Register GET /health/ready before the /api/* auth middleware so it is
public and reachable by K8s readinessProbe on port 3000 without auth.
On success → 200 {"status":"ready"}; on any DB/schema failure → 503
{"status":"degraded"}. Logs the pg error code; never leaks SQL in body.
A dropped schema (42P01) surfaces as non-200, closing the /health mask
that allowed the GRO-2678 incident to go undetected for ~43h.
Add health-ready.test.ts covering the 200 success path, 503 on schema
drop (42P01), and 503 on connection error; all assert no SQL leakage.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
- Register GET /api/readyz after /api/health in src/index.ts (before authMiddleware)
- Runs `getDb().select({id: staff.id}).from(staff).limit(1)` inside try/catch
- Success → 200 {"status":"ready"}; failure → 503 {"status":"degraded","check":"db"}
- Raw driver error is console.error'd but never leaked to the response body
- /health and /api/health are unchanged (K8s probes remain DB-less per CTO decision)
- New unit tests: mock getDb to resolve → assert 200/ready; throw → assert 503/degraded
- Updated UAT_PLAYBOOK.md §4.0 with TC-API-0.2 and TC-API-0.3
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Root cause: drizzle-orm's HWM-based migrate() marked migrations as applied in the ledger but rolled back schema changes when it detected a dirty state. Replacing with execSync("pnpm exec drizzle-kit migrate") matches the K8s migrate-schema Job, ensuring reset leaves a fully-migrated schema before seeding. Fixes reset-demo-data CronJob failures and resolves seed-test-data relation-does-not-exist reconcile loop.
QA approved: gb_lint (PR #229, head f7f90a7)
drizzle-orm's migrate() has a high-water-mark (HWM) bug: on a fresh DB,
migration 0000 sets the watermark to 2026-03-17. Migrations 0001, 0003,
0010, 0011 have stale 2025-era `when` timestamps and are silently
skipped. Migration 0003 (recurring_series) is the blocker — its skip
leaves `recurring_series`, `appointments.series_id`, and
`appointments.series_index` missing. A downstream migration inside
migrate()'s single Postgres transaction then fails, rolling back
everything including 0000's `staff` and `services` tables.
Replace the drizzle-orm migrate() call with `pnpm exec drizzle-kit
migrate` (hash-based). drizzle-kit applies every unhashed migration
regardless of `when` ordering, matching the K8s migrate Job exactly.
Merge origin/main into uat to resolve divergence (main was 27 commits
ahead). Only real conflict: UAT_PLAYBOOK.md — kept uat's §4.19 Boot
Resilience and §4.20 OOBE clients-from-auth test cases; excluded
.mcp.json and trigger-uat-1779751324.txt from merge result (CTO review
5042 — these must not land on main).
Co-Authored-By: Paperclip <noreply@paperclip.ing>
AbortSignal.timeout(5000) in initAuth's discovery fetch races with
vitest's 5000ms default test timeout, causing intermittent failures on CI push
runs. Stub fetch to return ok:false instantly so tests complete in <1s.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Previously the initAuth() function set authInitPromise to the async IIFE's
Promise but never cleared it on rejection. The index.ts retry loop called
initAuth() up to 10 times, but on attempt 2+ the check
`if (authInitPromise) { await authInitPromise; return; }` would immediately
re-throw the original rejection without performing a real retry.
Fix: wrap the final `await authInitPromise` in try/catch and reset
authInitPromise = null on error. This allows the index.ts retry loop to
create a fresh attempt on each call after a failure.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Previously the initAuth() function set authInitPromise to the async IIFE's
Promise but never cleared it on rejection. The index.ts retry loop called
initAuth() up to 10 times, but on attempt 2+ the check
`if (authInitPromise) { await authInitPromise; return; }` would immediately
re-throw the original rejection without performing a real retry.
Fix: wrap the final `await authInitPromise` in try/catch and reset
authInitPromise = null on error. This allows the index.ts retry loop to
create a fresh attempt on each call after a failure.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Merges feature/GRO-2652-boot-econnreset-resilience into dev. CI passed (run #440). Fixes PROD CrashLoopBackOff caused by bare top-level await initAuth() + unhandled ECONNRESET in auth_provider_config DB query.
Gitea OCI registry sporadically rejects cache blob writes with
"error writing layer blob: unknown" — this fails the build step
even though the image itself was pushed successfully. Adding
ignore-error=true makes cache write failures non-fatal so the
build proceeds regardless of registry-side cache issues.
Fixes recurring CI failure on Build and push Seed image step.
Previously: await initAuth() was a top-level ESM await. Any boot-time
ECONNRESET from Postgres propagated as an uncaught module-evaluation
error and killed the process before the server even started.
Now:
- serve() starts immediately so /health and public routes are available
- initAuth() runs in a retry loop (up to 10 attempts, exponential
500 ms → 30 s); a permanent failure degrades to 503 on auth routes
(existing catch-to-503 in authRouter) rather than crashing the pod
Co-Authored-By: Paperclip <noreply@paperclip.ing>
The auth_provider_config DB query at boot has no error handling;
a transient ECONNRESET causes authInitPromise to reject, propagating
to the top-level await initAuth() and crashing the process (exit 1).
Add up to 5 retry attempts with exponential backoff (1 s, 2 s, 4 s, 8 s)
so a single connection reset does not abort initialization.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
The OOBE flow on the web portal calls this endpoint to create a fresh
`clients` row bound to the Better Auth user's email when the SSO
bridge returns 404. Returns 201 on success, 409 if a client with that
email already exists (portal-selection case), 401/503 on auth issues,
400 on invalid body.
The OOBE success path navigates the user back to `/` and lets the
existing `session-from-auth` re-bridge; the new client is now
resolvable by email, so the bridge mints a real portal session.
Tests cover: 401 (no session), 400 (zod), 201 + persisted values
(name trimmed, optional fields normalized to null), 409 (existing
client or unique-constraint race), 503 (auth not configured).
Paired with the web PR on `feature/2357-p2-sso-to-oobe-routing`.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
(cherry picked from commit cdeebec021)
The OOBE flow on the web portal calls this endpoint to create a fresh
`clients` row bound to the Better Auth user's email when the SSO
bridge returns 404. Returns 201 on success, 409 if a client with that
email already exists (portal-selection case), 401/503 on auth issues,
400 on invalid body.
The OOBE success path navigates the user back to `/` and lets the
existing `session-from-auth` re-bridge; the new client is now
resolvable by email, so the bridge mints a real portal session.
Tests cover: 401 (no session), 400 (zod), 201 + persisted values
(name trimmed, optional fields normalized to null), 409 (existing
client or unique-constraint race), 503 (auth not configured).
Paired with the web PR on `feature/2357-p2-sso-to-oobe-routing`.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Resolves conflicts in UAT_PLAYBOOK.md, src/routes/portal.ts, and
src/__tests__/portal.test.ts (dev side wins — GRO-2342 changes are
the only diff in scope). Carries forward GRO-2139 reset.ts advisory
lock + GRO-2294 infra mcp trigger that were merged to dev but not
yet promoted to uat.
- src/routes/portal.ts: GET /portal/appointments now populates
service: {id, name} on both the synthetic waitlist card and the
appointment card (was {id} only). Same shape, no portal change
required.
- src/__tests__/portal.test.ts: services mock + TC-API-8.20 GRO-2342
assertions on the synthetic waitlist card service name.
- UAT_PLAYBOOK.md: TC-API-8.20 (GRO-2342) appended; TC-API-8.19
(GRO-2319) retained verbatim.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Cosmetic follow-up to GRO-2319 (Phase 4 review by CTO). The synthetic
waitlist card on GET /portal/appointments returned service: {id} only,
so the portal fell back to the literal 'Service' label. CMPO spec did
not call for a service name on the waitlist card, but populating the
real name is non-urgent and closes the cosmetic gap.
- src/routes/portal.ts: include a services SELECT (in addition to
pets and staff) covering both appointment and waitlist serviceIds.
serviceMap feeds a service.name lookup. The synthetic waitlist
card's service object is now {id, name} — same shape the
appointments join returns — so the portal renders the real name.
The appointments join also gains a name (consistent shape, no
regression for the existing path).
- src/__tests__/portal.test.ts: mock the services table and assert
service: {id, name} on both the synthetic waitlist card and the
appointment card.
- UAT_PLAYBOOK.md: TC-API-8.20 covering the waitlist card service
name (TC-API-8.19 retained verbatim for the original GRO-2319
surfacing contract).
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Serialize the entire db:reset chain (DROP → migrate → seed) inside one withSeedAdvisoryLock callback so a concurrent same-PRNG seeder cannot interleave and collide on invoices_pkey. Pool sized max:6 (1 reserved for the lock + work headroom) to avoid the connection-starvation deadlock the CTO caught. Verified with three end-to-end live db:reset runs against a throwaway Postgres.
cc @cpfarhood
| TC-API-0.3 | DB-touching readiness check — response body safe | GET /api/readyz and inspect body | Body contains only `status` and (on error) `check` fields — no raw SQL, driver messages, or stack traces |
> **Note (GRO-1544):** Health endpoint registered on `api` basePath before auth middleware at `/api/health`. The old path `/health` was incorrect (routed to web pod via HTTPRoute `/*` rule).
> **Note (GRO-2678):** `/api/readyz` is a separate DB-touching endpoint for monitoring. It is intentionally NOT used for K8s liveness/readiness probes — those remain DB-less to avoid pod cycling on transient DB blips.
### 4.1 Authentication
@@ -439,6 +442,38 @@ Both use the stops' stored `latitude`/`longitude` in `stopOrder`: **origin = fir
| TC-API-18.10 | Groomer cannot export another's route | As groomer, export a route owned by a different groomer | 403 Forbidden (`groomers may only access their own route`) |
| TC-API-18.11 | Receptionist denied | As **receptionist**, export any route | 403 Forbidden (role not permitted) |
Verifies the API process does not crash on transient boot-time DB connection resets and that auth routes degrade gracefully until initialization succeeds.
| TC | Test Case | Steps | Expected Result |
|----|-----------|-------|-----------------|
| TC-API-19.1 | Health endpoint available before auth init | 1. Deploy the image (or restart the api pod)<br>2. `GET /health` immediately (within first 2 s of pod start) | 200 `{"status":"ok"}` — server accepts requests before `initAuth()` completes |
| TC-API-19.2 | Auth routes return 503 when auth not yet initialized | 1. Temporarily set `OIDC_ISSUER` to an unreachable host so `initAuth()` keeps retrying<br>2. `POST /api/auth/sign-in/email` during the retry window | 503 `{"error":"Authentication not configured"}` — process stays alive, does not exit |
| TC-API-19.3 | Pod does not crash on first-attempt DB reset | 1. Review pod restart count after normal deployment<br>2. Confirm `kubectl get pod -n groombook` shows `RESTARTS: 0` (or same as before deploy) for the new pod | No new restarts — ECONNRESET causes retry, not process exit |
| TC-API-19.4 | DB query retry log lines visible | After deploy, `kubectl logs -n groombook <api-pod>` | If any DB retry occurred, log lines matching `[auth] DB query attempt N failed` are present; on clean boot no retry lines appear |
| TC-API-19.5 | Auth init retry log lines visible | When auth init fails and retries, check pod logs | Log lines matching `[auth] initAuth attempt N failed` present; process continues; no `process.exit` |
| TC-API-19.6 | Auth succeeds after transient DB hiccup | 1. Allow pod to retry until DB is available<br>2. `POST /api/auth/sign-in/email` with valid credentials after init succeeds | 200 with session cookie — auth recovers without pod restart |
| TC-API-19.7 | Normal sign-in still works end-to-end | Follow TC-WEB-SSO-3 (SSO sign-in) on UAT | Successful sign-in, staff list visible — no regression from resilience changes |
| TC-API-19.8 | Public routes unaffected during auth retry | While auth is retrying (TC-API-19.2 setup), `GET /api/branding` | 200 with branding data — public routes bypass auth and serve normally |
### 4.20 Portal OOBE — Create Client from Auth (GRO-2359)
Verifies the `POST /api/portal/clients-from-auth` endpoint that creates a new `clients` row for a first-time SSO user (out-of-box-experience registration). This endpoint requires a valid Better Auth session but does NOT require a portal session; it is the pre-portal step in the new-user OOBE flow.
| TC | Test Case | Steps | Expected Result |
|----|-----------|-------|-----------------|
| TC-API-20.1 | Successful client creation | 1. Sign in via SSO to obtain a Better Auth session<br>2. `POST /api/portal/clients-from-auth` with `{ "name": "Test User" }` | 201 `{ "id": "<uuid>", "name": "Test User", "email": "<sso-email>" }` — new `clients` row created |
| TC-API-20.2 | All optional fields accepted | `POST /api/portal/clients-from-auth` with `{ "name": "Test User", "phone": "555-1234", "address": "1 Main St", "notes": "VIP" }` (authenticated) | 201 with `id`, `name`, `email`; row in DB has all four fields |
| TC-API-20.3 | Invalid body — missing name | `POST /api/portal/clients-from-auth` with `{}` (authenticated) | 400 (Zod validation failure); no row created |
| TC-API-20.4 | Invalid body — empty name | `POST /api/portal/clients-from-auth` with `{ "name": "" }` (authenticated) | 400 — name must be at least 1 character |
| TC-API-20.5 | No session — 401 | `POST /api/portal/clients-from-auth` with a valid body but **no** Better Auth session cookie | 401 `{ "error": "Unauthorized" }` |
| TC-API-20.6 | Existing email — 409 | 1. Create a client row whose email matches the signed-in SSO user's email<br>2. `POST /api/portal/clients-from-auth` as that user | 409 `{ "error": "A customer record with this email already exists" }` — no duplicate row |
| TC-API-20.7 | Auth not configured — 503 | Temporarily disable auth (e.g., point `OIDC_ISSUER` to an invalid host) so `getAuth()` throws<br>2. `POST /api/portal/clients-from-auth` | 503 `{ "error": "Authentication not configured" }` — graceful degradation |
| TC-API-20.8 | Concurrent insert race — 409 | Simulate two near-simultaneous requests from the same SSO user (before any row exists) | At most one request returns 201; the other returns 409 — no duplicate row, no 500 |
## Pass/Fail Criteria
**Pass:**
@@ -460,3 +495,4 @@ Both use the stops' stored `latitude`/`longitude` in `stopOrder`: **origin = fir
## Update Policy
Any PR that changes user-facing behaviour MUST update this file. Test cases must be added, modified, or removed to reflect the new behaviour. The PR description must reference which playbook section was updated (e.g., "Updated UAT_PLAYBOOK.md §4.4 — new appointment rescheduling flow").
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.