/api/readyz is a public monitoring endpoint like /api/health.
app.basePath("/api") + api.use("*", authMiddleware) applies to all
/api/* paths regardless of route registration order (GRO-2692 UAT
failure: Hono basePath middleware bypass).
Co-Authored-By: Paperclip <noreply@paperclip.ing>
/api/readyz was returning 401 Unauthorized in UAT (GRO-2692 Shedward
regression). The authMiddleware for /api/* did not whitelist /api/readyz,
causing every unauthenticated probe request to be rejected.
Add /api/readyz to the bypass condition alongside /api/auth/* and
/api/health. Add readyz-auth-bypass.test.ts confirming the route passes
through the middleware without auth and that non-whitelisted paths still
block (503 when auth not configured).
Closes GRO-2692 regression (UAT fail).
CTO architectural ruling (PR #235 review): K8s readiness/liveness probes
must remain DB-less to prevent transient DB blips from cycling pods. The
/api/readyz endpoint (GRO-2687) is the canonical DB-health signal for the
monitoring layer and satisfies the GRO-2678 detection-gap requirement.
Removes:
- GET /health/ready route from src/index.ts
- src/__tests__/health-ready.test.ts
Refs GRO-2689, GRO-2687, GRO-2678
Register GET /health/ready before the /api/* auth middleware so it is
public and reachable by K8s readinessProbe on port 3000 without auth.
On success → 200 {"status":"ready"}; on any DB/schema failure → 503
{"status":"degraded"}. Logs the pg error code; never leaks SQL in body.
A dropped schema (42P01) surfaces as non-200, closing the /health mask
that allowed the GRO-2678 incident to go undetected for ~43h.
Add health-ready.test.ts covering the 200 success path, 503 on schema
drop (42P01), and 503 on connection error; all assert no SQL leakage.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
- Register GET /api/readyz after /api/health in src/index.ts (before authMiddleware)
- Runs `getDb().select({id: staff.id}).from(staff).limit(1)` inside try/catch
- Success → 200 {"status":"ready"}; failure → 503 {"status":"degraded","check":"db"}
- Raw driver error is console.error'd but never leaked to the response body
- /health and /api/health are unchanged (K8s probes remain DB-less per CTO decision)
- New unit tests: mock getDb to resolve → assert 200/ready; throw → assert 503/degraded
- Updated UAT_PLAYBOOK.md §4.0 with TC-API-0.2 and TC-API-0.3
Co-Authored-By: Paperclip <noreply@paperclip.ing>
drizzle-orm's migrate() has a high-water-mark (HWM) bug: on a fresh DB,
migration 0000 sets the watermark to 2026-03-17. Migrations 0001, 0003,
0010, 0011 have stale 2025-era `when` timestamps and are silently
skipped. Migration 0003 (recurring_series) is the blocker — its skip
leaves `recurring_series`, `appointments.series_id`, and
`appointments.series_index` missing. A downstream migration inside
migrate()'s single Postgres transaction then fails, rolling back
everything including 0000's `staff` and `services` tables.
Replace the drizzle-orm migrate() call with `pnpm exec drizzle-kit
migrate` (hash-based). drizzle-kit applies every unhashed migration
regardless of `when` ordering, matching the K8s migrate Job exactly.
AbortSignal.timeout(5000) in initAuth's discovery fetch races with
vitest's 5000ms default test timeout, causing intermittent failures on CI push
runs. Stub fetch to return ok:false instantly so tests complete in <1s.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Previously the initAuth() function set authInitPromise to the async IIFE's
Promise but never cleared it on rejection. The index.ts retry loop called
initAuth() up to 10 times, but on attempt 2+ the check
`if (authInitPromise) { await authInitPromise; return; }` would immediately
re-throw the original rejection without performing a real retry.
Fix: wrap the final `await authInitPromise` in try/catch and reset
authInitPromise = null on error. This allows the index.ts retry loop to
create a fresh attempt on each call after a failure.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Previously the initAuth() function set authInitPromise to the async IIFE's
Promise but never cleared it on rejection. The index.ts retry loop called
initAuth() up to 10 times, but on attempt 2+ the check
`if (authInitPromise) { await authInitPromise; return; }` would immediately
re-throw the original rejection without performing a real retry.
Fix: wrap the final `await authInitPromise` in try/catch and reset
authInitPromise = null on error. This allows the index.ts retry loop to
create a fresh attempt on each call after a failure.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Merges feature/GRO-2652-boot-econnreset-resilience into dev. CI passed (run #440). Fixes PROD CrashLoopBackOff caused by bare top-level await initAuth() + unhandled ECONNRESET in auth_provider_config DB query.
2026-08-05 09:01:12 +00:00
10 changed files with 378 additions and 7 deletions
| TC-API-0.3 | DB-touching readiness check — response body safe | GET /api/readyz and inspect body | Body contains only `status` and (on error) `check` fields — no raw SQL, driver messages, or stack traces |
> **Note (GRO-1544):** Health endpoint registered on `api` basePath before auth middleware at `/api/health`. The old path `/health` was incorrect (routed to web pod via HTTPRoute `/*` rule).
> **Note (GRO-2678):** `/api/readyz` is a separate DB-touching endpoint for monitoring. It is intentionally NOT used for K8s liveness/readiness probes — those remain DB-less to avoid pod cycling on transient DB blips.
### 4.1 Authentication
@@ -455,6 +458,22 @@ Verifies the API process does not crash on transient boot-time DB connection res
| TC-API-19.7 | Normal sign-in still works end-to-end | Follow TC-WEB-SSO-3 (SSO sign-in) on UAT | Successful sign-in, staff list visible — no regression from resilience changes |
| TC-API-19.8 | Public routes unaffected during auth retry | While auth is retrying (TC-API-19.2 setup), `GET /api/branding` | 200 with branding data — public routes bypass auth and serve normally |
### 4.20 Portal OOBE — Create Client from Auth (GRO-2359)
Verifies the `POST /api/portal/clients-from-auth` endpoint that creates a new `clients` row for a first-time SSO user (out-of-box-experience registration). This endpoint requires a valid Better Auth session but does NOT require a portal session; it is the pre-portal step in the new-user OOBE flow.
| TC | Test Case | Steps | Expected Result |
|----|-----------|-------|-----------------|
| TC-API-20.1 | Successful client creation | 1. Sign in via SSO to obtain a Better Auth session<br>2. `POST /api/portal/clients-from-auth` with `{ "name": "Test User" }` | 201 `{ "id": "<uuid>", "name": "Test User", "email": "<sso-email>" }` — new `clients` row created |
| TC-API-20.2 | All optional fields accepted | `POST /api/portal/clients-from-auth` with `{ "name": "Test User", "phone": "555-1234", "address": "1 Main St", "notes": "VIP" }` (authenticated) | 201 with `id`, `name`, `email`; row in DB has all four fields |
| TC-API-20.3 | Invalid body — missing name | `POST /api/portal/clients-from-auth` with `{}` (authenticated) | 400 (Zod validation failure); no row created |
| TC-API-20.4 | Invalid body — empty name | `POST /api/portal/clients-from-auth` with `{ "name": "" }` (authenticated) | 400 — name must be at least 1 character |
| TC-API-20.5 | No session — 401 | `POST /api/portal/clients-from-auth` with a valid body but **no** Better Auth session cookie | 401 `{ "error": "Unauthorized" }` |
| TC-API-20.6 | Existing email — 409 | 1. Create a client row whose email matches the signed-in SSO user's email<br>2. `POST /api/portal/clients-from-auth` as that user | 409 `{ "error": "A customer record with this email already exists" }` — no duplicate row |
| TC-API-20.7 | Auth not configured — 503 | Temporarily disable auth (e.g., point `OIDC_ISSUER` to an invalid host) so `getAuth()` throws<br>2. `POST /api/portal/clients-from-auth` | 503 `{ "error": "Authentication not configured" }` — graceful degradation |
| TC-API-20.8 | Concurrent insert race — 409 | Simulate two near-simultaneous requests from the same SSO user (before any row exists) | At most one request returns 201; the other returns 409 — no duplicate row, no 500 |
## Pass/Fail Criteria
**Pass:**
@@ -476,3 +495,4 @@ Verifies the API process does not crash on transient boot-time DB connection res
## Update Policy
Any PR that changes user-facing behaviour MUST update this file. Test cases must be added, modified, or removed to reflect the new behaviour. The PR description must reference which playbook section was updated (e.g., "Updated UAT_PLAYBOOK.md §4.4 — new appointment rescheduling flow").
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.