feat(api): add DB-touching /health/ready readiness probe (GRO-2678)
CI / Lint & Typecheck (pull_request) Successful in 19s
CI / Test (pull_request) Successful in 21s
CI / Build & Push Docker Images (pull_request) Successful in 3m25s

Register GET /health/ready before the /api/* auth middleware so it is
public and reachable by K8s readinessProbe on port 3000 without auth.
On success → 200 {"status":"ready"}; on any DB/schema failure → 503
{"status":"degraded"}. Logs the pg error code; never leaks SQL in body.
A dropped schema (42P01) surfaces as non-200, closing the /health mask
that allowed the GRO-2678 incident to go undetected for ~43h.

Add health-ready.test.ts covering the 200 success path, 503 on schema
drop (42P01), and 503 on connection error; all assert no SQL leakage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
This commit is contained in:
Flea Flicker
2026-08-09 09:21:56 +00:00
parent a164c24e8e
commit 7679fada0a
2 changed files with 113 additions and 0 deletions
+11
View File
@@ -65,6 +65,17 @@ app.use(
app.get("/health", (c) => c.json({ status: "ok" }));
// /api/health: used by Gateway HTTPRoute (/api/* → API pod)
app.get("/api/health", (c) => c.json({ status: "ok" }));
// /health/ready: DB-touching readiness probe — K8s removes pod from endpoints when schema is dropped (GRO-2689)
app.get("/health/ready", async (c) => {
try {
await getDb().select({ id: staff.id }).from(staff).limit(1);
return c.json({ status: "ready" }, 200);
} catch (err) {
const pgCode = (err as Record<string, unknown>).code ?? "unknown";
console.error("[health/ready] DB check failed:", pgCode);
return c.json({ status: "degraded" }, 503);
}
});
// /api/readyz: DB-touching deep health check consumed by monitoring (not K8s probes)
// Distinct from /health so a dropped schema triggers an alert without cycling pods (GRO-2678)
app.get("/api/readyz", async (c) => {