fix(GRO-2652): boot ECONNRESET resilience — retry-with-backoff in initAuth, server-first startup

Merges feature/GRO-2652-boot-econnreset-resilience into dev. CI passed (run #440). Fixes PROD CrashLoopBackOff caused by bare top-level await initAuth() + unhandled ECONNRESET in auth_provider_config DB query.
This commit was merged in pull request #222.
This commit is contained in:
2026-08-05 09:01:12 +00:00
4 changed files with 60 additions and 195 deletions
+1 -187
View File
File diff suppressed because one or more lines are too long
+16
View File
@@ -439,6 +439,22 @@ Both use the stops' stored `latitude`/`longitude` in `stopOrder`: **origin = fir
| TC-API-18.10 | Groomer cannot export another's route | As groomer, export a route owned by a different groomer | 403 Forbidden (`groomers may only access their own route`) |
| TC-API-18.11 | Receptionist denied | As **receptionist**, export any route | 403 Forbidden (role not permitted) |
### 4.19 Boot Resilience — ECONNRESET Recovery (GRO-2652)
Verifies the API process does not crash on transient boot-time DB connection resets and that auth routes degrade gracefully until initialization succeeds.
| TC | Test Case | Steps | Expected Result |
|----|-----------|-------|-----------------|
| TC-API-19.1 | Health endpoint available before auth init | 1. Deploy the image (or restart the api pod)<br>2. `GET /health` immediately (within first 2 s of pod start) | 200 `{"status":"ok"}` — server accepts requests before `initAuth()` completes |
| TC-API-19.2 | Auth routes return 503 when auth not yet initialized | 1. Temporarily set `OIDC_ISSUER` to an unreachable host so `initAuth()` keeps retrying<br>2. `POST /api/auth/sign-in/email` during the retry window | 503 `{"error":"Authentication not configured"}` — process stays alive, does not exit |
| TC-API-19.3 | Pod does not crash on first-attempt DB reset | 1. Review pod restart count after normal deployment<br>2. Confirm `kubectl get pod -n groombook` shows `RESTARTS: 0` (or same as before deploy) for the new pod | No new restarts — ECONNRESET causes retry, not process exit |
| TC-API-19.4 | DB query retry log lines visible | After deploy, `kubectl logs -n groombook <api-pod>` | If any DB retry occurred, log lines matching `[auth] DB query attempt N failed` are present; on clean boot no retry lines appear |
| TC-API-19.5 | Auth init retry log lines visible | When auth init fails and retries, check pod logs | Log lines matching `[auth] initAuth attempt N failed` present; process continues; no `process.exit` |
| TC-API-19.6 | Auth succeeds after transient DB hiccup | 1. Allow pod to retry until DB is available<br>2. `POST /api/auth/sign-in/email` with valid credentials after init succeeds | 200 with session cookie — auth recovers without pod restart |
| TC-API-19.7 | Normal sign-in still works end-to-end | Follow TC-WEB-SSO-3 (SSO sign-in) on UAT | Successful sign-in, staff list visible — no regression from resilience changes |
| TC-API-19.8 | Public routes unaffected during auth retry | While auth is retrying (TC-API-19.2 setup), `GET /api/branding` | 200 with branding data — public routes bypass auth and serve normally |
## Pass/Fail Criteria
**Pass:**
+22 -2
View File
@@ -292,14 +292,34 @@ api.route("/search", searchRouter);
api.route("/buffer-rules", bufferRulesRouter);
api.route("/routes", routesRouter);
// Start the HTTP server first so /health and public routes are available immediately.
// Auth initialization runs afterward with retry — a transient DB ECONNRESET at boot
// must not crash the process (GRO-2652). Auth routes return 503 until initAuth succeeds.
const port = Number(process.env.PORT ?? 3000);
await initAuth();
console.log(`API server listening on port ${port}`);
const server = serve({ fetch: app.fetch, port });
console.log(`API server listening on port ${port}`);
// Start background reminder scheduler (runs every minute to check for upcoming appointments)
startReminderScheduler();
let initAttempt = 0;
while (true) {
try {
await initAuth();
break;
} catch (err) {
initAttempt++;
const delay = Math.min(2 ** initAttempt * 500, 30_000);
console.error(`[auth] initAuth attempt ${initAttempt} failed: ${err}`);
if (initAttempt >= 10) {
console.error("[auth] auth init permanently failed — auth endpoints will serve 503");
break;
}
console.error(`[auth] retrying in ${delay}ms`);
await new Promise((r) => setTimeout(r, delay));
}
}
function shutdown() {
console.log("Shutting down gracefully...");
// SIGTERM/SIGINT → server.close() → callback → process.exit(0)
+17 -2
View File
@@ -124,13 +124,28 @@ export async function initAuth(): Promise<void> {
return;
}
// Step 1: Try to load config from DB
// Step 1: Try to load config from DB, with retry-with-backoff for transient ECONNRESET (GRO-2652).
// A single connection reset during boot must not abort initialization.
const db = getDb();
const [dbConfig] = await db
let dbQueryRows: (typeof authProviderConfig.$inferSelect)[] = [];
let dbAttempt = 0;
while (true) {
try {
dbQueryRows = await db
.select()
.from(authProviderConfig)
.where(eq(authProviderConfig.enabled, true))
.limit(1);
break;
} catch (err) {
dbAttempt++;
if (dbAttempt >= 5) throw err;
const delay = Math.min(1000 * 2 ** (dbAttempt - 1), 8_000);
console.warn(`[auth] DB query attempt ${dbAttempt} failed (${err}), retrying in ${delay}ms`);
await new Promise((r) => setTimeout(r, delay));
}
}
const [dbConfig] = dbQueryRows;
let providerConfig: {
providerId: string;