fix(GRO-2652): boot ECONNRESET resilience — retry-with-backoff in initAuth, server-first startup
Merges feature/GRO-2652-boot-econnreset-resilience into dev. CI passed (run #440). Fixes PROD CrashLoopBackOff caused by bare top-level await initAuth() + unhandled ECONNRESET in auth_provider_config DB query.
This commit was merged in pull request #222.
This commit is contained in:
+1
-187
File diff suppressed because one or more lines are too long
@@ -439,6 +439,22 @@ Both use the stops' stored `latitude`/`longitude` in `stopOrder`: **origin = fir
|
||||
| TC-API-18.10 | Groomer cannot export another's route | As groomer, export a route owned by a different groomer | 403 Forbidden (`groomers may only access their own route`) |
|
||||
| TC-API-18.11 | Receptionist denied | As **receptionist**, export any route | 403 Forbidden (role not permitted) |
|
||||
|
||||
|
||||
### 4.19 Boot Resilience — ECONNRESET Recovery (GRO-2652)
|
||||
|
||||
Verifies the API process does not crash on transient boot-time DB connection resets and that auth routes degrade gracefully until initialization succeeds.
|
||||
|
||||
| TC | Test Case | Steps | Expected Result |
|
||||
|----|-----------|-------|-----------------|
|
||||
| TC-API-19.1 | Health endpoint available before auth init | 1. Deploy the image (or restart the api pod)<br>2. `GET /health` immediately (within first 2 s of pod start) | 200 `{"status":"ok"}` — server accepts requests before `initAuth()` completes |
|
||||
| TC-API-19.2 | Auth routes return 503 when auth not yet initialized | 1. Temporarily set `OIDC_ISSUER` to an unreachable host so `initAuth()` keeps retrying<br>2. `POST /api/auth/sign-in/email` during the retry window | 503 `{"error":"Authentication not configured"}` — process stays alive, does not exit |
|
||||
| TC-API-19.3 | Pod does not crash on first-attempt DB reset | 1. Review pod restart count after normal deployment<br>2. Confirm `kubectl get pod -n groombook` shows `RESTARTS: 0` (or same as before deploy) for the new pod | No new restarts — ECONNRESET causes retry, not process exit |
|
||||
| TC-API-19.4 | DB query retry log lines visible | After deploy, `kubectl logs -n groombook <api-pod>` | If any DB retry occurred, log lines matching `[auth] DB query attempt N failed` are present; on clean boot no retry lines appear |
|
||||
| TC-API-19.5 | Auth init retry log lines visible | When auth init fails and retries, check pod logs | Log lines matching `[auth] initAuth attempt N failed` present; process continues; no `process.exit` |
|
||||
| TC-API-19.6 | Auth succeeds after transient DB hiccup | 1. Allow pod to retry until DB is available<br>2. `POST /api/auth/sign-in/email` with valid credentials after init succeeds | 200 with session cookie — auth recovers without pod restart |
|
||||
| TC-API-19.7 | Normal sign-in still works end-to-end | Follow TC-WEB-SSO-3 (SSO sign-in) on UAT | Successful sign-in, staff list visible — no regression from resilience changes |
|
||||
| TC-API-19.8 | Public routes unaffected during auth retry | While auth is retrying (TC-API-19.2 setup), `GET /api/branding` | 200 with branding data — public routes bypass auth and serve normally |
|
||||
|
||||
## Pass/Fail Criteria
|
||||
|
||||
**Pass:**
|
||||
|
||||
+22
-2
@@ -292,14 +292,34 @@ api.route("/search", searchRouter);
|
||||
api.route("/buffer-rules", bufferRulesRouter);
|
||||
api.route("/routes", routesRouter);
|
||||
|
||||
// Start the HTTP server first so /health and public routes are available immediately.
|
||||
// Auth initialization runs afterward with retry — a transient DB ECONNRESET at boot
|
||||
// must not crash the process (GRO-2652). Auth routes return 503 until initAuth succeeds.
|
||||
const port = Number(process.env.PORT ?? 3000);
|
||||
await initAuth();
|
||||
console.log(`API server listening on port ${port}`);
|
||||
const server = serve({ fetch: app.fetch, port });
|
||||
console.log(`API server listening on port ${port}`);
|
||||
|
||||
// Start background reminder scheduler (runs every minute to check for upcoming appointments)
|
||||
startReminderScheduler();
|
||||
|
||||
let initAttempt = 0;
|
||||
while (true) {
|
||||
try {
|
||||
await initAuth();
|
||||
break;
|
||||
} catch (err) {
|
||||
initAttempt++;
|
||||
const delay = Math.min(2 ** initAttempt * 500, 30_000);
|
||||
console.error(`[auth] initAuth attempt ${initAttempt} failed: ${err}`);
|
||||
if (initAttempt >= 10) {
|
||||
console.error("[auth] auth init permanently failed — auth endpoints will serve 503");
|
||||
break;
|
||||
}
|
||||
console.error(`[auth] retrying in ${delay}ms`);
|
||||
await new Promise((r) => setTimeout(r, delay));
|
||||
}
|
||||
}
|
||||
|
||||
function shutdown() {
|
||||
console.log("Shutting down gracefully...");
|
||||
// SIGTERM/SIGINT → server.close() → callback → process.exit(0)
|
||||
|
||||
+17
-2
@@ -124,13 +124,28 @@ export async function initAuth(): Promise<void> {
|
||||
return;
|
||||
}
|
||||
|
||||
// Step 1: Try to load config from DB
|
||||
// Step 1: Try to load config from DB, with retry-with-backoff for transient ECONNRESET (GRO-2652).
|
||||
// A single connection reset during boot must not abort initialization.
|
||||
const db = getDb();
|
||||
const [dbConfig] = await db
|
||||
let dbQueryRows: (typeof authProviderConfig.$inferSelect)[] = [];
|
||||
let dbAttempt = 0;
|
||||
while (true) {
|
||||
try {
|
||||
dbQueryRows = await db
|
||||
.select()
|
||||
.from(authProviderConfig)
|
||||
.where(eq(authProviderConfig.enabled, true))
|
||||
.limit(1);
|
||||
break;
|
||||
} catch (err) {
|
||||
dbAttempt++;
|
||||
if (dbAttempt >= 5) throw err;
|
||||
const delay = Math.min(1000 * 2 ** (dbAttempt - 1), 8_000);
|
||||
console.warn(`[auth] DB query attempt ${dbAttempt} failed (${err}), retrying in ${delay}ms`);
|
||||
await new Promise((r) => setTimeout(r, delay));
|
||||
}
|
||||
}
|
||||
const [dbConfig] = dbQueryRows;
|
||||
|
||||
let providerConfig: {
|
||||
providerId: string;
|
||||
|
||||
Reference in New Issue
Block a user