Compare commits

..

6 Commits

Author SHA1 Message Date
gb_flea af1af5c211 ci: trigger CI run for ignore-error=true cache-to fix 2026-08-05 08:56:30 +00:00
Flea Flicker 713e8bf415 ci: add ignore-error=true to all cache-to targets
Gitea OCI registry sporadically rejects cache blob writes with
"error writing layer blob: unknown" — this fails the build step
even though the image itself was pushed successfully. Adding
ignore-error=true makes cache write failures non-fatal so the
build proceeds regardless of registry-side cache issues.

Fixes recurring CI failure on Build and push Seed image step.
2026-08-05 08:54:00 +00:00
Flea Flicker 8d42bfffc9 test(GRO-2652): add boot resilience UAT test cases (TC-API-19.x)
CI / Test (pull_request) Successful in 25s
CI / Lint & Typecheck (pull_request) Successful in 28s
CI / Build & Push Docker Images (pull_request) Successful in 56s
New section 4.19 verifies:
- /health available before initAuth completes
- auth routes return 503 (not crash) during init retry window
- pod restart count stays stable after deploy
- retry log lines emitted correctly
- auth recovers after transient DB hiccup

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-05 08:22:00 +00:00
Flea Flicker 095efb1cca fix(GRO-2652): start server before initAuth; retry auth init on ECONNRESET
Previously: await initAuth() was a top-level ESM await. Any boot-time
ECONNRESET from Postgres propagated as an uncaught module-evaluation
error and killed the process before the server even started.

Now:
- serve() starts immediately so /health and public routes are available
- initAuth() runs in a retry loop (up to 10 attempts, exponential
  500 ms → 30 s); a permanent failure degrades to 503 on auth routes
  (existing catch-to-503 in authRouter) rather than crashing the pod

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-05 08:21:14 +00:00
Flea Flicker 632dadd064 fix(GRO-2652): retry DB query in initAuth on transient ECONNRESET
The auth_provider_config DB query at boot has no error handling;
a transient ECONNRESET causes authInitPromise to reject, propagating
to the top-level await initAuth() and crashing the process (exit 1).

Add up to 5 retry attempts with exponential backoff (1 s, 2 s, 4 s, 8 s)
so a single connection reset does not abort initialization.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-05 08:17:31 +00:00
Flea Flicker dace2c4e66 fix(GRO-2586): enforce trusted-origins allowlist on Better Auth CORS responses (#219)
CI / Lint & Typecheck (push) Successful in 19s
CI / Test (push) Successful in 21s
CI / Build & Push Docker Images (push) Successful in 1m4s
fix(GRO-2586): enforce trusted-origins allowlist on Better Auth CORS responses

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-06-26 13:35:59 +00:00
4 changed files with 60 additions and 195 deletions
+1 -187
View File
File diff suppressed because one or more lines are too long
+16
View File
@@ -439,6 +439,22 @@ Both use the stops' stored `latitude`/`longitude` in `stopOrder`: **origin = fir
| TC-API-18.10 | Groomer cannot export another's route | As groomer, export a route owned by a different groomer | 403 Forbidden (`groomers may only access their own route`) |
| TC-API-18.11 | Receptionist denied | As **receptionist**, export any route | 403 Forbidden (role not permitted) |
### 4.19 Boot Resilience — ECONNRESET Recovery (GRO-2652)
Verifies the API process does not crash on transient boot-time DB connection resets and that auth routes degrade gracefully until initialization succeeds.
| TC | Test Case | Steps | Expected Result |
|----|-----------|-------|-----------------|
| TC-API-19.1 | Health endpoint available before auth init | 1. Deploy the image (or restart the api pod)<br>2. `GET /health` immediately (within first 2 s of pod start) | 200 `{"status":"ok"}` — server accepts requests before `initAuth()` completes |
| TC-API-19.2 | Auth routes return 503 when auth not yet initialized | 1. Temporarily set `OIDC_ISSUER` to an unreachable host so `initAuth()` keeps retrying<br>2. `POST /api/auth/sign-in/email` during the retry window | 503 `{"error":"Authentication not configured"}` — process stays alive, does not exit |
| TC-API-19.3 | Pod does not crash on first-attempt DB reset | 1. Review pod restart count after normal deployment<br>2. Confirm `kubectl get pod -n groombook` shows `RESTARTS: 0` (or same as before deploy) for the new pod | No new restarts — ECONNRESET causes retry, not process exit |
| TC-API-19.4 | DB query retry log lines visible | After deploy, `kubectl logs -n groombook <api-pod>` | If any DB retry occurred, log lines matching `[auth] DB query attempt N failed` are present; on clean boot no retry lines appear |
| TC-API-19.5 | Auth init retry log lines visible | When auth init fails and retries, check pod logs | Log lines matching `[auth] initAuth attempt N failed` present; process continues; no `process.exit` |
| TC-API-19.6 | Auth succeeds after transient DB hiccup | 1. Allow pod to retry until DB is available<br>2. `POST /api/auth/sign-in/email` with valid credentials after init succeeds | 200 with session cookie — auth recovers without pod restart |
| TC-API-19.7 | Normal sign-in still works end-to-end | Follow TC-WEB-SSO-3 (SSO sign-in) on UAT | Successful sign-in, staff list visible — no regression from resilience changes |
| TC-API-19.8 | Public routes unaffected during auth retry | While auth is retrying (TC-API-19.2 setup), `GET /api/branding` | 200 with branding data — public routes bypass auth and serve normally |
## Pass/Fail Criteria
**Pass:**
+22 -2
View File
@@ -292,14 +292,34 @@ api.route("/search", searchRouter);
api.route("/buffer-rules", bufferRulesRouter);
api.route("/routes", routesRouter);
// Start the HTTP server first so /health and public routes are available immediately.
// Auth initialization runs afterward with retry — a transient DB ECONNRESET at boot
// must not crash the process (GRO-2652). Auth routes return 503 until initAuth succeeds.
const port = Number(process.env.PORT ?? 3000);
await initAuth();
console.log(`API server listening on port ${port}`);
const server = serve({ fetch: app.fetch, port });
console.log(`API server listening on port ${port}`);
// Start background reminder scheduler (runs every minute to check for upcoming appointments)
startReminderScheduler();
let initAttempt = 0;
while (true) {
try {
await initAuth();
break;
} catch (err) {
initAttempt++;
const delay = Math.min(2 ** initAttempt * 500, 30_000);
console.error(`[auth] initAuth attempt ${initAttempt} failed: ${err}`);
if (initAttempt >= 10) {
console.error("[auth] auth init permanently failed — auth endpoints will serve 503");
break;
}
console.error(`[auth] retrying in ${delay}ms`);
await new Promise((r) => setTimeout(r, delay));
}
}
function shutdown() {
console.log("Shutting down gracefully...");
// SIGTERM/SIGINT → server.close() → callback → process.exit(0)
+21 -6
View File
@@ -124,13 +124,28 @@ export async function initAuth(): Promise<void> {
return;
}
// Step 1: Try to load config from DB
// Step 1: Try to load config from DB, with retry-with-backoff for transient ECONNRESET (GRO-2652).
// A single connection reset during boot must not abort initialization.
const db = getDb();
const [dbConfig] = await db
.select()
.from(authProviderConfig)
.where(eq(authProviderConfig.enabled, true))
.limit(1);
let dbQueryRows: (typeof authProviderConfig.$inferSelect)[] = [];
let dbAttempt = 0;
while (true) {
try {
dbQueryRows = await db
.select()
.from(authProviderConfig)
.where(eq(authProviderConfig.enabled, true))
.limit(1);
break;
} catch (err) {
dbAttempt++;
if (dbAttempt >= 5) throw err;
const delay = Math.min(1000 * 2 ** (dbAttempt - 1), 8_000);
console.warn(`[auth] DB query attempt ${dbAttempt} failed (${err}), retrying in ${delay}ms`);
await new Promise((r) => setTimeout(r, delay));
}
}
const [dbConfig] = dbQueryRows;
let providerConfig: {
providerId: string;