diff --git a/docs/shutdown-gpf-analysis.md b/docs/shutdown-gpf-analysis.md index 6c9df74..9c04b3b 100644 --- a/docs/shutdown-gpf-analysis.md +++ b/docs/shutdown-gpf-analysis.md @@ -78,3 +78,35 @@ candidate fix (pending core confirmation of #1): after `log.info("relay stopped cleanly")`, bypass the racy libc teardown with an explicit `_exit(0)` — the process is terminating anyway and all durable state (event log, pg) is already flushed by the component deinits that ran above. + +## RESOLUTION (2026-06-14) + +**core symbolized** (operator, `docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md`): +the crash is a "third door" — neither `free` nor the exit machinery itself, but +a **worker thread** faulting inside glibc locale parsing +(`__GI_____strtold_l_internal`) at full speed. the all-threads histogram shows +all ~2,800 backend threads alive at crash, ~1,066 clustered in glibc +getaddrinfo/DNS (`add_region`/`getifaddrs`/`strverscmp`) — host workers +mid-reconnect during teardown. mechanism = suspect #1 confirmed: at SIGTERM, +glibc exit-time teardown frees shared libc state (locale data) while live +backend threads are still grinding DNS/locale paths; one dereferences the freed +state first and takes the SIGSEGV (kernel logs a GPF). not in zlay/zat code. + +**fix implemented** (`main.zig`): `main()` now calls `runRelay()` (the former +body, with all its defers — so the event-log flush in `DiskPersist.deinit` and +the pg disconnect still run), then issues the raw **`exit_group(2)`** syscall +(`std.os.linux.exit_group(0)`) instead of returning into glibc `exit(3)`. +exit_group terminates all threads atomically in-kernel, so no thread can fault +on freed locale data and no atexit/TLS/locale teardown runs at all. + +- deliberately **not** `std.process.exit` — under `link_libc` it calls libc + `exit(3)`, i.e. the exact teardown we're skipping. +- the operator's secondary idea (stop host reconnects before exit) is + **unnecessary with exit_group**: the atomic in-kernel kill gives no thread the + chance to run, reconnecting or not. kept minimal. + +**validation**: can't be proven locally (teardown path, real thread scale). +builds clean (native + `x86_64-linux-gnu` ReleaseSafe), `zig build test` green. +needs an operator canary/restart: SIGTERM the pod and confirm (a) no SIGILL/GPF +in the kernel journal, and (b) the event log is intact across restart (seq +continuity) — proving the flush still ran before exit_group. diff --git a/src/main.zig b/src/main.zig index 84454f8..56bd14d 100644 --- a/src/main.zig +++ b/src/main.zig @@ -153,6 +153,21 @@ const MetricsServer = struct { }; pub fn main() !void { + try runRelay(); + // clean shutdown reached. runRelay's defers (incl. the event-log flush in + // DiskPersist.deinit) have already run, so all durable state is persisted. + // now skip glibc's exit-time teardown (atexit handlers / locale free): with + // ~2,800 Io.Threaded backend threads still live (the backends are never + // deinit'd), that teardown races a worker mid-getaddrinfo/strtold and GPFs on + // every shutdown (libc.so.6+0x50f). the raw exit_group(2) syscall terminates + // all threads atomically in-kernel before any of them can fault — note we + // must NOT use std.process.exit here, which calls libc exit(3) (running the + // very atexit/locale teardown we're avoiding) when linked against libc. see + // docs/shutdown-gpf-analysis.md. + if (builtin.os.tag == .linux) std.os.linux.exit_group(0) else std.process.exit(0); +} + +fn runRelay() !void { // exp-002: optional GPA wrapper for leak detection. // build with -Duse_gpa=true to enable. on clean shutdown (SIGTERM), // GPA logs every allocation that was never freed, with stack traces.