shutdown GPF — source-level analysis (2026-06-12) #
staged hypothesis for the deterministic GPF at libc.so.6+0x50f on every
shutdown (operator handoff 2026-06-11). this is a read-only code audit; it does
not change teardown. the decisive next step needs the kept core specimen
symbolized — see the bottom.
what was audited #
the standing lead was "unjoined per-host subscriber threads touch the persist
layer (dp) after teardown starts → use-after-free." i traced every fiber/thread
owner and the defer order. that lead is not supported by the code — teardown
is actually well-ordered.
defer execution order at main() return (LIFO) #
registered: bc.deinit(219) → val.deinit(223) → dp.deinit(232) →
slurper.deinit(323) → broadcast_future.cancel(350) → metrics_future.cancel(394)
→ server.deinit(407) → server_future.cancel(416).
executes in reverse, so: server_future → server → metrics_future → broadcast_future → slurper.deinit → dp.deinit → val.deinit → bc.deinit. the explicit shutdown block (main.zig:423–454) runs before any defer and joins gc/db-workers/host_ops/backfiller/cleaner + cancels the three futures.
every fiber owner cancel-and-awaits before freeing its state #
- slurper workers + frame pool (
slurper.deinit, slurper.zig:559): cancels all per-hostrunWorkerfutures (579), thenframe_pool.shutdown()/deinit()joins the 16 pool threads (582). runs beforedp/val/bc.deinit. ordered. - broadcaster consumers (
bc.deinit→Consumer.shutdown, broadcaster.zig:499):f.cancel(io)awaits thewriteLoopfiber beforeallocator.destroy(consumer). ordered (this is the "Consumer.shutdown loose end" — it is actually joined). - validator resolvers (
val.deinit, validator.zig:98): cancels allresolver_futuresbeforecache.deinit(). ordered. - the three main futures:
Future.cancelis idempotent (Io.zig:1191 — nullsany_futureon first call, returns cached result after). the double-cancel (explicit block + defer) is therefore safe, not a UAF. ruled out.
Future.cancel = "cancel request + await", so each of the above blocks until the
task returns. conclusion: there is no simple ordering-UAF in component teardown.
remaining suspects (for core symbolication to disambiguate) #
- process-exit with ~2,800 live threads (leading). the three
Io.Threadedbackends (backend,pool_io_backend,debug_threaded_io) are never deinit'd — nobackend.deinit()anywhere. their thread-pool workers are still alive (idle on condvars) whenmain()returns into glibcexit(). glibcexit()(atexit handlers →__call_tls_dtors→_IO_cleanup→ stdio flush) is not safe against running threads; if a pool thread holds the malloc arena lock or a stdio lock, cleanup faults. a deterministic same-instruction GPF in low libc is consistent with this. - cancel vs a non-cancellable blocked syscall at scale. a
runWorker/writeLoopfiber blocked in a TLSread(not a cancellation point) — cancel awaits, but interaction at ~2,800 fibers is worth ruling out. - cross-backend resource ownership — a fiber on
iowhose freed resource lived onpool_io. the dual-Io setup is flagged in-code as a historical crash class (main.zig:191).
decisive next step (needs the core) #
symbolize libc.so.6+0x50f against the kept specimen
(/var/crash/core.zlay.1781157122.1040178, binary ReleaseSafe-b49a95a in
buildah /opt/zlay):
# on the node, with the matching unstripped binary:
gdb /opt/zlay/zlay /var/crash/core.zlay.1781157122.1040178 -batch \
-ex 'bt' -ex 'info registers rip' -ex 'thread apply all bt'
- fault in
free/malloc/_int_free→ allocator UAF/double-free; re-audit component deinits for a double-free, not an ordering issue. - fault in
__call_tls_dtors/_IO_cleanup/pthread_*→ suspect #1 confirmed: the fix is to either join the backends before return or (simpler, and common for servers) callstd.posix.exit/_exitonce cleanup is done to skip glibc's atexit/TLS/stdio teardown that races with the live pool threads.
candidate fix (pending core confirmation of #1): after log.info("relay stopped cleanly"), bypass the racy libc teardown with an explicit _exit(0) — the
process is terminating anyway and all durable state (event log, pg) is already
flushed by the component deinits that ran above.
RESOLUTION (2026-06-14) #
core symbolized (operator, docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md):
the crash is a "third door" — neither free nor the exit machinery itself, but
a worker thread faulting inside glibc locale parsing
(__GI_____strtold_l_internal) at full speed. the all-threads histogram shows
all ~2,800 backend threads alive at crash, ~1,066 clustered in glibc
getaddrinfo/DNS (add_region/getifaddrs/strverscmp) — host workers
mid-reconnect during teardown. mechanism = suspect #1 confirmed: at SIGTERM,
glibc exit-time teardown frees shared libc state (locale data) while live
backend threads are still grinding DNS/locale paths; one dereferences the freed
state first and takes the SIGSEGV (kernel logs a GPF). not in zlay/zat code.
fix implemented (main.zig): main() now calls runRelay() (the former
body, with all its defers — so the event-log flush in DiskPersist.deinit and
the pg disconnect still run), then issues the raw exit_group(2) syscall
(std.os.linux.exit_group(0)) instead of returning into glibc exit(3).
exit_group terminates all threads atomically in-kernel, so no thread can fault
on freed locale data and no atexit/TLS/locale teardown runs at all.
- deliberately not
std.process.exit— underlink_libcit calls libcexit(3), i.e. the exact teardown we're skipping. - the operator's secondary idea (stop host reconnects before exit) is unnecessary with exit_group: the atomic in-kernel kill gives no thread the chance to run, reconnecting or not. kept minimal.
validation: can't be proven locally (teardown path, real thread scale).
builds clean (native + x86_64-linux-gnu ReleaseSafe), zig build test green.
needs an operator canary/restart: SIGTERM the pod and confirm (a) no SIGILL/GPF
in the kernel journal, and (b) the event log is intact across restart (seq
continuity) — proving the flush still ran before exit_group.