atproto relay in zig zlay.waow.tech
relay zig atproto
zlay docs shutdown-gpf-analysis.md
6.5 kB
Markdown

shutdown GPF — source-level analysis (2026-06-12) #

staged hypothesis for the deterministic GPF at libc.so.6+0x50f on every shutdown (operator handoff 2026-06-11). this is a read-only code audit; it does not change teardown. the decisive next step needs the kept core specimen symbolized — see the bottom.

what was audited #

the standing lead was "unjoined per-host subscriber threads touch the persist layer (dp) after teardown starts → use-after-free." i traced every fiber/thread owner and the defer order. that lead is not supported by the code — teardown is actually well-ordered.

defer execution order at main() return (LIFO) #

registered: bc.deinit(219) → val.deinit(223) → dp.deinit(232) → slurper.deinit(323) → broadcast_future.cancel(350) → metrics_future.cancel(394) → server.deinit(407) → server_future.cancel(416).

executes in reverse, so: server_future → server → metrics_future → broadcast_future → slurper.deinit → dp.deinit → val.deinit → bc.deinit. the explicit shutdown block (main.zig:423–454) runs before any defer and joins gc/db-workers/host_ops/backfiller/cleaner + cancels the three futures.

every fiber owner cancel-and-awaits before freeing its state #

  • slurper workers + frame pool (slurper.deinit, slurper.zig:559): cancels all per-host runWorker futures (579), then frame_pool.shutdown()/deinit() joins the 16 pool threads (582). runs before dp/val/bc.deinit. ordered.
  • broadcaster consumers (bc.deinit → Consumer.shutdown, broadcaster.zig:499): f.cancel(io) awaits the writeLoop fiber before allocator.destroy(consumer). ordered (this is the "Consumer.shutdown loose end" — it is actually joined).
  • validator resolvers (val.deinit, validator.zig:98): cancels all resolver_futures before cache.deinit(). ordered.
  • the three main futures: Future.cancel is idempotent (Io.zig:1191 — nulls any_future on first call, returns cached result after). the double-cancel (explicit block + defer) is therefore safe, not a UAF. ruled out.

Future.cancel = "cancel request + await", so each of the above blocks until the task returns. conclusion: there is no simple ordering-UAF in component teardown.

remaining suspects (for core symbolication to disambiguate) #

  1. process-exit with ~2,800 live threads (leading). the three Io.Threaded backends (backend, pool_io_backend, debug_threaded_io) are never deinit'd — no backend.deinit() anywhere. their thread-pool workers are still alive (idle on condvars) when main() returns into glibc exit(). glibc exit() (atexit handlers → __call_tls_dtors → _IO_cleanup → stdio flush) is not safe against running threads; if a pool thread holds the malloc arena lock or a stdio lock, cleanup faults. a deterministic same-instruction GPF in low libc is consistent with this.
  2. cancel vs a non-cancellable blocked syscall at scale. a runWorker/ writeLoop fiber blocked in a TLS read (not a cancellation point) — cancel awaits, but interaction at ~2,800 fibers is worth ruling out.
  3. cross-backend resource ownership — a fiber on io whose freed resource lived on pool_io. the dual-Io setup is flagged in-code as a historical crash class (main.zig:191).

decisive next step (needs the core) #

symbolize libc.so.6+0x50f against the kept specimen (/var/crash/core.zlay.1781157122.1040178, binary ReleaseSafe-b49a95a in buildah /opt/zlay):

# on the node, with the matching unstripped binary:
gdb /opt/zlay/zlay /var/crash/core.zlay.1781157122.1040178 -batch \
  -ex 'bt' -ex 'info registers rip' -ex 'thread apply all bt'
  • fault in free/malloc/_int_free → allocator UAF/double-free; re-audit component deinits for a double-free, not an ordering issue.
  • fault in __call_tls_dtors / _IO_cleanup / pthread_* → suspect #1 confirmed: the fix is to either join the backends before return or (simpler, and common for servers) call std.posix.exit/_exit once cleanup is done to skip glibc's atexit/TLS/stdio teardown that races with the live pool threads.

candidate fix (pending core confirmation of #1): after log.info("relay stopped cleanly"), bypass the racy libc teardown with an explicit _exit(0) — the process is terminating anyway and all durable state (event log, pg) is already flushed by the component deinits that ran above.

RESOLUTION (2026-06-14) #

core symbolized (operator, docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md): the crash is a "third door" — neither free nor the exit machinery itself, but a worker thread faulting inside glibc locale parsing (__GI_____strtold_l_internal) at full speed. the all-threads histogram shows all ~2,800 backend threads alive at crash, ~1,066 clustered in glibc getaddrinfo/DNS (add_region/getifaddrs/strverscmp) — host workers mid-reconnect during teardown. mechanism = suspect #1 confirmed: at SIGTERM, glibc exit-time teardown frees shared libc state (locale data) while live backend threads are still grinding DNS/locale paths; one dereferences the freed state first and takes the SIGSEGV (kernel logs a GPF). not in zlay/zat code.

fix implemented (main.zig): main() now calls runRelay() (the former body, with all its defers — so the event-log flush in DiskPersist.deinit and the pg disconnect still run), then issues the raw exit_group(2) syscall (std.os.linux.exit_group(0)) instead of returning into glibc exit(3). exit_group terminates all threads atomically in-kernel, so no thread can fault on freed locale data and no atexit/TLS/locale teardown runs at all.

  • deliberately not std.process.exit — under link_libc it calls libc exit(3), i.e. the exact teardown we're skipping.
  • the operator's secondary idea (stop host reconnects before exit) is unnecessary with exit_group: the atomic in-kernel kill gives no thread the chance to run, reconnecting or not. kept minimal.

validation: can't be proven locally (teardown path, real thread scale). builds clean (native + x86_64-linux-gnu ReleaseSafe), zig build test green. needs an operator canary/restart: SIGTERM the pod and confirm (a) no SIGILL/GPF in the kernel journal, and (b) the event log is intact across restart (seq continuity) — proving the flush still ran before exit_group.