From 85c0c32636812f795b6837547daa31e73b539bee Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Sun, 14 Jun 2026 15:58:26 -0500 Subject: [PATCH] =?UTF-8?q?docs:=20record=20musl/mimalloc=20verdict=20?= =?UTF-8?q?=E2=80=94=20closed,=20prod=20stays=20glibc?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit the canary isolation matrix resolved the static-binary investigation: - "musl breaks RocksDB" is retired — musl builds/links/runs SIGILL-free on 0.16 at thread scale (was a 0.15 codegen artifact). - mimalloc is ruled out for zlay's ~2,800-thread model — RSS climbs unbounded (v2/v3/purge-forced/even under glibc); the glibc+mimalloc row isolated the allocator as the variable. so a static-musl deploy is blocked on the allocator (glibc malloc's page-return is what keeps prod flat and is glibc-specific), and there's no need anyway. - musl-investigation.md: full matrix + verdict + dead-end allocator list + where the revivable -Duse_mimalloc harness lives (musl branch, not merged). - CLAUDE.md / deployment.md: correct the stale "must use glibc / musl breaks RocksDB" notes to the evidence-backed verdict. - design.md: note mimalloc evaluated + ruled out in the memory model. - relocate the three 2026-06-12/14 operator handoffs into docs/handoffs/. Co-Authored-By: Claude Opus 4.8 --- CLAUDE.md | 2 +- docs/deployment.md | 2 +- docs/design.md | 5 +- ...FF-2026-06-12-canary-purge-test-verdict.md | 69 ++++++++++++++++ ...NDOFF-2026-06-12-gpf-core-symbolization.md | 78 +++++++++++++++++++ ...HANDOFF-2026-06-14-canary-final-verdict.md | 68 ++++++++++++++++ docs/musl-investigation.md | 62 +++++++++++++++ 7 files changed, 283 insertions(+), 3 deletions(-) create mode 100644 docs/handoffs/HANDOFF-2026-06-12-canary-purge-test-verdict.md create mode 100644 docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md create mode 100644 docs/handoffs/HANDOFF-2026-06-14-canary-final-verdict.md create mode 100644 docs/musl-investigation.md diff --git a/CLAUDE.md b/CLAUDE.md index c302e57..5628ffb 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -7,7 +7,7 @@ attempt shelved — see docs/evented-attempt.md). ## before pushing - `zig fmt --check .` and `zig build test` -- MUST use `-Dtarget=x86_64-linux-gnu` for production (musl breaks RocksDB) +- production is `-Dtarget=x86_64-linux-gnu` (glibc malloc's per-thread arenas + page-return are load-bearing for RSS at ~2,800 threads). the "musl breaks RocksDB" belief is **retired** — musl builds and runs SIGILL-free on 0.16 — but mimalloc (the only static-musl allocator we found viable) leaks under our thread model, so prod stays glibc. see [docs/musl-investigation.md](docs/musl-investigation.md) - ReleaseFast has a known double-free — do not use ## deploy diff --git a/docs/deployment.md b/docs/deployment.md index 92fec42..6c5f441 100644 --- a/docs/deployment.md +++ b/docs/deployment.md @@ -26,7 +26,7 @@ the full `Dockerfile` exists for CI/standalone builds but is slow on Mac (cross- ### build flags -- `-Dtarget=x86_64-linux-gnu` — **must use glibc**, not musl. zig 0.15's C++ codegen for musl produces illegal instructions in RocksDB's LRU cache. +- `-Dtarget=x86_64-linux-gnu` — production target, glibc. glibc malloc (per-thread arenas + `madvise` page-return + `malloc_trim`) is load-bearing for RSS at ~2,800 threads. **NB:** the old "musl breaks RocksDB / illegal instructions" claim was a zig 0.15 artifact and is **retired** — on 0.16 musl builds, links, and runs SIGILL-free at thread scale (canary, 2026-06). musl stays off prod only because the one static-musl allocator we validated as buildable — mimalloc — leaks RSS under our thread model (v2/v3/purge-forced/even under glibc). full matrix in [musl-investigation.md](musl-investigation.md). - `-Dcpu=baseline` — required when building inside Docker/QEMU (not needed for `zlay-publish-remote` since it builds natively). - `-Doptimize=ReleaseSafe` — safety checks on, optimizations on. production default since 2026-03-05. previously caused OOM (see [incident-2026-03-04.md](incident-2026-03-04.md)) — resolved by the frame pool moving heavy work off reader threads. diff --git a/docs/design.md b/docs/design.md index 5f67875..b2282bb 100644 --- a/docs/design.md +++ b/docs/design.md @@ -108,7 +108,10 @@ the deepest paths (crypto, CBOR, DB) now run on pool workers. **allocator**: `std.heap.c_allocator` (libc malloc). glibc has per-thread arenas and `madvise`-based page return. the general-purpose allocator (GPA) is a debug allocator that never returns freed pages — unsuitable for -long-running servers. +long-running servers. **mimalloc was evaluated (2026-06) and ruled out** — it +does not return memory to the OS under zlay's ~2,800-thread-heap workload, RSS +climbs unbounded (v2/v3/purge-forced/even under glibc). glibc malloc's +page-return is what keeps RSS flat; see [musl-investigation.md](musl-investigation.md). **shared TLS CA bundle**: loaded once by the slurper, passed to all ~2,750 subscriber connections via `config.ca_bundle`. without this, each websocket.zig diff --git a/docs/handoffs/HANDOFF-2026-06-12-canary-purge-test-verdict.md b/docs/handoffs/HANDOFF-2026-06-12-canary-purge-test-verdict.md new file mode 100644 index 0000000..c01954e --- /dev/null +++ b/docs/handoffs/HANDOFF-2026-06-12-canary-purge-test-verdict.md @@ -0,0 +1,69 @@ +# operator readout 2026-06-12 — purge test says leak-shaped, not retention + +ran your env lever. short version: the climb survives `MIMALLOC_PURGE_DELAY=0`, +so by your own discriminator this is not v2.3.2 retention — it's +live-allocation growth. details and the curve below; the canary is being left +to run into its 4Gi OOM deliberately (contained, and the re-climb after +restart is the reproducibility check). + +## banked first: criterion #1 is settled + +zero SIGILL/traps across the entire canary life — 13h+ on the first pod +(15.4M frames, ramped to 2,076 hosts ≈ your reference scale, RocksDB indexing +live + listReposByCollection paging), and the restarted pod continued clean. +the year-old "musl breaks RocksDB" rule is retired with prejudice. musl is +viable independent of the allocator question. + +## the purge test (your lever #1) + +applied `MIMALLOC_PURGE_DELAY=0` + `MIMALLOC_PURGE_DECOMMITS=1` at ~14:30Z +(pod restart; host table + index persist on PVC, so it reconnected all +~2,076 hosts within minutes — the curve below is at full thread scale nearly +from t-zero, minimal ramp contamination). + +container working set, stable 2,074–2,076 hosts throughout: + +``` +14:32 1735 Mi (restart + reconnect churn) +15:02 2363 +15:32 2640 <- slope decelerating here, looked like it might flatten +~18:00 3421 <- alert line; slope re-accelerated to ~ +300 Mi/h +``` + +- with purging forced on, RSS ≈ live allocation. a monotonic, accelerating + climb at constant host count is not explainable by per-thread-heap + retention. +- for reference the first pod (default purge delay) reached ~2.6 Gi at the + same host count — the purged pod is *worse* at matched wall-clock, which + is consistent with "the growth was never retained-free memory." +- glibc prod comparison: flat 2.6 Gi for 30h+ at 2,976 hosts on the same + network load. + +## interpretation (labeled) + +**confirmed:** growth is in live allocations (or memory mimalloc cannot +purge), not purge-delay retention. cold-index build alone is hard to credit +for an *accelerating* slope at hour 16+ of indexing. + +**hypothesis, untested:** whether this is (a) a genuine allocation leak that +only manifests on the musl/mimalloc build (e.g. a path where glibc build +behavior differs), (b) the known v2.x multithreaded RSS pathology being +worse than "retention" in practice (issue #1111's reports do describe +unbounded-looking growth), or (c) mimalloc arena behavior `PURGE_DELAY` +doesn't reach. your staged v3 bump discriminates (b) cheaply; (a) would need +heap diffing on the canary. + +## state + suggested next + +- canary left running into its 4Gi OOM on purpose. post-OOM it restarts + itself (k8s) and i'll capture whether the curve reproduces — if yes, we + have a reliable repro harness for whatever round two you want. +- your call per your own framing: deploy the staged v3 bump for one more + round (cheap, the repro harness is standing), or park the track with + "SIGILL dead / RSS leak-shaped on v2.3.2 even with purge=0" — both + conclusions are now cleanly documented either way. +- if you want allocation-level data from the canary before deciding: + MIMALLOC_SHOW_STATS on exit is free to enable on the next restart; say the + word. + +— operator diff --git a/docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md b/docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md new file mode 100644 index 0000000..9e5092e --- /dev/null +++ b/docs/handoffs/HANDOFF-2026-06-12-gpf-core-symbolization.md @@ -0,0 +1,78 @@ +# operator reply 2026-06-12 — shutdown-GPF core symbolized: third door, but your mechanism + +ran your exact gdb triage against the kept core +(`core.zlay.1781157122.1040178`, b49a95a binary extracted from buildah +storage to `/var/crash/zlay-b49a95a`). the answer is **neither of your two +decision branches** — but it confirms the *mechanism* of your leading +suspect (live backend threads racing glibc exit-time teardown). + +## crashing thread (verbatim) + +``` +Program terminated with signal SIGSEGV, Segmentation fault. +#0 0x00007d2a302b150f in __GI_____strtold_l_internal (nptr=, + endptr=0x0, group=, loc=) + at ../stdlib/strtod_l.c:1093 +#1 0x00007d2a249c3250 in ?? () +#2 0x00007d23990aa6c0 in ?? () +#3 0x0000000000000001 in ?? () +rip 0x7d2a302b150f <__GI_____strtold_l_internal+4287> +``` + +- not `free`/`_int_free` → not an allocator double-free +- not `__call_tls_dtors`/`_IO_cleanup`/`pthread_*` → not the exit machinery + crashing on its own +- it's a **worker thread doing live parsing work** — locale-aware + string-to-float (`strtold_l`, note the `loc=` argument) — faulting at full + speed. the unwind above frame 0 is garbage, consistent with the thread's + surroundings being torn down while it ran. + +## all-threads histogram (top frame, 2,800+ threads at death) + +``` +1947 Io.Threaded.Syscall.finish <- idle pool threads, parked +1066 add_region <- glibc internal + 3 __GI___strverscmp + 2 __GI_____strtold_l_internal <- incl. the crasher + 2 getifaddrs_internal + 1 findidx <- glibc locale collation internal + 1 __GI_memcpy + 1 math.rotl__anon_33657 + 1 Io.Threaded.unlockStderr +``` + +## reading (confirmed vs hypothesis) + +**confirmed:** +- all ~2,800 backend threads were alive at crash time — your "Io.Threaded + backends never deinit'd" observation is visible in the core itself. +- the fault is inside glibc locale/parse internals on a worker thread, not + in zlay/zat code and not in exit machinery. + +**hypothesis (consistent, not proven):** the 1,066-thread `add_region` + +`getifaddrs_internal` + `strverscmp`/`findidx` cluster is glibc +getaddrinfo/DNS machinery — i.e. at SIGTERM, host workers were mid-reconnect +(likely a reconnect burst as connections dropped during teardown), grinding +through DNS/locale paths, when exit-time cleanup freed shared libc state +(locale data being the prime candidate given where the crasher died). one +thread dereferenced it first and took the SIGSEGV (kernel logs it as a GPF). + +**implication:** your `_exit(0)`-after-cleanup direction kills this class +outright — no libc teardown under running threads. independently, stopping +host reconnect attempts before process exit would shrink the racing +population from ~1,000 to ~0, which matters if any cleanup must still run. +both are your calls; the evidence supports either or both. + +## housekeeping + +- core + extracted binary stay in `/var/crash` until you say done + (36 GB total — tell me when I can reclaim). +- stall-window logs (06-11 16:15–17:40Z): **unrecoverable** — at prod log + volume the node retains ~20 minutes of rotated pod logs and nothing + upstream collects them. if the stall investigation needs log evidence, + we'd need it to recur with collection in place; say the word and I'll add + a lightweight log sink first. +- canary: still zero traps, settled-RSS read coming after its first quiet + day, then the collection-backfill hammer. + +— operator diff --git a/docs/handoffs/HANDOFF-2026-06-14-canary-final-verdict.md b/docs/handoffs/HANDOFF-2026-06-14-canary-final-verdict.md new file mode 100644 index 0000000..8a14b6e --- /dev/null +++ b/docs/handoffs/HANDOFF-2026-06-14-canary-final-verdict.md @@ -0,0 +1,68 @@ +# operator final readout 2026-06-14 — canary track closed: mimalloc is the variable, musl is clean + +the isolation step resolved it. **the RSS growth is mimalloc, not musl.** +SIGILL-dead is banked permanently. recommendation: park mimalloc; musl itself +is viable for the shutdown-GPF/static-binary goals if paired with a different +allocator. full matrix below. + +## the matrix (RSS at ~2,030–2,080 hosts, same firehose load) + +| build | libc | allocator | RSS behavior | verdict | +|---|---|---|---|---| +| prod | glibc | glibc malloc | **flat 2.6 GiB, 30h+** | baseline | +| canary | musl | mimalloc v2.3.2 (default) | climb → 2.6 GiB+ slow | fails | +| canary | musl | mimalloc v2.3.2, PURGE_DELAY=0 | climb ~+300 MiB/h | retention ruled out | +| canary | musl | mimalloc v3.3.2 (default) | climb → 4 GiB OOM crashloop | version not the cause | +| **canary** | **glibc** | **mimalloc v3.3.2** | **climb → 3.4 GiB+ and rising** | **isolates the cause** | + +the last row is the discriminator: holding the allocator at mimalloc and +swapping musl→glibc, the climb **persists**. holding the allocator at glibc +malloc (prod), RSS is flat. the only variable that tracks the climb is +mimalloc. + +(this pod warm-started at full host scale off the persisted host table, so it +hit 3.4 GiB in ~1h rather than the multi-hour cold ramp the earlier pods +showed — faster, but the same monotonic shape and the same trajectory toward +OOM. restarts=0 at time of writing; left running, but the verdict doesn't +need the OOM — the climb at matched scale against flat glibc-malloc prod is +the result.) + +## what this means + +- **musl-libc: exonerated.** it was never the variable. the static-musl + build is sound — RocksDB compiles/links/runs clean, and (criterion #1, + banked permanently) **zero SIGILL across every musl pod's life** at thread + scale. the year-old "musl breaks RocksDB" belief is retired with evidence. +- **mimalloc: implicated.** it does not return memory to the OS under + zlay's ~2,800-thread-heap workload — not v2, not v3, not with immediate + purge, not under glibc. your #1111 hypothesis was a good, specific call; + it just turned out the v3 fix doesn't cover whatever path our thread model + hits. this is consistent with mimalloc's per-thread-heap design meeting a + pathological thread count, but i haven't proven the mechanism (MIMALLOC_SHOW_STATS + was armed but exit-time output didn't survive the kill path; a deliberate + SIGTERM read is available if you want the allocator's own accounting before + fully closing this). + +## recommendation + +park the mimalloc track. two clean futures for musl if/when it's worth +revisiting, both independent of this result: +1. static musl + a different allocator (jemalloc segfaulted on musl per your + notes; tcmalloc or a tuned glibc-static are the remaining candidates) — + only worth it if the static-binary/`_exit` story needs it. +2. nothing — glibc prod is flat, cheap, and fine; the SIGILL finding is + logged for whenever it matters. + +net from the whole exercise: one false belief retired (musl/RocksDB), one +allocator ruled out for our workload (mimalloc), prod untouched throughout, +and the higher-priority tracks (shutdown-GPF fix, ingestion-stall) are still +where the real urgency is. + +## housekeeping + +- canary infra (own pg pod, 2 PVCs, deployment) is still up as a repro + harness. say the word and i tear it down to reclaim the node resources, or + leave it for a SHOW_STATS read / future allocator trial. +- GPF core + b49a95a binary still in /var/crash pending your GPF work. + +— operator diff --git a/docs/musl-investigation.md b/docs/musl-investigation.md new file mode 100644 index 0000000..5d45fdd --- /dev/null +++ b/docs/musl-investigation.md @@ -0,0 +1,62 @@ +# musl + static-binary investigation — CLOSED (2026-06-14) + +investigation into shipping zlay as a fully static `FROM scratch` binary +(runs on any distro, tiny image, hermetic). **outcome: closed. prod stays +glibc + glibc-malloc.** two durable results below. + +## verdict + +1. **the "musl breaks RocksDB" belief is retired.** it was a zig 0.15 codegen + artifact (illegal instructions in RocksDB's LRU cache). on **0.16, musl + builds, links, and runs clean** — zero SIGILL across every musl canary pod's + life at thread scale (~2,000–2,080 hosts, 15M+ frames, RocksDB indexing live + + `listReposByCollection` paging). musl-libc is sound for our use. +2. **mimalloc is ruled out for zlay's workload.** it does not return memory to + the OS under our ~2,800-thread-heap model — RSS climbs unbounded. confirmed + across v2.3.2, v3.3.2, `MIMALLOC_PURGE_DELAY=0`, and even under glibc. + +the blocker for a static-musl deploy is therefore the **allocator**, not musl: +glibc malloc's per-thread arenas + `madvise` page-return are what keep prod RSS +flat (~2.5 GiB at ~2,950 hosts), and that's glibc-specific. no validated +static-musl allocator reproduces it. + +## the isolation matrix (operator canary, RSS at ~2,030–2,080 hosts, same load) + +| libc | allocator | RSS behavior | verdict | +|---|---|---|---| +| glibc | glibc malloc (prod) | **flat 2.6 GiB, 30h+** | baseline | +| musl | mimalloc v2.3.2 (default) | slow climb past 2.6 GiB | fails | +| musl | mimalloc v2.3.2, PURGE_DELAY=0 | climb ~+300 MiB/h | retention ruled out | +| musl | mimalloc v3.3.2 (default) | climb → 4 GiB OOM | version not the cause | +| **glibc** | **mimalloc v3.3.2** | **climb → 3.4 GiB+ rising** | **isolates the cause** | + +the last row is the discriminator: holding the allocator at mimalloc and +swapping musl→glibc, the climb persists; holding the allocator at glibc malloc, +RSS is flat. the only variable that tracks the climb is mimalloc. (mimalloc +#1111 reports a v2.x multithreaded RSS pathology fixed in v3 — but v3 still +climbs here, so our thread model hits a path the v3 fix doesn't cover. mechanism +not proven; `MIMALLOC_SHOW_STATS` didn't survive the kill path.) + +## allocator options for a static-musl binary (all dead-ends, for the record) + +- **mimalloc** — ruled out (above). +- **jemalloc** — segfaults under musl. +- **musl's own malloc** — no per-thread arenas; known-poor at our thread count; + no reason to expect glibc-like page return. untested, not pursued. +- **tcmalloc / static-glibc** — untested candidates; only worth trying if a + static binary becomes a genuine need. + +## what was built (the revivable harness) + +the `-Duse_mimalloc=true` build option (compiles mimalloc's `static.c` with +`MI_MALLOC_OVERRIDE`, `MI_LIBC_MUSL` on musl) lives on branch +**`musl-mimalloc-static`** — pushed, **not merged** (we stay glibc). the three +glibc-malloc calls (`mallinfo`/`malloc_info`/`malloc_trim`) are gated on +`isGnuLibC()` there. the operator left the canary infra (isolated pg + PVCs + +deployment) up as a revivable repro harness. + +## if revisited + +only worth it if a static binary becomes a real need (e.g. a deploy constraint). +the path would be static-musl + tcmalloc (or a tuned static-glibc), measured the +same way. otherwise: nothing — glibc prod is flat, cheap, and fine. -- 2.51.2