From 8c08a4f54d7f80f6086bc3410152bf7d882d90da Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Thu, 6 Aug 2026 02:04:38 -0500 Subject: [PATCH] docs: drop the retracted stack-aliasing claim MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit the 08-04 handoff's own correction retracts it — the park-shot detector compared snapshots at different park depths, so its corruption reports were unfounded. yesterday's docs pass propagated the stale claim from the operator brief; sequential per-address connect stays as a deliberate behavioral choice, not a bug workaround. Co-Authored-By: Claude Fable 5 --- AGENTS.md | 3 +-- CLAUDE.md | 3 +-- docs/design.md | 7 ++++--- 3 files changed, 6 insertions(+), 7 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 5942e10..7bd0e48 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -13,7 +13,6 @@ loops, no sub-ms sleep spins (see docs/handoffs/HANDOFF-2026-08-05-zio-canary.md - `zig fmt --check .` and `zig build test` - production is `-Dtarget=x86_64-linux-gnu` (glibc malloc's per-thread arenas + page-return are load-bearing for RSS at ~2,800 threads under the Threaded fallback). the "musl breaks RocksDB" belief is **retired** — musl builds and runs SIGILL-free on 0.16 — but mimalloc (the only static-musl allocator we found viable) leaks under our thread model, so prod stays glibc. see [docs/musl-investigation.md](docs/musl-investigation.md) - ReleaseFast has a known double-free — do not use -- don't reintroduce Io.Group usage on hot paths at fleet scale — open zio stack-aliasing bug under group churn (see docs/handoffs/HANDOFF-2026-08-04-zio-backend.md) ## deploy @@ -31,5 +30,5 @@ in-pod + a /metrics dump) — rolling back deletes the pod and the evidence. - [docs/gotchas.md](docs/gotchas.md) — zig/pg.zig/rocksdb-zig/deploy traps - [docs/incident-2026-03-04.md](docs/incident-2026-03-04.md) — ReleaseSafe RSS analysis - [docs/evented-attempt.md](docs/evented-attempt.md) — Evented backend attempt and why we reverted -- [docs/handoffs/HANDOFF-2026-08-04-zio-backend.md](docs/handoffs/HANDOFF-2026-08-04-zio-backend.md) — getting zio to fleet scale (eight fixes + open aliasing bug) +- [docs/handoffs/HANDOFF-2026-08-04-zio-backend.md](docs/handoffs/HANDOFF-2026-08-04-zio-backend.md) — getting zio to fleet scale (eight fixes; the stack-aliasing diagnosis was retracted as a detector artifact) - [docs/handoffs/HANDOFF-2026-08-05-zio-canary.md](docs/handoffs/HANDOFF-2026-08-05-zio-canary.md) — the sub-ms sleep root cause and canary #2 verdict diff --git a/CLAUDE.md b/CLAUDE.md index 5942e10..7bd0e48 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -13,7 +13,6 @@ loops, no sub-ms sleep spins (see docs/handoffs/HANDOFF-2026-08-05-zio-canary.md - `zig fmt --check .` and `zig build test` - production is `-Dtarget=x86_64-linux-gnu` (glibc malloc's per-thread arenas + page-return are load-bearing for RSS at ~2,800 threads under the Threaded fallback). the "musl breaks RocksDB" belief is **retired** — musl builds and runs SIGILL-free on 0.16 — but mimalloc (the only static-musl allocator we found viable) leaks under our thread model, so prod stays glibc. see [docs/musl-investigation.md](docs/musl-investigation.md) - ReleaseFast has a known double-free — do not use -- don't reintroduce Io.Group usage on hot paths at fleet scale — open zio stack-aliasing bug under group churn (see docs/handoffs/HANDOFF-2026-08-04-zio-backend.md) ## deploy @@ -31,5 +30,5 @@ in-pod + a /metrics dump) — rolling back deletes the pod and the evidence. - [docs/gotchas.md](docs/gotchas.md) — zig/pg.zig/rocksdb-zig/deploy traps - [docs/incident-2026-03-04.md](docs/incident-2026-03-04.md) — ReleaseSafe RSS analysis - [docs/evented-attempt.md](docs/evented-attempt.md) — Evented backend attempt and why we reverted -- [docs/handoffs/HANDOFF-2026-08-04-zio-backend.md](docs/handoffs/HANDOFF-2026-08-04-zio-backend.md) — getting zio to fleet scale (eight fixes + open aliasing bug) +- [docs/handoffs/HANDOFF-2026-08-04-zio-backend.md](docs/handoffs/HANDOFF-2026-08-04-zio-backend.md) — getting zio to fleet scale (eight fixes; the stack-aliasing diagnosis was retracted as a detector artifact) - [docs/handoffs/HANDOFF-2026-08-05-zio-canary.md](docs/handoffs/HANDOFF-2026-08-05-zio-canary.md) — the sub-ms sleep root cause and canary #2 verdict diff --git a/docs/design.md b/docs/design.md index bc4a16e..cc51158 100644 --- a/docs/design.md +++ b/docs/design.md @@ -252,9 +252,10 @@ at ~240 MiB. resource limits: 8 GiB memory, 1 GiB request, 1000m CPU. > the default build and rollback path. full record: > [handoffs/HANDOFF-2026-08-04-zio-backend.md](handoffs/HANDOFF-2026-08-04-zio-backend.md) > and [handoffs/HANDOFF-2026-08-05-zio-canary.md](handoffs/HANDOFF-2026-08-05-zio-canary.md). -> known debt: an open zio stack-aliasing bug under `Io.Group` churn — groups are -> off the hot path (sequential per-address connect); don't reintroduce them at -> scale until it's closed. +> note: subscriber connect is now sequential per-address (no `Io.Group`, no +> v4/v6 racing) on both backends — a deliberate behavioral change kept from +> this work. an earlier "stack aliasing under group churn" diagnosis was +> retracted as a detector artifact (see the 08-04 handoff's correction). - zig 0.16's `Io` (io_uring/kqueue) would replace OS threads with fibers (~2,800 → ~35), scaling toward 100K+ hosts per process. - **attempted and shelved twice** — `Io.Evented` ran in production, was reverted -- 2.51.2