diff --git a/docs/design.md b/docs/design.md index b2282bb..9050b07 100644 --- a/docs/design.md +++ b/docs/design.md @@ -242,13 +242,50 @@ at ~240 MiB. resource limits: 8 GiB memory, 1 GiB request, 1000m CPU. ### long-term (async I/O) - zig 0.16's `Io` (io_uring/kqueue) would replace OS threads with fibers (~2,800 → ~35), scaling toward 100K+ hosts per process. -- **attempted and shelved** — the `Io.Evented` backend hit 8 cross-Io crash - classes plus a fiber-context-switch GPF; stdlib `Io.Evented` still doesn't - implement networking on 0.16. see [evented-attempt.md](evented-attempt.md). +- **attempted and shelved twice** — `Io.Evented` ran in production, was reverted + on a misdiagnosis, was re-enabled, then shelved for good at `e6cdf84` + (2026-04-09): 8 cross-Io crash classes, a fiber-context-switch GPF, and an + untraced ~10-15% coverage regression. stdlib `Io.Evented` still doesn't + implement networking (verified against 0.16.0 release, `Uring.zig:774-789`). + see [evented-attempt.md](evented-attempt.md). - 2026-06 reassessment: the relevant path is no longer "wait for stdlib Evented" - but a spike on **zio** (a third-party complete `std.Io` impl with its own - stack-switching, sidestepping the stdlib fiber bug). still gated on a real - evaluation; the dual-Io mutex hazard and DbRequestQueue bridge carry over. + but **[zio](https://tangled.org/zzstoatzz.io/zio)** — an ordinary package + implementing the `std.Io` vtable, no compiler fork and no stdlib patch. it is + an active project with zlay named as its first intended consumer; see its + `docs/adoption-guide.md`, `benchmarks.md`, and `fiber-gauntlet-2026-07-09.md`. + not yet wired into zlay. +- zio differs from the shelved `Io.Evented` attempt in two ways that matter + here: its core is a **readiness loop** (epoll/kqueue driving per-connection + state machines) with no stack switching — so the ReleaseSafe GPF class does + not apply — and it is a **hybrid**, wrapping an inner `Io.Threaded` and + delegating everything it does not implement. graceful degradation rather than + the all-or-nothing dual-Io segregation that made the Evented attempt fragile. +- **`pg.Pool` Io-agnosticism is NOT a prerequisite** — measured 2026-07-31, not + reasoned. a Threaded `pg.Pool` driven from zio fibers was correct at every + size tried (64/256/1024 concurrent callers, up to 20,480 ops, plus repeat + runs): zero failures, zero wrong results, no crash. the Evented-era corruption + class did not reproduce. zio's futex ops do wake both domains, as its rule 2 + claims. + - but calling the pool *from fibers* serializes everything. with a starved + pool (8 conns, 256 callers, 20ms queries): Threaded 7.6s / 337 ops-sec vs + zio-fibers 58.5s / 44 ops-sec. max acquire-wait under zio was 0.02ms — + fibers never contend, because each one blocks the whole loop thread for its + query. 2,560 ops × 20ms ≈ 51s serial, against 58.5s observed. + - so the fix is not a smarter pool; it is **not doing DB work on fibers**. + the two-io split plus `DbRequestQueue` that zlay already has is the correct + architecture, and should be kept rather than dissolved. + - caveats: macOS/kqueue and Debug, not linux/epoll under ReleaseSafe; longest + run ~60s, so slow-onset faults are not excluded. +- **remaining open questions**, in order: + 1. the ~10-15% coverage regression from the Evented era is still unexplained. + it was never root-caused, so it cannot be assumed specific to `Io.Uring`. + 2. re-run the above on linux under ReleaseSafe, and soak for longer than a + minute, before trusting the safety result for production. + 3. validate zio's throughput numbers against zlay's actual socket workload — + its benchmarks are its own; `Io.Threaded` remains the reference column. +- cheaper intermediate step that needs none of the above: the reader-thread + multiplexing on the near-term list, which cuts most of the ~2,800 threads + without introducing a second Io type. ## deliberate divergences from indigo diff --git a/docs/evented-attempt.md b/docs/evented-attempt.md index afc36be..c5371a8 100644 --- a/docs/evented-attempt.md +++ b/docs/evented-attempt.md @@ -1,16 +1,18 @@ -# Evented backend attempt (2026-03 to 2026-04) +# Evented backend attempt (2026-04-03 to 2026-04-10) -> **Current status:** shelved. zlay uses `Io.Threaded`, and the Docker build no -> longer applies `patches/uring-networking.patch`. This document is retained as -> a historical record of the experiment. +> **Current status:** shelved, permanently as of `e6cdf84` (2026-04-09). zlay +> uses `Io.Threaded` (`src/main.zig:62`), and the Docker build no longer applies +> `patches/uring-networking.patch`. This document is a historical record. log of getting zlay running on `Io.Evented` (io_uring fibers) instead of `Io.Threaded` (one OS thread per task). goal: drop from ~2,800 threads to ~35. -after 28 commits fixing cross-Io issues and a brief revert to Threaded, we -discovered the production SIGSEGV was a websocket handshake bug (TCP split -mid-CRLF, fixed in websocket.zig `9ac64da`), not a fiber context-switch -issue. Evented backend is back, now running ReleaseSafe. +Evented was shelved **twice**. the first revert (`42e1019`) was a +misdiagnosis: the production SIGSEGV turned out to be a websocket handshake bug +(TCP split mid-CRLF, fixed in websocket.zig `9ac64da`), not fiber machinery, so +Evented was re-enabled the same day (`02434de`). it was shelved again four days +later (`e6cdf84`) — that decision stuck, and for different reasons. see +[why it was shelved for good](#why-it-was-shelved-for-good). ## what we built @@ -61,7 +63,31 @@ ReleaseSafe case. this is a zig stdlib bug — the inline asm that swaps rsp/rbp/rip between fiber contexts corrupts state under certain optimizer configurations. -## upstream status (checked 2026-04-05) +## why it was shelved for good + +the second shelving (`e6cdf84`, 2026-04-09) was not triggered by a crash. per +that commit, three things together: + +1. the 8 cross-Io crash classes above — all fixed, but the architecture stayed + fragile +2. the ReleaseSafe GPF from the zig codegen bug +3. **a persistent ~10-15% coverage degradation that nobody could trace** + +(3) is the one that matters most for any future attempt, and it is the least +understood. `Io.Threaded` restored the proven 0.15 model — thread-per-PDS at +99%+ coverage — and the regression went away. no root cause was ever found, so +it is not known whether it was specific to `Io.Uring`, to the networking patch, +or inherent to running ~2,800 subscribers as fibers. **any fiber-based retry +must explain this first**; a backend that fixes only the GPF does not address it. + +what switching back bought, per the same commit: the entire cross-Io problem +class vanished, ReleaseSafe worked again, DNS worked natively, and the uring +networking patch became inert. the cost was ~2,800 OS threads instead of ~35 +fibers. + +## upstream status + +**checked 2026-04-05** (against dev.3059/dev.3091): - `fiber.zig`: **unchanged** between dev.3059 (our pin) and dev.3091 (latest) - uring networking: **still fully stubbed** — all six `*Unavailable` functions @@ -69,6 +95,20 @@ configurations. to be done before they can be used reliably" (HN, feb 2026) - `Io.Threaded` is the recommended production backend +**re-checked 2026-07-30** (against the pinned 0.16.0 release, `zig env` +`.std_dir`) — no movement: + +- `Io.Evented` still resolves to `Io.Uring` on linux (`std/Io.zig:31-36`) +- `Uring.zig` networking **still fully stubbed**: `netListenIp`, `netConnectIp`, + `netSend`, `netRead`, `netLookup` and the rest are `*Unavailable` + (`std/Io/Uring.zig:774-789`). `Threaded.zig:1916-1961` implements them. +- so `patches/uring-networking.patch` and the Threaded DNS fallback remain + load-bearing for any Evented path +- `std.Io` itself is a plain vtable interface (`std/Io.zig:26`, `VTable` at + `:51`) with no instability markers — but it is **113 function pointers** wide + (fs, process spawn, mmap, clocks, random, futex, net). that is the bar for any + complete third-party implementation. + ## what we kept - `patches/uring-networking.patch` — the networking implementation, for @@ -79,14 +119,40 @@ configurations. ## timeline -- 2026-03-08: first Evented deploy (`39134d1`) -- 2026-03-08 to 2026-04-04: 28 commits fixing cross-Io crashes -- 2026-04-05: reverted to Threaded after SIGSEGV blamed on fiber machinery -- 2026-04-05: discovered real cause — websocket handshake TCP split bug -- 2026-04-05: fixed websocket.zig (`9ac64da`), re-enabled Evented + ReleaseSafe +commit dates from `git log`, not reconstructed: -## fallback - -if Evented + ReleaseSafe hits the repro GPF in production, switch to -ReleaseFast in the Dockerfile. if that also crashes, flip `Backend` back -to `Io.Threaded` in `src/main.zig:58` — one-line change. +- 2026-04-03: `6c83f11` Evented backend via patched Uring.zig networking +- 2026-04-03: `39134d1` evented backend with single ordered publication path +- 2026-04-03 to 2026-04-04: 28 commits fixing cross-Io crashes, ending + `c3bc3be` (DbRequestQueue) +- 2026-04-05: `42e1019` reverted to Threaded after SIGSEGV blamed on fibers +- 2026-04-05: discovered real cause — websocket handshake TCP split bug +- 2026-04-05: `02434de` fixed websocket.zig (`9ac64da`), re-enabled Evented +- 2026-04-09: `e6cdf84` **switched back to Threaded and shelved Evented** — + final, and the reason the rest of this doc is history +- 2026-04-10: `80eca78` cleaned up stale Evented-era comments in the codebase + +## if someone tries fibers again + +the backend swap itself is one line — `const Backend = Io.Threaded;` +(`src/main.zig:62`), with `Backend.init` already branched at `:199-203`. the +seam survived the revert: `pool_io_backend` (`:211`) is still a separate +Threaded runtime, documented in-place as "redundant since Backend is also +Threaded, but harmless". so the scaffolding is cheap to reuse. + +what is *not* cheap, and what a swap does not fix: + +- **`pg.Pool` is not Io-agnostic.** this is the root of the cross-Io rule and + the source of three separate heap-corruption commits. `docs/notes.md` lists + making it Io-agnostic as the follow-up that "would eliminate cross-Io issues". + that is the real prerequisite. +- **the XRPC/admin hole was never closed** — API handlers ran on fibers and + queried the Threaded `pg.Pool`. see `docs/notes.md`, "known remaining". +- **the ~10-15% coverage regression is unexplained.** see above. + +a third-party `std.Io` implementation with its own stack-switching (e.g. the +**zio** spike noted in [design.md](design.md)) would plausibly avoid the +`fiber.zig` GPF, since that bug is in stdlib inline asm. it does not address any +of the three items above, and it inherits the stubbed-networking problem unless +it implements the net vtable itself. **zio has never been tried in this repo** — +as of 2026-07-30 it exists only as a note in design.md. diff --git a/docs/notes.md b/docs/notes.md index dc50c7f..539aae6 100644 --- a/docs/notes.md +++ b/docs/notes.md @@ -1,8 +1,17 @@ # zlay 0.16 migration — status and known issues -last updated: 2026-04-04 -stable production build: `a931853` (zig 0.15) -latest 0.16 build: `b433403` (not yet deployed — includes all fixes through crash 8) +last updated: 2026-07-30 (migration content below is from 2026-04-04) + +**the 0.16 migration is done and deployed.** zlay runs 0.16.0 release on +`Io.Threaded`; current state is zlay 0.0.5 at `863a6ef`. this document is +retained for the eight crash write-ups and the cross-Io rule, which remain the +best record of why the architecture looks the way it does — but read the +per-section notes: anything describing `Io.Evented` as live is history, shelved +at `e6cdf84` (2026-04-09). see [evented-attempt.md](evented-attempt.md). + +historical context, at time of writing: +- stable production build: `a931853` (zig 0.15) +- latest 0.16 build: `b433403` (not yet deployed — all fixes through crash 8) ## what happened @@ -22,7 +31,7 @@ futex wait/wake access fiber-local `Thread.current()` state. but zlay's frame pool (`thread_pool.zig`) spawns plain `std.Thread` workers — they don't have fiber state. any `Io.Mutex` operation from a plain thread segfaulted. -**fix**: force `Backend = Io.Threaded` in `src/main.zig:55`. threaded futex +**fix**: force `Backend = Io.Threaded` in `src/main.zig:62`. threaded futex uses direct kernel syscalls that work from any execution context. all 0.16 API benefits are preserved; `io.concurrent` under Threaded spawns real OS threads. @@ -30,6 +39,12 @@ benefits are preserved; `io.concurrent` under Threaded spawns real OS threads. cross-boundary modules a dedicated `Threaded` sync_io for mutex operations. then Evented can be re-enabled for the network I/O layer. +> **2026-07 update**: this "future path" was tried and abandoned. Evented ran in +> production twice and was shelved for good at `e6cdf84` (2026-04-09) — not for +> this crash, which was fixed, but for an untraced ~10-15% coverage regression. +> the durable fix is making `pg.Pool` Io-agnostic, not re-enabling Evented. +> see [evented-attempt.md](evented-attempt.md). + ### crash 2: pool acquire panic under load (fixed in `e5ed0d1`) **symptom**: `unreachable` panic in `Io.Event.waitTimeout` after 30-60s of @@ -251,9 +266,16 @@ summary: ### concurrency architecture -three layers, shaped by the cross-Io constraint: +> **as of `e6cdf84` (2026-04-09) `Backend = Io.Threaded` (`src/main.zig:62`), +> so there are no fibers.** `io.concurrent()` under Threaded spawns real OS +> threads, and the cross-Io hazard below is dormant — every runtime in the +> process is Threaded. the layering is kept because it still describes what runs +> where, and because it is the shape any future fiber attempt has to reproduce. + +three layers, originally shaped by the cross-Io constraint: -**Evented fibers (`io.concurrent` on `Io.Evented`)** +**`io.concurrent` on the primary backend** (was Evented fibers, now Threaded +OS threads) - upstream PDS subscribers (read loops, ping loops) - downstream consumer write loops - DID resolver loops (validator) @@ -264,7 +286,8 @@ three layers, shaped by the cross-Io constraint: **plain `std.Thread` with `pool_io` (`Io.Threaded`)** - GC loop — uses `DiskPersist.mutex` + `pg.Pool` (crash 8) - resyncer — uses `DiskPersist` + HTTP client (crash 6) -- these MUST NOT run as Evented fibers (see "the cross-Io rule") +- these MUST NOT run as fibers (see "the cross-Io rule") — still true if fibers + ever return, even though `pool_io` is redundant today **CPU-bound ordered processing → explicit `std.Thread` workers** - `thread_pool.zig`: `workers[host_id % N]` ensures per-key FIFO ordering @@ -272,13 +295,23 @@ three layers, shaped by the cross-Io constraint: - bounded backpressure: blocking submit when queue full → TCP backpressure - uses `Io.Mutex` / `Io.Condition` with `pool_io` (Threaded futex) -### dependency versions (current `b433403`) +### dependency versions + +current, from `build.zig.zon` at `863a6ef` (zlay 0.0.5): ``` -zat v0.3.0-alpha.16 (tangled.org) -websocket.zig 80c6434 (github, master) +zat v0.3.20 (tangled.org) +websocket.zig v0.1.11 (tangled.org) pg.zig 5ce2355 (github, dev branch) rocksdb-zig cdef67b (github) +zig 0.16.0 (release; no Uring networking patch applied) +``` + +historical, at `b433403` during the Evented attempt: + +``` +zat v0.3.0-alpha.16 (tangled.org) +websocket.zig 80c6434 (github, master) zig 0.16.0-dev.3059+42e33db9d (patched Uring networking) ``` @@ -297,7 +330,14 @@ zig 0.16.0-dev.3059+42e33db9d (patched Uring networking) | `RELAY_MAX_EVENTS_GB` | 100 | max disk usage | | `DATABASE_URL` | postgres://relay:relay@localhost:5432/relay | Postgres connection | -## what needs to happen next +## what needed to happen next (2026-04, resolved) + +> items 1 and 2 are done — `b433403` shipped and 0.16 has been in production +> since. item 3 was **superseded**: XRPC/admin no longer cross Io boundaries, +> because there is only one Io type now (Threaded). the underlying cause — +> `pg.Pool` not being Io-agnostic — is untouched and is now the main +> prerequisite for any future fiber work. item 4's "Evented viability" line is +> closed; see [evented-attempt.md](evented-attempt.md). 1. **deploy `b433403`** — all eight crashes/fixes are in. the steady-state heap corruption (cross-Io GC + health checks) is fixed. broadcaster @@ -318,7 +358,11 @@ zig 0.16.0-dev.3059+42e33db9d (patched Uring networking) - fix: either run API handlers on pool_io, or make pg.Pool Io-agnostic 4. **follow-up work**: - - investigate Evented backend viability (frame workers → io.concurrent?) + - ~~investigate Evented backend viability (frame workers → io.concurrent?)~~ + — **closed 2026-04-09** (`e6cdf84`), see [evented-attempt.md](evented-attempt.md) - consider upstreaming the client write lock to karlseguin/websocket.zig - - consider upstreaming Uring networking patch (zig#31723) - - evaluate whether pg.Pool can be made Io-agnostic (would eliminate cross-Io issues) + - consider upstreaming Uring networking patch (zig#31723) — still relevant; + upstream `Uring.zig` networking is still stubbed as of 0.16.0 release + - **evaluate whether pg.Pool can be made Io-agnostic** (would eliminate + cross-Io issues) — still open, and the highest-leverage item here: it is + the prerequisite for any fiber-based backend, stdlib or third-party