zio #
a fiber-based std.Io backend for zig, built to run thousands of blocking-style
connections on one loop thread under ReleaseSafe — with all safety checks on.
zig 0.16 shipped Io.Threaded as the only production backend; Io.Evented
is experimental, has no networking, and its fiber machinery crashes under
ReleaseSafe. zio fills that gap as an ordinary package — no compiler fork, no
stdlib patch.
what it is #
zio.ZioBackend wraps an inner Io.Threaded and overrides the ops that
benefit from a fiber scheduler:
- concurrent/await/cancel spawn fibers on a single loop thread; cancellation reaches fibers parked on anything (io, futex, timer, await) promptly via an all-fibers interrupt registry
- net connect/accept/read/write park on poller readiness (epoll/kqueue, per-direction one-shot arms); ip4 and ip6 both connect non-blocking
- sleep goes to the sched timer heap
- fiber-aware futexWait/Wake makes
Io.Mutex/Condition/Queuecorrect across fiber and thread domains: fibers park in the sched table, foreign threads on the kernel word, wakes reach both - Io.Group members run as fibers; group await parks the caller, not the loop (but see caveats)
- netLookup spills getaddrinfo to a small pool of plain threads zio owns; the calling fiber parks until completion, so stack-resident result buffers cannot outlive their writers
everything else delegates to the inner Io.Threaded — graceful degradation
by construction. unmodified Io.net code runs on fibers. fiber stacks are
mmap'd with guard pages, so a fiber costs the pages it touches, not its
256 KiB reservation.
adoption notes: docs/adoption-guide.md
field results #
zio's first production consumer is an AT Protocol relay holding ~2,800 TLS websocket subscriptions plus full firehose validation. it has run in production on zio (zig 0.16, ReleaseSafe, x86_64-linux):
- RSS 1.33–1.45 GB against 2.27 GB flat on thread-per-connection — about a 40% cut on identical live load, with ~90 OS threads instead of ~2,800. measured at steady state after cache refill; a mid-refill reading of the same deployment showed ~45%, so quote the smaller number.
- the older scoreboard's ~200× RSS figure is raw TCP echo with no TLS and no per-connection state. treat it as a microbenchmark ceiling, not an expectation.
getting there took nine root-caused bugs. the one worth repeating to anyone building on a single-threaded fiber runtime:
a sleep that does not park is a busy-wait that owns your scheduler. zio's clock is millisecond-granular and durations were truncated, so
io.sleep(100us)returned without ever parking. the relay'swhile (!done) io.sleep(100us)database wait then spun the loop thread, starving every other fiber on it — accept loops went deaf while ingest kept trickling and the process looked healthy. durations now round up.
eight scheduler/backend bugs were found and fixed getting there; the debug
instrumentation that caught them (parked-stack scribble detector, list
membership invariants) lives in-tree — the detector is comptime-gated off
unless a root module opts in via pub const zio_debug_park_shot = true.
known caveats #
- single loop thread. every fiber's work is serialized there, so any per-event cost (TLS, fan-out to N consumers) is serialized too. this is the main scaling consideration when adopting zio; multi-loop (thread-per-core, one Sched per core) is design-supported but not wired.
Io.Groupat scale is lightly proven. the corruption once attributed to groups was reallyio.asyncfalling through to the inner Threaded backend and returning a foreign future that zio then reinterpreted as its own task (fixed). groups pass the suite and the storm bench, but no production workload leans on them yet.mainand thecanary/obs-2026-08-04branch have diverged and want reconciling: the canary deliberately reverted three unsoaked semantic changes (theasyncoverride, the DNS broadcast wake, a task-result stamp) to stay close to a certified build.mainkeeps them.- the DNS broadcast wake on
mainis O(fibers) per lookup completion. it is correct but wants replacing with a targeted wake — the principled fix is for zio to own the lookup rather than delegate toThreaded.netLookupand compensate for its kernel-only wake. - io_uring engine not started (default docker seccomp blocks io_uring).
testing #
zig build test -Doptimize=ReleaseSafe # unit + integration, both domains
zig build storm -Doptimize=ReleaseSafe # adversarial: 2k fibers, error
# propagation across parks, futex/
# cancel/lookup/socket churn, heap
# canaries, echo integrity
ReleaseSafe on real x86_64 hardware is the load-bearing configuration — the
historical fiber crashes were optimizer-dependent and invisible in Debug
(gauntlet write-up: docs/fiber-gauntlet-2026-07-09.md).
on macOS, guarded fiber stacks hang the test allocator (not production
paths) — run the suite on linux; an x86_64 container with -mcpu baseline
works (Rosetta misreports the host CPU).
repo tour #
src/backend.zig—ZioBackend: thestd.Iovtable splicesrc/fiber/sched.zig— the scheduler: ready queue, timers, futex table, inject queue (the only cross-thread surface), park-shot detectorsrc/fiber/vendored.zig— the gauntleted context-switch primitive (naked-function switch on x86_64; the aarch64.lrclobber fix)src/poller.zig— epoll/kqueue with per-direction one-shot armssrc/spill.zig— owned worker threads for blocking calls made on behalf of fiberssrc/net.zig— non-blocking socket helpers shared by the backend, the scheduler, and the demosbench/— live demo (demo), adversarial storm (storm)
origins #
zio came out of running ~2,800 production websocket subscriptions on both
Io.Threaded and Io.Evented and watching the experimental backend fail in
ways that forced a retreat to thread-per-connection. the write-ups that led
here (retrospective, stdlib empirics, plan) live with that project's docs;
the short version is in docs/benchmarks.md.