about things notes.zzstoatzz.io
notes
notes languages ziglang io patterns.md
12 kB

patterns #

practical patterns for using std.Io in real applications.

backend selection #

const std = @import("std");
const Io = std.Io;

const Backend = if (Io.Evented != void) Io.Evented else Io.Threaded;
var backend: Backend = undefined;

pub fn main() !void {
    const allocator = std.heap.smp_allocator;

    if (Backend == Io.Threaded) {
        backend = Io.Threaded.init(allocator, .{});
    } else {
        try Backend.init(&backend, allocator, .{});
    }
    const io = backend.io();

    // pass io to your app
    try app(io, allocator);
}

production: Io.Evented has known bugs as of 0.16.0-dev.3059 — see below. the code is identical between backends — just swap the init.

argument ordering: where does io go relative to allocator? #

no rule is written down, so when both Allocator and Io show up in one signature you have to decide. what std actually does is a useful prior. as of 0.16.0, signatures taking both lean hard toward allocator first — ~104 allocator-before-io vs ~4 io-before, with the shape:

self (if a method) → allocator → io → operation args.

pub fn init(gpa: Allocator, io: Io, ...) ...        // std.Build.Fuzz.init
pub fn init(mr: *MultiReader, gpa: Allocator, io: Io, ...) void  // Io/File/MultiReader
fn load(gpa: Allocator, io: Io, elf_file: Io.File, ...) ...      // debug/ElfFile

the handful that lead with io are free ...Alloc helpers whose whole job is io and that allocate only the return value (process.currentPathAlloc(io, allocator)) — io is the subject there, allocator a trailing "…and put the result here." so the lean tracks intent: for a stateful constructor that allocates and owns an object, the allocator that backs it reads as the more fundamental of the two and goes first.

it's a lean, not a law (the 4 exceptions are real, and conventions drift — re-run the count on a newer std before leaning on it). a plausible counter-read is "io is an ambient capability like a context handle, so put it first"; std today just doesn't go that way for constructor-shaped functions. we followed the prevailing lean in our websocket.zig/http.zig forks (WorkerState.init(allocator, io, config)); karl's upstream uses (io, allocator) — both compile, they just can't be mixed in one graph (see ../websocket/).

re-check it against your own std:

STD=$(dirname $(which zig))/../lib/std   # or wherever your zig lib lives
rg -U -o '\b\w+: ?(mem\.)?Allocator ?, ?io: ?(std\.)?Io\b' "$STD" | wc -l  # allocator-first
rg -U -o '\bio: ?(std\.)?Io ?, ?\w+: ?(mem\.)?Allocator\b' "$STD" | wc -l  # io-first

Threaded InitOptions #

Io.Threaded.init(allocator, opts) accepts:

Io.Threaded.init(allocator, .{
    .stack_size = 8 * 1024 * 1024,     // default: 16MB (std.Thread.SpawnConfig.default_stack_size)
    .async_limit = .{ .value = 8 },    // default: CPU count - 1
    .concurrent_limit = .{ .value = 64 }, // default: .unlimited
});
option default what it does
stack_size 16MB per-thread stack. affects all spawned threads.
async_limit CPU - 1 bounded pool for io.async(). overflow runs task inline.
concurrent_limit .unlimited pool for io.concurrent(). overflow returns error.ConcurrencyUnavailable.

the thread explosion lesson #

with default concurrent_limit = .unlimited, every io.concurrent() call that outlives its parent creates a permanent OS thread. a relay connecting to 2,000+ hosts with 2 concurrent tasks each (read loop + ping loop) creates ~4,000 threads at 16MB stack = 64GB virtual memory.

mitigations:

  1. set concurrent_limit to a bounded value
  2. set stack_size to what you actually need (8MB is plenty for I/O tasks)
  3. reduce concurrent tasks per unit of work — merge ping into read loop (1 task per host, not 2)
  4. use Io.Group for lifecycle management — cancel all subscribers on shutdown

under Evented, io.concurrent() creates fibers (cheap userspace stacks). the 2-tasks-per-host architecture is fine there. Threaded InitOptions let you bound the damage until Evented is production-ready.

debug_io override #

std.Options.debug_io is backed by a single-threaded instance. using it for application I/O silently serializes everything (measured: an event pipeline dropped from ~60 to ~4 events/s after migrating to Io.Mutex on debug_io — coral, 2026).

override in your root source file:

var app_threaded_io: Io.Threaded = undefined;
pub const std_options_debug_threaded_io: ?*Io.Threaded = &app_threaded_io;

pub fn main() !void {
    app_threaded_io = Io.Threaded.init(allocator, .{});
    // now all std.Options.debug_io usage gets the real threaded instance
}

this works because the Io struct holds a pointer to Threaded — the pointer is stable even though the data is undefined at comptime.

or just pass io explicitly — create Io.Threaded in main, call .io(), thread it through functions. avoids globals entirely but more invasive.

long-lived task lifecycle #

replacing std.Thread.spawn with io.concurrent for I/O-bound loops:

// old pattern
self.thread = try std.Thread.spawn(.{ .stack_size = 8 * 1024 * 1024 }, runLoop, .{self});
// ... later:
if (self.thread) |t| t.join();

// new pattern
self.future = try io.concurrent(runLoop, .{self});
// ... later:
_ = self.future.cancel(io);

the task function should exit cleanly on cancellation:

fn runLoop(self: *Self) void {
    while (!self.shouldStop()) {
        self.io.sleep(Io.Duration.fromMilliseconds(100), .awake) catch break;
        // ... work ...
    }
}

io.sleep() is a cancellation point. when future.cancel(io) is called, sleep returns error.Canceled. the catch break exits the loop.

reset reconnect backoff on success #

when a long-lived task wraps a flaky connection with exponential backoff, reset the backoff inside the successful connect path (after handshake, not just on init). otherwise a long-running connection that briefly drops inherits the maxed-out backoff from earlier failures and wastes minutes reconnecting. the bug is invisible until you've been running for hours.

// in the persistent connect-and-stream loop:
fn connectAndStream(self: *Worker) !void {
    var client = try connect();
    defer client.close();
    self.backoff_ms = 1000;  // ← reset HERE, not in init
    while (...) { ... work ... }
}

cancel vs await #

  • cancel(io) — requests cancellation + blocks until done. returns the task's result.
  • await(io) — just blocks until done. no cancellation request.
  • both are idempotent and consume the future.
  • both are NOT threadsafe — only call from the parent task.

managing dynamic task sets with Group #

for a dynamic set of long-lived tasks (e.g., subscriber connections):

var subscribers: Io.Group = .init;

// spawn subscribers as they're discovered
for (hosts) |host| {
    subscribers.concurrent(io, runSubscriber, .{host, io}) catch {
        log.warn("concurrent limit reached for {s}", .{host});
        continue;
    };
}

// on shutdown — cancel all at once
subscribers.cancel(io);

Group resources per task are freed when that task returns, not when the group is awaited. safe for long-lived groups where tasks come and go.

std.net moved to Io.net #

const net = Io.net;

// connecting
const host_name = try net.HostName.init(host);
const stream = try host_name.connect(io, port, .{});

// listening
var addr = try net.IpAddress.parse("::", port);
var server = try net.IpAddress.listen(&addr, io, .{ .reuse_address = true });
defer server.deinit(io);

// accepting
const stream = try server.accept(io);

// reading/writing (need wrapper)
var reader = net.Stream.Reader.init(stream, io, &read_buf);
var writer = net.Stream.Writer.init(stream, io, &write_buf);

net.Stream no longer has direct read/writeAll. use Stream.Reader/Stream.Writer.

Evented production experience #

what running many thousands of long-lived connections on stock Io.Evented (0.16.0-dev.3059) teaches (evidence: an AT Protocol relay, ~2,800 PDS connections).

fiber contextSwitch GPF under ReleaseSafe #

stock Io.Evented fibers crash under ReleaseSafe (GPF inside std.Io.fiber.contextSwitch), forcing ReleaseFast. root cause and fixes — two clobber-list bugs in the inline asm — are in fibers.md; until stdlib ships them, a stock Evented deployment is a ReleaseFast deployment.

cross-backend bridging (Evented fibers ↔ Threaded workers) #

the Io interface is backend-agnostic, but you cannot mix execution contexts. Evented fibers cannot safely lock a Threaded mutex — the scheduler accesses thread-local state that doesn't exist in the fiber context.

pattern: bridge with a lock-free MPSC queue using atomics:

[Evented fibers] --atomics→ [ring buffer] --wake→ [Threaded worker pool]

Evented subscriber fibers enqueue work items via atomic CAS. a bounded set of Threaded workers dequeue and execute (e.g., postgres queries). no mutex crossing between backends.

this is the "DbRequestQueue" pattern — decouples the hot networking path (Evented) from blocking I/O (database) that can't run in fibers.

safety checks matter more under Evented #

under Threaded/ReleaseSafe, a bounds error panics with a stack trace pointing to the exact line. under Evented/ReleaseFast (forced by the GPF bug), the same error silently corrupts memory and manifests as a SIGSEGV minutes or hours later with no useful diagnostic.

example: a websocket library assumed \r\n always arrives in a single TCP read; when TCP splits mid-CRLF the next slice has start > end. under ReleaseSafe that is an immediate panic naming the exact line; under ReleaseFast it was silent corruption → SIGSEGV every 30-90 min across ~2,800 connections, and only switching back to Threaded/ReleaseSafe produced the stack trace that identified the bug.

lesson: when forced into ReleaseFast by the fiber GPF, you lose the single most valuable debugging tool zig provides. any bug that would be trivially caught by bounds checking becomes a production mystery.

thread count: Evented vs Threaded #

backend OS threads subscriber tasks RSS
Threaded (ReleaseSafe) ~2,830 ~2,830 ~1.9 GiB
Evented (ReleaseFast) ~47 ~2,830 ~1.2 GiB

Evented runs the same ~2,800 subscriber tasks on ~47 OS threads (bounded worker pool + io_uring event loop). RSS is lower partly due to fewer thread stacks and partly due to ReleaseFast stripping safety metadata.

uring networking patch #

Io.Uring ships with networking functions stubbed out as *Unavailable (return error.NetworkDown). to use Evented for real networking, you need to patch Uring.zig to implement netListenIp, netAccept, netConnectIp, netSend, netRead, netWrite using io_uring opcodes (ACCEPT, CONNECT, SENDMSG, READV, etc.).

note: bind and listen use sync syscalls because IORING_OP_BIND / IORING_OP_LISTEN require kernel 6.11+. DNS resolution (netLookup) is also not patched — subscribers resolve hostnames through a Threaded pool_io fallback.

timing: std.time.Timer is gone #

wall-clock/monotonic timing goes through the io's clock now:

const t0 = std.Io.Clock.awake.now(io);
// ... work ...
const ns = t0.durationTo(std.Io.Clock.awake.now(io)).toNanoseconds();

Io.Timestamp.durationTo returns Io.Duration (has .toNanoseconds() directly); the Clock.Timestamp variant wraps it in .raw. easy to mix up — the compile error names the missing field.

sources #

  • coral — debug_io serialization incident (60 to 4 events/s), 2026