about things notes.zzstoatzz.io
notes
notes languages ziglang page-allocator-granularity.md
6.0 kB

page_allocator granularity and "small-alloc leak amplification" #

written against zig 0.16.

Why dropping a []Value backing on std.heap.page_allocator once a second turns into a ~0.6 MiB/min RSS leak.

the symptom #

prefect-server (zig 0.16, alpine + musl, idle):

RSS slope:                       +10,886 B/s   (+0.62 MiB/min, ~7 GiB/hour at sustained 200 RPS)
total broker-consumer iter rate:  2.92/s
bytes per iteration:              +3,556 B     (~3.5 KiB/iter)

Three broker consumers were each polling redis once per second (XAUTOCLAIM + XREADGROUP with a 1-second block). Every iteration, the redis client library produced an array-shaped response, which was being persistently held.

zero-length allocations are free #

A suspect that can be ruled out immediately: alloc(T, 0). The standard Allocator interface short-circuits zero-byte requests before they ever reach the backing allocator:

// std/mem/Allocator.zig:295-298
fn allocBytesWithAlignment(...) Error![*]align(alignment.toByteUnits()) u8 {
    if (byte_count == 0) {
        const ptr = comptime alignment.backward(math.maxInt(usize));
        return @as([*]align(alignment.toByteUnits()) u8, @ptrFromInt(ptr));
    }
    const byte_ptr = self.rawAlloc(byte_count, alignment, return_address) orelse return error.OutOfMemory;
    ...
}

alloc(T, 0) returns a synthetic max-aligned-down pointer. No mmap, no page consumed. The zero-length call is essentially free (verified in std/mem/Allocator.zig:289-303 on zig 0.16), so do not contort code to avoid it.

the actual footgun #

Look at std/heap/PageAllocator.zig:109:

const page_aligned_len = mem.alignForward(usize, n, page_size);

Every non-zero allocation is rounded up to a page. A 72-byte request (alloc(Value, 3) with @sizeOf(Value) == 24) consumes a full 4 KiB page on most linux targets, 16 KiB on macOS-arm64.

Combine that with a code path that allocates small structures through page_allocator and doesn't reliably free them — for example, a protocol-client array result tree where the leaves alias another buffer (see per-command-arena.md for the redis case) — and the kernel charges you a page per allocation, regardless of how little payload you asked for.

Arithmetic from the prefect-server hunt:

~1 XAUTOCLAIM/sec/consumer
  × 3 consumers
  × 1 array response per call
  × ~1 page (4 KiB on aarch64-linux) per array response, never freed
≈ 12 KiB/sec → ~0.7 MiB/min   (we measured +0.62 MiB/min)

The cost is page-granular even for tiny allocations. Wasting bytes is fine; wasting pages many times per second is what scales into a real leak.

what zig 0.16 says about this #

The release notes explicitly call out page_allocator as a workaround, not a default:

When upgrading code, if you find yourself without access to an Io instance, you can get one like this:

var threaded: Io.Threaded = .init_single_threaded;
const io = threaded.io();

This works as long as you don't need task-level concurrency, however, it is a non-ideal workaround — like reaching for std.heap.page_allocator when you need an Allocator and do not have one. Instead, it is better to accept an Io parameter if you need one [...].

— Zig 0.16 release notes, "IO as an Interface"

The same release notes also call out ThreadSafeAllocator as "an anti-pattern" and document that ArenaAllocator is now thread-safe and lock-free, which is exactly what makes the per-command-arena pattern in per-command-arena.md compose cleanly inside a stateful client without needing an Io or mutex.

three rules #

  1. alloc(T, 0) is fine. Sentinel pointer, no syscall. Do not contort code to avoid it.
  2. page_allocator.alloc(SmallStruct, N) is not fine when called often. Every request consumes a full page minimum. If those allocations aren't paired with reliable frees, RSS grows page-granularly.
  3. For "many small allocations with one lifetime" → ArenaAllocator on top of page_allocator. The arena keeps the pages allocated and reuses them, so steady-state memory plateaus. arena.reset(.retain_capacity) pairs naturally with protocol command boundaries (see per-command-arena.md). On zig 0.16 the arena is lock-free, so it's safe to embed in a single-threaded client struct without extra synchronization.

For a general-purpose default allocator, std.heap.smp_allocator (or std.heap.DebugAllocator during a leak hunt) is the right reach. page_allocator is the substrate, not the application allocator.

how to spot this empirically #

Two metrics make this kind of leak attributable:

  • prefect_process_rss_anon_bytes (from /proc/self/status's RssAnon) — total heap, includes every glibc per-thread arena and every mmap'd page_allocator page. The source-of-truth gauge for "is RSS growing."
  • An iteration counter on the suspect loop (in the prefect-server case: prefect_broker_consumer_iterations_*_total, three flat per-topic counters).

Compute (RSS slope) / (iteration rate) over an idle window. If it lands at a multiple of the page size, you're paying page granularity for small allocations that aren't getting reclaimed. The fix is allocator shape, not call-site cleanup.

sources #

  • ~/.local/share/zigup/0.16.0/files/lib/std/mem/Allocator.zig:289-303 — zero-byte short-circuit
  • ~/.local/share/zigup/0.16.0/files/lib/std/heap/PageAllocator.zig:109 — alignForward to page
  • Zig 0.16 release notes — heap.ArenaAllocator Becomes Thread-Safe and Lock-Free, heap.ThreadSafe Allocator Removed
  • per-command-arena.md — the pattern that fixes this in the redis client
  • prefect-server's docs/experiments.md (exp-001 through exp-005) and docs/leak-investigation-plan.md — the empirical write-up that produced this note