jetstream v2 in zig stream.waow.tech
stream docs benchmarks.md
7.4 kB
Markdown

Benchmarks #

just bench runs a ReleaseSafe binary over the hot path of each subsystem and prints ns/op plus a throughput figure. It is a local, comparative tool: the numbers are only meaningful against another run on the same machine, and only the ratio between runs should be quoted. Nothing here is a promotion gate — docs/semantic-parity.md owns that.

Why these, and not others #

The set mirrors upstream Jetstream V2's own benchmark surface, which is a useful prior: those are the paths they chose to measure after running the system. Upstream has 21 benchmarks in five places — segment/block_bench_test.go, segment/seal_bench_test.go, internal/ingest/writer_bench_test.go, internal/subscribe/coldfanout_bench_test.go, and internal/client/decode*_bench_test.go — covering block encode/decode, append, flush, seal, reader open, bloom lookup, the writer under backfill and live shapes, cold fan-out, and client decode.

Each of ours names the upstream benchmark it corresponds to, so a divergence in shape is visible rather than accidental.

Bench Subsystem Upstream counterpart
archive_append_live ingest→storage BenchmarkWriterLiveShape
archive_append_backfill ingest→storage BenchmarkWriterBackfillShape
archive_seal storage BenchmarkSeal
sealed_parse storage BenchmarkReaderOpen
sealed_parse_unchecked storage BenchmarkReaderOpenNoVerify
sealed_read_block storage BenchmarkDecodeBlockSealed
bloom_lookup storage BenchmarkBlockBloom
wire_encode_v1 / wire_encode_v2 serve BenchmarkDecodeSegmentEvent* (mirror side)
cold_replay_batch serve BenchmarkColdFanout

The table is exactly what just bench prints — ten rows, in that order. If the binary prints a different set, this table is stale.

Four upstream benchmarks deliberately have no counterpart yet:

  • BenchmarkEncodeColumns / BenchmarkEncodeBlock and BenchmarkDecodeColumns / BenchmarkDecodeBlock — bare block encode/decode in isolation. Ours are covered transitively: archive_append_* drives encode and sealed_read_block drives decode, both through the real call path rather than a synthetic block. Worth splitting out only if a regression lands in encode/decode specifically and the append/read numbers are too coarse to localize it.
  • BenchmarkDeleteCompactionSyntheticArchive — compaction is measured, but a synthetic-archive harness is a bigger fixture than this file should own.
  • BenchmarkFlushToTmpfs — isolates filesystem cost from encode cost; worth adding when a flush regression is actually suspected.

Reading the output #

ns/op is wall time per operation. The throughput column is derived from it, not measured independently, and its unit is named per row — rows/s, frames/s, lookups/s, parses/s. Deliberately no bytes/sec figure for the parse benches: opening a sealed segment reads the header and block index, not the file body, so dividing file size by parse time would not measure a throughput the code achieves.

The harness allocates through c_allocator, which is what src/main.zig uses. Benchmarking on DebugAllocator measures the allocator instead: wire_encode_v1 reads 7.19 µs/op there versus 496 ns on c_allocator — 14x apart.

A sample run, for shape rather than as a target (Apple M-series, 2026-07-25, ReleaseSafe):

archive_append_live        ingest->storage       245 ns    4.08M rows/s
archive_append_backfill    ingest->storage       125 ns    7.97M rows/s
archive_seal               storage              6.17 ms    3.24M rows/s
sealed_parse               storage              1.39 us  717.51K parses/s
sealed_parse_unchecked     storage                27 ns   37.62M parses/s
sealed_read_block          storage               144 us   23.18M rows/s
bloom_lookup               storage                23 ns   43.13M lookups/s
wire_encode_v1             serve                 496 ns    2.02M frames/s
wire_encode_v2             serve                 701 ns    1.43M frames/s
cold_replay_batch          serve                 733 ns    1.36M rows/s

The one ratio worth carrying in your head: checksum-verifying a sealed segment open costs ~51x an unchecked one. That is why the cold path opens with parseUnchecked and the admission-critical paths do not, and why upstream splits ReaderOpen from ReaderOpenNoVerify.

("Admission-critical" means a path whose result decides whether a build is allowed to ship — see docs/semantic-parity.md. Those paths pay for the checksum because a silently corrupt segment read there would let a broken artifact through. The cold read path is serving traffic that a client can retry, so it takes the 51x instead.)

The harness is deliberately simple — a fixed iteration count, no warmup discipline beyond a short prime, no statistical treatment. It catches order-of-magnitude regressions, not differences of a few percent. For precision, use a dedicated measurement.

Head-to-head vs upstream (2026-08-08, Apple M5 Pro, one run each) #

Same machine, same afternoon: zig build bench -Doptimize=ReleaseSafe here, go test -bench at upstream f29815c. Single runs — treat as ratios with ~30-40% run variance, not as absolute truth.

path stream upstream (Go) ratio parity
writer, live shape 3.20M rows/s 2.13M events/s (sync) ~1.5x clean
writer, backfill shape 5.30M rows/s 4.43M events/s (async8, its best) ~1.2x clean
sealed block decode 15.0M rows/s 6.25M rows/s (654µs / 4096-event block) ~2.4x clean
block bloom lookup 35 ns 630 ns ~18x clean; includes our fixed-4096 bloom, see memory-ceiling note
cold fan-out, 1 consumer 849K rows/s 281K frames/s ~3x shapes differ slightly (their bench spans consumer counts)
seal 3.04M rows/s (seal step only) 1.22M events/s (append+flush+seal per op) directionally ahead NOT clean — different op boundaries
reader open 123 ns 21.2 µs — NOT comparable: ours parses resident bytes, theirs opens and reads a file

Context that keeps these honest:

  • Upstream jetstream V2 ingests through atmos (jcalabro/atmos v0.2.14), which atproto-bench measures directly against zat. The SDK gap is path-dependent: CID-verified decode is zat 249K vs atmos 91K frames/s (~2.7x), but the combined relay hot path — decode + CID verify + commit signature verify — is zat 19.7K vs atmos 18.1K (~1.1x), because ECDSA is ~95% of the frame budget and every implementation pays it. On the signature-verifying ingest path the language advantage mostly washes out; it survives on the sig-free paths (format work, serving, scans), which is what the table above measures — the zat-vs-their-Go format comparison is the block-decode row (2.4x), via atproto-bench/go-jss driving their segment package on the same sealed-segment shape.
  • The live network runs ~200-450 events/s. Every row above is in the millions/s: CPU is not the live tail's bottleneck for either implementation. Where the margin actually pays: replay serving, cold fan-out, compaction/rebuild scans, and box size for the same workload.
  • Nothing here is a promotion gate; parity rows marked NOT clean need a matched-shape bench before being quoted as a ratio.