Benchmarks #
just bench runs a ReleaseSafe binary over the hot path of each subsystem and
prints ns/op plus a throughput figure. It is a local, comparative tool: the
numbers are only meaningful against another run on the same machine, and only
the ratio between runs should be quoted. Nothing here is a promotion gate —
docs/semantic-parity.md owns that.
Why these, and not others #
The set mirrors upstream Jetstream V2's own benchmark surface, which is a
useful prior: those are the paths they chose to measure after running the
system. Upstream has 21 benchmarks in five places —
segment/block_bench_test.go, segment/seal_bench_test.go,
internal/ingest/writer_bench_test.go,
internal/subscribe/coldfanout_bench_test.go, and
internal/client/decode*_bench_test.go — covering block encode/decode, append,
flush, seal, reader open, bloom lookup, the writer under backfill and live
shapes, cold fan-out, and client decode.
Each of ours names the upstream benchmark it corresponds to, so a divergence in shape is visible rather than accidental.
| Bench | Subsystem | Upstream counterpart |
|---|---|---|
archive_append_live |
ingest→storage | BenchmarkWriterLiveShape |
archive_append_backfill |
ingest→storage | BenchmarkWriterBackfillShape |
archive_seal |
storage | BenchmarkSeal |
sealed_parse |
storage | BenchmarkReaderOpen |
sealed_parse_unchecked |
storage | BenchmarkReaderOpenNoVerify |
sealed_read_block |
storage | BenchmarkDecodeBlockSealed |
bloom_lookup |
storage | BenchmarkBlockBloom |
wire_encode_v1 / wire_encode_v2 |
serve | BenchmarkDecodeSegmentEvent* (mirror side) |
cold_replay_batch |
serve | BenchmarkColdFanout |
The table is exactly what just bench prints — ten rows, in that order. If
the binary prints a different set, this table is stale.
Four upstream benchmarks deliberately have no counterpart yet:
BenchmarkEncodeColumns/BenchmarkEncodeBlockandBenchmarkDecodeColumns/BenchmarkDecodeBlock— bare block encode/decode in isolation. Ours are covered transitively:archive_append_*drives encode andsealed_read_blockdrives decode, both through the real call path rather than a synthetic block. Worth splitting out only if a regression lands in encode/decode specifically and the append/read numbers are too coarse to localize it.BenchmarkDeleteCompactionSyntheticArchive— compaction is measured, but a synthetic-archive harness is a bigger fixture than this file should own.BenchmarkFlushToTmpfs— isolates filesystem cost from encode cost; worth adding when a flush regression is actually suspected.
Reading the output #
ns/op is wall time per operation. The throughput column is derived from it,
not measured independently, and its unit is named per row — rows/s, frames/s,
lookups/s, parses/s. Deliberately no bytes/sec figure for the parse benches:
opening a sealed segment reads the header and block index, not the file body,
so dividing file size by parse time would not measure a throughput the code
achieves.
The harness allocates through c_allocator, which is what src/main.zig uses.
Benchmarking on DebugAllocator measures the allocator instead:
wire_encode_v1 reads 7.19 µs/op there versus 496 ns on c_allocator — 14x
apart.
A sample run, for shape rather than as a target (Apple M-series, 2026-07-25, ReleaseSafe):
archive_append_live ingest->storage 245 ns 4.08M rows/s
archive_append_backfill ingest->storage 125 ns 7.97M rows/s
archive_seal storage 6.17 ms 3.24M rows/s
sealed_parse storage 1.39 us 717.51K parses/s
sealed_parse_unchecked storage 27 ns 37.62M parses/s
sealed_read_block storage 144 us 23.18M rows/s
bloom_lookup storage 23 ns 43.13M lookups/s
wire_encode_v1 serve 496 ns 2.02M frames/s
wire_encode_v2 serve 701 ns 1.43M frames/s
cold_replay_batch serve 733 ns 1.36M rows/s
The one ratio worth carrying in your head: checksum-verifying a sealed segment
open costs ~51x an unchecked one. That is why the cold path opens with
parseUnchecked and the admission-critical paths do not, and why upstream
splits ReaderOpen from ReaderOpenNoVerify.
("Admission-critical" means a path whose result decides whether a build is
allowed to ship — see docs/semantic-parity.md. Those paths pay for the
checksum because a silently corrupt segment read there would let a broken
artifact through. The cold read path is serving traffic that a client can
retry, so it takes the 51x instead.)
The harness is deliberately simple — a fixed iteration count, no warmup discipline beyond a short prime, no statistical treatment. It catches order-of-magnitude regressions, not differences of a few percent. For precision, use a dedicated measurement.
Head-to-head vs upstream (2026-08-08, Apple M5 Pro, one run each) #
Same machine, same afternoon: zig build bench -Doptimize=ReleaseSafe here,
go test -bench at upstream f29815c. Single runs — treat as ratios with
~30-40% run variance, not as absolute truth.
| path | stream | upstream (Go) | ratio | parity |
|---|---|---|---|---|
| writer, live shape | 3.20M rows/s | 2.13M events/s (sync) | ~1.5x | clean |
| writer, backfill shape | 5.30M rows/s | 4.43M events/s (async8, its best) | ~1.2x | clean |
| sealed block decode | 15.0M rows/s | 6.25M rows/s (654µs / 4096-event block) | ~2.4x | clean |
| block bloom lookup | 35 ns | 630 ns | ~18x | clean; includes our fixed-4096 bloom, see memory-ceiling note |
| cold fan-out, 1 consumer | 849K rows/s | 281K frames/s | ~3x | shapes differ slightly (their bench spans consumer counts) |
| seal | 3.04M rows/s (seal step only) | 1.22M events/s (append+flush+seal per op) | directionally ahead | NOT clean — different op boundaries |
| reader open | 123 ns | 21.2 µs | — | NOT comparable: ours parses resident bytes, theirs opens and reads a file |
Context that keeps these honest:
- Upstream jetstream V2 ingests through atmos (
jcalabro/atmos v0.2.14), which atproto-bench measures directly against zat. The SDK gap is path-dependent: CID-verified decode is zat 249K vs atmos 91K frames/s (~2.7x), but the combined relay hot path — decode + CID verify + commit signature verify — is zat 19.7K vs atmos 18.1K (~1.1x), because ECDSA is ~95% of the frame budget and every implementation pays it. On the signature-verifying ingest path the language advantage mostly washes out; it survives on the sig-free paths (format work, serving, scans), which is what the table above measures — the zat-vs-their-Go format comparison is the block-decode row (2.4x), viaatproto-bench/go-jssdriving theirsegmentpackage on the same sealed-segment shape. - The live network runs ~200-450 events/s. Every row above is in the millions/s: CPU is not the live tail's bottleneck for either implementation. Where the margin actually pays: replay serving, cold fan-out, compaction/rebuild scans, and box size for the same workload.
- Nothing here is a promotion gate; parity rows marked NOT clean need a matched-shape bench before being quoted as a ratio.