# Benchmarks `just bench` runs a ReleaseSafe binary over the hot path of each subsystem and prints ns/op plus a throughput figure. It is a **local, comparative** tool: the numbers are only meaningful against another run on the same machine, and only the ratio between runs should be quoted. Nothing here is a promotion gate — `docs/semantic-parity.md` owns that. ## Why these, and not others The set mirrors upstream Jetstream V2's own benchmark surface, which is a useful prior: those are the paths they chose to measure after running the system. Upstream has 21 benchmarks in five places — `segment/block_bench_test.go`, `segment/seal_bench_test.go`, `internal/ingest/writer_bench_test.go`, `internal/subscribe/coldfanout_bench_test.go`, and `internal/client/decode*_bench_test.go` — covering block encode/decode, append, flush, seal, reader open, bloom lookup, the writer under backfill and live shapes, cold fan-out, and client decode. Each of ours names the upstream benchmark it corresponds to, so a divergence in shape is visible rather than accidental. | Bench | Subsystem | Upstream counterpart | |---|---|---| | `archive_append_live` | ingest→storage | `BenchmarkWriterLiveShape` | | `archive_append_backfill` | ingest→storage | `BenchmarkWriterBackfillShape` | | `archive_seal` | storage | `BenchmarkSeal` | | `sealed_parse` | storage | `BenchmarkReaderOpen` | | `sealed_parse_unchecked` | storage | `BenchmarkReaderOpenNoVerify` | | `sealed_read_block` | storage | `BenchmarkDecodeBlockSealed` | | `bloom_lookup` | storage | `BenchmarkBlockBloom` | | `wire_encode_v1` / `wire_encode_v2` | serve | `BenchmarkDecodeSegmentEvent*` (mirror side) | | `cold_replay_batch` | serve | `BenchmarkColdFanout` | The table is exactly what `just bench` prints — ten rows, in that order. If the binary prints a different set, this table is stale. Four upstream benchmarks deliberately have no counterpart yet: - `BenchmarkEncodeColumns` / `BenchmarkEncodeBlock` and `BenchmarkDecodeColumns` / `BenchmarkDecodeBlock` — bare block encode/decode in isolation. Ours are covered transitively: `archive_append_*` drives encode and `sealed_read_block` drives decode, both through the real call path rather than a synthetic block. Worth splitting out only if a regression lands in encode/decode specifically and the append/read numbers are too coarse to localize it. - `BenchmarkDeleteCompactionSyntheticArchive` — compaction is measured, but a synthetic-archive harness is a bigger fixture than this file should own. - `BenchmarkFlushToTmpfs` — isolates filesystem cost from encode cost; worth adding when a flush regression is actually suspected. ## Reading the output `ns/op` is wall time per operation. The throughput column is derived from it, not measured independently, and its unit is named per row — rows/s, frames/s, lookups/s, parses/s. Deliberately no bytes/sec figure for the parse benches: opening a sealed segment reads the header and block index, not the file body, so dividing file size by parse time would not measure a throughput the code achieves. The harness allocates through `c_allocator`, which is what `src/main.zig` uses. Benchmarking on `DebugAllocator` measures the allocator instead: `wire_encode_v1` reads 7.19 µs/op there versus 496 ns on `c_allocator` — 14x apart. A sample run, for shape rather than as a target (Apple M-series, 2026-07-25, ReleaseSafe): ``` archive_append_live ingest->storage 245 ns 4.08M rows/s archive_append_backfill ingest->storage 125 ns 7.97M rows/s archive_seal storage 6.17 ms 3.24M rows/s sealed_parse storage 1.39 us 717.51K parses/s sealed_parse_unchecked storage 27 ns 37.62M parses/s sealed_read_block storage 144 us 23.18M rows/s bloom_lookup storage 23 ns 43.13M lookups/s wire_encode_v1 serve 496 ns 2.02M frames/s wire_encode_v2 serve 701 ns 1.43M frames/s cold_replay_batch serve 733 ns 1.36M rows/s ``` The one ratio worth carrying in your head: checksum-verifying a sealed segment open costs ~51x an unchecked one. That is why the cold path opens with `parseUnchecked` and the admission-critical paths do not, and why upstream splits `ReaderOpen` from `ReaderOpenNoVerify`. ("Admission-critical" means a path whose result decides whether a build is allowed to ship — see `docs/semantic-parity.md`. Those paths pay for the checksum because a silently corrupt segment read there would let a broken artifact through. The cold read path is serving traffic that a client can retry, so it takes the 51x instead.) The harness is deliberately simple — a fixed iteration count, no warmup discipline beyond a short prime, no statistical treatment. It catches order-of-magnitude regressions, not differences of a few percent. For precision, use a dedicated measurement. ## Head-to-head vs upstream (2026-08-08, Apple M5 Pro, one run each) Same machine, same afternoon: `zig build bench -Doptimize=ReleaseSafe` here, `go test -bench` at upstream `f29815c`. Single runs — treat as ratios with ~30-40% run variance, not as absolute truth. | path | stream | upstream (Go) | ratio | parity | |---|---|---|---|---| | writer, live shape | 3.20M rows/s | 2.13M events/s (sync) | ~1.5x | clean | | writer, backfill shape | 5.30M rows/s | 4.43M events/s (async8, its best) | ~1.2x | clean | | sealed block decode | 15.0M rows/s | 6.25M rows/s (654µs / 4096-event block) | ~2.4x | clean | | block bloom lookup | 35 ns | 630 ns | ~18x | clean; includes our fixed-4096 bloom, see memory-ceiling note | | cold fan-out, 1 consumer | 849K rows/s | 281K frames/s | ~3x | shapes differ slightly (their bench spans consumer counts) | | seal | 3.04M rows/s (seal step only) | 1.22M events/s (append+flush+seal per op) | directionally ahead | NOT clean — different op boundaries | | reader open | 123 ns | 21.2 µs | — | NOT comparable: ours parses resident bytes, theirs opens and reads a file | Context that keeps these honest: - Upstream jetstream V2 ingests through atmos (`jcalabro/atmos v0.2.14`), which atproto-bench measures directly against zat. The SDK gap is path-dependent: CID-verified decode is zat 249K vs atmos 91K frames/s (~2.7x), but the combined relay hot path — decode + CID verify + commit signature verify — is zat 19.7K vs atmos 18.1K (~1.1x), because ECDSA is ~95% of the frame budget and every implementation pays it. On the signature-verifying ingest path the language advantage mostly washes out; it survives on the sig-free paths (format work, serving, scans), which is what the table above measures — the zat-vs-their-Go format comparison is the block-decode row (2.4x), via `atproto-bench/go-jss` driving their `segment` package on the same sealed-segment shape. - The live network runs ~200-450 events/s. Every row above is in the millions/s: CPU is not the live tail's bottleneck for either implementation. Where the margin actually pays: replay serving, cold fan-out, compaction/rebuild scans, and box size for the same workload. - Nothing here is a promotion gate; parity rows marked NOT clean need a matched-shape bench before being quoted as a ratio.