jetstream v2 in zig stream.waow.tech
stream docs compaction-cost.md
8.4 kB
Markdown

why delete compaction is slow #

Measured on production (hel1, cx43, 8 vCPU) on 2026-08-10, against build 22bf62c, while the archive held 6,584 sealed segments / ~23.2B events.

The short version: a pass rewrites the whole archive to apply the deletes of a much smaller window, and nothing in the design bounds that. Everything else here is a multiplier on top of it.

the measurements #

quantity value how
sealed segments 6,584 /status
events per segment 3.33M median (1.8M–3.6M) listSegments, 100-segment sample
segment size ~277 MB getSegment
uncompacted window 100.9M events ≈ 30 segments tip 23,236,818,914 − watermark 23,135,908,535
segments rewritten per pass 6,584 pass.zig passes all of metas to the rewrite
live tombstones 3.69M entries / 472 MB jetstream_compaction_tombstone_set_entries
watermark lag 2.3 days jetstream_compaction_watermark_lag_seconds
observed rate 3.4–3.5 segments/min segments_examined_total sampled over hours
box state while compacting 56% idle, iowait 1.1%, stream at 245% CPU top, 8 cores

So: ~220x amplification (6,584 rewritten to apply the deletes found in ~30), CPU-bound on zstd, not disk.

why it is O(archive) and not O(new data) #

A delete is retroactive. A tombstone written today can target a record written any time in the past, so "what is new" does not bound "what must be rewritten." The fold is over the new window; the rewrite is over everything.

The mechanism meant to cut that down is the candidate-DID bloom prefilter, and at this scale it does nothing:

  • Snapshot.candidateDids returns null once distinct DIDs exceed bloom_narrow_max_dids (100,000). With 3.69M tombstones it is never close.
  • Raising the cap would not help and would cost memory. The test is "does this segment contain any tombstoned DID", but a segment holds 3.3M events from a large slice of the network, so with millions of candidate DIDs essentially every segment matches. Expected tombstoned rows per segment ≈ 3.69M / 6,584 ≈ 560 — no segment is clean.
  • The real mismatch: the prefilter is DID-level while the unit of deletion is (did, collection, rkey). A precise-enough skip test is the only thing that makes this sublinear.

Upstream is identical. Verified against bluesky-social/jetstream at the pinned f29815c: applyCompactionChunk fans out over all sealed segments, and compact_deletes.go:342 does the same candidateDIDs = nil past the same 100k default. This amplification is inherited, not a porting mistake.

the feedback loop that made it permanent #

This is the part that matters for prevention. Until 2026-08-10 the watermark was saved once per chunk, and a chunk was the whole archive (bounded only by a 32M tombstone cap against a 3.7M live set). At 3.4 segments/min a pass needs far longer than the interval between process restarts, so:

restart → refold → rewrite for hours → killed by the next deploy → watermark never moves → backlog grows → more tombstones → slower pass → repeat.

jetstream_compaction_passes_total sat at 0 across two binaries and 4+ days while segments_examined_total climbed the whole time. The pass was never stuck; it was never finishing. A counter that only increments on completion cannot distinguish those two, which is why this went unexplained twice.

what was fixed, and what remains #

Fixed 2026-08-10:

  • 7136ebb — commit the watermark every 64 segments instead of once per chunk. Each segment is still rewritten once; only commit granularity changes, so a restart costs at most a batch instead of everything. Breaks the loop above.
  • b3ea690 — reuse a block's compressed frame when it drops nothing. The port recompressed every block unconditionally and discarded the result when the segment turned out clean; upstream's segment/rewrite.go keeps the source frame. ~810 blocks per segment, ~50% of them untouched, and zstd compression is the expensive direction.

Still open, in the order I would do them:

  1. Rewrite workers. Production runs 2. Upstream's default and ours is min(NumCPU, 8); the 2 is deployment config, on a box measured 56% idle. Memory is the real constraint — each worker holds a segment's output buffer.
  2. A precise skip test. Per-segment or per-block filters keyed on the record identity rather than the DID, so old segments can be rejected without being decompressed. This is the only fix that changes the complexity class.
  3. Bound the backlog, not just the pass. See below.

how to not get here again #

The failure was not slowness; it was slowness that compounded unobserved. Watch for the loop, not the symptom:

  • Alert on jetstream_compaction_watermark_lag_seconds. It is the honest measure of exposure: everything witnessed inside that window still serves deleted rows, because nothing in the read path consults the tombstone set — rewriting the segment is what removes a row.
  • Alert on pass duration approaching the interval. Once a pass cannot finish within the interval, the backlog only grows, and every restart is a full reset of the work.
  • Treat passes_total == 0 with a climbing segments_examined_total as in-progress, not idle. Opposite responses.
  • The tombstone set is also a memory cost: 3.69M entries held 472 MB of a 15.6 GB box, and it only shrinks when the watermark advances.

2026-08-17 re-measurement: the bottleneck moved #

The fix that this doc's measurements predate (b3ea690, frame reuse for blocks that drop nothing — landed the same day this doc was written) changed the answer. Re-measured on the same production box during the first completed 4h-interval pass:

quantity 2026-08-10 (pre frame-reuse) 2026-08-17
observed rate 3.4 segments/min ~20 segments/min
segments rewritten 6,584 of 6,584 5,008 of 5,252 examined (244 clean)
bytes moved per pass (not measured) ~1.7 TB read + 1.34 TB written
rows dropped per pass — 2.3M (~0.01% of the archive)
bottleneck CPU (zstd recompress) disk

The disk evidence: a single-stream cold read of one 270 MB segment runs ~105 MB/s while the pass's 4 workers are active, and the pass aggregate (~3 TB moved in ~4h15m) is ~200 MB/s sustained — /data is a network-attached cloud volume, and that is its ceiling under this load. zstd no longer dominates: with tombstones scattered thin (hundreds of dropped rows across ~1,600 blocks per segment), most blocks reuse their compressed frame verbatim.

Consequences:

  • Pass duration (5h27m measured, first completed pass 2026-08-17 06:16Z: 6,358/6,666 rewritten, 3.56M rows dropped) exceeds the 4h interval — but the loop is sequential and the timer resets AFTER a pass completes, so the real cadence is pass+interval ≈ 9.9h per cycle (~55% duty), not continuous. Upstream's loop behaves identically at this archive size — shared design cost, not a config error. The interval stays at upstream's 4h default (decided 2026-08-16; do not re-litigate).
  • Age-tiered skipping is falsified by the archive's demographics (first completed pass's rewritten-age histogram): only 84 rewritten segments were younger than 7 days; 6,274 sat in the 7–30d bucket. Steady state seals only ~6-7 segments/day, so the archive is overwhelmingly the backfill-era mass and that is where tombstoned rows live. Tiering by age would exempt almost nothing.
  • More rewrite workers would not help this box: the volume is saturated at 4. (--compaction-rewrite-workers=4 vs upstream's default 8 is now a non-difference here.)
  • The O(archive) amplification analysis above still holds exactly — the fold is O(new data), the rewrite is O(archive BYTES THROUGH THE DISK), and frame reuse cut the CPU multiplier, not the byte volume. The sublinear path remains a skip test precise enough that untouched segments are never read — see the FP-union caveat before assuming blooms get there (a candidate set of ~800K record keys unions bloom false positives: at fp=1e-6, expected ~0.8 false hits per segment — every segment matches again; exact per-segment indexes or a threshold-triggered lazy rewrite are the honest design space).