# why delete compaction is slow Measured on production (hel1, cx43, 8 vCPU) on 2026-08-10, against build `22bf62c`, while the archive held 6,584 sealed segments / ~23.2B events. The short version: **a pass rewrites the whole archive to apply the deletes of a much smaller window**, and nothing in the design bounds that. Everything else here is a multiplier on top of it. ## the measurements | quantity | value | how | | --- | --- | --- | | sealed segments | 6,584 | `/status` | | events per segment | 3.33M median (1.8M–3.6M) | `listSegments`, 100-segment sample | | segment size | ~277 MB | `getSegment` | | uncompacted window | 100.9M events ≈ **30 segments** | tip 23,236,818,914 − watermark 23,135,908,535 | | segments rewritten per pass | **6,584** | `pass.zig` passes all of `metas` to the rewrite | | live tombstones | 3.69M entries / 472 MB | `jetstream_compaction_tombstone_set_entries` | | watermark lag | 2.3 days | `jetstream_compaction_watermark_lag_seconds` | | observed rate | **3.4–3.5 segments/min** | `segments_examined_total` sampled over hours | | box state while compacting | 56% idle, iowait 1.1%, stream at 245% CPU | `top`, 8 cores | So: **~220x amplification** (6,584 rewritten to apply the deletes found in ~30), CPU-bound on zstd, not disk. ## why it is O(archive) and not O(new data) A delete is retroactive. A tombstone written today can target a record written any time in the past, so "what is new" does not bound "what must be rewritten." The fold is over the new window; the rewrite is over everything. The mechanism meant to cut that down is the candidate-DID bloom prefilter, and at this scale it does nothing: - `Snapshot.candidateDids` returns `null` once distinct DIDs exceed `bloom_narrow_max_dids` (100,000). With 3.69M tombstones it is never close. - Raising the cap would not help and would cost memory. The test is *"does this segment contain any tombstoned DID"*, but a segment holds 3.3M events from a large slice of the network, so with millions of candidate DIDs essentially every segment matches. Expected tombstoned rows per segment ≈ 3.69M / 6,584 ≈ **560** — no segment is clean. - The real mismatch: the prefilter is **DID-level** while the unit of deletion is `(did, collection, rkey)`. A precise-enough skip test is the only thing that makes this sublinear. **Upstream is identical.** Verified against `bluesky-social/jetstream` at the pinned `f29815c`: `applyCompactionChunk` fans out over all `sealed` segments, and `compact_deletes.go:342` does the same `candidateDIDs = nil` past the same 100k default. This amplification is inherited, not a porting mistake. ## the feedback loop that made it permanent This is the part that matters for prevention. Until 2026-08-10 the watermark was saved **once per chunk**, and a chunk was the whole archive (bounded only by a 32M tombstone cap against a 3.7M live set). At 3.4 segments/min a pass needs far longer than the interval between process restarts, so: restart → refold → rewrite for hours → killed by the next deploy → watermark never moves → backlog grows → more tombstones → slower pass → repeat. `jetstream_compaction_passes_total` sat at **0 across two binaries and 4+ days** while `segments_examined_total` climbed the whole time. The pass was never stuck; it was never finishing. A counter that only increments on completion cannot distinguish those two, which is why this went unexplained twice. ## what was fixed, and what remains Fixed 2026-08-10: - `7136ebb` — commit the watermark every 64 segments instead of once per chunk. Each segment is still rewritten once; only commit granularity changes, so a restart costs at most a batch instead of everything. Breaks the loop above. - `b3ea690` — reuse a block's compressed frame when it drops nothing. The port recompressed every block unconditionally and discarded the result when the segment turned out clean; upstream's `segment/rewrite.go` keeps the source frame. ~810 blocks per segment, ~50% of them untouched, and zstd compression is the expensive direction. Still open, in the order I would do them: 1. **Rewrite workers.** Production runs 2. Upstream's default and ours is `min(NumCPU, 8)`; the 2 is deployment config, on a box measured 56% idle. Memory is the real constraint — each worker holds a segment's output buffer. 2. **A precise skip test.** Per-segment or per-block filters keyed on the record identity rather than the DID, so old segments can be rejected without being decompressed. This is the only fix that changes the complexity class. 3. **Bound the backlog, not just the pass.** See below. ## how to not get here again The failure was not slowness; it was slowness that compounded unobserved. Watch for the loop, not the symptom: - **Alert on `jetstream_compaction_watermark_lag_seconds`.** It is the honest measure of exposure: everything witnessed inside that window still serves deleted rows, because nothing in the read path consults the tombstone set — rewriting the segment is what removes a row. - **Alert on pass duration approaching the interval.** Once a pass cannot finish within the interval, the backlog only grows, and every restart is a full reset of the work. - **Treat `passes_total == 0` with a climbing `segments_examined_total` as in-progress, not idle.** Opposite responses. - The tombstone set is also a memory cost: 3.69M entries held 472 MB of a 15.6 GB box, and it only shrinks when the watermark advances. ## 2026-08-17 re-measurement: the bottleneck moved The fix that this doc's measurements predate (b3ea690, frame reuse for blocks that drop nothing — landed the same day this doc was written) changed the answer. Re-measured on the same production box during the first completed 4h-interval pass: | quantity | 2026-08-10 (pre frame-reuse) | 2026-08-17 | | --- | --- | --- | | observed rate | 3.4 segments/min | ~20 segments/min | | segments rewritten | 6,584 of 6,584 | 5,008 of 5,252 examined (244 clean) | | bytes moved per pass | (not measured) | ~1.7 TB read + 1.34 TB written | | rows dropped per pass | — | 2.3M (~0.01% of the archive) | | bottleneck | CPU (zstd recompress) | **disk** | The disk evidence: a single-stream cold read of one 270 MB segment runs ~105 MB/s while the pass's 4 workers are active, and the pass aggregate (~3 TB moved in ~4h15m) is ~200 MB/s sustained — `/data` is a network-attached cloud volume, and that is its ceiling under this load. zstd no longer dominates: with tombstones scattered thin (hundreds of dropped rows across ~1,600 blocks per segment), most blocks reuse their compressed frame verbatim. Consequences: - **Pass duration (5h27m measured, first completed pass 2026-08-17 06:16Z: 6,358/6,666 rewritten, 3.56M rows dropped) exceeds the 4h interval — but the loop is sequential and the timer resets AFTER a pass completes**, so the real cadence is pass+interval ≈ 9.9h per cycle (~55% duty), not continuous. Upstream's loop behaves identically at this archive size — shared design cost, not a config error. The interval stays at upstream's 4h default (decided 2026-08-16; do not re-litigate). - **Age-tiered skipping is falsified by the archive's demographics** (first completed pass's rewritten-age histogram): only 84 rewritten segments were younger than 7 days; 6,274 sat in the 7–30d bucket. Steady state seals only ~6-7 segments/day, so the archive is overwhelmingly the backfill-era mass and that is where tombstoned rows live. Tiering by age would exempt almost nothing. - **More rewrite workers would not help this box**: the volume is saturated at 4. (`--compaction-rewrite-workers=4` vs upstream's default 8 is now a non-difference here.) - The O(archive) amplification analysis above still holds exactly — the fold is O(new data), the rewrite is O(archive BYTES THROUGH THE DISK), and frame reuse cut the CPU multiplier, not the byte volume. The sublinear path remains a skip test precise enough that untouched segments are never read — see the FP-union caveat before assuming blooms get there (a candidate set of ~800K record keys unions bloom false positives: at fp=1e-6, expected ~0.8 false hits per segment — every segment matches again; exact per-segment indexes or a threshold-triggered lazy rewrite are the honest design space).