why delete compaction is slow #
Measured on production (hel1, cx43, 8 vCPU) on 2026-08-10, against build
22bf62c, while the archive held 6,584 sealed segments / ~23.2B events.
The short version: a pass rewrites the whole archive to apply the deletes of a much smaller window, and nothing in the design bounds that. Everything else here is a multiplier on top of it.
the measurements #
| quantity | value | how |
|---|---|---|
| sealed segments | 6,584 | /status |
| events per segment | 3.33M median (1.8M–3.6M) | listSegments, 100-segment sample |
| segment size | ~277 MB | getSegment |
| uncompacted window | 100.9M events ≈ 30 segments | tip 23,236,818,914 − watermark 23,135,908,535 |
| segments rewritten per pass | 6,584 | pass.zig passes all of metas to the rewrite |
| live tombstones | 3.69M entries / 472 MB | jetstream_compaction_tombstone_set_entries |
| watermark lag | 2.3 days | jetstream_compaction_watermark_lag_seconds |
| observed rate | 3.4–3.5 segments/min | segments_examined_total sampled over hours |
| box state while compacting | 56% idle, iowait 1.1%, stream at 245% CPU | top, 8 cores |
So: ~220x amplification (6,584 rewritten to apply the deletes found in ~30), CPU-bound on zstd, not disk.
why it is O(archive) and not O(new data) #
A delete is retroactive. A tombstone written today can target a record written any time in the past, so "what is new" does not bound "what must be rewritten." The fold is over the new window; the rewrite is over everything.
The mechanism meant to cut that down is the candidate-DID bloom prefilter, and at this scale it does nothing:
Snapshot.candidateDidsreturnsnullonce distinct DIDs exceedbloom_narrow_max_dids(100,000). With 3.69M tombstones it is never close.- Raising the cap would not help and would cost memory. The test is "does this segment contain any tombstoned DID", but a segment holds 3.3M events from a large slice of the network, so with millions of candidate DIDs essentially every segment matches. Expected tombstoned rows per segment ≈ 3.69M / 6,584 ≈ 560 — no segment is clean.
- The real mismatch: the prefilter is DID-level while the unit of deletion
is
(did, collection, rkey). A precise-enough skip test is the only thing that makes this sublinear.
Upstream is identical. Verified against bluesky-social/jetstream at the
pinned f29815c: applyCompactionChunk fans out over all sealed segments,
and compact_deletes.go:342 does the same candidateDIDs = nil past the same
100k default. This amplification is inherited, not a porting mistake.
the feedback loop that made it permanent #
This is the part that matters for prevention. Until 2026-08-10 the watermark was saved once per chunk, and a chunk was the whole archive (bounded only by a 32M tombstone cap against a 3.7M live set). At 3.4 segments/min a pass needs far longer than the interval between process restarts, so:
restart → refold → rewrite for hours → killed by the next deploy → watermark never moves → backlog grows → more tombstones → slower pass → repeat.
jetstream_compaction_passes_total sat at 0 across two binaries and 4+ days
while segments_examined_total climbed the whole time. The pass was never
stuck; it was never finishing. A counter that only increments on completion
cannot distinguish those two, which is why this went unexplained twice.
what was fixed, and what remains #
Fixed 2026-08-10:
7136ebb— commit the watermark every 64 segments instead of once per chunk. Each segment is still rewritten once; only commit granularity changes, so a restart costs at most a batch instead of everything. Breaks the loop above.b3ea690— reuse a block's compressed frame when it drops nothing. The port recompressed every block unconditionally and discarded the result when the segment turned out clean; upstream'ssegment/rewrite.gokeeps the source frame. ~810 blocks per segment, ~50% of them untouched, and zstd compression is the expensive direction.
Still open, in the order I would do them:
- Rewrite workers. Production runs 2. Upstream's default and ours is
min(NumCPU, 8); the 2 is deployment config, on a box measured 56% idle. Memory is the real constraint — each worker holds a segment's output buffer. - A precise skip test. Per-segment or per-block filters keyed on the record identity rather than the DID, so old segments can be rejected without being decompressed. This is the only fix that changes the complexity class.
- Bound the backlog, not just the pass. See below.
how to not get here again #
The failure was not slowness; it was slowness that compounded unobserved. Watch for the loop, not the symptom:
- Alert on
jetstream_compaction_watermark_lag_seconds. It is the honest measure of exposure: everything witnessed inside that window still serves deleted rows, because nothing in the read path consults the tombstone set — rewriting the segment is what removes a row. - Alert on pass duration approaching the interval. Once a pass cannot finish within the interval, the backlog only grows, and every restart is a full reset of the work.
- Treat
passes_total == 0with a climbingsegments_examined_totalas in-progress, not idle. Opposite responses. - The tombstone set is also a memory cost: 3.69M entries held 472 MB of a 15.6 GB box, and it only shrinks when the watermark advances.
2026-08-17 re-measurement: the bottleneck moved #
The fix that this doc's measurements predate (b3ea690, frame reuse for blocks that drop nothing — landed the same day this doc was written) changed the answer. Re-measured on the same production box during the first completed 4h-interval pass:
| quantity | 2026-08-10 (pre frame-reuse) | 2026-08-17 |
|---|---|---|
| observed rate | 3.4 segments/min | ~20 segments/min |
| segments rewritten | 6,584 of 6,584 | 5,008 of 5,252 examined (244 clean) |
| bytes moved per pass | (not measured) | ~1.7 TB read + 1.34 TB written |
| rows dropped per pass | — | 2.3M (~0.01% of the archive) |
| bottleneck | CPU (zstd recompress) | disk |
The disk evidence: a single-stream cold read of one 270 MB segment runs
~105 MB/s while the pass's 4 workers are active, and the pass aggregate
(~3 TB moved in ~4h15m) is ~200 MB/s sustained — /data is a
network-attached cloud volume, and that is its ceiling under this load.
zstd no longer dominates: with tombstones scattered thin (hundreds of
dropped rows across ~1,600 blocks per segment), most blocks reuse their
compressed frame verbatim.
Consequences:
- Pass duration (5h27m measured, first completed pass 2026-08-17 06:16Z: 6,358/6,666 rewritten, 3.56M rows dropped) exceeds the 4h interval — but the loop is sequential and the timer resets AFTER a pass completes, so the real cadence is pass+interval ≈ 9.9h per cycle (~55% duty), not continuous. Upstream's loop behaves identically at this archive size — shared design cost, not a config error. The interval stays at upstream's 4h default (decided 2026-08-16; do not re-litigate).
- Age-tiered skipping is falsified by the archive's demographics (first completed pass's rewritten-age histogram): only 84 rewritten segments were younger than 7 days; 6,274 sat in the 7–30d bucket. Steady state seals only ~6-7 segments/day, so the archive is overwhelmingly the backfill-era mass and that is where tombstoned rows live. Tiering by age would exempt almost nothing.
- More rewrite workers would not help this box: the volume is
saturated at 4. (
--compaction-rewrite-workers=4vs upstream's default 8 is now a non-difference here.) - The O(archive) amplification analysis above still holds exactly — the fold is O(new data), the rewrite is O(archive BYTES THROUGH THE DISK), and frame reuse cut the CPU multiplier, not the byte volume. The sublinear path remains a skip test precise enough that untouched segments are never read — see the FP-union caveat before assuming blooms get there (a candidate set of ~800K record keys unions bloom false positives: at fp=1e-6, expected ~0.8 false hits per segment — every segment matches again; exact per-segment indexes or a threshold-triggered lazy rewrite are the honest design space).