diff --git a/docs/compaction-cost.md b/docs/compaction-cost.md index d0264a7..feb7f1f 100644 --- a/docs/compaction-cost.md +++ b/docs/compaction-cost.md @@ -103,3 +103,46 @@ for the loop, not the symptom: in-progress, not idle.** Opposite responses. - The tombstone set is also a memory cost: 3.69M entries held 472 MB of a 15.6 GB box, and it only shrinks when the watermark advances. + +## 2026-08-17 re-measurement: the bottleneck moved + +The fix that this doc's measurements predate (b3ea690, frame reuse for +blocks that drop nothing — landed the same day this doc was written) +changed the answer. Re-measured on the same production box during the +first completed 4h-interval pass: + +| quantity | 2026-08-10 (pre frame-reuse) | 2026-08-17 | +| --- | --- | --- | +| observed rate | 3.4 segments/min | ~20 segments/min | +| segments rewritten | 6,584 of 6,584 | 5,008 of 5,252 examined (244 clean) | +| bytes moved per pass | (not measured) | ~1.7 TB read + 1.34 TB written | +| rows dropped per pass | — | 2.3M (~0.01% of the archive) | +| bottleneck | CPU (zstd recompress) | **disk** | + +The disk evidence: a single-stream cold read of one 270 MB segment runs +~105 MB/s while the pass's 4 workers are active, and the pass aggregate +(~3 TB moved in ~4h15m) is ~200 MB/s sustained — `/data` is a +network-attached cloud volume, and that is its ceiling under this load. +zstd no longer dominates: with tombstones scattered thin (hundreds of +dropped rows across ~1,600 blocks per segment), most blocks reuse their +compressed frame verbatim. + +Consequences: + +- **Pass duration (~4.5h+) exceeds the 4h interval**, so the steady + compactor refires immediately on completion: compaction is effectively + continuous. Upstream's loop behaves identically at this archive size — + shared design cost, not a config error. The interval stays at + upstream's 4h default (decided 2026-08-16; do not re-litigate). +- **More rewrite workers would not help this box**: the volume is + saturated at 4. (`--compaction-rewrite-workers=4` vs upstream's + default 8 is now a non-difference here.) +- The O(archive) amplification analysis above still holds exactly — the + fold is O(new data), the rewrite is O(archive BYTES THROUGH THE DISK), + and frame reuse cut the CPU multiplier, not the byte volume. The + sublinear path remains a skip test precise enough that untouched + segments are never read — see the FP-union caveat before assuming + blooms get there (a candidate set of ~800K record keys unions bloom + false positives: at fp=1e-6, expected ~0.8 false hits per segment — + every segment matches again; exact per-segment indexes or a + threshold-triggered lazy rewrite are the honest design space).