From 1fb55f2aa74bc1320256fe1d691898dcca536dfe Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Thu, 30 Jul 2026 07:22:05 -0500 Subject: [PATCH] docs: repair's steady ERROR stream is expected; and worker_count*2 is now 3 bounds live-repair.md verifies clean -- worker_count 32, pending_capacity 2048, limiter_capacity 16_384, and a five-token bucket with 12s refill (five per minute, burst five) all match ingest/repair.zig. Two things it does not say. Repair logs one ERROR line per failed attempt (repair.zig:297) and the limiter allows five attempts per DID per minute, so a small permanently broken set of accounts emits errors forever. Measured on experiment 6: 141 "repair worker failed" lines in 10 minutes from only 19 distinct DIDs -- GetRepoFailed 94, RepairAuthenticationFailed 39, RepairRateLimited 6, AttemptTimeout 2. That is the specified behaviour, not an incident. Added the numbers and the rule that follows: watch distinct DIDs, not line rate, because line rate scales with the retry allowance rather than the problem. Second: the "64-job bounded queue" is queue_capacity = worker_count * 2. That expression is now the bound in three places -- backfill dispatch, the live scheduler's per-DID pending capacity, and this repair pool -- each written in prose as a bare "64" as if chosen locally. It is the same expression whose interaction with a 100,000-repo batch caused the experiment 4 deadlock, so raising worker_count moves all three at once. Noted here and in live-scheduler.md. Also expands Atmos on first use. Co-Authored-By: Claude Opus 5 (1M context) --- docs/live-repair.md | 42 ++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 40 insertions(+), 2 deletions(-) diff --git a/docs/live-repair.md b/docs/live-repair.md index a17d7f0..cdc14db 100644 --- a/docs/live-repair.md +++ b/docs/live-repair.md @@ -1,7 +1,45 @@ # Live Sync 1.1 repair -Stream follows the pinned Atmos/Jetstream V2 repair contract rather than -inventing a second recovery model. +Stream follows the pinned Atmos/Jetstream V2 repair contract — Atmos being the +Go atproto library upstream Jetstream V2 runs — rather than inventing a second +recovery model. + +Verified against `ingest/repair.zig` on 2026-07-30: `worker_count = 32`, +`pending_capacity = 2048`, `limiter_capacity = 16_384`, and a token bucket of +five with a 12-second refill (five per minute, burst five) all exist as +described. + +**The "64-job bounded queue" is `worker_count * 2`, not an independent 64.** +That expression is now the bound in three separate places — the backfill +dispatch queue, the live scheduler's per-DID pending capacity, and this repair +pool — and each is written in prose as a bare "64" as though it were chosen +there. It is the same expression whose interaction with a 100,000-repo batch +caused the experiment 4 deadlock (`invariants.md`, "a bounded producer must not +outrun its consumer"). Anyone raising `worker_count` moves all three bounds at +once. + +## Expected log noise, so you do not chase it + +Repair logs **one `ERROR` line per failed attempt** (`repair.zig:297`), and the +limiter permits five attempts per DID per minute, so a small permanently-broken +set of accounts produces a steady error stream indefinitely. Measured during +experiment 6: + +``` +141 "repair worker failed" lines in 10 minutes + -- but only 19 distinct DIDs + 94 GetRepoFailed + 39 RepairAuthenticationFailed + 6 RepairRateLimited + 2 AttemptTimeout +``` + +~7 attempts per DID per 10 minutes against a handful of unreachable or +misconfigured PDSes. **This is the system working as specified, not an +incident.** The number to watch is the count of *distinct* DIDs, not the line +rate — line rate scales with the retry allowance, distinct DIDs with the actual +problem. `RepairAuthenticationFailed` in particular is a property of the remote +repository, not of Stream. ## Contract -- 2.51.2