Live Sync 1.1 repair #
Stream follows the pinned Atmos/Jetstream V2 repair contract — Atmos being the Go atproto library upstream Jetstream V2 runs.
Verified against ingest/repair.zig on 2026-07-30: worker_count = 32,
pending_capacity = 2048, limiter_capacity = 16_384, and a token bucket of
five with a 12-second refill (five per minute, burst five) all exist as
described.
The "64-job bounded queue" is worker_count * 2, not an independent 64.
The same expression bounds the backfill dispatch queue, the live scheduler's
per-DID pending capacity, and this repair pool (see invariants.md, "a bounded
producer must not outrun its consumer"). Raising worker_count moves all three
bounds at once.
Expected log noise, so you do not chase it #
Repair logs one ERROR line per failed attempt (repair.zig:297), and the
limiter permits five attempts per DID per minute, so a small permanently-broken
set of accounts produces a steady error stream indefinitely. Measured during
experiment 6:
141 "repair worker failed" lines in 10 minutes
-- but only 19 distinct DIDs
94 GetRepoFailed
39 RepairAuthenticationFailed
6 RepairRateLimited
2 AttemptTimeout
~7 attempts per DID per 10 minutes against a handful of unreachable or
misconfigured PDSes. This is specified behavior. The number to watch is the count of distinct DIDs, not the line
rate — line rate scales with the retry allowance, distinct DIDs with the actual
problem. RepairAuthenticationFailed in particular is a property of the remote
repository, not of Stream.
Contract #
- Recoverable chain, inversion, CAR-decode, duplicate-path, and op-CID
failures withhold the offending event and schedule authoritative
getRepothrough the configured relay. Signature failures, future revisions, oversized input, and outer/inner envelope disagreement bypass repair. - A
#syncdivergence first performs that fetch synchronously, with the DID lane released during network I/O, and queues one async retry only for a retryable fetch failure. A successful inline repair advances the original firehose cursor at the replacement set's durable boundary; the deferred retry is synthetic and carries no cursor. - For a divergent
#sync, Atmos tests data-root divergence before comparing the outer and inner rev. An invalid outer rev therefore still repairs and durably advances verifier chain state, after which the ingest gate drops the original tombstone and every replacement row as onelive/invalid_revevent. No invalid material reaches the archive. - The pool is 32 workers behind a 64-job bounded queue. A fetch owns a five-minute budget and does not hold the DID lane. Commits arriving during the fetch enter a 2,048-frame FIFO; overflow drops and reports the oldest.
- A bounded 16,384-DID token-bucket cache permits five repairs per minute with burst five. Rate-limit eviction only resets a DID's repair allowance; it never drops repository state.
- Apply re-authenticates the prepared complete repository, rejects a head older than the current chain, and rejects an equal revision with a different data CID. Pending commits replay against the fetched head under the same DID lane. Replay failures are reported and never recursively enqueue repair.
- Completion is registered under the DID lane, but cannot pass the archive
writer until its trigger ticket has completed. The durable row order is a
syncDID tombstone, the authoritativecreate_resyncmaterialization, then any valid pending commits. - A withheld trigger or pending commit does not advance the relay cursor by itself. As in Atmos, a later normally emitted upstream event may move the cursor watermark beyond those silent drops; the in-memory repair queue is not a crash journal. A repair completion is synthetic and carries no relay cursor. This matches the pinned upstream restart behavior.
- Archive sequence assignment and verifier/cursor/hosting metadata staging share one writer-lock transaction. Any fsync callback therefore sees the matching checkpoint already registered; tail publication remains gated on that same durable boundary.
- Complete CARs stream through
<data-dir>/repair-scratch/live, outside the bootstrap lifecycle tree. Startup removes any killed download before workers begin; successful and failed attempts delete their files. The separate namespace preserves the invariant thatbackfill/disappears at cutover.
Jetstream runs Atmos with its default hosting policy (HostingTrack), so
account state is persisted but does not gate commit or repair fetches.
Offline receipts #
zig build test requires no public network. The repair tests use loopback HTTP
and generated cryptographic material to exercise the production paths:
- a signed, complete repository CAR is streamed to scratch, copied into the bounded preparation allocator, its mmap is released, and the owned repo is completeness-checked, resolved through a real DID document, and verified;
- the HTTP body is held open while a same-DID encoded firehose frame enters the pending FIFO, proving network fetch does not own the apply lane;
- limiter, overflow, stale-head, contradictory-head, and trigger-ticket ordering boundaries are checked deterministically;
- a real synchronous-
#synccompletion is written throughConsumer.emitRepair, then read back as sealed JSS rows while its RocksDB chain checkpoint and v2 tail frames are asserted with its original relay cursor at the same durable sequence boundary; and - a divergent invalid-rev
#syncrepairs chain state but emits no archive rows, with the canonical live drop counter proving the downstream gate; and - verifier silent-drops leave the durable relay cursor unchanged until a later ordinary event is archived, matching Atmos's watermark behavior.