jetstream v2 in zig stream.waow.tech
stream docs live-repair.md
5.9 kB
Markdown

Live Sync 1.1 repair #

Stream follows the pinned Atmos/Jetstream V2 repair contract — Atmos being the Go atproto library upstream Jetstream V2 runs.

Verified against ingest/repair.zig on 2026-07-30: worker_count = 32, pending_capacity = 2048, limiter_capacity = 16_384, and a token bucket of five with a 12-second refill (five per minute, burst five) all exist as described.

The "64-job bounded queue" is worker_count * 2, not an independent 64. The same expression bounds the backfill dispatch queue, the live scheduler's per-DID pending capacity, and this repair pool (see invariants.md, "a bounded producer must not outrun its consumer"). Raising worker_count moves all three bounds at once.

Expected log noise, so you do not chase it #

Repair logs one ERROR line per failed attempt (repair.zig:297), and the limiter permits five attempts per DID per minute, so a small permanently-broken set of accounts produces a steady error stream indefinitely. Measured during experiment 6:

141 "repair worker failed" lines in 10 minutes
 -- but only 19 distinct DIDs
     94  GetRepoFailed
     39  RepairAuthenticationFailed
      6  RepairRateLimited
      2  AttemptTimeout

~7 attempts per DID per 10 minutes against a handful of unreachable or misconfigured PDSes. This is specified behavior. The number to watch is the count of distinct DIDs, not the line rate — line rate scales with the retry allowance, distinct DIDs with the actual problem. RepairAuthenticationFailed in particular is a property of the remote repository, not of Stream.

Contract #

  • Recoverable chain, inversion, CAR-decode, duplicate-path, and op-CID failures withhold the offending event and schedule authoritative getRepo through the configured relay. Signature failures, future revisions, oversized input, and outer/inner envelope disagreement bypass repair.
  • A #sync divergence first performs that fetch synchronously, with the DID lane released during network I/O, and queues one async retry only for a retryable fetch failure. A successful inline repair advances the original firehose cursor at the replacement set's durable boundary; the deferred retry is synthetic and carries no cursor.
  • For a divergent #sync, Atmos tests data-root divergence before comparing the outer and inner rev. An invalid outer rev therefore still repairs and durably advances verifier chain state, after which the ingest gate drops the original tombstone and every replacement row as one live/invalid_rev event. No invalid material reaches the archive.
  • The pool is 32 workers behind a 64-job bounded queue. A fetch owns a five-minute budget and does not hold the DID lane. Commits arriving during the fetch enter a 2,048-frame FIFO; overflow drops and reports the oldest.
  • A bounded 16,384-DID token-bucket cache permits five repairs per minute with burst five. Rate-limit eviction only resets a DID's repair allowance; it never drops repository state.
  • Apply re-authenticates the prepared complete repository, rejects a head older than the current chain, and rejects an equal revision with a different data CID. Pending commits replay against the fetched head under the same DID lane. Replay failures are reported and never recursively enqueue repair.
  • Completion is registered under the DID lane, but cannot pass the archive writer until its trigger ticket has completed. The durable row order is a sync DID tombstone, the authoritative create_resync materialization, then any valid pending commits.
  • A withheld trigger or pending commit does not advance the relay cursor by itself. As in Atmos, a later normally emitted upstream event may move the cursor watermark beyond those silent drops; the in-memory repair queue is not a crash journal. A repair completion is synthetic and carries no relay cursor. This matches the pinned upstream restart behavior.
  • Archive sequence assignment and verifier/cursor/hosting metadata staging share one writer-lock transaction. Any fsync callback therefore sees the matching checkpoint already registered; tail publication remains gated on that same durable boundary.
  • Complete CARs stream through <data-dir>/repair-scratch/live, outside the bootstrap lifecycle tree. Startup removes any killed download before workers begin; successful and failed attempts delete their files. The separate namespace preserves the invariant that backfill/ disappears at cutover.

Jetstream runs Atmos with its default hosting policy (HostingTrack), so account state is persisted but does not gate commit or repair fetches.

Offline receipts #

zig build test requires no public network. The repair tests use loopback HTTP and generated cryptographic material to exercise the production paths:

  1. a signed, complete repository CAR is streamed to scratch, copied into the bounded preparation allocator, its mmap is released, and the owned repo is completeness-checked, resolved through a real DID document, and verified;
  2. the HTTP body is held open while a same-DID encoded firehose frame enters the pending FIFO, proving network fetch does not own the apply lane;
  3. limiter, overflow, stale-head, contradictory-head, and trigger-ticket ordering boundaries are checked deterministically;
  4. a real synchronous-#sync completion is written through Consumer.emitRepair, then read back as sealed JSS rows while its RocksDB chain checkpoint and v2 tail frames are asserted with its original relay cursor at the same durable sequence boundary; and
  5. a divergent invalid-rev #sync repairs chain state but emits no archive rows, with the canonical live drop counter proving the downstream gate; and
  6. verifier silent-drops leave the durable relay cursor unchanged until a later ordinary event is archived, matching Atmos's watermark behavior.