jetstream v2 in zig stream.waow.tech
stream docs bootstrap-semantic-parity.md
20 kB
Markdown

Jetstream V2 bootstrap and lifecycle audit #

Reference:

  • upstream Jetstream f29815c391fc2644f8a3dd36b899fb3697dd1ea6
  • Atmos v0.2.14 (the Go atproto library upstream runs, not a Stream dependency)
  • Stream: the recorded basis hash is unresolvable — see below.

The Stream basis cannot be located. This file named Stream a345f4054df568e364aff08a9b1726275fc48e29 "plus the discovery changes and audit updates in this commit". That object does not exist in this repository (git cat-file cannot resolve it, and no ref contains it), so the basis is not reproducible.

This file was last substantively updated by 4f15a19 (2026-07-25, "adopt upstream's subsystem vocabulary, and fix two rotted audits"); treat that as the effective basis and git log -- docs/bootstrap-semantic-parity.md as the authority. Experiment 6 deployed 29705cb, and semantic-parity.md is based on eb95c14; neither matches this file.

The hash is not replaced because no audit has been performed against a later commit. Re-basing this audit is a separate deliberate act.

The status words have the meanings defined in semantic-parity.md. A configuration knob being wired correctly does not verify the behavior it controls.

Current decision #

The overall lifecycle remains blocked. The dispatch-batch completion defect, correctness-metadata fail-open reads, merge restart guard, and bootstrap-to-merging ordering, and post-bootstrap discovery now have production-boundary fixes and focused proof. Independent recovery, retry, live-encoding, and reconnect defects still prevent another deployment.

Initial enumeration and download #

Invariant Status Audited evidence Gap or required proof
listRepos page size and accumulation verified Stream requests protocol-sized pages, counts every entry toward a page-aligned configurable batch, and defaults the batch target to 100,000. The producer ledger tests cover page ordering and the zero/default value. This says nothing about completion durability inside the batch.
Eligible-repository shuffle verified The accumulated eligible jobs are shuffled immediately before dispatch. Deterministic helper tests establish that the production helper changes order. Statistical PDS distribution is an operational measurement, not part of this narrow row.
Initial worker admission verified Positive worker counts become physical concurrent getRepo requests; zero/omitted select 100. Held-request production tests observed 7 and 100 active requests. A 200-worker run was slower than 100 workers.
Request path through relay redirects verified Stream requests getRepo from the configured relay and follows redirects, retaining the final host for diagnostics. This matches pinned Atmos. Neither implementation knows the redirected PDS before the request, so redirected requests receive reactive rather than proactive per-host rate limiting. This is a shared limitation, not a Stream parity gap.
Repository-size exclusion verified No hard CAR-size rejection remains. The response streams to data-volume scratch and a repository may consume the configured preparation budget. Capacity exhaustion remains fatal by design and must be sized operationally.
Complete-CAR validation verified Preparation parses the CAR, verifies block CIDs, loads the MST, and walks referenced nodes/records before row emission. Truncation fixtures fail before rows are emitted. The claim is limited to the tested truncation/missing-block shapes.
Authoritative DID verified Emitted rows use the listRepos/selected DID even when the embedded commit DID differs. A hostile fixture exercises this. None known for this invariant.
Attempt deadline boundary verified One cancellable operation covers response streaming and CAR preparation; handler emission and durable writes are outside the five-minute network deadline. A deadline failure is terminal for that run. None known for this invariant.
Ordinary retry budget verified The initial attempt plus three ordinary transient retries, exponential base, jitter, and 30-second cap have mixed 503/429 fixtures. The production firehose reconnect defect is separate from getRepo retry.
Rate-limit retry budget verified 429 uses a separate twenty-retry budget and honors parsed server reset information with the bootstrap ceiling. Tests show 429 does not consume the ordinary budget. Pre-request per-PDS scheduling is unavailable after relay indirection, as it is upstream.
Terminal repository errors verified Structured RepoNotFound becomes complete with no rows; deactivated/suspended/takendown become unavailable. Focused tests inspect the durable status. None known for the classified responses.
Inactive repository during initial discovery verified A new inactive entry is stored as not_started/active=false and is not dispatched; later active flips preserve the row. Post-bootstrap discovery likewise preserves inactive unknown entries as failed retry-state rows. None known for this invariant.
Explicit selected backfill partial The selected DID list bypasses listRepos, preserves caller order, conflicts with max-repos, clears stale discovery state, resolves identity metadata, maintains the handle index, and uses the same writer-gated per-repository completion machinery as whole-network bootstrap. The exact current artifact has not rerun the full selected-repo receipt.
Max-repos debug mode partial The flag truncates selected work and avoids committing a resumable whole-network cursor. Phase and relay-cursor reads now fail closed. It remains a debug path and is not evidence for whole-network discovery or merge behavior.

Repository emission and durable completion #

Invariant Status Audited evidence Gap or required proof
Record path validation and drops verified Bad NSID/rkey/field-width records are dropped individually; other valid records from the repository survive. Counters and physical rows are asserted. None known for the tested shapes.
Repository revision watermark verified Backfill rows receive the repository head revision and successful completion stores the immutable backfill revision plus latest revision. Durability timing of that completion is blocked below.
1,024-row handler emission verified Prepared repositories emit bounded row batches and tests exercise multi-batch repositories. This is an emission/memory boundary, not a repository completion boundary.
Async compression ordering partial Four workers by default prepare real blocks concurrently; commit order is serialized; sync and async fixtures produce byte-equivalent JSS. Cancellation/OOM behavior at every queued/compressed/commit state has not been compared against upstream in this audit.
Archive before repository completion verified Each successful CAR records its final emitted archive sequence. The archive durability hook stages only completions whose final sequence is covered; a forced drain fails rather than advancing past an uncovered completion. Focused tests inspect the physical boundary and terminal state. The claim is limited to bootstrap repository completion; other metadata boundaries have their own rows.
Per-repository completion durability verified Stream now ports upstream's watermark queue: data-bearing and genuinely empty successful CARs queue independently, duplicate DIDs replace in place, RepoNotFound follows upstream's immediate terminal OnFail path, and captured completion time survives deferred durability. Re-run the process-level oracle before artifact admission; the production mechanism and focused boundary are established here.
Slow or hung straggler isolation verified A production-boundary test leaves one repository without a worker result, crosses a writer durability boundary, proves its completed sibling durable while the listRepos cursor is absent, fully reopens RocksDB, and proves only the interrupted repository moves to pending. The admission oracle must scale this invariant beyond a two-repository fixture and inject an actual process kill.
Completion metrics verified Queue count/depth change at queue/replacement; durable batch/repo counters change only after the synced repo-state batch; queue wait uses the captured completion time; RepoNotFound bypasses those queue metrics as upstream does. Prometheus counters are process-local by definition. Restart-stable progress comes from durable repository states, not counter continuity.
Resource exhaustion classification partial Preparation allocation failure is treated as lifecycle-fatal rather than a durable “repo too large” result. Archive startup now propagates iterator allocation failure rather than treating it as a torn tail. Fault coverage does not establish cleanup and restart behavior at each preparation/emission ownership transfer.

Cursors and lifecycle transitions #

Invariant Status Audited evidence Gap or required proof
Relay listRepos cursor checkpoint verified Stream atomically writes the relay cursor and last non-empty discovery cursor only after the dispatch batch drains. Final empty cursor handling has unit coverage, while the straggler/reopen test proves individual repository completion can become durable with the page cursor still absent. The 100,000-entry unit is checkpoint/re-enumeration cadence, not repository completion durability.
Relay firehose cursor read verified Durable cursor writes remain coupled to archive metadata batches. Reads now propagate RocksDB errors and require upstream's [version=1][uint64 LE] encoding, rejecting wrong width, unknown version, and values above MaxInt64; writes reject negative cursors. Focused tests exercise every rejection, and a startup migration converts Stream's prior valid 8-byte encoding without accepting other malformed values. This row does not close the independent Zat reconnect failure.
listRepos cursor read verified loadRelayListReposCursor propagates RocksDB errors and distinguishes a missing key. Initial enumeration and merge discovery retain dynamically owned cursors without a local length ceiling; real-process receipts walk 4 KiB cursors through both paths. None known for this invariant.
Lifecycle phase read verified Phase writes are synced. Reads distinguish a missing key from RocksDB failure and reject every value outside bootstrap, merging, and steady_state; focused tests inject both corruption and a real Store read error. The production orchestrator uses the same fallible read before any fresh-directory write. Persisted phase-entry timing remains a separate gap below.
Bootstrap → merging commit point verified After the backfill future drains successfully, Stream syncs phase=merging while bootstrap-live capture is still running, then requests capture shutdown and performs the archive cutover. This is the pinned upstream ordering. Named process-abort seams immediately before the write and immediately after it/before live stop both recover on the same disk through the lifecycle oracle. This row establishes the durable phase boundary, not the separate persisted phase-entry timestamp/status contract.
Merging → steady_state commit point partial Stream drains, compacts, reconciles the manifest, discovers, removes the source tree, syncs the data directory, deletes cursors, publishes seq metadata, then writes steady_state. Phase reads fail closed, and the restart guard distinguishes only FileNotFound from every other filesystem error before taking the cleanup-complete path. Named write-side crash tests cover several seams. Exact persisted phase-entry timing remains absent, and other merge/discovery rows remain partial or blocked.
Exact persisted phase timing verified phase/entered_at and backfill/timing/{started_at,completed_at} are persisted, and the transition that closes backfill writes the phase and both endpoints in one synced batch. Absent stays absent rather than becoming zero, and a backwards clock clamps to zero instead of yielding a negative duration. The status summary does not yet render these fields; the durable side is done.

Bootstrap-live capture and merge #

Invariant Status Audited evidence Gap or required proof
Concurrent bootstrap-live attachment partial The production lifecycle starts a separate capture archive, crash/oracle fixtures send live events during bootstrap, and the merging phase is now committed before capture is asked to stop. The firehose dependency can still terminate during reconnect sleep.
Bootstrap-live failure propagation partial The lifecycle now observes an unexpectedly completed capture future and treats it as fatal. HttpConnectionClosing can still escape Zat's reconnect loop and terminate the process; the exact dependency behavior needs a regression test on Linux.
Merge source existence guard verified Only FileNotFound selects the cleanup-complete restart path. Other filesystem errors propagate without cursor deletion or phase advancement; focused tests exercise a non-directory source and the normal absent-source restart. This narrow guard does not prove every source-file read in the merge walker; those remain covered by their own rows.
Merge source cursor read verified Successful source completion atomically commits upstream's [version=1][uint64 LE] cursor and latest-revision updates. Reads distinguish absence from Store failure and reject wrong width/version; focused tests inject each failure, and startup migrates Stream's prior valid 8-byte cursor. This row proves cursor decoding and commit atomicity, not every source-file read.
Merge row filtering verified Rows are dropped only when the repository is complete and the source revision is at/below its backfill watermark; missing lookups are cached. Kept/dropped fixtures exercise this. This predicate proof does not establish source-file recovery behavior.
Latest revision refresh partial Kept rev-bearing rows update latest revision in the same batch as the source cursor while preserving the backfill revision. Missing repo rows are skipped; this matches the intended defensive behavior, but corrupt cursor state can replay updates and rows.
Pending-repository pass verified The pending pass is upstream's RunPendingRepoRetryPass: the same bounded retry runner with eligible_status flipped to pending, so it inherits the worker pool, per-host gate, computed backoff and 429 host parking. A test asserts a failed pending repository gets upstream's RecordRetryFailure bookkeeping and that a failed sibling is not selected. None known.
Merge-tail compaction and manifest reconcile partial Destination sealing precedes delete/update compaction; manifest reconciliation happens before serving is enabled. Physical rewrite tests cover named write failures; compaction watermark, merge source admission, and active-tail startup recovery fail closed. Exact cutover and post-startup rewrite failure coverage remain independent blockers.
Post-bootstrap discovery verified It resumes from the last non-empty bootstrap cursor, writes every previously unknown active or inactive DID as failed while preserving the relay's active flag, follows dynamically owned cursors, rejects cursor loops, and propagates relay/store errors before cleanup. A real-process relay fixture returns an inactive DID, a 4 KiB cursor, and a second DID; both durable account rows are inspected after steady-state admission. This row proves discovery completeness and failure behavior; the downstream retry runner remains independently blocked.
Cleanup durability verified On the successful path, the backfill tree is removed, the data directory is synced, and merge/discovery cursors are deleted in a synced RocksDB batch afterward. Restart-after-cleanup repeats the directory sync, and the source-existence guard fails closed on non-not-found errors. None known for this cleanup ordering invariant.

Steady failed-repository retry #

Invariant Status Audited evidence Gap or required proof
Global and per-host concurrency verified The implementation has real global workers and host gates; focused held-request tests measure configured/default 16 and 4 limits. The candidate set is still materialized eagerly.
Candidate scan memory verified Candidates stream straight off the store into a bounded queue sized to the worker pool, and host gates are created on demand, so neither the candidate set nor the host set is materialized. The merge's pending pass shares that runner. None known.
Final-host attribution verified The final post-redirect authority replaces prior attribution after a response; redirected tests inspect the durable host. Requests that fail before a response may remain unattributed, as upstream permits.
429 host parking partial A 429 delays only work assigned to that final host in focused tests. Stream persists host parking whereas upstream keeps it process-local. Allocation, RocksDB, and malformed-value errors all mean “not parked,” which can hammer a limited PDS. Decide whether persistence is an intentional divergence; in either case, fail loud on corrupted state.
Retry state and diagnostics verified Attempts, class, last error, next attempt, host aggregates, bounded recent samples, active/status counts, and success clearing are committed with repository transitions in focused restart tests. This does not prove the retry supervisor remains alive.
Retry subsystem failure verified A pass that fails for our own reasons latches a terminal error, counts retry_terminal_failures_total, and requests process shutdown — upstream returns the error into an errgroup that cancels steady state. Per-repo remote failures still become backoff and never reach that path; a real injected metadata read fault drives the test. None known.

Why the previous adversity gate passed #

The previous gate described the oracle as covering “after-repo-complete” crashes. In Stream, that crashpoint is executed once per result only after the full putMany succeeds. In the four-account restart world, every numbered crash therefore occurred after all four completions were already durable. It could not exercise the upstream state “repo A complete and durable while repo B from the same listRepos dispatch is still running.”

The replacement adversity matrix must include:

  1. N−1 durable completions plus one indefinitely held repository;
  2. a crash after each individual writer durability hook, before listRepos cursor advancement;
  3. phase, relay cursor, merge cursor, compaction watermark, and host-park read errors plus malformed bytes;
  4. unreadable versus genuinely absent bootstrap-live directories;
  5. inactive unknown repositories and a cursor longer than every local fixed buffer;
  6. retry-store failure with health/readiness assertions; and
  7. a process restart proving completed repositories are skipped without relying only on final-state convergence.