Jetstream V2 semantic parity audit #
This is a deployment gate. A row describes what the current implementation can prove today; its status reflects current evidence only.
Re-pin in progress (2026-08-12). Upstream renamed the v2 surface to
network.bsky.jetstream.subscribeEvents(proposal-0015 lexicon frames), renamedplanBackfilltoplanSnapshot, removedrecordCbor, added a v1 envelopecursor, and retrained the v2 zstd dictionary. Stream ported the new contract against pin289b0328c2e1a0ccf8c870cb45de0b2397de19fb; rows below still describe the f29815c-era audit and are being re-proved through the gate (the differential oracle and archive-client e2e now run upstream's d4dd2f0 client against Stream). One deliberate non-port: upstream's PDS-direct fleet bootstrap and its four backfill flags (environment_contract.pyupstream_unported) — Stream's bootstrap remains relay-getRepo until that design is ported on purpose.Startup ordering (2026-08-16, task #12): listeners now bind and accept before storage recovery, matching upstream's server-first startup (jetstreamd runtime.go registers all routes, then recovery runs behind lifecycle.IsSteadyState 503s on subscribe and the xrpcapi Ready hook). Stream's equivalent gates are hub.serving (lifecycle) and the new hub.storage_ready (Ready): both closed from before bind, opened by release stores after the stores are published. Storage-free surfaces (homepage, endpoints.json, healthz/readyz/metrics, a minimal /status) serve throughout, so a deploy presents 503s for seconds instead of connection-refused for minutes (measured 2.5 min on 2026-08-16).
Failed-repo retry routing (2026-08-15, task #11): the retry pass now ports
retryRunner.download(retry.go:425-438) exactly — stamped-PDS direct first; relay 302 fallback only on an unroutable stamp or a directisRepoNotFoundError; restamp guarded by upstream'scompletionHost != cand.PDS; foreign embedded DID rejected. One recorded data-model difference, not a behavioral one: upstream stamps bare authorities and dials them via atmos's https-only hardened client builder, while Stream stamps the DID document's#atproto_pdsURL —directBasereduces it toscheme://normalized-authoritybefore dialing, and restamps ashttps://authority, matching what upstream's builder would dial.
Audit basis #
- Stream:
eb95c14310ef18c7ab9edbcff991a04a560a7d86 - admitted artifact:
sha256:4300104135c6fd35bd13a264b744c5f842c41ee9a51e0cf1c57c9a59640551a2(receipts/eb95c14.json, 20/20 suites, built natively on linux/amd64) - upstream Jetstream:
f29815c391fc2644f8a3dd36b899fb3697dd1ea6 - Atmos:
v0.2.14— the Go atproto library upstream Jetstream V2 runs, and the reference for repair, hosting-policy and rate-limiting behaviour. It is upstream's dependency, not one of Stream's, so it does not appear inbuild.zig.zon; it is pinned here because rows compare against it. - Zat:
955e9ca9afb41d66d41e998803f6fcfc8fc4330f(v0.3.22) — Stream's own atproto library, and this pin does matchbuild.zig.zon. - audit date: 2026-07-28
The basis commit is behind the artifact that actually ran experiment 6. This audit is based on Stream
eb95c14, but the deployed image was29705cb(receipts/29705cb.json,sha256:0f45468c…). Fifteen commits separate them, and three change bootstrap behaviour this audit has rows for:e2ab25cdeleted the in-flight byte budget,fb3cea4made repository parsing concurrent, andfcc63f8broke the batch submit/consume deadlock.The basis is not rewritten to
29705cbhere because no audit has been performed against that commit. The bootstrap rows need re-confirmation against29705cb; re-basing the audit is a separate deliberate act.
Earlier receipts remain valid for the digests they cover — a receipt proves those suites passed against that image, and that stays true. "Admitted" means "earned a receipt", not "current", so confirm an artifact matches the commit you intend before deploying it.
The upstream checkout is exactly at the recorded pin. The untracked
simulator path in that checkout is not treated as upstream source.
Status vocabulary #
- verified: source comparison establishes the same narrow invariant and a test exercises the production boundary, including the relevant failure mode.
- partial: some named invariants are verified, but the surface has an unproved or non-equivalent boundary.
- blocked: a known semantic mismatch, data-loss risk, false-negative API result, or process-fatal remote-input path exists.
- intentional divergence: the behavior deliberately differs and must not be called parity.
- unverified: no known mismatch was found, but the available evidence is insufficient.
Only verified rows may contribute to a parity claim. Passing a test after the audited commit does not update this document automatically. The source, test, and this audit must be reviewed together.
Current decision #
No row is blocked. A whole-network backfill is feasible on the hardware we already use — ~8.7 TB, ~42 hours at our measured rate.
That row was briefly marked blocked on an estimate that was wrong by ~21x.
Repository size is heavy-tailed (median 10 KB, mean 0.45 MB, p99 10 MB) and
listRepos is ordered largest-first, so a sample of consecutive head entries
produced 9.45 MB per repo and a 40-day projection. Four hundred independent
random cursors gave 0.45 MB and 42 hours, which reconciles with upstream's
documented 16 hours on a single server. Correctness evidence is not
feasibility evidence; feasibility is measured separately.
A second correction: the whole-network bootstrap row was previously marked verified and was wrong.
Experiment 4 (27 July) ran 55 minutes and durably committed nothing: 12,686
repositories fetched and emitted, zero ever queued for completion. The cause
was a three-way circular wait in finishWorkBatch — the main fiber submitting
a whole 100,000-job batch into a worker_count * 2 queue before consuming any
result, workers blocked in Budget.acquire (which waits rather than failing),
and the budget released only by the consumption the main fiber could not reach.
Fixed in fcc63f8; reproduced and verified against bsky.network on a ccx23,
where the previous build handled 1,403 repositories with zero durable and the
fixed build holds queued == handled throughout.
The row had been marked verified on suites that could not reach the failure: the pinned simulator's repositories are ~4 KiB, so concurrent demand can never exceed a budget a single repository fits in. An evidence limit on a durability property is treated as unproven, not as absent.
Also closed since the 2026-07-24 audit: cold replay holes, the planBackfill
false-negative paths, listSegments drop-on-encode, ordinary live encoding,
retry-supervisor health, pending-repository orchestration, and the firehose
reconnect loop — the last a real process-killing race in the shared
websocket.zig fork, fixed there and shipped through zat v0.3.18.
This section is a summary of the table below and must be re-read whenever a row changes.
The detailed bootstrap audit is in bootstrap-semantic-parity.md.
Audited checklist #
| Surface | Status | What is actually established | Blocking or missing evidence |
|---|---|---|---|
| Whole-network feasibility | verified | Measured 2026-07-28 against bsky.network: ~19.4M repositories (cursor-space binary search x measured density), mean 0.45 MB and median 0.01 MB per repository from 400 independent uniformly-random cursors, so ~8.7 TB for a full pass. At the ~460 Mbps a CCX43 sustains that is ~42 hours; at 1 Gbps ~19. Upstream's documented ~16 hours needs ~1.2 Gbps, which matches their stated profile of a single server with a few TB of disk. | A first estimate of 183 TB and 40+ days was wrong by ~21x: it sampled consecutive listRepos entries, which are correlated, from the head of the list, which is the largest accounts. Repository size is heavy-tailed and listRepos is ordered worst-first, so early throughput understates the crawl by more than an order of magnitude. Sample independent random cursors. |
| Lifecycle phase machine | verified | Phase reads distinguish absence from RocksDB failure, reject unknown values, and use synced writes. Bootstrap writes merging after backfill drains while bootstrap-live is still running, exactly at the pinned upstream boundary; process-kill crashpoints exercise both sides of that write. Cleanup writes steady_state only after directory sync and cursor deletion. |
Exact persisted phase-entry timestamps are a separate status-surface gap; this row establishes transition state and ordering only. |
| Whole-network bootstrap | partial | listRepos pages form a page-aligned dispatch/checkpoint unit; eligible repositories are shuffled; download concurrency and CAR preparation are real. Successful repositories become durable independently at archive-writer durability boundaries. Cleanup at every preparation/emission ownership transfer is the parse arena's by construction, and an allocation-failure sweep asserts exhaustion always surfaces as OutOfMemory. The batch submit/consume deadlock below is fixed and the fix is confirmed against the real network. |
This row was marked verified and was wrong. finishWorkBatch submitted an entire 100,000-job batch before consuming a result, against a queue of worker_count * 2; with workers blocked in Budget.acquire (which waits rather than failing) and the budget released only by the main fiber consuming results, bootstrap deadlocked with zero repositories ever queued for completion. Experiment 4 ran 55 minutes and committed nothing. Fixed in fcc63f8, reproduced and verified on a ccx23 against bsky.network -- 1,403 handled / 0 durable before, queued == handled after. It stays partial because no offline suite reaches it: the pinned simulator's repositories are ~4 KiB, so concurrent demand cannot exceed a budget one repository fits in. Only a real-network run confirms this class. |
| Merge and post-bootstrap discovery | verified | Source rows are filtered by the backfill revision; source cursor and latest-revision updates share one synced batch (commitMergeSource, which refreshes latest_rev/updated_us while preserving the backfill rev — asserted by the real-merge test); merge cursor corruption/read errors fail closed; the restart guard treats only FileNotFound as cleanup-complete; discovery records active and inactive unknown repositories and follows arbitrary-length cursors with loop detection; successful cleanup is directory-synced. Every source-file path fails closed — readFileAlloc, checksum-verifying Sealed.parse, readBlock and the destination append all propagate, with no catch swallow — and a missing source segment raises SourceIndexGap rather than being skipped; a physical test deletes the middle of three sealed sources and asserts the error, where without the guard merge consumes 2 of 3 sources and reports success. The pending pass is now upstream's RunPendingRepoRetryPass exactly: the same bounded runner with eligible_status flipped to pending, so it inherits the worker pool, per-host gate, computed backoff and 429 host parking rather than reimplementing a weaker version. A test asserts a failed pending repository gets upstream's RecordRetryFailure bookkeeping — status failed, attempts and retry_count incremented, next_attempt_us in the future — and that a failed sibling is untouched, which the previous hand-rolled loop got wrong on the last two fields. |
No known divergence remains. The runner is shared with the steady retry loop, so that row's persisted-host-parking caveat applies here too; it is tracked there rather than duplicated. |
| Failed-repository healing | verified | A real global/per-host worker gate exists, final redirect hosts are recorded, and retry state is stored. Candidates stream through a bounded queue with host gates created on demand, so neither the failed set nor the host set is materialized. A pass that fails for our own reasons (store, archive, memory) now latches a terminal error, counts retry_terminal_failures_total, and requests process shutdown instead of sleeping until the next interval — matching upstream, whose errgroup cancels steady state. Per-repo remote failures still become backoff and never reach that path; a real read fault drives the test. Host parking now lives in the runner's memory for its lifetime, as upstream's hostParked map does, and never shortens a park — upstream's if old.After(until) return, which the persisted version lacked. |
None known. |
| JSS sealed format interoperability | verified | Upstream-produced sealed fixtures are parsed; Stream-produced sealed files are consumed by the pinned Go reader; header/footer, block index, blooms, collections, compression, and checksums have reciprocal fixtures. | This verifies sealed-format compatibility only. It does not verify startup recovery or replay completeness. |
| Archive startup recovery | verified | Startup validates the active header, walks complete frames, propagates read/decode/allocation failures, truncates and fsyncs only a framing-torn suffix, reconstructs block/event/sequence state, and resumes the same active file and segment index. Sealed high-water floors are accepted only through the checksum-verifying parser, including the empty-active/lower-sealed case. Physical tests prove repeated same-file restart, exact torn-tail truncation, byte-preserving failure on a complete corrupt frame, and rejection of a forged sealed max_seq. |
This row establishes startup recovery. Cold replay and post-startup rewrite paths remain separate surfaces. |
listSegments, getSegment, getBlock |
verified | The pinned Go client and direct HTTP tests cover normal responses, byte identity, checksums/ETags, conditional requests, ranges, cache headers, and storage-error responses through the production server. listSegments rendering was extracted so encode failure propagates; a fail-index sweep over the production handler proves every response is the exact baseline list or 5xx, never a truncated 200 and never a dropped request, so an internal failure cannot produce a silently short plan. |
The claim is limited to these three methods and the tested file states. |
planBackfill |
verified | Normal DID/collection/sequence filtering, whole-segment versus block mode, pagination, and configured limits match the pinned client fixtures. The planner is one-sided like upstream: a block whose collection bitmask row is absent, or a segment carrying no collection index, fails open instead of being pruned (upstream blockHasAnyCollection). A fail-index sweep over the whole production handler proves every response is either the exact baseline plan or 5xx, never a 200 with fewer segments, never a 400, and never a dropped request, so an internal failure cannot produce a silently short plan. |
Exact collection reporting (inspect-segment) deliberately stays fail-closed; that split matches upstream and is asserted separately. |
| Subscribe v1/v2 handshake and wire formats | verified | Cursor parsing, v1/v2 filter shapes, WebSocket framing, v1 deflate, v2 dictionary zstd, dictionary validators, size caps, pings, write deadlines, and shutdown close frames have production-path tests. | The two cross-references this row was waiting on are closed: cold replay propagates rather than skipping an unreadable segment, and remote-data encoding is isolated per row per wire. This row establishes the handshake and wire formats themselves; replay completeness is the cold-replay row. |
| Cold archive replay | verified | Normal sealed and active-prefix replay reaches the hot tail in order in existing fixtures. Timestamp cursor corruption is surfaced for the selected block. Neither cold walker will step over a manifest-listed segment it cannot open: both propagate, matching upstream walkSealedSegment, and physical tests delete a listed segment and assert nothing is emitted from the intact segment behind the hole. The rotation-seam no-progress invariant is now enforced: a cold pass that ends below the readable-log floor without advancing is retried once, and a second consecutive stall disconnects the subscriber with upstream's diagnostic rather than spinning, counted by stream_subscribe_cold_stalls_total. Reaching the floor is completion, not a stall, which a test pins so a caught-up subscriber is never disconnected. |
Stream walks manifest summaries linearly rather than resolving SegmentForSeq per pass, so the seam is retried at the subscriber loop rather than inside one replay call. The observable contract — bounded retry, then a loud failure — matches. |
| Live Sync 1.1 verification | verified | DID/signature checks, MST inversion, op-CID checks, replay/future-rev gates, durable chain/hosting updates, and several repair-trigger cases use real RocksDB, CARs, and sockets. Relay cursor corruption and read failures now fail closed. An accepted live row whose record the wire encoders cannot render no longer kills ingest: each wire is isolated per row like upstream's per-entry memoized encode error, the row stays durable, subscribers skip the empty body, and unencodable_record counts it. OutOfMemory still propagates so exhaustion cannot masquerade as a bad record. |
Retry-subsystem terminal failures now stop the service (retry_terminal_failures_total plus the shutdown flag main waits on), so that cross-reference is closed. The row covers live verification and repair triggering; whole-repo repair capacity is its own row. |
| Live scheduling and cursor durability | verified | Per-DID ordering, bounded pending work, worker admission, completion-order emission, archive-before-cursor write ordering, strict versioned cursor decoding, and cursor-read failure propagation are represented in production code and focused tests. | The reconnect race was a real process-killing defect in the shared websocket.zig fork — std.posix.setsockopt maps BADF/NOTSOCK/INVAL to unreachable, which is reachable on a connection socket — fixed there, shipped through zat v0.3.18, and pinned by a production-path regression that drives real abrupt disconnects over a loopback socket. A Zat pin bump re-opens this row. |
| Whole-repository repair/resync | verified | Fetched repositories are authenticated before replacement emission; valid repair emits sync then bounded replacement batches; repair-tail encoding errors are isolated per row. Capacity tests exercise 32 active plus 64 queued repairs. | Encoding isolation covers both repair-tail and ordinary live publication, and retry-subsystem terminal failures are tied to service health. Capacity is proved at 32 active plus 64 queued repairs. |
| Delete/update compaction | verified | Rewrite selection, survivor correctness, manifest-before-watermark ordering, fsync/rename cutpoints, bounded workers, and reason counters have physical-file tests. The versioned watermark now distinguishes absence from Store failure, rejects wrong width/version, and refuses to initialize over corrupt bytes; focused tests inject the production Store read fault. | This row does not close merge orchestration or cold-serving gaps. |
| Timestamp import | verified | CSV parsing, rule precedence, patch topology preservation, job persistence, restart, HTTP authentication, and write/fsync/rename power-loss cuts have concrete tests. | Public visibility after import rides the cold path, which is now verified; this row does not separately re-prove replay. |
| Status: hosts and accounts | verified | Durable host aggregates and account lookup/verification have focused production HTTP tests, including rate limiting and restart. Together with the segments, collections and summary views this is upstream's status surface rather than a subset of it. | Rate-limit windowing is Stream's own; upstream does not limit these routes. |
| Status: summary, collections, segments | verified | The field report at /status carries upstream's status information set: repository rollup with percent complete (BackfillStats), live cursors (LiveStats), retention with configured lookback and oldest retained seq (CursorLookbackStats), timestamp-import job state and phase (ImportInfo), and the durable lifecycle phase with its age and backfill duration (PhaseInfo). Segments and collections render from resident manifest metadata, cross-checked against listSegments. Absence is distinguished from zero throughout — no repository state does not render as "0 of 0 complete", and disabled cursor replay is reported rather than shown as no retention. A test pins the set so it cannot quietly shrink back to process-local counters. |
The ASCII river at GET / is an intentional aesthetic divergence and is not a status surface; upstream has no equivalent. Per-tree storage totals and a PebbleStats-equivalent metadata-store size are not surfaced. |
| Prometheus metric exposition | partial | Metrics are emitted in Prometheus syntax and focused tests show families increment at their intended boundaries. Backfill progress now has a restart-stable series: stream_backfill_repos_durable{status=...} is read from the metadata store at scrape time, so it continues across a restart instead of resetting — measured at 22 complete before a restart and 49 after, where the process-local gauge starts again from zero. Absence is distinguished from zero: no repository state emits no series rather than zero-of-zero. |
Producer semantics are proved for the progress families specifically, not for every family on the dashboard. Label equivalence with upstream and monotonicity are still unproved in general. |
| Grafana dashboard reuse | partial | The upstream dashboard is checksum-pinned and deliberate runtime-only panels are separated from upstream panels. The progress chart's "durably committed" series now queries the durable store-backed gauge; it previously queried jetstream_backfill_progress_completed, a process-local counter, whose value resets on restart; the durable gauge reflects the archive itself. A contract test pins which metric backs that legend. |
The remaining panels are still validated by family presence rather than producer semantics, so this row tracks the progress chart only. |
| Public/debug listeners | verified | Production-process tests establish route isolation, disabled debug binding, public WebSocket admission, readiness, and the historical combined listener mode. | This row does not establish lifecycle readiness or data completeness. |
| Graceful shutdown | verified | Listener admission stops, clients receive 1001 within configured budgets, blocked peers are interrupted, and a hung listRepos bootstrap exits cooperatively in the focused receipt. | Listener admission stops, clients receive 1001 within configured budgets, blocked peers are interrupted, and a hung listRepos bootstrap exits cooperatively. The exact deployed image now is admitted: shutdown-contract passes inside receipts/79ecfa5.json, bound to the published digest. Dependency reconnect failures are no longer process-fatal. |
| Logging | verified | Real processes exercise JSON/text selection, level filtering, default values, and invalid configuration. | Semantic parity is limited to the public controls and rendered fields, not byte-identical Go logging internals. |
| OpenTelemetry | partial | Provider configuration, propagation, sampling, batching, TLS/mTLS, and representative production spans have real collector tests. | “Every pinned production span” has not been independently re-audited in this pass. Treat the enumerated tested spans as evidence, not a blanket closure. |
| Inspect/version command surfaces | verified | version, sealed inspect-segment, active inspection, and representative inspect-all reports have fixtures and golden comparisons. The golden comparisons run inside zig build test, so they are current with every commit by construction rather than by a remembered manual pass, and that suite is bound to the admitted digest. |
Inspection is an offline surface and is not a substitute for the online status tabs — which now exist. |
| Differential oracle | partial | It provides event-log, final-state, restart, repair and public-client coverage against a pinned upstream simulator. The property most at risk — per-repository rather than per-batch completion, whose absence caused the July replay — is proved at the production boundary by focused tests: writer durability commits completed repositories independently of a dispatch straggler (a repository becomes durable while a sibling in the same dispatch batch is still incomplete) and durable repository completion survives reopen before listRepos cursor advance. |
The oracle itself cannot confirm that property end to end. Observing the difference requires a crash landing with some repositories durably complete and others not; in a 100-repo simulator the archive flushes once for the whole corpus, so every cut leaves either none or all of them durable and both designs look identical. A dispatch-scale kill test could not be made meaningful at this fixture size and was removed. This is a limit of the harness, not a divergence — Stream does not behave differently from upstream here. Only a real whole-network run, where in-flight flushes leave a partial prefix, confirms it end to end. |
| Strict power-loss oracle | verified | It exercises acknowledged-write reconstruction at named write/fsync/rename boundaries. Focused tests separately inject Store read failures for lifecycle phase, relay cursor, compaction watermark, and merge cursor; the merge restart guard also distinguishes source absence from filesystem failure. The completion reopen test proves completed-versus-interrupted repository classification around an unchanged listRepos cursor. It now also cuts power between the archive fsync that makes a repository's rows durable and the metadata commit that acknowledges completion — the durability-ordering seam the whole design rests on. Recovery reconstructs every acknowledged event and the repository simply completes again. | Missing manifest files during replay and arbitrary-length listRepos cursors are covered by physical tests elsewhere (cold replay deletes a manifest-listed segment; discovery uses 4 KiB cursors) rather than by a power cut, because neither is a durability boundary. |
| Exact artifact admission | verified | scripts/admit implements the procedure end to end and has now been run for real. Commit b04bb34 passed all 20 suites — both oracles included — and the receipt receipts/b04bb34.json binds those results to sha256:2f2bda74d78bd7c5596acc368a60946020ed1a6b2ab4759566d2196f2e7a71ec, which is published. admit verify admits that digest and refuses sha256:adc276…, the e1926f3 image that failed the July experiment; the experiment harness deploy.sh calls it before rollout. A receipt covers a digest, never a tag; a skipped suite blocks admission exactly like a failure. The artifact was built on a native linux/amd64 host rather than under emulation, so it is the same architecture as the deploy target rather than a cross-emulated approximation. tests/admission_contract.py pins the refusals. |
The receipt is bound to b04bb34; any later commit needs its own admit run and admit publish. |
Evidence that must not be overread #
required=104 present=104 missing=0means only that the scrape exposed every family named by the dashboard.- A reciprocal JSS fixture proves sealed-format compatibility, not crash recovery or replay continuity.
- Final-state convergence does not prove bounded replay work or the timing of individual durable completions.
- A successful restart after a named crashpoint proves only that the crashpoint was placed at a safe boundary. It does not prove the boundary matches upstream.
- A four- or 104-account simulator cannot by itself establish behavior at a 100,000-entry dispatch/checkpoint boundary. Per-repository completion needs an explicit held-sibling invariant, not merely a smaller happy-path batch.
Admission requirements #
A new whole-network experiment is forbidden until all blocked rows are resolved and all partial rows that can affect durability, completeness, availability, or operator truth are either verified or explicitly accepted as divergences.
The admission command must:
- require a clean worktree and record the exact Stream, upstream, Atmos, Zat, and dashboard revisions;
- build the exact linux/amd64 artifact that will be deployed;
- run invariant-based negative tests, including metadata read corruption, per-repository completion crashes, replay holes, persistent encode errors, and transient firehose disconnects;
- run the positive interoperability and configuration receipts;
- publish by digest and emit one signed/immutable receipt binding every result to that digest; and
- make deployment reject any other digest.
The previous 4d163e8 receipt remains historical evidence for that artifact.
It is not a current admission and must not be copied forward.
archive key ring (2026-08-17, extension) #
upstream OSS has no archive auth at all — its only bearer surface is the
single timestamp-import token (xrpcapi/auth.go). Bluesky's hosted
instances gate and meter archives in a proprietary gateway in front
(2 MB/s per key by default; the gateway exists to adjust limits per key —
alex.bsky.team, 2026-08-17). stream's single --archive-api-key was
already an extension mirroring that hosted behavior; the key ring
(--archive-api-keys-file, serve/api_keys.zig) extends it into an
in-process approximation of the gateway: named keys (name:key[:mbps]
lines), per-key token-bucket byte budgets (429 + retry semantics the SDK
already honors), and revocation by deleting a line (mtime-based reload,
no restart). the fleet key remains valid and unmetered. rationale:
metering protects compaction/live-delivery from uncapped archive drains
on a ~200 MB/s volume; revocation avoids rotating the fleet credential
per consumer.
in-process admission caps (2026-08-21, extension) #
upstream keeps connection admission at the edge: docs/README.md §7 puts
per-IP limits on subscriber counts, live-tail events, and segment
downloads at the CDN and origin proxy, explicitly calling in-process
limits "an implementation detail of the Bluesky-hosted instance" — the
binary itself only bounds per-request cost (filter caps, list limits,
byte-budgeted caches) plus one per-IP fixed-window limiter guarding
web repo actions. stream fronts with a single Caddy (no CDN), so two
caps that need process knowledge live in-process as opt-in extensions
(--max-subscribers, --max-cold-readers; 0/unset = unbounded,
preserving upstream behavior exactly):
- max_subscribers: global subscriber ceiling; excess upgrades close 1013 ("at capacity") before a Subscriber is allocated.
- max_cold_readers: concurrent cold-replay bound. only the process knows a cursor is about to cost archive reads; over-limit passes return .contended (20 ms pause + retry — backpressure, not an error; never counts toward the rotation-seam stall disconnect). known limits, accepted at write time: slots are global (one client with N connections can hold every slot — per-consumer fairness needs subscriber identity, which the open websocket does not have), and the retry sleep assumes the thread-per-subscriber serving model.
request-rate admission (per-IP/per-subnet) stays at the proxy layer per upstream doctrine: deploy/caddy builds Caddy with mholt/caddy-ratelimit and deploy/site/Caddyfile defines the zones (ws connects, keyed xrpc, anonymous http).