Live scheduler parity #
Stream's live path follows the scheduler and cursor behavior in pinned Atmos v0.2.14 — the Go atproto library upstream Jetstream V2 runs, and the reference these numbers come from — which Jetstream V2 configures with its defaults. The constants below are matched to upstream.
Verified against ingest/pipeline.zig on 2026-07-30: worker_count = 32,
result_capacity = 4096, batch_size = 50, batch_timeout = 500 ms, and the
jetstream_livestream_verify_queue_dropped_events_total counter all exist as
described.
The 32 and the 64 are one constant, not two. Per-DID pending capacity is
key_queue_capacity = worker_count * 2, so "32 workers" and "64 pending events"
are not independent knobs — changing worker_count changes the drop-oldest
threshold with it (see invariants.md, "a bounded producer must not outrun its
consumer").
"Cutover" below does not mean what it means elsewhere. In this file
"graceful cutover" is the shutdown handover — stop admission, drain, tear down
workers. In deployment-runbook.md, upstream-bootstrap-spec.md and the other
lifecycle docs, cutover is the phase transition out of merge into steady state
where serving stops returning 503.
Contract #
- Thirty-two workers execute at most thirty-two active DID chains. A new DID backpressures the websocket dispatch path until a worker is available.
- A worker owns one DID chain until it is empty. Events for that DID execute in arrival order; other DIDs can complete and reach the writer first.
- Each active DID may retain 64 pending events. The 65th pending arrival
removes the oldest pending event and appends the newest. The displaced seq
leaves the inflight set and increments the canonical
jetstream_livestream_verify_queue_dropped_events_totalcounter. - Results are delivered in completion order in batches of 50 or after 500 ms. Completed-result buffering is bounded to Atmos's 4,096 entries; a full result channel backpressures the worker that owns that DID chain. A malformed frame, verifier error, replay, or repair-withheld event may be collected silently; it does not manufacture a non-empty delivery batch.
- After a successful non-empty batch, the relay cursor becomes the greater of
its prior value and
min(inflight)-1. If nothing remains inflight, it uses the largest seq in that batch. The pipeline is seeded from the durable cursor at startup, so a replayed lower batch cannot regress it. The cursor callback runs only after all event callbacks in the batch succeed. - Archive rows, verifier state, replay ratchets, and relay cursor retain their existing fsync coupling. Cursor staging or durable metadata failure is lifecycle-fatal.
- Graceful cutover stops admission, drains every accepted frame and ordered repair completion, then tears down worker tasks. Abrupt cancellation owns and frees worker-local and writer-local jobs rather than abandoning them.
Offline receipts #
zig build test uses no public network. Scheduler tests encode real firehose
control frames and use the production verifier with a temporary RocksDB
store. Holding the verifier's hosting-state lane supplies deterministic
pressure:
- a second DID completes while the first is blocked;
- a second event for the blocked DID cannot pass the first;
- all 32 active DID chains occupy the pool and the 33rd submit remains blocked until one completes;
- 65 pending events retain seqs 3 through 66, proving seq 2 was the exact drop-oldest victim and left the inflight set; and
- cursor calculation returns the smallest inflight seq minus one, then the batch maximum once the set is empty.
These receipts complement the loopback HTTP and signed-CAR repair tests in
live-repair.md; neither suite contacts a public PDS, relay, or PLC service.