jetstream v2 in zig stream.waow.tech
stream docs live-scheduler.md
3.8 kB
Markdown

Live scheduler parity #

Stream's live path follows the scheduler and cursor behavior in pinned Atmos v0.2.14 — the Go atproto library upstream Jetstream V2 runs, and the reference these numbers come from — which Jetstream V2 configures with its defaults. The constants below are matched to upstream.

Verified against ingest/pipeline.zig on 2026-07-30: worker_count = 32, result_capacity = 4096, batch_size = 50, batch_timeout = 500 ms, and the jetstream_livestream_verify_queue_dropped_events_total counter all exist as described.

The 32 and the 64 are one constant, not two. Per-DID pending capacity is key_queue_capacity = worker_count * 2, so "32 workers" and "64 pending events" are not independent knobs — changing worker_count changes the drop-oldest threshold with it (see invariants.md, "a bounded producer must not outrun its consumer").

"Cutover" below does not mean what it means elsewhere. In this file "graceful cutover" is the shutdown handover — stop admission, drain, tear down workers. In deployment-runbook.md, upstream-bootstrap-spec.md and the other lifecycle docs, cutover is the phase transition out of merge into steady state where serving stops returning 503.

Contract #

  • Thirty-two workers execute at most thirty-two active DID chains. A new DID backpressures the websocket dispatch path until a worker is available.
  • A worker owns one DID chain until it is empty. Events for that DID execute in arrival order; other DIDs can complete and reach the writer first.
  • Each active DID may retain 64 pending events. The 65th pending arrival removes the oldest pending event and appends the newest. The displaced seq leaves the inflight set and increments the canonical jetstream_livestream_verify_queue_dropped_events_total counter.
  • Results are delivered in completion order in batches of 50 or after 500 ms. Completed-result buffering is bounded to Atmos's 4,096 entries; a full result channel backpressures the worker that owns that DID chain. A malformed frame, verifier error, replay, or repair-withheld event may be collected silently; it does not manufacture a non-empty delivery batch.
  • After a successful non-empty batch, the relay cursor becomes the greater of its prior value and min(inflight)-1. If nothing remains inflight, it uses the largest seq in that batch. The pipeline is seeded from the durable cursor at startup, so a replayed lower batch cannot regress it. The cursor callback runs only after all event callbacks in the batch succeed.
  • Archive rows, verifier state, replay ratchets, and relay cursor retain their existing fsync coupling. Cursor staging or durable metadata failure is lifecycle-fatal.
  • Graceful cutover stops admission, drains every accepted frame and ordered repair completion, then tears down worker tasks. Abrupt cancellation owns and frees worker-local and writer-local jobs rather than abandoning them.

Offline receipts #

zig build test uses no public network. Scheduler tests encode real firehose control frames and use the production verifier with a temporary RocksDB store. Holding the verifier's hosting-state lane supplies deterministic pressure:

  1. a second DID completes while the first is blocked;
  2. a second event for the blocked DID cannot pass the first;
  3. all 32 active DID chains occupy the pool and the 33rd submit remains blocked until one completes;
  4. 65 pending events retain seqs 3 through 66, proving seq 2 was the exact drop-oldest victim and left the inflight set; and
  5. cursor calculation returns the smallest inflight seq minus one, then the batch maximum once the set is empty.

These receipts complement the loopback HTTP and signed-CAR repair tests in live-repair.md; neither suite contacts a public PDS, relay, or PLC service.