jetstream v2 in zig stream.waow.tech
stream docs configuration.md
16 kB
Markdown

Configuration reference #

What each runtime control does and what its defaults are. This is the operator reference; configuration-parity.md audits how each control is proved, and semantic-parity.md is the deployment gate. Serve flags are declared in internal/runtime/cli.zig; offline subcommands are dispatched from src/main.zig. Deploy notes live in deploy/README.md.

Flag and environment names are derived from the options struct at comptime: a field backfill_workers becomes --backfill-workers and JETSTREAM_BACKFILL_WORKERS via flagName / envName. Adding a field to the struct defines both spellings automatically.

Stream uses Jetstream V2's canonical names. Every canonical flag also reads the same JETSTREAM_* environment variable as the pinned upstream command. Explicit CLI values take precedence over environment values. Unknown JETSTREAM_* names fail the process before command execution, which catches misspelled production settings; Kubernetes's JETSTREAM_APP_* service variables and upstream's JETSTREAM_ORACLE_* and JETSTREAM_SIM_* test namespaces are exempt.

--relay-url drives both the live WebSocket and the HTTP repository clients. --plc-url configures the DID resolver. The older split flags remain accepted for existing local harnesses.

Intentional divergences #

Two controls deliberately differ from upstream and are recorded here per the configuration audit:

  • The deprecated cursor block-index cache flag is a compatibility no-op. Stream accepts signed values and ignores them; its sealed indexes are always manifest-resident. Accepting the flag keeps existing command lines working.
  • There are no /debug/pprof/* routes. Zig has no pprof runtime; process metrics on the debug listener serve that purpose.

Listeners #

stream serve --addr=:8080 --debug-addr=127.0.0.1:6060

The public listener owns /, /status, /subscribe, /subscribe-v2, and the archive XRPC methods. The optional debug listener owns /healthz, /readyz, and /metrics; an empty or omitted --debug-addr binds no debug socket.

/readyz means the configured accept loops are running — not that bootstrap has reached steady state. The flag-only --port=6008 form binds a single combined listener and is retained for existing deployments and offline harnesses; new deployments should use serve --addr and keep the debug listener private.

Shutdown #

--shutdown-timeout=5s
--client-drain-timeout=10s

On SIGINT or SIGTERM, Stream stops listener admission, sends WebSocket close code 1001 to all subscribers concurrently (a backpressured client cannot delay the others), drains cooperative subscribers until the client budget expires, and shuts down the public and debug HTTP listeners from the same instant while the accepted ingest pipeline reaches its durable boundary.

Matching upstream, a non-positive HTTP shutdown timeout selects 30 seconds. A non-positive client-drain timeout is treated as already expired: teardown is immediate and delivery of the close frame is best-effort.

Logging #

--log-level accepts debug/info/warn/error and their upstream aliases; --log-format selects json (default) or text. Filtering is runtime-configurable in ReleaseSafe builds and applies to Stream and its Zig dependencies. Root flags are persistent: both stream --log-level=debug serve and stream serve --log-level=debug work.

Subscribers #

The writer-owned hot log retains 256 MiB by default and never evicts rows above the durable archive watermark; tune with --subscribe-read-log-retention-bytes.

The slow-client policy is configurable with --subscribe-slow-window (default 60s) and --subscribe-slow-min-rate (default 5 log events/second). A client is disconnected only when both limits are continuously violated while it is more than 100,000 events behind.

Cold subscribers share a 64 MiB decoded-block cache (--subscribe-block-cache-bytes) covering decoded JSS events, memoized v1/v2 JSON, and the corresponding dictionary-zstd frames. Compaction and timestamp rewrites invalidate the affected segment generation.

Pulls scan at most 1,024 raw log entries at a time (--subscribe-read-batch). The bound applies before filters and v1 skip rules in both hot and durable reads, so selective subscriptions still yield regularly and slow-client accounting follows raw network progress.

Replay window #

Subscriber replay implements upstream's 36-hour default lookback. A v1 sequence cursor older than the retained window is clamped to the first eligible sealed segment; v2 rejects the same cursor before the WebSocket upgrade with a cursor too old response. Timestamp cursors are clamped for both protocols. --cursor-lookback=72h changes the window; zero makes both endpoints pure-live.

Archive serving #

Sealed downloads use upstream's cache policy. The default is Cache-Control: public, no-cache; a positive Go-style duration such as --segment-cache-max-age=15m is rounded up to whole seconds in max-age. Use a positive max-age only when compaction is disabled or the cache lifetime is comfortably shorter than the compaction interval, because compaction can replace a segment while retaining its name.

planBackfill exposes upstream's controls and defaults:

--plan-max-dids=1000
--plan-max-collections=25
--plan-max-entries=100000
--plan-whole-segment-threshold=0.75

Zero disables the corresponding cap for non-empty DID or collection filters. For --plan-max-entries, zero means one unbounded page. The threshold must be greater than zero and at most one.

Compaction #

--compaction-interval=4h
--compaction-tombstone-cap=32000000
--compaction-rewrite-workers=0

The interval accepts a non-negative Go-style duration; zero disables all compaction. The tombstone cap bounds each durable fold/rewrite chunk; zero is upstream's unlimited sentinel. Rewrite workers bound concurrent sealed-segment rewrites; zero selects min(CPU count, 8). Reaching the cap wakes the compactor immediately, subject to upstream's 30-second between-pass floor.

Failed-repository healing #

--failed-repo-retry-interval=4h
--failed-repo-retry-workers=16
--failed-repo-retry-host-workers=4
--failed-repo-retry-max-delay=168h

Interval and maximum delay accept non-negative Go-style durations; a zero interval disables the background loop. Zero worker values select upstream's 16-global/4-per-host defaults; zero maximum delay selects seven days. Transient failures use upstream's capped exponential delay plus uniform [0, delay/2) jitter.

Attempts, error class/message, final host, retry count, and next attempt are durable in RocksDB. Every repo transition also updates a persistent host/<bucket> diagnostic aggregate in the same synced batch: current lifecycle/active counts, cumulative failure classes, and the five most recent bounded error samples survive restart without scanning the whole-network repo keyspace.

Bootstrap #

--backfill
--backfill-workers=100
--backfill-batch-size=100000
--backfill-async-flush-workers=4
--skip-merge-discovery

Worker count controls concurrent network downloads; zero selects the 100-worker default. Batch size counts every listRepos entry, including inactive and already-complete repositories: Stream accumulates whole 1,000-entry protocol pages until it reaches the target, then shuffles eligible repositories before dispatch. Requests go through the relay and follow its redirect, so batch composition does not control which PDS is contacted. Zero selects the 100,000-entry default.

Worker count is the only concurrency bound, matching upstream. Preparation copies the downloaded CAR into the worker's own parse arena and releases the scratch mmap before emission, so each repository's bytes are owned, and repositories prepare and emit concurrently. --backfill-async-flush-workers bootstrap-only workers detach complete JSS blocks, compress concurrently, and commit/fsync in preparation order; zero selects synchronous compression.

--backfill-max-inflight-bytes was removed in an earlier revision and is no longer accepted.

Downloads need temporary free space on --data-dir equal to the concurrently in-flight CARs. Scratch files are deleted after each attempt and cleared on restart. Bootstrap owns <data-dir>/backfill/repo-scratch; steady live and failed-repo healing use <data-dir>/repair-scratch/{live,failed-repo}, so repair cannot recreate lifecycle state after cutover.

Resource exhaustion is lifecycle-fatal. It is never persisted as a repo-level success or failure, and never classified as malformed remote data; see invariants.md.

Backfill progress, active workers, queue depth, and current/peak transient bytes are exported on /metrics. Successful CARs record their final emitted archive sequence, and writer durability commits every covered repository independently: the dispatch batch controls listRepos cursor-checkpoint cadence, but a slow sibling does not withhold already-durable repository completion. The authoritative lifecycle analysis is bootstrap-semantic-parity.md.

Discovery and selected repositories #

--skip-merge-discovery omits the post-merge listRepos rescan. Matching upstream, Stream enables it automatically for --max-backfill-repos and --backfill-repos; a normal full crawl still replays the last page so accounts created during bootstrap enter durable retry state. Enabling it for a whole-network crawl omits accounts created during bootstrap.

The durable account row retains the distinct initial-backfill and latest known revisions, update time, declared handle, PDS endpoint, and reserved record/byte counters. The debug-only explicit --backfill-repos=did:plc:example[,did:web:example.com] path processes the requested DIDs serially, resolves their real DID documents, and removes stale full-network merge-discovery state. Ordinary whole-network bootstrap does not issue per-repo identity requests.

handle/<normalized> is maintained atomically with every repo transition, and each merge-source commit advances its cursor together with latest-revision refreshes.

Account verification #

Account lookup resolves handles through the configured identity directory, reconstructs the repository from sealed JSS segments, the active writer, pending rows, and bootstrap live segments at a rotation-safe boundary, then compares the canonical local MST root with the PDS commit root obtained through Sync 1.1 getLatestCommit and getBlocks. Verification runs automatically for a found account.

Matching upstream, expensive repo actions are limited to four per source IP per minute with a bounded 4,096-source map. Trusted operator deployments can pass --disable-repo-action-rate-limits.

Rebloom sweep #

--rebloom-sweep

Stream-only. Runs a one-shot background sweep at startup that right-sizes per-block DID bloom regions in segments sealed before the cardinality-sizing fix, reusing each segment's data region byte for byte and refreshing the resident manifest entry as it goes (resident memory falls while the sweep runs). Already-right-sized segments are skipped, so the sweep is idempotent. Requires --compaction-interval=0: sealed-file rewriting is single-writer, and the process refuses to start the sweep alongside an enabled compactor. Progress is exported as stream_rebloom_segments_total and stream_rebloom_bytes_reclaimed_total and logged every 100 segments.

Timestamp import #

Disabled to remote callers unless a non-empty bearer token is configured. Stage plain CSVs under the confined import directory, submit a relative path, then poll the returned job id:

--timestamp-import-token=<secret>
--archive-api-key=<secret>
--timestamp-import-dir=/srv/stream/imports

POST /xrpc/network.bsky.jetstream.importTimestamps
GET  /xrpc/network.bsky.jetstream.getImportStatus?job=<id>

Both routes require Authorization: Bearer <secret>. Put them behind TLS at the reverse proxy; Stream serves plain HTTP internally. Empty, missing, and bad credentials all produce the same 401. Submitted paths are canonicalized through symlinks and must resolve to regular files inside the configured directory.

archive key ring #

--archive-api-keys-file=<path> (JETSTREAM_ARCHIVE_API_KEYS_FILE) adds named, metered archive credentials beside the fleet key:

# name:key[:mbps]     mbps omitted or 0 = unmetered
evelyn:some-long-random-secret:8
ci:another-secret
  • keys authorize the archive endpoints exactly like --archive-api-key (which stays valid and unmetered)
  • mbps is a per-key token bucket (30s burst; burst_seconds in serve/api_keys.zig) charged on served bytes; a key in deficit gets 429 RateLimitExceeded until it refills — the SDK's fetch retry already backs off on 429
  • revocation = delete the line. the file's mtime is rechecked at most once per second on the request path; no restart, no signal
  • an unreadable/missing file revokes every ring key (fail closed) and logs a warning; the fleet key is unaffected

admission caps #

--max-subscribers=<n> (JETSTREAM_MAX_SUBSCRIBERS) and --max-cold-readers=<n> (JETSTREAM_MAX_COLD_READERS); non-positive or unset = unbounded (the upstream-faithful default — upstream bounds these at its CDN/proxy, docs/semantic-parity.md).

  • at the subscriber ceiling, new upgrades close with 1013 ("at capacity; retry later") before a subscriber is allocated; counted in stream_subscribe_capacity_rejects_total
  • max-cold-readers bounds concurrent cold-replay passes (a cursor below the resident tail reads the archive). an over-limit pass waits 20 ms and retries — the subscriber stays connected and delivery is delayed, not dropped; counted in stream_subscribe_cold_contended_total
  • request-rate limits per client IP are NOT here — they live in the caddy layer (deploy/site/Caddyfile; see docs/deploying.md "Admission limits" for the layer table and current values)

slow upstream reconnect #

--upstream-slow-min-rate=<frames/s> (JETSTREAM_UPSTREAM_SLOW_MIN_RATE); non-positive or unset = off (the upstream-faithful default — atmos reconnects only when the socket closes or errors).

  • the live consumer counts relay frames in a 60s window per connection. a window below the floor drops the connection and reconnects from the cursor, so nothing is skipped; counted in stream_upstream_slow_reconnects_total
  • each slow verdict doubles the next window (60s up to 16m) and a healthy window resets it, so a relay that is slow on every connection costs one reconnect per 16 minutes, not one per minute
  • it complements the fixed 20s idle timeout, which only sees total silence. on 2026-10-01 relay1.us-east.bsky.network delivered ~18 frames/s on some connections for up to 80 minutes; other connections to the same host were at full rate, which is why a reconnect is worth trying
  • set the floor well under the network's quietest hour, and leave it off against a simulator or a small relay

OpenTelemetry #

Stream installs the same process-wide tracing provider as pinned Jetstream V2. Tracing is a no-op unless OTEL_EXPORTER_OTLP_ENDPOINT or OTEL_EXPORTER_OTLP_TRACES_ENDPOINT is set. When configured, completed spans are batched and exported as OTLP/HTTP protobuf; the generic endpoint receives the /v1/traces suffix and the trace-specific endpoint is used verbatim. --otel-service-name and OTEL_SERVICE_NAME set service.name; the build version supplies service.version. Full endpoint, header, gzip, insecure, timeout, custom-CA, client-certificate, batch, and sampler controls are in opentelemetry.md.

The container includes the libcurl 4 runtime, used only when an OTLP client certificate and key are configured. Ordinary HTTPS and custom-CA-only exports use Zig's native TLS path and do not load libcurl.