jetstream v2 in zig stream.waow.tech
README.md

deploying stream #

This file is the deploy procedure. It deliberately does not track which experiment is running — that state belongs in ../docs/full-network-experiment-2026-07.md, which is maintained as runs happen. An earlier revision asserted that "no replacement experiment should be provisioned" while experiment 6 was in fact running on an admitted artifact, which is exactly the rot that got deployment state removed from the top-level README.

The retired live-only canary no longer runs on the Indigo relay node.

The long-term production target remains a separate host with redundant local storage; passing a cloud-volume experiment is evidence for sizing and throughput, not approval to colocate Stream with Indigo.

The previously deployed e1926f3 image is historical evidence only. It was not the artifact covered by the earlier admission receipt and it failed during the experiment. Do not reuse its digest as a promotion candidate. That gap — an image no receipt covered, with nothing to refuse it — is what scripts/admit now closes.

just package        # ReleaseSafe x86_64-linux-gnu binary in zig-out/bin/
just admit run      # clean tree -> every suite -> exact image -> one receipt
just admit publish  # push that image, pin the receipt to its registry digest
just admit verify sha256:...   # deployment gate; non-zero unless admitted

Nothing may be deployed whose digest just admit verify does not accept. A receipt covers exactly one image digest, never a tag, and a suite recorded as skipped blocks admission just as a failure does — a partial run must never read as a full one. just admission-contract proves those refusals offline.

Do not run just admit run as the deploy gate. It is fine for debugging one suite, but it is invisible to the concurrency limit that keeps two gates off the same machine — which we learned by running two at once and throttling both (see ../docs/deploying.md). The sanctioned entry point is the stream-admission Prefect deployment in the sibling repo zzstoatzz.io/my-prefect-server, which runs the identical scripts/admit run on bare metal (heavypad) with concurrency_limit: 1. .tangled/workflows/admission.yml only triggers it.

Concretely, this repo's deploy path depends on that repo for:

what where
the gate runner (flow + deployment) my-prefect-server: flows/stream_admission.py, prefect.yaml
the machine it runs on heavypad, via Prefect's home-pool (a process pool)
serialisation the deployment's concurrency_limit: 1 (ENQUEUE)
toolchain on that box zig, go, just, uv, docker — go/just live in the shared nix profile

If that deployment is gone or its worker is offline, pushes will not produce receipts and nothing will be deployable. That is deliberate: no receipt, no deploy.

required environment (docs/lessons-from-zlay.md #8, #9) #

  • MALLOC_ARENA_MAX=4 — the highest-leverage glibc RSS knob for thread-heavy services (per-thread arena fragmentation)
  • public traffic must target --addr; probes and Prometheus must target the private --debug-addr (/healthz, /readyz, and /metrics)
  • immutable image tags (git SHA); confirm what's running via the stream_build_info{git_sha,optimize} metric

flags for production #

serve
--addr=:8080
--debug-addr=127.0.0.1:6060
--relay-url=https://relay1.us-east.bsky.network
--upstream-slow-min-rate=50
--plc-url=https://plc.directory
--data-dir=/data
--backfill
--backfill-workers=100
--backfill-async-flush-workers=4
--subscribe-read-log-retention-bytes=268435456
--subscribe-block-cache-bytes=67108864
--subscribe-read-batch=1024
--subscribe-slow-window=60s
--subscribe-slow-min-rate=5

Omitting or emptying --debug-addr disables the debug socket. Do not expose it through the public ingress. /readyz proves both configured accept loops are running; use /status and metrics for ingest/bootstrap health. Run just listener-contract before packaging to prove split routing and the disabled-debug case offline.

The production host target is 16 cores, 32–64 GB RAM, and at least 3.84 TB usable mirrored SSD/NVMe. Keep RocksDB and the archive on the same RAID1 XFS filesystem: segment fsync precedes the RocksDB batch that acknowledges those bytes, and the unified filesystem is reflink-snapshotted for backup.

The defaults reserve at most 8 GiB for transient backfill allocations while admitting 100 concurrent disk-backed downloads. CAR preparation is serialized, copies each complete CAR into that budget, and releases the scratch mmap before concurrent emission. This avoids partial-index deadlock while prepared repositories still emit concurrently. Four bootstrap-only workers compress detached JSS blocks and commit them in order. The pod still needs headroom for RocksDB, archive/compaction buffers, live capture, thread stacks, zstd frames, and libc; do not set the container limit equal to the backfill budget.

promotion gate #

The complete gate is ../docs/semantic-parity.md. In particular, a promotion must test individual repository durability inside a large dispatch batch, metadata read corruption, unreadable archive files, persistent live encoding failures, and transient firehose disconnects. The receipt must bind those results to the exact deployed image digest. Throughput, backup restore, client interoperability, and alerts remain necessary, but they cannot waive a blocked semantic row.

alert rules #

prometheus/rules.yml is the in-repo source for the Prometheus alert rules; the box's Prometheus reads its own copy at /opt/stream-experiment/rules.yml (mounted to /etc/prometheus/rules.yml, listed under rule_files in prometheus.yml), so a rule change here is not live until that copy is replaced and Prometheus restarted (docker compose restart prometheus; the container has no --web.enable-lifecycle, so there is no reload endpoint, and up -d does not restart a container whose compose definition is unchanged). Write the box copy in place (cat new > rules.yml): the file is a single-file bind mount. Wired up 2026-08-29; before that the box had no rules loaded at all. Validate with promtool check rules deploy/prometheus/rules.yml (e.g. via docker run --rm -v "$PWD/deploy/prometheus:/rules:ro" --entrypoint promtool prom/prometheus check rules /rules/rules.yml).

paging #

Since 2026-10-02 an alertmanager service in the box's compose stack (prom/alertmanager:v0.33.0, pinned by digest) posts critical and warning alerts to Discord; info alerts are dropped. prometheus/alertmanager.yml and prometheus/prometheus.yml are the in-repo sources for the box copies at /opt/stream-experiment/. The compose file itself lives only on the box.

The webhook URL is the one zlay's Alertmanager uses (the zlay-discord-webhook secret in the zlay cluster's monitoring namespace). It is on the box as /opt/stream-experiment/discord-webhook, owner 65534, mode 400, and is read through webhook_url_file. Rotating the webhook means replacing both copies. To prove the route end to end:

docker compose exec -T alertmanager amtool --alertmanager.url=http://127.0.0.1:9093 \
  alert add StreamPagingTest severity=warning job=stream
docker compose exec -T alertmanager wget -qO- http://127.0.0.1:9093/metrics \
  | grep 'alertmanager_notifications.*discord'

dashboard gate #

grafana/jetstream-upstream.json is the byte-exact dashboard from the pinned Jetstream V2 reference. grafana/stream.json is the deployable artifact, generated deterministically from that source. The Go runtime row is adapted to honest process metrics, and one explicit Stream-only row exposes scheduler tickets, unsettled stage gaps, and terminal failures during bootstrap and steady state; every upstream protocol/lifecycle panel remains unchanged. Run just dashboard-test, then just process-metrics-contract, then just dashboard-contract <metrics-url-or-saved-scrape>. These checks are network-independent when given a saved scrape or the loopback harness and fail when a queried family is absent. They verify dashboard structure and metric presence only; they are not a promotion decision.