# deploying stream **This file is the deploy *procedure*. It deliberately does not track which experiment is running** — that state belongs in [`../docs/full-network-experiment-2026-07.md`](../docs/full-network-experiment-2026-07.md), which is maintained as runs happen. An earlier revision asserted that "no replacement experiment should be provisioned" while experiment 6 was in fact running on an admitted artifact, which is exactly the rot that got deployment state removed from the top-level README. The retired live-only canary no longer runs on the Indigo relay node. The long-term production target remains a separate host with redundant local storage; passing a cloud-volume experiment is evidence for sizing and throughput, not approval to colocate Stream with Indigo. The previously deployed `e1926f3` image is historical evidence only. It was not the artifact covered by the earlier admission receipt and it failed during the experiment. Do not reuse its digest as a promotion candidate. That gap — an image no receipt covered, with nothing to refuse it — is what `scripts/admit` now closes. ```sh just package # ReleaseSafe x86_64-linux-gnu binary in zig-out/bin/ just admit run # clean tree -> every suite -> exact image -> one receipt just admit publish # push that image, pin the receipt to its registry digest just admit verify sha256:... # deployment gate; non-zero unless admitted ``` Nothing may be deployed whose digest `just admit verify` does not accept. A receipt covers exactly one image digest, never a tag, and a suite recorded as `skipped` blocks admission just as a failure does — a partial run must never read as a full one. `just admission-contract` proves those refusals offline. **Do not run `just admit run` as the deploy gate.** It is fine for debugging one suite, but it is invisible to the concurrency limit that keeps two gates off the same machine — which we learned by running two at once and throttling both (see `../docs/deploying.md`). The sanctioned entry point is the `stream-admission` Prefect deployment in the sibling repo [`zzstoatzz.io/my-prefect-server`](https://tangled.org/zzstoatzz.io/my-prefect-server), which runs the identical `scripts/admit run` on bare metal (heavypad) with `concurrency_limit: 1`. `.tangled/workflows/admission.yml` only triggers it. Concretely, this repo's deploy path depends on that repo for: | what | where | |---|---| | the gate runner (flow + deployment) | `my-prefect-server: flows/stream_admission.py`, `prefect.yaml` | | the machine it runs on | heavypad, via Prefect's `home-pool` (a *process* pool) | | serialisation | the deployment's `concurrency_limit: 1` (ENQUEUE) | | toolchain on that box | zig, go, just, uv, docker — go/just live in the shared nix profile | If that deployment is gone or its worker is offline, pushes will not produce receipts and nothing will be deployable. That is deliberate: no receipt, no deploy. ## required environment (docs/lessons-from-zlay.md #8, #9) - `MALLOC_ARENA_MAX=4` — the highest-leverage glibc RSS knob for thread-heavy services (per-thread arena fragmentation) - public traffic must target `--addr`; probes and Prometheus must target the private `--debug-addr` (`/healthz`, `/readyz`, and `/metrics`) - immutable image tags (git SHA); confirm what's running via the `stream_build_info{git_sha,optimize}` metric ## flags for production ``` serve --addr=:8080 --debug-addr=127.0.0.1:6060 --relay-url=https://relay1.us-east.bsky.network --upstream-slow-min-rate=50 --plc-url=https://plc.directory --data-dir=/data --backfill --backfill-workers=100 --backfill-async-flush-workers=4 --subscribe-read-log-retention-bytes=268435456 --subscribe-block-cache-bytes=67108864 --subscribe-read-batch=1024 --subscribe-slow-window=60s --subscribe-slow-min-rate=5 ``` Omitting or emptying `--debug-addr` disables the debug socket. Do not expose it through the public ingress. `/readyz` proves both configured accept loops are running; use `/status` and metrics for ingest/bootstrap health. Run `just listener-contract` before packaging to prove split routing and the disabled-debug case offline. The production host target is 16 cores, 32–64 GB RAM, and at least 3.84 TB usable mirrored SSD/NVMe. Keep RocksDB and the archive on the same RAID1 XFS filesystem: segment fsync precedes the RocksDB batch that acknowledges those bytes, and the unified filesystem is reflink-snapshotted for backup. The defaults reserve at most 8 GiB for transient backfill allocations while admitting 100 concurrent disk-backed downloads. CAR preparation is serialized, copies each complete CAR into that budget, and releases the scratch mmap before concurrent emission. This avoids partial-index deadlock while prepared repositories still emit concurrently. Four bootstrap-only workers compress detached JSS blocks and commit them in order. The pod still needs headroom for RocksDB, archive/compaction buffers, live capture, thread stacks, zstd frames, and libc; do not set the container limit equal to the backfill budget. ## promotion gate The complete gate is `../docs/semantic-parity.md`. In particular, a promotion must test individual repository durability inside a large dispatch batch, metadata read corruption, unreadable archive files, persistent live encoding failures, and transient firehose disconnects. The receipt must bind those results to the exact deployed image digest. Throughput, backup restore, client interoperability, and alerts remain necessary, but they cannot waive a blocked semantic row. ## alert rules `prometheus/rules.yml` is the in-repo source for the Prometheus alert rules; the box's Prometheus reads its own copy at `/opt/stream-experiment/rules.yml` (mounted to `/etc/prometheus/rules.yml`, listed under `rule_files` in `prometheus.yml`), so a rule change here is not live until that copy is replaced and Prometheus restarted (`docker compose restart prometheus`; the container has no `--web.enable-lifecycle`, so there is no reload endpoint, and `up -d` does not restart a container whose compose definition is unchanged). Write the box copy in place (`cat new > rules.yml`): the file is a single-file bind mount. Wired up 2026-08-29; before that the box had no rules loaded at all. Validate with `promtool check rules deploy/prometheus/rules.yml` (e.g. via `docker run --rm -v "$PWD/deploy/prometheus:/rules:ro" --entrypoint promtool prom/prometheus check rules /rules/rules.yml`). ### paging Since 2026-10-02 an `alertmanager` service in the box's compose stack (`prom/alertmanager:v0.33.0`, pinned by digest) posts `critical` and `warning` alerts to Discord; `info` alerts are dropped. `prometheus/alertmanager.yml` and `prometheus/prometheus.yml` are the in-repo sources for the box copies at `/opt/stream-experiment/`. The compose file itself lives only on the box. The webhook URL is the one zlay's Alertmanager uses (the `zlay-discord-webhook` secret in the zlay cluster's `monitoring` namespace). It is on the box as `/opt/stream-experiment/discord-webhook`, owner 65534, mode 400, and is read through `webhook_url_file`. Rotating the webhook means replacing both copies. To prove the route end to end: ```sh docker compose exec -T alertmanager amtool --alertmanager.url=http://127.0.0.1:9093 \ alert add StreamPagingTest severity=warning job=stream docker compose exec -T alertmanager wget -qO- http://127.0.0.1:9093/metrics \ | grep 'alertmanager_notifications.*discord' ``` ## dashboard gate `grafana/jetstream-upstream.json` is the byte-exact dashboard from the pinned Jetstream V2 reference. `grafana/stream.json` is the deployable artifact, generated deterministically from that source. The Go runtime row is adapted to honest process metrics, and one explicit Stream-only row exposes scheduler tickets, unsettled stage gaps, and terminal failures during bootstrap and steady state; every upstream protocol/lifecycle panel remains unchanged. Run `just dashboard-test`, then `just process-metrics-contract`, then `just dashboard-contract `. These checks are network-independent when given a saved scrape or the loopback harness and fail when a queried family is absent. They verify dashboard structure and metric presence only; they are not a promotion decision.