deploying stream #
This file is the deploy procedure. It deliberately does not track which
experiment is running — that state belongs in
../docs/full-network-experiment-2026-07.md,
which is maintained as runs happen. An earlier revision asserted that "no
replacement experiment should be provisioned" while experiment 6 was in fact
running on an admitted artifact, which is exactly the rot that got deployment
state removed from the top-level README.
The retired live-only canary no longer runs on the Indigo relay node.
The long-term production target remains a separate host with redundant local storage; passing a cloud-volume experiment is evidence for sizing and throughput, not approval to colocate Stream with Indigo.
The previously deployed e1926f3 image is historical evidence only. It was
not the artifact covered by the earlier admission receipt and it failed during
the experiment. Do not reuse its digest as a promotion candidate. That gap —
an image no receipt covered, with nothing to refuse it — is what
scripts/admit now closes.
just package # ReleaseSafe x86_64-linux-gnu binary in zig-out/bin/
just admit run # clean tree -> every suite -> exact image -> one receipt
just admit publish # push that image, pin the receipt to its registry digest
just admit verify sha256:... # deployment gate; non-zero unless admitted
Nothing may be deployed whose digest just admit verify does not accept. A
receipt covers exactly one image digest, never a tag, and a suite recorded as
skipped blocks admission just as a failure does — a partial run must never
read as a full one. just admission-contract proves those refusals offline.
Do not run just admit run as the deploy gate. It is fine for debugging
one suite, but it is invisible to the concurrency limit that keeps two gates
off the same machine — which we learned by running two at once and throttling
both (see ../docs/deploying.md). The sanctioned entry point is the
stream-admission Prefect deployment in the sibling repo
zzstoatzz.io/my-prefect-server,
which runs the identical scripts/admit run on bare metal (heavypad) with
concurrency_limit: 1. .tangled/workflows/admission.yml only triggers it.
Concretely, this repo's deploy path depends on that repo for:
| what | where |
|---|---|
| the gate runner (flow + deployment) | my-prefect-server: flows/stream_admission.py, prefect.yaml |
| the machine it runs on | heavypad, via Prefect's home-pool (a process pool) |
| serialisation | the deployment's concurrency_limit: 1 (ENQUEUE) |
| toolchain on that box | zig, go, just, uv, docker — go/just live in the shared nix profile |
If that deployment is gone or its worker is offline, pushes will not produce receipts and nothing will be deployable. That is deliberate: no receipt, no deploy.
required environment (docs/lessons-from-zlay.md #8, #9) #
MALLOC_ARENA_MAX=4— the highest-leverage glibc RSS knob for thread-heavy services (per-thread arena fragmentation)- public traffic must target
--addr; probes and Prometheus must target the private--debug-addr(/healthz,/readyz, and/metrics) - immutable image tags (git SHA); confirm what's running via the
stream_build_info{git_sha,optimize}metric
flags for production #
serve
--addr=:8080
--debug-addr=127.0.0.1:6060
--relay-url=https://relay1.us-east.bsky.network
--upstream-slow-min-rate=50
--plc-url=https://plc.directory
--data-dir=/data
--backfill
--backfill-workers=100
--backfill-async-flush-workers=4
--subscribe-read-log-retention-bytes=268435456
--subscribe-block-cache-bytes=67108864
--subscribe-read-batch=1024
--subscribe-slow-window=60s
--subscribe-slow-min-rate=5
Omitting or emptying --debug-addr disables the debug socket. Do not expose
it through the public ingress. /readyz proves both configured accept loops
are running; use /status and metrics for ingest/bootstrap health. Run
just listener-contract before packaging to prove split routing and the
disabled-debug case offline.
The production host target is 16 cores, 32–64 GB RAM, and at least 3.84 TB usable mirrored SSD/NVMe. Keep RocksDB and the archive on the same RAID1 XFS filesystem: segment fsync precedes the RocksDB batch that acknowledges those bytes, and the unified filesystem is reflink-snapshotted for backup.
The defaults reserve at most 8 GiB for transient backfill allocations while admitting 100 concurrent disk-backed downloads. CAR preparation is serialized, copies each complete CAR into that budget, and releases the scratch mmap before concurrent emission. This avoids partial-index deadlock while prepared repositories still emit concurrently. Four bootstrap-only workers compress detached JSS blocks and commit them in order. The pod still needs headroom for RocksDB, archive/compaction buffers, live capture, thread stacks, zstd frames, and libc; do not set the container limit equal to the backfill budget.
promotion gate #
The complete gate is ../docs/semantic-parity.md. In particular, a promotion
must test individual repository durability inside a large dispatch batch,
metadata read corruption, unreadable archive files, persistent live encoding
failures, and transient firehose disconnects. The receipt must bind those
results to the exact deployed image digest. Throughput, backup restore, client
interoperability, and alerts remain necessary, but they cannot waive a blocked
semantic row.
alert rules #
prometheus/rules.yml is the in-repo source for the Prometheus alert rules;
the box's Prometheus reads its own copy at /opt/stream-experiment/rules.yml
(mounted to /etc/prometheus/rules.yml, listed under rule_files in
prometheus.yml), so a rule change here is not live until that copy is
replaced and Prometheus restarted (docker compose restart prometheus; the
container has no --web.enable-lifecycle, so there is no reload endpoint, and
up -d does not restart a container whose compose definition is unchanged).
Write the box copy in place (cat new > rules.yml): the file is a single-file
bind mount. Wired up 2026-08-29; before that the box had no rules loaded at
all. Validate with promtool check rules deploy/prometheus/rules.yml
(e.g. via docker run --rm -v "$PWD/deploy/prometheus:/rules:ro" --entrypoint promtool prom/prometheus check rules /rules/rules.yml).
paging #
Since 2026-10-02 an alertmanager service in the box's compose stack
(prom/alertmanager:v0.33.0, pinned by digest) posts critical and warning
alerts to Discord; info alerts are dropped. prometheus/alertmanager.yml
and prometheus/prometheus.yml are the in-repo sources for the box copies at
/opt/stream-experiment/. The compose file itself lives only on the box.
The webhook URL is the one zlay's Alertmanager uses (the
zlay-discord-webhook secret in the zlay cluster's monitoring namespace).
It is on the box as /opt/stream-experiment/discord-webhook, owner 65534,
mode 400, and is read through webhook_url_file. Rotating the webhook means
replacing both copies. To prove the route end to end:
docker compose exec -T alertmanager amtool --alertmanager.url=http://127.0.0.1:9093 \
alert add StreamPagingTest severity=warning job=stream
docker compose exec -T alertmanager wget -qO- http://127.0.0.1:9093/metrics \
| grep 'alertmanager_notifications.*discord'
dashboard gate #
grafana/jetstream-upstream.json is the byte-exact dashboard from the pinned
Jetstream V2 reference. grafana/stream.json is the deployable artifact,
generated deterministically from that source. The Go runtime row is adapted to
honest process metrics, and one explicit Stream-only row exposes scheduler
tickets, unsettled stage gaps, and terminal failures during bootstrap and
steady state; every upstream protocol/lifecycle panel remains unchanged. Run
just dashboard-test, then
just process-metrics-contract, then
just dashboard-contract <metrics-url-or-saved-scrape>. These checks are
network-independent when given a saved scrape or the loopback harness and fail
when a queried family is absent. They verify dashboard structure and metric
presence only; they are not a promotion decision.