deploying — routine code changes to the live box #
How to ship a commit to the running instance at stream.waow.tech, what each
stage costs, and what to expect while it happens. This is the everyday flow;
the whole-network experiment procedure (provisioning, watchdogs, cost
envelopes) is deployment-runbook.md.
The instance is one Hetzner box (stream-cx43, 89.167.122.160, hel1) running
docker compose at /opt/stream-experiment/. The stream image is pinned in
/opt/stream-experiment/.env as STREAM_IMAGE=atcr.io/zat.dev/stream:<sha>.
where the gate runs (this repo is not self-contained) #
Admission runs on heavypad, as the stream-admission Prefect deployment
defined in the sibling repo
zzstoatzz.io/my-prefect-server
(flows/stream_admission.py, prefect.yaml). Prefect's home-pool is a
process pool, so a flow run there is a bare-metal run on a 32-core box — no
VM tax, no per-run guest setup.
That deployment is the only sanctioned entry point, because it is the only
one carrying concurrency_limit: 1. On 2026-08-09 two gates ran on the same
box at once (a spindle microVM guest and a bare-metal run) and crossed the
worker cgroup's memory.high 534,646 times, silently throttling both. Running
./scripts/admit run by hand still works and is fine for debugging a single
suite, but it is invisible to that limit — do not use it as the deploy gate.
.tangled/workflows/admission.yml therefore only triggers the deployment
(POST /deployments/<id>/create_flow_run — the endpoint that sets
Scheduled; POST /flow_runs/ would create an inert Pending run). It needs
the PREFECT_API_URL and PREFECT_API_AUTH_STRING repo secrets.
the pipeline, in order #
| stage | how | notes |
|---|---|---|
| 1. land the change | commit + push to main | ci.yml runs fmt + debug/releasesafe tests; admission below is the deploy gate |
| 2. admission | push triggers the stream-admission deployment, or trigger it by hand with sha + build_image=true |
bare metal on heavypad; serialised by concurrency_limit=1 |
| 3. publish (optional) | ./scripts/admit publish [sha] |
pins the registry digest into the receipt. Skippable: the registry has been over quota, and deployment verifies by image id instead |
| 4. cut over | flip STREAM_IMAGE in /opt/stream-experiment/.env, docker compose up -d stream |
seconds to swap |
| 5. restart recovery | (automatic) | ~4–5 min to serving + dialing upstream; see below |
Notes on the stages:
- Admission is the gate, not CI. A receipt (
receipts/<sha>.json) binds one full-suite run to one image digest;scripts/admit verify <digest>refuses anything without a fully-passing receipt. Never point.envat an image thatadmit verifyrejects. - The registry is not on the critical path.
22bf62cshipped while atcr was over quota by verifying the receipt's recordedlocal_idagainst the image id on the box.docker save | docker loadpreserves image ids, so a gate that built on heavypad can still be deployed by transferring it. - Uncommitted trees are refused (
require_clean_tree) — the image tag is the commit, by construction. - Partial runs cannot deploy.
only/skiprecord every unrun suite asskipped, andverifytreatsskippedexactly like a failure;build_image=falsewrites no receipt at all.
what to expect at restart #
- The process replays its cursor from the durable watermark; with
delete-compaction enabled, startup first rebuilds the live tombstone set.
Since
fba93d0(header-level watermark skip, upstream spec §3.4) that rebuild reads only segment headers below the watermark and costs seconds to minutes, not hours. Before that fix a restart paid a ~3h full-archive scan with the tail deaf throughout — if a restart ever takes that long again, the skip has regressed; check the rebuild progress lines in the logs. - Subscribers are dropped at the restart. zat-based consumers rotate to their next configured host (official jetstreams) and come back on their own schedule — they only return to stream when their current host drops them. Do not conclude the deploy broke a consumer because it is camping on a fallback host.
- The live tail replays the full gap on reconnect; the archive serves throughout (the ~30s of docker swap excepted).
- No consumer bounce is needed any more. The cold→hot seam wedge that
made cursor-subscribers deliver in bursts ("jerky, 0↔40/s") was fixed in
fca6cc7; the22bf62ccutover confirmed it — coral ran 12 clean samples at 20–42/s withjetstream_subscribe_cold_reads_totalfrozen, and no consumer was restarted. If burstiness returns, that counter climbing after catch-up is the tell, and it is a regression, not an expected cost.
post-deploy verification (do all four) #
curl -s https://stream.waow.tech/metrics | grep stream_build_info— the sha must be the one you shipped.upstream seqadvancing across two reads ~10s apart (an honest live-tail check;state livealone only means the process is serving).- Delivery cadence:
tools/loadgenone connection with--report-s=1for ~60s — per-second deltas, no silent seconds. Run it once live-attach and once with--rewind-s=10(cursor resume — the path that broke 2026-08-08). - Grafana (
grafana.stream.waow.tech): rate, RSS, and dropped-events panels at their usual shapes.
Then commit the ops note: what shipped, why, what was verified. The operator
journal is docs/ops-changelog.md in the relay repo.
when a stage goes wrong #
| symptom | do |
|---|---|
| admission suite fails | fix it; there is no skip lever — a skipped verdict blocks admission like a failure |
admit verify refuses at cutover |
you are deploying an image with no passing receipt; go back to stage 2 |
| restart takes hours, tail deaf | the tombstone-rebuild header skip regressed (see above); the archive still serves — decide rollback vs. wait using the rebuild progress logs |
| consumers stay on fallback hosts | expected; bounce a consumer if you need it back on stream to verify a fix |
| box unreachable mid-deploy | the old container keeps running unless compose already swapped; verify with docker ps before assuming an outage |
operator reference (moved from the root HANDOFF, 2026-08-18) #
Access. Prod box: ssh root@89.167.122.160, compose at
/opt/stream-experiment/. heavypad: ssh stoat@heavypad over Tailscale (a
hang means re-auth locally, not a dead host); non-login shells need
PATH=/home/stoat/.local/bin:/nix/var/nix/profiles/default/bin:/usr/local/bin:/usr/bin:/bin
and DOCKER_HOST=unix:///var/run/docker.sock. heavypad is on the home
LAN: no firehose probes, no bulk transfers through it — probe from the
prod box instead. Upstream clone lives at
~/github.com/bluesky-social/jetstream; keep it at the pin
(re-pin deliberately in campaigns, never chase upstream main).
Triggering the gate. Always pass explicit parameters — the stored deployment defaults are a smoke test, not a contract:
just -f ~/tangled.org/zzstoatzz.io/my-prefect-server/justfile \
prefect deployment run "stream-admission/stream-admission" \
--param sha=<sha> --param 'only=[]' --param 'skip=[]' --param build_image=true
The gate self-heals its fixtures: ensure_simulator recycles a >12h/wedged
simulator (its getRepo CARs outgrow the oracle window), and
ensure_upstream reads the pin from docs/upstream-harness.md's header.
For local iteration on oracle failures,
STREAM_UPSTREAM_REPO=~/github.com/bluesky-social/jetstream uv run python tests/differential_oracle.py gives ~90s cycles against the gate's ~25min.
Shipping the image. Preferred: registry push (needs a live
docker-credential-atcr login session on the pushing machine). Fallback,
used whenever registry auth is stale — checksummed tar relay, never an
unverified pipe (a dropped docker save stream loads "successfully" with
missing layers via content-store dedupe):
ssh stoat@heavypad '... docker save <tag> > image.tar && sha256sum image.tar'
scp heavypad:image.tar . && verify sha && scp to box && verify sha && docker load
heavypad has no ssh key to the box, so the relay goes through the
workstation. Deploy = flip STREAM_IMAGE in /opt/stream-experiment/.env,
docker compose up -d stream. Rollback = flip it back.
Restart semantics. All maintenance timers are boot-relative: every restart resets the compaction interval AND the 24h failed-repo retry timer. Schedule deploys around pending watch points, or accept moving them. Startup holds data paths at 503 for ~3–4 minutes (storage gate); the site, /status, and metrics stay up, and one relay-eval window overlapping the gate will read near-0% — cosmetic and self-explaining.
Archive keys. Fleet key: STREAM_ARCHIVE_API_KEY in the box .env
(never in git, unmetered). Consumer keys: name:key[:mbps] lines in
/opt/stream-experiment/archive-keys.conf (0600, mounted ro at /keys/)
— mint by appending a line, retune by editing the mbps, revoke by
deleting the line; the ring reloads on file mtime within a second, no
restart. See docs/configuration.md and docs/semantic-parity.md.
Admission limits. Three layers; each limit lives at the layer that can see what it bounds. Source of truth for current values, per layer:
| layer | bounds | read current values | change |
|---|---|---|---|
caddy zones (deploy/site/Caddyfile) |
request rates per client IP/subnet | grep -B2 -A6 'zone ' deploy/site/Caddyfile |
edit repo copy, just caddy (reload; severs nothing) |
key ring (/opt/stream-experiment/archive-keys.conf) |
archive bytes/s per key | curl -s https://stream.waow.tech/metrics | grep rate_bps |
edit the box file (see Archive keys below) |
| stream process | subscriber count, concurrent cold replays | grep JETSTREAM_MAX /opt/stream-experiment/.env (absent/0 = unbounded) |
set JETSTREAM_MAX_SUBSCRIBERS / JETSTREAM_MAX_COLD_READERS in the box .env, docker compose up -d stream |
Values as first deployed (2026-08-21, untested guesses — retune from metrics, not from this table):
| control | value | basis |
|---|---|---|
| ws connect attempts | 30 / 5m per /24 | reconnect loops are rare in healthy clients |
| keyed xrpc requests | 15,000 / min per IP | a 32 MB/s consumer ≈ 150 getBlock/s (~214 KB avg block); floor is highest observed legit rate + headroom |
| anonymous http | 240 / min per /24 | static site + probes; also absorbs 401 spray |
| max subscribers | unset (unbounded) | pending first load test |
| max cold readers | unset (unbounded) | pending first load test |
Retuning inputs: stream_archive_key_* series (per-key request/byte
rates), stream_subscribe_capacity_rejects_total,
stream_subscribe_cold_contended_total, caddy logs (429 counts per
zone). When a limit fires against legitimate traffic, the fix is a new
row here with a new date and basis, not a silent edit.
The caddy image is custom (stream-caddy:rl, deploy/caddy/Dockerfile:
caddy 2.10 + mholt/caddy-ratelimit) — stock caddy has no rate
limiter. Image rebuild: docker build -t stream-caddy:rl -f caddy.Dockerfile . on the box, then docker compose up -d caddy; a
recreate severs live subscriber websockets once (clients reconnect).
Caddyfile-only changes go through just caddy (reload; severs
nothing). Cap semantics and upstream contrast:
docs/semantic-parity.md.
Namespace doctrine (settled). We serve upstream's
network.bsky.jetstream.* NSIDs on purpose — the NSID names Bluesky's
schema contract, not the server. Stream-only extensions, if ever added,
get dev.zat.stream.*. Rationale: docs/product-copy-context.md.
Strata playground #
just strata publishes https://stream.waow.tech/strata/
independently of the homepage, Worker, and Stream binary. It builds the
strata/ source (run bun install there once for its pinned esbuild) and
fetches one public summary snapshot during the deploy. Caddy serves only static files, including gzip sidecars. No backend,
credential, archive sampling endpoint, or scheduled refresh is deployed.
Historical sampling stays local; the public playground offers the map, search,
profile avatars, and canvas pan/zoom. The former live-event feed has been
removed; the playground opens no WebSocket connections.
The storage comparison uses strata/storage-report.json, embedded in the
static summary at deploy time. Refresh it manually with
uv run strata/inspect-storage.py before deploying when needed; its scope and
limits are in strata-storage-discovery.md.
That bounded inspection reads file listings, headers, and indexes using the
existing archive credential locally; it does not download record bodies.
Deploying does not run the inspection. Visitors trigger no archive reads or
Microcosm API calls; UFOs integration is a contextual external link.
The homepage footer links to /strata/. The former /_strata-preview/ route
redirects there, including asset paths. Existing /strata/api endpoints still
proxy to the summary Worker for compatibility. The route definition is in
deploy/site/Caddyfile; normal updates need no Caddy reload. just strata-preview
remains an alias for just strata.
Uploads remain in site/.strata-preview-releases/<UTC timestamp> on the existing
host, selected by the atomically replaced site/strata symlink. Rollback selects
an earlier release with the same temporary-symlink/rename operation in
scripts/deploy-strata-preview. The historical release-directory and script
names are retained; the public URL and deployment output are /strata/.
Profile recognition is also embedded in the static summary.
python3 strata/refresh-identities.py refreshes the shared catalog manually, using batches of 25 handles with a 100-request ceiling, a pause between
batches, and no retry loop. Successful matches expire after seven days; missing
profiles after one day. It stops on an HTTP error, including rate limiting, and
preserves completed batches. Deploy does not refresh this catalog. Visitors
make no profile API calls; avatar images use their source's HTTP cache and a
bounded canvas drawing cache. No avatar blobs are stored on the Stream disk.
Slow-connection loading #
The public HTML embeds only source totals and recognition metadata. The full record-type catalogue is grouped by namespace and loaded once on demand when a visitor selects a source or searches. Counts and record types remain exact. The build bundles the scripts with the sibling repo's installed esbuild, and content-hashes script, stylesheet, and catalogue filenames. Caddy caches these immutable assets for one year; prior asset versions are retained during release switches so a cached page continues to work. HTML is fresh for 60 seconds, with stale-while-revalidate and stale-if-error allowances. This is HTTP caching, not a guarantee of offline availability. Detail requests time out after 15 seconds and can be retried by selecting a source or searching again.
For a local build without deploying, run python3 scripts/deploy-strata-preview --output-dir /tmp/strata-build. --overview PATH reuses a saved overview instead
of fetching the public summaries. Avatar resolution and archive inspection are
still manual; deployment adds no service or scheduled job.