jetstream v2 in zig stream.waow.tech
stream docs gotchas.md
21 kB
Markdown

Gotchas #

Accepted limitations and known traps: things that look like bugs but are deliberate, and behavior that is easy to misread. Modelled on upstream Jetstream V2's specs/gotchas.md.

If something here starts costing more than it saves, it belongs in docs/semantic-parity.md as a real gap instead.

Deliberate behavior that looks wrong #

--backfill-only reaches the sealed frontier, not the archive tip. The segment-download API serves sealed segments; the active segment is not part of it. A backfill client stopping short of the tip is correct — an ordinary subscriber replays sealed plus active durable rows and reaches everything. The archive contract asserts both sides of that boundary on purpose.

--max-backfill-repos caps dispatch per run, not the corpus. A restart under that flag picks up a fresh slice rather than replaying, because completed repositories are skipped and the cap counts only dispatched ones.

A skipped suite blocks admission exactly like a failure. scripts/admit records a suite that did not run rather than omitting it, and verify refuses the digest. A partial run must never read as a full one.

The lookback floor is computed over sealed segments only. With every fixture timestamp in the past, the floor clamps to the freshest sealed segment's min_seq. Startup resumes the active segment in place rather than sealing it, so the floor is not the active segment's first seq.

Progress metrics are process-local; durable counts come from the store. jetstream_backfill_completed_total counts work done by this process. To ask what is durably complete, read the status page's host aggregates. Confusing the two makes a restart look like it redid everything.

Traps in the harness, not the system #

A killed power-loss run leaves its NBD device claimed. The next run dies with nbd: nbdN already in use while /sys/block/nbdN still reports size=0 and no pid, which reads like a Stream failure and is not. Point STREAM_POWERLOSS_DEVICE at another device (nbd0..nbd3). Repeated cycles can wedge the Docker VM outright; restarting it is the repair.

The simulator grows for as long as it is up. It accumulates world state continuously — multiple GB over a long session — and fixtures assume a small world. A large one makes bootstrap outlast fixed timeouts in the harness; a very small one makes bootstrap finish before a transient gauge can be sampled. Reset it before a long run (rm -rf its data dir, restart), and note that scripts/admit runs the two simulator-backed oracles first for this reason.

Do not sample a transient to prove a phase happened. Whether jetstream_orchestrator_phase is observable at all depends on how long bootstrap takes, which depends on simulator size. Assert the cumulative transition counter instead; it cannot be missed by arriving late.

One injected fault can surface under two names. The relay/cursor store fault races to the process boundary either directly as InjectedStoreFault or wrapped by the archive append that triggered it as ArchiveAppendFailed. Both are correct and both exit non-zero. Pin the invariant — a fatal cause reaches main — not one side of the race.

A recycled Hetzner IP fails SSH as a host-key attack warning. Tear down an experiment and the next one can be handed an address a previous run used — 65.109.0.222 came back as a build host having been experiment 3's watchdog, and SSH refused with REMOTE HOST IDENTIFICATION HAS CHANGED. It is ordinary address reuse, not an interception.

This is not just a manual-SSH nuisance: it broke a real provision. The watchdog landed on a build host's recycled address, arm_watchdog.sh could not connect, and the failed-apply trap tore the run down. StrictHostKeyChecking=accept-new admits an unknown host but hard-fails a changed one, so it does not help here. provision.sh and arm_watchdog.sh now scrub the host keys of the servers they just created, which is safe precisely because those servers are seconds old.

Build the deploy artifact natively. Emulating linux/amd64 on an arm64 workstation turns an 8-minute compile into 40-plus and yields a cross-emulated approximation of the thing that ships. STREAM_BUILD_HOST points admit at a native amd64 daemon; the image is streamed back so the push uses local credentials and the build host never receives them.

Repository size is heavy-tailed and listRepos is ordered largest-first. Median 10 KB, mean 0.45 MB, p99 10 MB, max 69 MB (400 independent random cursors). The first 100,000 listRepos entries average ~10 MB — roughly 21x the network mean. Two consequences:

  • Never sample consecutive entries. They share a PDS and a creation cohort, so forty consecutive repos is one sample, not forty. Use independent random cursors.
  • Never extrapolate early crawl throughput. The first hour is the heaviest 0.5% of the network. Repos/s rises by more than an order of magnitude after the head; bytes/s is the more reliable health signal.

The pinned simulator's repositories are ~4 KiB; real ones vary by 4 orders of magnitude. Measured against bsky.network: 15, 3.4, 2.2, 15, 4.6 and 37 MB in one sample. Any failure that depends on real repository sizes — memory pressure, worker contention, bandwidth saturation — is structurally unreachable offline. The batch submit/consume deadlock survived 20 passing suites for this reason.

Reproducing that class costs about €0.03: a small cloud box crawling bsky.network showed the deadlock in 90 seconds. Stream the image over SSH (docker image save | ssh 'docker image load') since the registry needs auth, and write /etc/docker/daemon.json with explicit DNS first or BuildKit cannot resolve. Reach for this whenever a hypothesis involves anything the fixture cannot express; do not conclude from a green suite that the condition does not occur.

Open defects #

FIXED 2026-07-31 (stream d9c4f90, zat 0.3.23). Accept-Encoding: identity on the getRepo hot path. fetched_car.zig overrode the request encoding to identity, so every getRepo call pulled uncompressed CARs. Kept here because the fix itself has a trap. Measured 2026-07-30 against real PDSes: they honour gzip and return 2.5x smaller bodies (they do not honour zstd — it silently falls back to identity).

Upstream does not do this. It sets no Accept-Encoding and never sets DisableCompression, so Go's net/http transparently requests and decodes gzip, giving it 2.5x compression on this path.

The override paired with a decoding constraint. Content-Encoding is transport-level, so the decoded body is byte-identical and verification is unaffected — but Zig's response.reader() returns the raw body and does not decompress. Only response.readerDecompressing() does. The override was therefore correct as paired with that reader: deleting it alone would feed gzip bytes straight into the CAR parser and break every fetch.

The real fix is both together — accept compression and switch to readerDecompressing with a std.compress.flate.max_window_len buffer. Both body-reading sites need it: the CAR stream, and classifyErrorResponse, which matches on the error body's text and would silently fail every pattern against a gzipped body, misclassifying a known error as unknown.

The override entered in c52765b.

Throughput impact is small. Measured at concurrency 100, gzip gave 136.7 repos/s against ~120 for identity — about 1.14x — because neither the test nor production is bandwidth-bound (production runs 41 MB/s against 150+ available). The main effect of the fix is pulling 2.5x less data from other operators' PDSes across millions of requests. During that measurement a third run drew 4,337 errors against ~160 in the first two, consistent with throttling.

The fix's own trap: the transfer buffer, not just the decompress buffer. readerDecompressing(transfer_buffer, &decompress, decompress_buffer) takes two buffers. The decompress buffer needs std.compress.flate.max_window_len; the transfer buffer feeds the decompressor its compressed input, and a zero-length one does not return an error — it aborts the process on the first gzip response. The first version of this fix (c8c8c5d) passed &.{} there and would have aborted on the first PDS that honoured the request. Caught by a loopback gzip test in zat, which failed the same way; both sites now pass an 8 KiB transfer buffer.

zat had the mirror-image bug: it passed a real transfer buffer and a zero-length decompress buffer, so gzip never worked there; the // disable gzip - zig stdlib issue comment above it attributed the failure to the stdlib, but the cause was the empty buffer. Fixed in 0.3.23, which also bounds stalled transfers and replays connections closed before they answer.

The status page divides by the wrong denominator. /status reports repository progress as complete / total, where total is the enumeration count — how far listRepos has walked — not the network size. On experiment 6 at hour 57 it read:

repositories
  total          3,300,002
  complete       3,142,519 (95%)

The real network is ~5.7 M repositories (indigo's account_repo), so actual coverage was ~58%. The enumeration total grows by 100,000 every batch and never settles, so this percentage climbs toward 100% for the whole run regardless of progress. Report the raw counts, or divide by a denominator that does not move.

StreamProcessRestarted has never fired, across 35 restarts. The rule is:

changes(process_start_time_seconds{job="stream"}[5m]) > 0

The batch-boundary restart takes the process down for ~4.5 minutes of startup, during which the scrape target is gone and process_start_time_seconds has no samples at all. It reappears with a new value, so within any 5-minute window changes() typically sees a single sample and evaluates to 0. The rule cannot observe a restart that outlasts its own window.

What did fire is StreamTargetDown (up{job="stream"} == 0), at 8 sample points in 24 h. Restarts are visible in alerting, but through that rule and only partially. Widening the changes() window past the startup duration, or keying off resets() on a counter that survives, would make the intended rule work.

Fixed 2026-08-22 in deploy/prometheus/rules.yml (the in-repo source for the rules; the box copy must be refreshed by hand). The rule is now

process_start_time_seconds{job="stream"} > time() - 900

which is true from the first post-startup scrape for 15 minutes and then clears — one firing per restart regardless of scrape interval or startup length. Since a single batch-boundary restart is by design, the same file adds StreamCrashLoop (changes(process_start_time_seconds[1h]) >= 2) and StreamLiveTailStalled (time() - jetstream_livestream_last_seen_upstream_event_timestamp_seconds > 600, guarded by > 0), the latter because up == 1 never distinguished a wedged live tail from a healthy one. None of these has a notification route yet.

HttpConnectionClosing escapes to main and ends the process. First observed twice in experiment 6 (2026-07-28 09:56 and 13:24 UTC, roughly every 3.5 h), each exit immediately preceded by:

{"level":"ERROR","msg":"HttpConnectionClosing","component":"main"}

component: "main" is logging.zig renaming Zig's default log scope, which is what the runtime uses when pub fn main returns an error — so the error escaped every handler and unwound out of main. A remote closing an HTTP connection is ordinary and must never be process-fatal; that is the same invariant as the setsockopt fix, one layer up.

ingest.zig already carries a regression test for this class ("abrupt upstream disconnects reconnect instead of ending the consumer"), written after an earlier production exit, and it passes. So the live consumer loop is covered and something else is not — most likely the bootstrap-phase HTTP client (getRepo/listRepos/DID resolution), which the test does not exercise. Finding the uncovered call site is the work.

Impact is bounded because durability holds: each restart resumed and the durable counter never went backwards. Cost is ~60 s of crawl plus a firehose reconnect per occurrence, about 17 over a 60-hour run.

The interval is not stable — it accelerates as the crawl progresses. By hour 51 of experiment 6 the gap between process starts had shrunk roughly 4x, and monotonically rather than noisily (30 starts retained in the logs):

gaps (h): 4.3 3.5 3.3 3.5 3.6 3.5 2.9 ... 1.5 1.4 1.2 1.1 1.0 0.9 0.9 0.9 1.0 1.3 1.1
first half median 2.89 h  ->  second half median 1.09 h

So "every 3.5 h" is the early rate, and quoting it understates the defect on a long run. If you are estimating exposure, measure the recent gaps rather than reusing this number:

docker compose logs --no-log-prefix -t stream | grep "subscribe controls:" | cut -c1-19

(subscribe controls: is the first line of a fresh process, so counting it counts process starts — see the RestartCount caveat below for why that is not the same as Docker's restart count.)

FIXED (2026-07-31, not yet deployed): it was a keep-alive idle race on list_client. engine.zig creates one long-lived zat.XrpcClient for listRepos and zat's transport defaults to keep_alive = true. A batch takes ~40 minutes to drain, during which listRepos is never called, so the pooled connection sits idle until the relay closes it. The next page reuses that dead socket, receiveHead reads zero bytes, and HttpConnectionClosing unwinds to main. That is why the interval equals the batch-drain time exactly: the connection is idle for precisely one batch.

ade3d5d fixed the identical race on the getRepo path by disabling keep-alive; list_client never got the same treatment. listReposPage now retries up to four times on a connection that died before answering — safe by construction, since zero bytes received means nothing was applied.

The restart is not what advances the dispatch loop. The dispatch loop is paging: while (true) and walks listRepos continuously in one process; progress comes from the durable cursor, not from the restart. The restart was crash recovery.

The interval is not a timer at all — it is one failure per batch boundary. Experiment 6 ran --backfill-batch-size=100000. Discovery admits 100,000 repos, the crawl drains them, and the process dies at the boundary where the next listRepos page is fetched. So the interval is simply batch size ÷ crawl throughput:

throughput 100,000 ÷ rate matches observed
~8 repos/s (early) 3.5 h the original "every 3.5 h"
~27 repos/s (hour 51) ~62 min the "accelerated" ~1 h

Nothing is degrading. The crawl got ~3x faster, so it hits the same batch-boundary bug ~3x more often. The shrinking interval is not progressive deterioration.

This also means the restart is load-bearing by accident: each fresh process re-runs discovery and admits the next 100,000, so the run advances because Docker restarts it. jetstream_backfill_discovered_total reads exactly 100000 on a long-lived process for this reason — it is per-process and capped, not a network total.

Why this matters beyond lost crawl time. The bounded-impact argument above was measured during bootstrap, where progress is per-repo and restart-safe. A death during merge is also covered, and that is verified rather than assumed: merge.zig fsyncs dst rows before advancing the cursor, commitMergeSource writes the cursor and rev updates in one commitDurable batch, and tests/powerloss_oracle.py kills the process at after-merge-dst-flush-before-source-commit on a real NBD device, asserting no duplicate recovered events. Power loss is strictly harsher than a clean exit.

The fix. HttpConnectionClosing is raised by std.http.Reader.receiveHead when a peer sends zero bytes of headers before closing — std's own Build/WebServer.zig treats it as => return, i.e. an ordinary event. Both http.Client and http.Server include it via http.Reader.HeadError, so the error type alone does not tell you which side hung up; find the call site rather than inferring it. Zero bytes received also makes the request unambiguously safe to retry — nothing was answered — so the correct handling at a client call site is to retry once on a fresh connection instead of propagating. Any bare try ....receiveHead(...) on a pooled keep-alive connection is exposed to this race by construction.

Measuring it: do not read docker inspect ... ExitCode on a running container. It reports 0 regardless, which reads as a clean exit and is simply the absence of an exit. Use the process-start timestamps in the log, or RestartCount, and note that those two count different things — Docker's policy restarts versus actual process starts.

Divergences that were never decided #

Host parking is no longer persisted. Persistence entered in a commit whose message does not mention parking. Upstream keeps parking in the runner's memory, and the persisted version lacked upstream's rule that a park never shortens. When checking whether a divergence was a decision, git log -S on the identifier usually settles it.

Dependency lessons #

websocket.zig's client set socket timeouts through std.posix.setsockopt, which maps BADF/NOTSOCK/INVAL to unreachable. Those are reachable on a connection socket — the peer resets, or another thread closes the fd — and unreachable cannot be caught, so an upstream dropping the connection killed the process from inside the handshake. Fixed in the fork; the server side had already taken the same fix. Any dependency that treats a syscall's race-condition errno as impossible is a process-kill waiting to happen.

The pins move together. zat pins websocket.zig too, so bumping Stream's websocket pin alone yields two module copies and file exists in modules 'build' and 'build0'. Order: websocket → zat → Stream.

"state live" on /status does not mean fresh content is flowing. Check seq advance across a window and median create-rkey TID age — never time_us, and never a single sample.

relay-eval under-reads during catch-up bursts, and every deploy costs one near-0% eval window (the measurement overlaps the startup storage gate). A dedicated probe from the prod box is the tiebreak; the eval's API is documented at relay-eval.waow.tech/llms.txt.

Compaction passes cause 1–2s live-delivery lulls (measured 2026-08-17: ten >1s gaps in 10 minutes mid-pass, max 2.08s; zero above 0.4s while the compactor idles). Consumers with strict sub-5s read timeouts will misclassify these as dead connections — the relay-eval collector did exactly that until its 2026-08-17 fix.

docker logs vanish on container recreate; Prometheus on the box survives restarts and is the durable record.

tangled short-sha links need /commits/<sha> — the 404 page returns HTTP 200, so a broken link looks alive to checkers.

The static site deploys separately from the binary (just site); a green gate does not update index/stats/llms.txt.

The edge rate limiter counts your own probes. /metrics, /status, and llms.txt checks from one workstation share that IP's anonymous 240/min bucket — a verification loop can 429 itself and read as an outage. Health checks from the box itself (localhost:6060) bypass caddy entirely; prefer those in scripts.

Caddy container recreate ≠ caddy reload. Recreate (image change) severs every proxied websocket once; reload (Caddyfile change) severs nothing (stream_close_delay). Don't pay the recreate blip for a config-only change.

listSegments lags a compaction rewrite by up to minutes. The rewrite renames the new file over the old one first and refreshes the manifest after (compact/pass.zig, on_rewritten), so a reader can fetch bytes whose header checksum is newer than the checksum listSegments (and the ETag) advertise. On a live-era segment with many DIDs the gap between the tmp file's last write and the rename was ~6 minutes (segment 6720, 2026-08-29). Consumers that verify checksums should treat a mismatch as "moving", skip, and re-list later — strata's flow does exactly that — rather than retry in a tight loop.