Gotchas #
Accepted limitations and known traps: things that look like bugs but are
deliberate, and behavior that is easy to misread. Modelled on upstream Jetstream
V2's specs/gotchas.md.
If something here starts costing more than it saves, it belongs in
docs/semantic-parity.md as a real gap instead.
Deliberate behavior that looks wrong #
--backfill-only reaches the sealed frontier, not the archive tip. The
segment-download API serves sealed segments; the active segment is not part of
it. A backfill client stopping short of the tip is correct — an ordinary
subscriber replays sealed plus active durable rows and reaches everything. The
archive contract asserts both sides of that boundary on purpose.
--max-backfill-repos caps dispatch per run, not the corpus. A restart
under that flag picks up a fresh slice rather than replaying, because
completed repositories are skipped and the cap counts only dispatched ones.
A skipped suite blocks admission exactly like a failure. scripts/admit
records a suite that did not run rather than omitting it, and verify refuses
the digest. A partial run must never read as a full one.
The lookback floor is computed over sealed segments only. With every
fixture timestamp in the past, the floor clamps to the freshest sealed
segment's min_seq. Startup resumes the active segment in place rather than
sealing it, so the floor is not the active segment's first seq.
Progress metrics are process-local; durable counts come from the store.
jetstream_backfill_completed_total counts work done by this process. To ask
what is durably complete, read the status page's host aggregates. Confusing the
two makes a restart look like it redid everything.
Traps in the harness, not the system #
A killed power-loss run leaves its NBD device claimed. The next run dies
with nbd: nbdN already in use while /sys/block/nbdN still reports size=0
and no pid, which reads like a Stream failure and is not. Point
STREAM_POWERLOSS_DEVICE at another device (nbd0..nbd3). Repeated cycles
can wedge the Docker VM outright; restarting it is the repair.
The simulator grows for as long as it is up. It accumulates world state
continuously — multiple GB over a long session — and fixtures assume a small
world. A large one makes bootstrap outlast fixed timeouts in the harness;
a very small one makes bootstrap finish before a transient gauge can be
sampled. Reset it before a long run (rm -rf its data dir, restart), and note
that scripts/admit runs the two simulator-backed oracles first for this
reason.
Do not sample a transient to prove a phase happened. Whether
jetstream_orchestrator_phase is observable at all depends on how long
bootstrap takes, which depends on simulator size. Assert the cumulative
transition counter instead; it cannot be missed by arriving late.
One injected fault can surface under two names. The relay/cursor store
fault races to the process boundary either directly as InjectedStoreFault or
wrapped by the archive append that triggered it as ArchiveAppendFailed. Both
are correct and both exit non-zero. Pin the invariant — a fatal cause reaches
main — not one side of the race.
A recycled Hetzner IP fails SSH as a host-key attack warning. Tear down an
experiment and the next one can be handed an address a previous run used —
65.109.0.222 came back as a build host having been experiment 3's watchdog,
and SSH refused with REMOTE HOST IDENTIFICATION HAS CHANGED. It is ordinary
address reuse, not an interception.
This is not just a manual-SSH nuisance: it broke a real provision. The watchdog
landed on a build host's recycled address, arm_watchdog.sh could not connect,
and the failed-apply trap tore the run down. StrictHostKeyChecking=accept-new
admits an unknown host but hard-fails a changed one, so it does not help
here. provision.sh and arm_watchdog.sh now scrub the host keys of the
servers they just created, which is safe precisely because those servers are
seconds old.
Build the deploy artifact natively. Emulating linux/amd64 on an arm64
workstation turns an 8-minute compile into 40-plus and yields a cross-emulated
approximation of the thing that ships. STREAM_BUILD_HOST points admit at a
native amd64 daemon; the image is streamed back so the push uses local
credentials and the build host never receives them.
Repository size is heavy-tailed and listRepos is ordered largest-first.
Median 10 KB, mean 0.45 MB, p99 10 MB, max 69 MB (400 independent random
cursors). The first 100,000 listRepos entries average ~10 MB — roughly 21x
the network mean. Two consequences:
- Never sample consecutive entries. They share a PDS and a creation cohort, so forty consecutive repos is one sample, not forty. Use independent random cursors.
- Never extrapolate early crawl throughput. The first hour is the heaviest 0.5% of the network. Repos/s rises by more than an order of magnitude after the head; bytes/s is the more reliable health signal.
The pinned simulator's repositories are ~4 KiB; real ones vary by 4 orders of magnitude. Measured against bsky.network: 15, 3.4, 2.2, 15, 4.6 and 37 MB in one sample. Any failure that depends on real repository sizes — memory pressure, worker contention, bandwidth saturation — is structurally unreachable offline. The batch submit/consume deadlock survived 20 passing suites for this reason.
Reproducing that class costs about €0.03: a small cloud box crawling
bsky.network showed the deadlock in 90 seconds. Stream the image over SSH
(docker image save | ssh 'docker image load') since the registry needs auth,
and write /etc/docker/daemon.json with explicit DNS first or BuildKit cannot
resolve. Reach for this whenever a hypothesis involves anything the fixture
cannot express; do not conclude from a green suite that the condition does not
occur.
Open defects #
FIXED 2026-07-31 (stream d9c4f90, zat 0.3.23). Accept-Encoding: identity
on the getRepo hot path. fetched_car.zig overrode the request encoding to
identity, so every getRepo call pulled uncompressed CARs. Kept here because
the fix itself has a trap.
Measured 2026-07-30 against real PDSes: they honour gzip and return 2.5x
smaller bodies (they do not honour zstd — it silently falls back to
identity).
Upstream does not do this. It sets no Accept-Encoding and never sets
DisableCompression, so Go's net/http transparently requests and decodes
gzip, giving it 2.5x compression on this path.
The override paired with a decoding constraint. Content-Encoding is
transport-level, so the decoded body is byte-identical and verification is
unaffected — but Zig's response.reader() returns the raw body and does not
decompress. Only response.readerDecompressing() does. The override was
therefore correct as paired with that reader: deleting it alone would feed gzip
bytes straight into the CAR parser and break every fetch.
The real fix is both together — accept compression and switch to
readerDecompressing with a std.compress.flate.max_window_len buffer. Both
body-reading sites need it: the CAR stream, and classifyErrorResponse, which
matches on the error body's text and would silently fail every pattern
against a gzipped body, misclassifying a known error as unknown.
The override entered in c52765b.
Throughput impact is small. Measured at concurrency 100, gzip gave 136.7 repos/s against ~120 for identity — about 1.14x — because neither the test nor production is bandwidth-bound (production runs 41 MB/s against 150+ available). The main effect of the fix is pulling 2.5x less data from other operators' PDSes across millions of requests. During that measurement a third run drew 4,337 errors against ~160 in the first two, consistent with throttling.
The fix's own trap: the transfer buffer, not just the decompress buffer.
readerDecompressing(transfer_buffer, &decompress, decompress_buffer) takes
two buffers. The decompress buffer needs
std.compress.flate.max_window_len; the transfer buffer feeds the
decompressor its compressed input, and a zero-length one does not return an
error — it aborts the process on the first gzip response. The first version
of this fix (c8c8c5d) passed &.{} there and would have aborted on the first
PDS that honoured the request. Caught by a loopback gzip test in zat, which
failed the same way; both sites now pass an 8 KiB transfer buffer.
zat had the mirror-image bug: it passed a real transfer buffer and a
zero-length decompress buffer, so gzip never worked there; the
// disable gzip - zig stdlib issue comment above it attributed the failure to
the stdlib, but the cause was the empty buffer. Fixed in 0.3.23, which also
bounds stalled transfers and replays connections closed before they answer.
The status page divides by the wrong denominator. /status reports
repository progress as complete / total, where total is the enumeration
count — how far listRepos has walked — not the network size. On experiment 6
at hour 57 it read:
repositories
total 3,300,002
complete 3,142,519 (95%)
The real network is ~5.7 M repositories (indigo's account_repo), so actual
coverage was ~58%. The enumeration total grows by 100,000 every batch and never
settles, so this percentage climbs toward 100% for the whole run regardless of
progress. Report the raw counts, or divide by a denominator that does not move.
StreamProcessRestarted has never fired, across 35 restarts. The rule is:
changes(process_start_time_seconds{job="stream"}[5m]) > 0
The batch-boundary restart takes the process down for ~4.5 minutes of startup,
during which the scrape target is gone and process_start_time_seconds has no
samples at all. It reappears with a new value, so within any 5-minute window
changes() typically sees a single sample and evaluates to 0. The rule cannot
observe a restart that outlasts its own window.
What did fire is StreamTargetDown (up{job="stream"} == 0), at 8 sample
points in 24 h. Restarts are visible in alerting, but through that rule and
only partially. Widening the changes() window past the startup duration, or keying
off resets() on a counter that survives, would make the intended rule work.
Fixed 2026-08-22 in deploy/prometheus/rules.yml (the in-repo source for the
rules; the box copy must be refreshed by hand). The rule is now
process_start_time_seconds{job="stream"} > time() - 900
which is true from the first post-startup scrape for 15 minutes and then
clears — one firing per restart regardless of scrape interval or startup
length. Since a single batch-boundary restart is by design, the same file adds
StreamCrashLoop (changes(process_start_time_seconds[1h]) >= 2) and
StreamLiveTailStalled (time() - jetstream_livestream_last_seen_upstream_event_timestamp_seconds > 600,
guarded by > 0), the latter because up == 1 never distinguished a wedged
live tail from a healthy one. None of these has a notification route yet.
HttpConnectionClosing escapes to main and ends the process. First
observed twice in experiment 6 (2026-07-28 09:56 and 13:24 UTC, roughly every
3.5 h), each exit immediately preceded by:
{"level":"ERROR","msg":"HttpConnectionClosing","component":"main"}
component: "main" is logging.zig renaming Zig's default log scope, which
is what the runtime uses when pub fn main returns an error — so the error
escaped every handler and unwound out of main. A remote closing an HTTP
connection is ordinary and must never be process-fatal; that is the same
invariant as the setsockopt fix, one layer up.
ingest.zig already carries a regression test for this class ("abrupt upstream
disconnects reconnect instead of ending the consumer"), written after an
earlier production exit, and it passes. So the live consumer loop is covered
and something else is not — most likely the bootstrap-phase HTTP client
(getRepo/listRepos/DID resolution), which the test does not exercise.
Finding the uncovered call site is the work.
Impact is bounded because durability holds: each restart resumed and the durable counter never went backwards. Cost is ~60 s of crawl plus a firehose reconnect per occurrence, about 17 over a 60-hour run.
The interval is not stable — it accelerates as the crawl progresses. By hour 51 of experiment 6 the gap between process starts had shrunk roughly 4x, and monotonically rather than noisily (30 starts retained in the logs):
gaps (h): 4.3 3.5 3.3 3.5 3.6 3.5 2.9 ... 1.5 1.4 1.2 1.1 1.0 0.9 0.9 0.9 1.0 1.3 1.1
first half median 2.89 h -> second half median 1.09 h
So "every 3.5 h" is the early rate, and quoting it understates the defect on a long run. If you are estimating exposure, measure the recent gaps rather than reusing this number:
docker compose logs --no-log-prefix -t stream | grep "subscribe controls:" | cut -c1-19
(subscribe controls: is the first line of a fresh process, so counting it
counts process starts — see the RestartCount caveat below for why that is
not the same as Docker's restart count.)
FIXED (2026-07-31, not yet deployed): it was a keep-alive idle race on
list_client. engine.zig creates one long-lived zat.XrpcClient for
listRepos and zat's transport defaults to keep_alive = true. A batch takes
~40 minutes to drain, during which listRepos is never called, so the pooled
connection sits idle until the relay closes it. The next page reuses that dead
socket, receiveHead reads zero bytes, and HttpConnectionClosing unwinds to
main. That is why the interval equals the batch-drain time exactly: the
connection is idle for precisely one batch.
ade3d5d fixed the identical race on the getRepo path by disabling keep-alive;
list_client never got the same treatment. listReposPage now retries up to
four times on a connection that died before answering — safe by construction,
since zero bytes received means nothing was applied.
The restart is not what advances the dispatch loop. The dispatch loop is
paging: while (true) and walks listRepos continuously in one process;
progress comes from the durable cursor, not from the restart. The restart was
crash recovery.
The interval is not a timer at all — it is one failure per batch boundary.
Experiment 6 ran --backfill-batch-size=100000. Discovery admits 100,000 repos,
the crawl drains them, and the process dies at the boundary where the next
listRepos page is fetched. So the interval is simply batch size ÷ crawl
throughput:
| throughput | 100,000 ÷ rate | matches observed |
|---|---|---|
| ~8 repos/s (early) | 3.5 h | the original "every 3.5 h" |
| ~27 repos/s (hour 51) | ~62 min | the "accelerated" ~1 h |
Nothing is degrading. The crawl got ~3x faster, so it hits the same batch-boundary bug ~3x more often. The shrinking interval is not progressive deterioration.
This also means the restart is load-bearing by accident: each fresh process
re-runs discovery and admits the next 100,000, so the run advances because
Docker restarts it. jetstream_backfill_discovered_total reads exactly 100000
on a long-lived process for this reason — it is per-process and capped, not a
network total.
Why this matters beyond lost crawl time. The bounded-impact argument above
was measured during bootstrap, where progress is per-repo and restart-safe. A
death during merge is also covered, and that is verified rather than assumed:
merge.zig fsyncs dst rows before advancing the cursor, commitMergeSource
writes the cursor and rev updates in one commitDurable batch, and
tests/powerloss_oracle.py kills the process at
after-merge-dst-flush-before-source-commit on a real NBD device, asserting no
duplicate recovered events. Power loss is strictly harsher than a clean exit.
The fix. HttpConnectionClosing is raised by std.http.Reader.receiveHead
when a peer sends zero bytes of headers before closing — std's own
Build/WebServer.zig treats it as => return, i.e. an ordinary event. Both
http.Client and http.Server include it via http.Reader.HeadError, so the
error type alone does not tell you which side hung up; find the call site rather
than inferring it. Zero bytes received also makes the request unambiguously
safe to retry — nothing was answered — so the correct handling at a client call
site is to retry once on a fresh connection instead of propagating. Any bare
try ....receiveHead(...) on a pooled keep-alive connection is exposed to this
race by construction.
Measuring it: do not read docker inspect ... ExitCode on a running
container. It reports 0 regardless, which reads as a clean exit and is simply
the absence of an exit. Use the process-start timestamps in the log, or
RestartCount, and note that those two count different things — Docker's
policy restarts versus actual process starts.
Divergences that were never decided #
Host parking is no longer persisted. Persistence entered in a commit whose
message does not mention parking. Upstream keeps parking in the runner's
memory, and the persisted version lacked upstream's rule that a park never
shortens. When checking whether a divergence was a decision, git log -S on
the identifier usually settles it.
Dependency lessons #
websocket.zig's client set socket timeouts through std.posix.setsockopt,
which maps BADF/NOTSOCK/INVAL to unreachable. Those are reachable on
a connection socket — the peer resets, or another thread closes the fd — and
unreachable cannot be caught, so an upstream dropping the connection killed
the process from inside the handshake. Fixed in the fork; the server side had
already taken the same fix. Any dependency that treats a syscall's
race-condition errno as impossible is a process-kill waiting to happen.
The pins move together. zat pins websocket.zig too, so bumping Stream's
websocket pin alone yields two module copies and
file exists in modules 'build' and 'build0'. Order: websocket → zat → Stream.
"state live" on /status does not mean fresh content is flowing. Check
seq advance across a window and median create-rkey TID age — never
time_us, and never a single sample.
relay-eval under-reads during catch-up bursts, and every deploy costs one near-0% eval window (the measurement overlaps the startup storage gate). A dedicated probe from the prod box is the tiebreak; the eval's API is documented at relay-eval.waow.tech/llms.txt.
Compaction passes cause 1–2s live-delivery lulls (measured 2026-08-17: ten >1s gaps in 10 minutes mid-pass, max 2.08s; zero above 0.4s while the compactor idles). Consumers with strict sub-5s read timeouts will misclassify these as dead connections — the relay-eval collector did exactly that until its 2026-08-17 fix.
docker logs vanish on container recreate; Prometheus on the box
survives restarts and is the durable record.
tangled short-sha links need /commits/<sha> — the 404 page returns
HTTP 200, so a broken link looks alive to checkers.
The static site deploys separately from the binary (just site); a
green gate does not update index/stats/llms.txt.
The edge rate limiter counts your own probes. /metrics, /status, and llms.txt checks from one workstation share that IP's anonymous 240/min bucket — a verification loop can 429 itself and read as an outage. Health checks from the box itself (localhost:6060) bypass caddy entirely; prefer those in scripts.
Caddy container recreate ≠ caddy reload. Recreate (image change)
severs every proxied websocket once; reload (Caddyfile change) severs
nothing (stream_close_delay). Don't pay the recreate blip for a
config-only change.
listSegments lags a compaction rewrite by up to minutes. The rewrite
renames the new file over the old one first and refreshes the manifest after
(compact/pass.zig, on_rewritten), so a reader can fetch bytes whose header
checksum is newer than the checksum listSegments (and the ETag) advertise.
On a live-era segment with many DIDs the gap between the tmp file's last write
and the rename was ~6 minutes (segment 6720, 2026-08-29). Consumers that verify
checksums should treat a mismatch as "moving", skip, and re-list later —
strata's flow does exactly that — rather than retry in a tight loop.