diff --git a/README.md b/README.md index edb67c6..55e1362 100644 --- a/README.md +++ b/README.md @@ -8,6 +8,7 @@ two full-network [ATProto](https://atproto.com) relays on independent Hetzner Cl |---|---|---| | **firehose** | `wss://relay.waow.tech` | `wss://zlay.waow.tech` | | **jetstream** | `wss://jetstream.waow.tech/subscribe` ([sidecar](https://github.com/bluesky-social/jetstream)) | — | +| **Jetstream V2 canary** | [`wss://stream.waow.tech/subscribe`](https://stream.waow.tech/) ([status](https://stream.waow.tech/status)) | — | | **collectiondir** | [lightrail](https://tangled.org/microcosm.blue/lightrail) sidecar on same endpoint | built-in (inspired by lightrail) | | **health** | [`relay.waow.tech/xrpc/_health`](https://relay.waow.tech/xrpc/_health) | [`zlay.waow.tech/_health`](https://zlay.waow.tech/_health) | | **metrics** | [`relay-metrics.waow.tech`](https://relay-metrics.waow.tech) | [`zlay-metrics.waow.tech`](https://zlay-metrics.waow.tech) | @@ -53,6 +54,7 @@ consumes the simplified JSON firehose via [jetstream](https://github.com/bluesky # point at a different jetstream instance ./scripts/jetstream --url wss://jetstream1.us-east.bsky.network +./scripts/jetstream --url wss://stream.waow.tech ``` ## relay-eval @@ -68,6 +70,7 @@ source: [`relay-eval/`](relay-eval/) ├── indigo/ # Go relay (indigo) — justfile, deploy configs, terraform ├── zlay/ # zig relay (zlay) — justfile, deploy configs, terraform ├── relay-eval/ # firehose comparison tool — zig, hetzner, systemd +├── stream/ # public Zig Jetstream V2 canary deployment ├── shared/deploy/ # helm values shared by both deployments ├── scripts/ # uv scripts — firehose, jetstream, backfill ├── docs/ # architecture, deployment guide, backfill @@ -81,6 +84,7 @@ just indigo deploy # deploy Go relay just indigo status # check Go relay pods just zlay deploy # deploy zig relay just zlay status # check zig relay pods +just --justfile stream/justfile status # check Stream canary just --list # see all available recipes ``` diff --git a/STREAM-HANDOFF.md b/STREAM-HANDOFF.md index f64d062..52cec78 100644 --- a/STREAM-HANDOFF.md +++ b/STREAM-HANDOFF.md @@ -3,7 +3,7 @@ > for the relay engineer. written 2026-07-12 by the stream build session. > source: [zat.dev/stream](https://tangled.org/zat.dev/stream) (local: `~/tangled.org/zat.dev/stream`) -> **deployment update, 2026-07-13:** ReleaseSafe `b94eaee` is live at +> **deployment update, 2026-07-13:** ReleaseSafe `bff8139` is live at > `wss://stream.waow.tech/subscribe` on the indigo k3s cluster, capped at > 2 CPU / 2 GiB and writing to a dedicated 500 GB Hetzner volume. This is the > live-only shakedown: no `--backfill`, compaction and failed-repo retry are @@ -30,10 +30,11 @@ upstream, and real-network scale validation. - **live path**: firehose consume → verify (PLC key cache, rotation-aware) → convert → hot tail fan-out + archive append. v1 wire parity + v2. - **bootstrap lifecycle**: `bootstrap → merging → steady_state`, durable - phase file, crash-safe at every commit boundary. full listRepos crawl via - the relay, concurrent live capture into a throwaway seq space, merge with - upstream's rev-watermark filter, serving 503-gated until steady_state. -- **compaction**: record + DID tombstones, chunked passes, watermark file, + phase in the shared metadata database, crash-safe at every commit boundary. + full listRepos crawl via the relay, concurrent live capture into a throwaway + seq space, merge with upstream's rev-watermark filter, serving 503-gated + until steady_state. +- **compaction**: record + DID tombstones, chunked passes, database watermark, atomic segment rewrites (tmp+fsync+rename, block topology preserved), merge-tail pass before first serve, steady 4h loop with cap trigger. - **repo healing** (this week): @@ -46,9 +47,9 @@ upstream, and real-network scale validation. 429 parks the host) - **oracle rig**: `just oracle` — SIGABRT-crashes the real binary at 8 named lifecycle seams (`--crashpoint=`), restarts, asserts convergence to a - healthy steady state (serving ungated, seq-monotonic replay, phase file, + healthy steady state (serving ungated, seq-monotonic replay, durable phase, cleanup done, interrupted repos repaired). all 8 pass. -- **tests**: ~70 zig unit/integration tests (`zig build test` + `zig build`), +- **tests**: 77 zig unit/integration tests (`zig build test` + `zig build`), python e2e (`just e2e`), oracle (`just oracle`). all green at handoff. verified e2e on the simulator this session: full crawl → merge → merge-tail @@ -60,7 +61,9 @@ pass resyncing a repo live next to the running consumer. - `just publish-docker` in the stream repo: fresh-clone docker build → `atcr.io/zat.dev/stream:`. ReleaseSafe, x86_64-linux-gnu (glibc — do NOT switch to musl; and never ReleaseFast, zlay lesson). - single static-ish binary, no postgres, no rocksdb — just a disk volume. + single static-ish binary with RocksDB embedded through `rocksdb-zig`; no + external database service. Archive segments and `/data/meta.rocksdb` share + the persistent disk. - runtime: `MALLOC_ARENA_MAX=4` is set in the image (defense in depth if you also set it in helm values). - k8s shape: same bjw-s app-template pattern as `zlay/deploy/zlay-values.yaml` @@ -117,17 +120,16 @@ restart (the oracle proves each boundary). `--backfill`: live-only consume for a day. watch throughput, memory, verify drops. this is the cheapest possible shakedown. 2. wipe, redeploy with `--backfill` against the chosen relay. bootstrap is - 503-gated; watch `repos.log` growth and the phase transitions in logs. + 503-gated; watch RocksDB backfill counts and the phase transitions in logs. 3. once steady, enable compaction + retry intervals (defaults are fine). 4. run relay-eval / coverage comparison against the go jetstream sidecar. ## known gaps + honest risks (please read before sizing) -- **throughput is unmeasured off-laptop.** the consumer, verifier, and - archive keep up with the simulator easily; nobody has pointed this at a - real full-network firehose. the verifier resolves PLC docs synchronously - on cache miss — cold-start against the real network will be - resolution-heavy (zlay's old cold-start lesson; the LRU holds 250k keys). +- **bootstrap throughput remains unmeasured at full-network scale.** the live + consumer, verifier, and archive are running against the real full-network + firehose. the verifier resolves PLC docs synchronously on cache miss, so a + cold bootstrap may still be resolution-heavy (the LRU holds 250k keys). - **backfill engine is serial** (one download at a time; upstream uses 100 workers). fine for the simulator, not for a multi-million-repo crawl. this is the single biggest remaining engineering item and it's isolated @@ -139,14 +141,11 @@ restart (the oracle proves each boundary). after the active segment seals, not instantly. data is never lost; it's a latency divergence from upstream. fix candidates are noted in `retry.zig`'s module doc. -- **retry backoff is in-memory** — a restart retries all failed repos once - immediately. harmless at small counts; if the failed set is huge this is - a thundering herd against the relay after every deploy. -- **duplicate-row window**: a crash between an archive block flush and the - repos.log flush can re-download a repo and duplicate its `create` rows - (same did/rkey/rev twice). at-least-once semantics, consumers must - tolerate; upstream commits both atomically in pebble, we can't. window is - one listRepos page. +- **metadata durability now follows upstream's topology**: one shared + RocksDB database stores repo state/counts, retry deadlines and host parks, + relay/listRepos/merge cursors, phase, and compaction/sequence watermarks. + Synced batches follow segment fsync, so metadata cannot publish bytes that + are not durable. The legacy cursor/phase/watermark files are migration-only. - **compaction skips** upstream's DID-bloom prefilter + worker fan-out (documented in `compact/pass.zig`) — passes are correct but do more I/O than upstream on large archives. diff --git a/docs/ops-changelog.md b/docs/ops-changelog.md index abad447..cf4452e 100644 --- a/docs/ops-changelog.md +++ b/docs/ops-changelog.md @@ -71,6 +71,26 @@ legitimate recovery cannot be killed by the steady-state liveness policy. Afterward the pod had zero restarts, the public harness passed 20/20, and a 20-second sample delivered 46.8 verified events/s with stable pipeline gaps. +The canary was then advanced to `bff8139`, replacing the provisional metadata +files and in-memory backfill retry state with one embedded RocksDB database at +`/data/meta.rocksdb`. Synced WAL batches now cover relay/listRepos/merge +cursors, sequence watermarks, lifecycle phase, compaction watermark, repo +status/counts, retry deadlines, and host rate-limit parks. Archive publication +is ordered as segment fsync followed by the corresponding metadata batch; an +injected commit-failure test proves a failed batch neither publishes nor +strands the durable suffix. Existing cursor/phase/watermark files migrate only +when the database key is absent and remain on disk for rollback. + +Validation included 77 Zig tests, the complete public-protocol e2e suite, all +eight crash-oracle seams, and an intentional public pod replacement. The first +RocksDB boot imported legacy relay cursor `837308931`; after ingesting further, +the replacement opened the existing database and resumed at `837316726` +without re-import. Public Stream delivery produced 199 events in five seconds, +health remained 200, drops remained zero, and the replacement pod used about +22 MiB RSS with 465 GiB free. The colocated relay snapshot remained healthy +at 6.5 GiB/12 GiB with no broadcast backlog; both Go Jetstream and Stream +delivered real events after the rollout. + --- ## 2026-06-16 diff --git a/stream/README.md b/stream/README.md index 9601815..b937a00 100644 --- a/stream/README.md +++ b/stream/README.md @@ -11,10 +11,16 @@ cluster; no existing clients or production routes were moved to it. - v2 endpoint: `wss://stream.waow.tech/subscribe-v2` - phone field report: `https://stream.waow.tech/status` - health and metrics: `https://stream.waow.tech/healthz`, `/metrics` -- image: `atcr.io/zat.dev/stream:b94eaee`, ReleaseSafe, Zig 0.16 +- image: `atcr.io/zat.dev/stream:bff8139`, ReleaseSafe, Zig 0.16 - upstream: the real `relay.waow.tech` firehose - archive: Hetzner volume `106343283`, 500 GB, mounted at `/mnt/stream-data` on the indigo node and `/data` in the pod +- metadata: one RocksDB database at `/data/meta.rocksdb`; synced WAL batches + atomically publish archive sequence watermarks, the covered relay cursor, + lifecycle state, backfill repo status/counts, retries, and merge/compaction + cursors. Segment bytes are fsynced before the corresponding batch commits. +- migration: pre-RocksDB cursor/phase/watermark files are imported only when + their database keys are absent and are retained as rollback material - limits: 2 CPU, 2 GiB memory - mode: live-only; backfill is off, compaction and failed-repo retry intervals are both zero @@ -46,6 +52,11 @@ just --justfile stream/justfile deploy The remote build uses Zig 0.16 and ReleaseSafe, imports the immutable image directly into k3s containerd, and pins it against image GC. +RocksDB is compiled into the image through `rocksdb-zig`; there is no external +database service. The database and archive must remain on the same persistent +volume. A restart should log `metadata store opened` and `resuming from +persisted cursor`; a repeated `imported legacy relay cursor` is an anomaly. + ## rollback and teardown Fast rollback leaves the archive and DNS intact: