diff --git a/docs/full-network-experiment-2026-07.md b/docs/full-network-experiment-2026-07.md index 2b99506..1c80d8a 100644 --- a/docs/full-network-experiment-2026-07.md +++ b/docs/full-network-experiment-2026-07.md @@ -444,3 +444,118 @@ Repos/s doubles while bytes/repo halves, because the crawl is working down the size distribution from the heaviest accounts. **Early throughput always understates the run.** Judge health by sustained bandwidth, not repos/s, until the crawl is past the head. + +### …and past the head, that advice inverts + +Extending the same table to hour 55 shows why the bandwidth rule has a shelf +life: + +| elapsed | repos/s | note | +|---|---|---| +| 0.5 h | 3.43 | 11.09 MB/repo | +| 1.4 h | 6.71 | 6.11 MB/repo | +| 30 h | 8.1 | still in the head | +| 45 h | 23.6 | | +| 52 h | 31.9 | | +| 55 h | **45.0** | ~1.1 MB/repo | + +Once the crawl is past the heavy accounts it becomes **request-bound, not +bytes-bound**: repos/s keeps climbing while bytes/s *falls*, because each +request now returns a small repository. A bandwidth-based projection therefore +reads **backwards** in this regime — it shows the run slowing down while it is +in fact 13x faster than at hour 1. Judge by `complete` against a known +denominator instead. + +### The denominator is not the metric sum + +`sum(stream_backfill_repos_durable)` is **enumeration progress, not the network +size.** It grows in exact +100,000 steps, one per batch, and went 132,000 → +3,200,002 over the run without ever settling. Dividing `complete` by it produced +a confident, meaningless "92.7% done" at a point that was really under half. + +The independent number, from the indigo relay's own database +(`account_repo`, across 2,530 hosts): **~5,716,765 repos**, with 5,778,477 +active accounts. Use that as the denominator. At hour 55, `complete` was +3,035,833 — **53.1%**. + +Also: the `pending` bucket froze at exactly **103,077** and never moved again. +It is inert, not in-flight, so adding it to "remaining" overstates the work. +`complete` is the only monotonic bucket. + +### The envelope moved twice, in place + +48 h as provisioned → 72 h → **120 h** (`expires_at_unix = 1785646907`, +Sun 2026-08-02 05:01 UTC), because the crawl needed roughly 24 h more when only +20 h remained. Done by writing `/etc/stream-watchdog/deadline` on the running +watchdog, which `stream-experiment-expire` re-reads every timer tick — **no +Terraform, so no watchdog replace.** Changing `expires_at_unix` alters +cloud-init and *replaces* the enforcer; a previous attempt was killed mid-replace +by a command timeout and destroyed the watchdog outright. Terraform state still +holds the old value; anything that replaces the watchdog restores the stale +baked-in deadline. `arm_watchdog.sh` only *verifies* a deadline, it does not set +one. + +### The process restarts every ~45 minutes, by design-accident + +`RestartCount` reached 32 by hour 55. The cause is neither the firehose nor OOM: +it is the **batch boundary**. With `--backfill-batch-size=100000`, the process +drains its batch, dies fetching the next `listRepos` page with +`HttpConnectionClosing` escaping to `main`, and Docker restarts it — whereupon +discovery admits the next 100,000. **The restart is load-bearing by accident: +the run advances because it keeps crashing.** + +The interval is therefore `batch size ÷ crawl rate`, which is why the originally +documented "every 3.5 h" became ~45 min: the crawl got faster, not worse. + +| crawl rate | interval | +|---|---| +| ~8 repos/s (early) | 3.5 h | +| ~45 repos/s (hour 55) | ~45 min | + +Exactly **one** real oom-kill occurred in the whole run (2026-07-29 06:06, +62.9 GB anon-rss). Every other restart has the boundary signature. Grep `dmesg` +for a literal `Out of memory: Killed process` before ever calling it OOM. +Recovery is clean every time — the active segment resumes at its last valid byte +and `complete` never regresses. + +### Resident blooms are the memory ceiling, and a recurring startup tax + +Measured at hour 55 from the sealed segment headers: + +| | | +|---|---| +| sealed segments | 4,767 | +| total blocks | 5,041,001 | +| archived events | **16,728,092,842** | +| per-block bloom at capacity 4096 | 8,409 bytes each | +| resident bloom bytes | **39.5 GiB** | +| process RSS | 44.3 GB (41.3 GiB) → **~96% is blooms** | + +`segment_writer.zig` passes `max_events_per_block` (4096) as the footer +`bloom_capacity`, so every per-block bloom marshals to 8,409 bytes however few +DIDs it holds. Upstream `f02919c` (2026-07-10) right-sizes these and reports the +same symptom — *">90% of server heap (44.3 GiB full-network)"* — with a median +per-block unique-DID count of **1–3**, making the sizing ~1000x too large. +Upstream states the fix is **not a format change**. Stream has not adopted it; +see `jss-seal-spec.md`. + +It is paid twice, because `manifest.zig` keeps that metadata resident and must +load it before crawling resumes: + +| startup latency (process start → backfill begins) | | +|---|---| +| early in the run | 2.0 min median | +| by hour 55 | **4.5 min** median | + +With a restart every ~45 min, that is **~10% of wall-clock time in startup**, +growing with the archive. It also revises the restart bug's cost upward: not +"~60 s of crawl per occurrence" but ~4.5 min. The two defects compound, and +either fix recovers most of it. + +### State at hour 55 + +`phase=1` still — bootstrap has never completed, so merge and cutover remain +unexercised at full-network scale, and `jetstream_orchestrator_phase_transitions_total` +does not yet exist as a series. 3,035,833 of ~5.72 M complete, 1.2 TB written of +2.5 TB, disk filling at 12 GB/h (108 h to full, not a constraint), spend ~€55 at +€1.01/h.