diff --git a/README.md b/README.md index 4e5b3c0..aa10963 100644 --- a/README.md +++ b/README.md @@ -14,7 +14,10 @@ the relay. Consumer docs, including endpoints and cursor semantics, are at `/subscribe-v2` (sequence cursors + record CBOR), with collection/DID filters, cursor replay, and zstd - **the archive**: `listSegments`, `getSegment`, `getBlock`, `planBackfill` — - gap-free recovery over sealed `jss` segments byte-compatible with upstream + gap-free recovery over sealed `jss` segments byte-compatible with upstream. + A working consumer (plan → fetch → decode → sqlite) is + [`examples/backfill-collection-sqlite.py`](examples/backfill-collection-sqlite.py); + best suited to dense collections and bulk drains (block-granular retrieval) - **whole-network bootstrap** with concurrent live capture, deterministic merge, and crash-safe resume - **Sync 1.1 commit verification** and continuing repair of repositories that diff --git a/deploy/site/llms.txt b/deploy/site/llms.txt index 17adb9f..8c4a3ed 100644 --- a/deploy/site/llms.txt +++ b/deploy/site/llms.txt @@ -34,9 +34,11 @@ - `GET /xrpc/network.bsky.jetstream.getSegment?name=` — download a sealed segment (~277 MB each; param is `name`, not `segment`). - `GET /xrpc/network.bsky.jetstream.getBlock?segment=&blockIndex=` — - one compressed block, if you don't want a whole segment. -- `GET /xrpc/network.bsky.jetstream.getZstdDictionary` — the v2 compression - dictionary. + one compressed block (a plain zstd frame — no dictionary needed), if you + don't want a whole segment. columnar layout: docs/jss-format-v1.md in the + source repo; a working python decoder ships in the repo's examples/. +- `GET /xrpc/network.bsky.jetstream.getZstdDictionary` — the v2 wire + compression dictionary (subscribe-v2 only; not used by archive blocks). - `POST /xrpc/network.bsky.jetstream.planBackfill` — plan an archive range download. @@ -54,6 +56,15 @@ - `GET /metrics` — Prometheus text format (jetstream_* and stream_* series), the raw source behind /stats. +### choosing archive vs other tools + +The archive's retrieval unit is a ~4096-event block: fetching a block with +one matching row costs the whole block. It excels at dense collections, +whole-network drains, per-DID history, and deleted/historical records — the +network's memory. For a SPARSE collection's current records (measured up to +~1,300x read amplification), prefer a collection directory such as +lightrail.microcosm.blue plus per-repo listRecords, then follow live here. + ## Sharp edges - trust model, precisely: /subscribe-v2's record CBOR is **CID-checkable**