From cf3edbce3d0da1fa22372395e3ecdf6ee2bf84a7 Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Sat, 26 Sep 2026 00:25:33 -0500 Subject: [PATCH] docs: bring Strata's discovery notes over with its source strata-discovery.md (2026-09-08 research pass on Strata's purpose) and strata-storage-discovery.md (2026-09-09 archive storage measurement) were untracked in the strata repo. The storage note documents strata/inspect-storage.py, which now lives here; its script path is updated and deploying.md links it. The rejected reading-app study stays in the strata repo's spikes/. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Ray8ErT7urWdq3aT9Ka87z --- docs/deploying.md | 3 +- docs/strata-discovery.md | 302 +++++++++++++++++++++++++++++++ docs/strata-storage-discovery.md | 77 ++++++++ 3 files changed, 381 insertions(+), 1 deletion(-) create mode 100644 docs/strata-discovery.md create mode 100644 docs/strata-storage-discovery.md diff --git a/docs/deploying.md b/docs/deploying.md index d1536b5..f9b738c 100644 --- a/docs/deploying.md +++ b/docs/deploying.md @@ -210,7 +210,8 @@ removed; the playground opens no WebSocket connections. The storage comparison uses `strata/storage-report.json`, embedded in the static summary at deploy time. Refresh it manually with -`uv run strata/inspect-storage.py` before deploying when needed. +`uv run strata/inspect-storage.py` before deploying when needed; its scope and +limits are in [strata-storage-discovery.md](strata-storage-discovery.md). That bounded inspection reads file listings, headers, and indexes using the existing archive credential locally; it does not download record bodies. Deploying does not run the inspection. Visitors trigger no archive reads or diff --git a/docs/strata-discovery.md b/docs/strata-discovery.md new file mode 100644 index 0000000..066bb63 --- /dev/null +++ b/docs/strata-discovery.md @@ -0,0 +1,302 @@ +# discovery before representation + +Research pass, 2026-09-08. This is context for deciding Strata's purpose, +not an approved experience or implementation plan. Operational measurements +below retain their original dates; this pass did not inspect the production +machine or establish current drain completeness. + +## what archival Jetstream enables + +Upstream's design describes a compressed network archive and live stream for +building AppViews and doing network analysis. Its explicit non-goals include +arbitrary queries and point lookups: the archive is a replay cache for range +scans. The local upstream reference is `bluesky-social/jetstream` at `289b032` +(2026-08-13), `docs/README.md` sections 1–2, not a claim about latest upstream. + +The important change for a builder is being able to choose what to consume +after capture began. They can backfill a view, rebuild it, or catch up after +being offline, then continue consuming live. The SDK owns the replay-to-live +seam; the application owns durable progress, idempotency, and folding markers. + +Stream's product context records the architectural motivation as: + + origin PDS network -> durable archive -> ephemeral application view + +Rebuilding an application need not mean recrawling all origin PDSes. Bobbin's +hydration from its Hydrant twin is an attributed example of this architecture, +not evidence that Bobbin consumes Stream. + +## what “history” can honestly mean + +- Bootstrap retrieves repository state available when fetched. It cannot + reconstruct every mutation preceding capture. +- Compaction removes superseded and deleted materialization rows. Markers + survive. Available replay is not an immutable recording of every old value. +- Witnessed time is the instance's observation time, distinct from record + creation. Backfill, catch-up, resync, and timestamp imports matter to its + interpretation. Archive order is not a network-wide chronology of creation. +- Coverage belongs to an instance, its enumeration and capture, and what was + retrievable. Counts of namespaces are not verified counts of independent + apps; counts of event rows are not counts of distinct current records. +- Sequence cursors belong to one instance. Cross-host recovery is approximate + and carries different guarantees from replay within one archive. + +Some older notes use “full event history” and “whole network” more broadly. +The owning compaction and client contracts qualify those phrases. + +## uses with evidence + +| use | evidence | qualification | +| --- | --- | --- | +| rebuild a filtered view, then follow live | SDK unified subscribe contract | application must fold markers and persist progress | +| replay a particular broadcast's chat | Streamplace example: 17 messages from 633 blocks | retained records only; archive read amplification is visible even in a tiny result | +| analyze or materialize a corpus | network-backfill note records a Bluesky archive-to-ClickHouse drain around 17h | dated, attributed measurement; not a current Stream benchmark | +| inspect archive composition cheaply | Strata's footer reader: about 1.3 KB in two requests per sampled segment | counts reveal composition, not content or historical user activity | +| query the long tail interactively | Strata namespace-partitioned Parquet and DuckDB proof | derived dataset; coverage and reconciliation are separate responsibilities | + +The August 17 retrieval comparison is particularly useful: one sparse +collection returned 17,537 archive rows for 497 MB versus 13,145 current rows +for 3.6 MB through directory plus PDS reads. These answer different questions. +The archive's benefit includes one retrieval source and retained events; it +does not guarantee the cheapest way to fetch any small collection. + +## operational lessons relevant to a playground + +1. Query cost follows blocks touched, not result size. Sparse scattered rows + can be expensive; estimates and bounded reads must precede bulk work. +2. Bootstrap layout and mature live layout differ. Per-DID locality degrades + as new events arrive interleaved; fresh-archive demos can mislead. +3. Compaction changes old files and counts. Strata's footer summaries reconcile + checksums, but its drain documentation still leaves selective re-drain open. +4. Receiving events does not establish freshness. During catch-up the service + can emit old content rapidly. Process uptime is not archive age or coverage. +5. Shared serving resources are finite. The August 17 compaction measurement + found disk saturation; Strata's drain subsequently gained CPU/memory caps. +6. The shelved backup proposal is evidence of a limit, not unfinished scope: + restore work grows with the archive while upstream replay retention is finite. + +## what Strata has, and what it has not established + +Original discussion: replay-use heatmap -> inspect the archive externally -> +progressive exploration of the long tail. The explicit constraint was to avoid +adding in-process telemetry purely for the playground. + +Current Strata summarizes storage slices, and separately queries extracted +event metadata. The extracted rows omit record bodies. Neither the heatmap nor +the Parquet path currently demonstrates an application's replay-to-live seam. +The first idea—visualizing what other consumers replay—requires demand data +different from archive composition. Footer counts cannot answer it. + +Read-only checks on September 8 found `truncated: true` on both `fm.teal` and +`sh.tangled` partition listings. The Worker returns one listing page and the UI +ignores truncation. Thus “every record an app has ever written” is unsupported +by both the available-history semantics and this concrete retrieval limit. + +## discovery across the other projects + +The first pass overweighted archive mechanics. Reading Typeahead, pub-search, +and plyr.fm changes the working interpretation: the archive is valuable as +material from which someone can construct a useful view of the network. +Those projects make that concrete in three different ways. + +### Typeahead: who becomes findable + +Typeahead is community actor search, including identities outside Bluesky. +Its current ingester consumes `com.atproto.sync.subscribeRepos` directly; +`services/src/ingest.zig` applies a collection allowlist locally. It is not +currently an archival Jetstream consumer. + +The August birds.place incident is a particularly useful case. An account +writing only `place.birds.*` records was invisible to the previous allowlist, +and identity resolution also failed. The current allowlist includes that +namespace. The broader lexicon-driven discovery proposal remains a design, +not something to count as shipped. Capturing a person's records, recognizing +their identity, extracting a profile, and making them searchable are separate +steps. A person without a Bluesky profile is still a network participant. + +The September 5 recovery retrospective reports 377 restored active accounts +and checks their search visibility after synchronization. That is a meaningful +outcome for recovery; a count of replayed frames alone would not establish it. +The report does not establish lossless historical recovery. + +Typeahead also separates costly index construction from serving: durable actor +data becomes an independently built search artifact, with recent changes in +an overlay. Its application index snapshot is different from a Jetstream +archive snapshot. Both can be called “snapshot,” but they serve different jobs. + +Read: `docs/discovery-and-identity.md`, `docs/collection-discovery.md`, +`docs/architecture.md`, `docs/retros/2026-09-05-ingester-recovery-closeout.md`, +`docs/retros/2026-05-23-standard-site-backfill.md`, current ingest allowlist, +README, AGENTS, and recent commit history. Live actor search for `zzstoatzz.io` +returned nate and two other matching handles during this pass. + +### pub-search: a corpus across publishing applications + +Pub-search provides the strongest direct example of archival Jetstream +replacing bespoke ingestion infrastructure. The August cutover adopted Stream; +the following unified-client change uses archive replay followed by live +events, and the separate ingester was decommissioned. The current consumer is +`backend/src/ingest/jetstream.zig`, with nine selected collections and a +persisted sequence. The README still describes the retired ingester, so the +cutover document, implementation, and commit history take precedence. + +Its product is one long-form corpus across publishing platforms. Several +platforms share `site.standard.document`; namespace is not application +identity. Even within that schema, extracting readable content depends on +platform conventions, including structured Leaflet blocks versus flat text. +Deduplication also changes the meaning of volume: the August cleanup removed +27,076 stale republish rows without removing distinct logical documents. + +The July corpus audit distinguishes genuinely missing writing, valid dedupe, +metadata-only records, excluded sources, and unreachable PDSes. Its 15,174 +genuine misses were primarily a historical backfill gap. Those are dated +findings, not a current completeness assessment. The lesson for Strata is +that “present in the archive,” “indexed,” and “available to discover” are +different statements. Pub-search also honors discovery preferences in its +discovery surfaces while allowing direct retrieval of a known public URI. + +The Atlas already demonstrates progressive exploration of network content: +regions, narrower clusters, and individual writing. Its labels derive from +semantic neighborhoods, not collection names. The two-dimensional layout is +a display projection; clustering uses a separate higher-dimensional +embedding. The live map displayed 69,013 documents during this pass. That +number describes this displayed dataset, not total network writing. + +Its agent surfaces make the same point without a map: curated operations +return an author's writing, related posts across platforms, and recommendations +with snippets. They package something a reader can understand and act on. + +Read: `docs/jetstream-cutover.md`, `docs/reconciliation.md`, +`docs/content-extraction.md`, `docs/corpus-audit-2026-07-20.md`, +`docs/atlas.md`, `docs/agent-surfaces.md`, the architecture portion of +`docs/snapshot-pipeline.md`, current consumer code, README, and cutover/recent +commit history. Older counts and API claims in overview documents require +care: these documents contain revisions and some contradictory older text. + +### plyr.fm: records become something you can listen to + +Plyr's public data model includes tracks, likes, timed comments, playlists, +and profiles. It also writes listening records under `fm.teal.*`. One +application spans multiple namespaces; another client can write records in +plyr's collections. Private application data is outside this public substrate. + +The current Python consumer uses the older live Jetstream interface with +`wantedCollections` and a timestamp cursor. It gates processing on known +artists locally and dispatches durable tasks. It is not evidence of an +already-shipped whole-network archival music rebuild. Own uploads can be +finalized when their PDS write echoes back; records written through another +client can enter the view for a known artist without an upload pending row. + +The dead-audioURL retrospective shows why readable records are only part of +the experience. Two records referenced unavailable CDN media while their PDS +blobs survived. A useful track needs playable media, not merely JSON. Later +replay work added revision protection after stale updates walked track +metadata backward. The recent player work also distinguishes a published +track's identity from an audio content hash shared by multiple uploads. + +Quiet collections expose another interpretive trap: no recent events need +not mean an unhealthy stream. Plyr detects missing echoes of known writes; +it has encountered a connected host silently omitting its collections. +These incidents concern the interfaces and hosts used at those dates, not +a blanket finding against today's v2 service. + +Read: README, STATUS vision and recent work, public lexicon overview, +`docs/internal/architecture/jetstream-ingest.md`, +`docs/internal/retrospectives/2026-06-30-dead-audiourl-ingest.md`, +frontend search/collection documentation, and the current Python consumer's +event processing and cursor handling. Inspected the live track listing and +its artist, tag, queue, and sharing context without triggering playback. + +## the current public Jetstream documentation + +Read the prose and code examples of all seven pages in the Jetstream section +on September 8, plus the linked Relay and Consuming the firehose guides: + +- [Jetstream](https://bsky.network/docs/jetstream/): filtered decoded events; + live, replay, and snapshot as related consumption modes. +- [SDK](https://bsky.network/docs/jetstream-sdk/): typed handlers, cancellation, + checkpointing, and application indexing examples. +- [Network Replay](https://bsky.network/docs/jetstream-replay/): bounded + snapshot or replay followed by live, planning, downloading, metering, and + recovery at the seam. Snapshotting is a section here, not another page. +- [Webhooks](https://bsky.network/docs/jetstream-webhooks/): turn a selected + stream into notifications relevant to a watchlist or topic. +- [Analytics](https://bsky.network/docs/jetstream-analytics/): materialize a + community's posts into DuckDB to ask questions about its activity. +- [Streaming moderation](https://bsky.network/docs/jetstream-coop/): translate + documents and authors into reviewable items and moderation workflows. +- [Self-hosting](https://bsky.network/docs/jetstream-self-host/): operate an + archive, including bootstrap, storage, retention, and serving costs. +- [Relay](https://bsky.network/docs/relay/) and + [Consuming the firehose](https://bsky.network/docs/consuming-the-firehose/): + distinguish raw verified repository events from downstream interpretation. + +These are current public documentation, unlike the August local upstream +checkout. The overview now recommends v2 for new projects and documents public +v2 endpoints; descriptions of the work as merely unannounced or preproduction +are stale. The HTTP reference link was located, but its client-rendered schemas +were not readable through the web text fetch and are not claimed as reviewed. + +The examples repeatedly go from selected records to a concrete outcome: an +alert, a queryable dataset, an index, a review item. The defining archival +addition is that a consumer can start after data was written, populate its +view, and keep it updated through one consumption flow. Tutorials simplify +some application concerns; the real projects show the extra work around +eligibility, content interpretation, identity, media, and durable processing. + +“Point-in-time copy” still requires care: a bounded archive read does not +restore superseded values already removed by compaction. Reading and folding +available events is distinct from an immutable historical database. Public +copy about replaying history must be read alongside the compaction contract. + +## synthesis before choosing an experience + +The three projects supply concrete evidence for a working hypothesis: +Strata has shown the archive's size and organization, but has not yet made the +possibility of building new views from shared network data tangible. Its +metadata-only drilldown stops before the person, piece of writing, or song +that gives those records meaning. This is an interpretation of the research, +not a settled brief or an instruction to clone any of these interfaces. + +Useful distinctions to preserve in subsequent design work: + +- A namespace groups record types; it does not identify one application. +- Events, distinct records, logical works, and people are different units. +- An archive supplies retained data; a consumer supplies interpretation and + product-specific inclusion, identity, relationships, and presentation. +- Backfill, application index construction, and live freshness are separate + stages even when the end user experiences one coherent product. +- “Can I find/read/play this?” is a stronger demonstration than “rows arrived.” + +## questions to settle before designing + +- Who should leave with a changed understanding: an app builder, a person + exploring the ecosystem, or an archive operator? +- Which capability should they have experienced: rebuilding, replaying, + inspecting a corpus, or understanding archive use? These can coexist, but + their evidence and success criteria differ. +- Which real example best demonstrates that capability on available data? + Verify its completeness, cost, and semantic limits before choosing its form. +- What freshness and coverage can Strata's derived dataset actually support? + Check drain completion, missing partitions, pagination, reconciliation, and + continued ingestion before promising exhaustive or live results. + +## source map + +- Stream: `docs/product-copy-context.md`, `docs/feature-surface.md`, + `docs/upstream-compaction-spec.md`, `docs/experiment-6-serving-receipt.md`, + `docs/experiment-6-handoff.md`, `docs/compaction-cost.md`, `docs/gotchas.md`, + and the opening status of `docs/backup-plan.md`. +- Jetstream SDK: `docs/client.md`, `docs/failover.md`, + `examples/streamplace_chat.md`. +- Notes: `protocols/atproto/{firehose,network-backfill,record-retrieval,serving-event-streams}.md` + and `storage/columnar-files.md`. Measurements are dated evidence, not SLAs. +- Strata: `docs/how-it-works.md`, `src/index.ts`, `src/ui.js`, and original + Stream conversation `31544ae6-0053-47db-a6a4-d0726cb701d1` (August 28–31). +- Upstream design: local `~/github.com/bluesky-social/jetstream/docs/README.md` + at the revision noted above. The checkout was not refreshed; current public + website documentation was read separately as listed above. +- Project source roots for the cross-project reading above: + `~/tangled.org/zzstoatzz.io/typeahead`, + `~/tangled.org/zzstoatzz.io/pub-search`, and + `~/tangled.org/zzstoatzz.io/plyr.fm`. diff --git a/docs/strata-storage-discovery.md b/docs/strata-storage-discovery.md new file mode 100644 index 0000000..4c8ce1b --- /dev/null +++ b/docs/strata-storage-discovery.md @@ -0,0 +1,77 @@ +# Storage discovery — 2026-09-09 + +The measured sealed-file listing totals **1,831,276,429,146 bytes (1.83 TB)** +across **6,959 files**. This is the sum of file lengths, not allocated filesystem +blocks, total host disk use, database size, or a measure of reclaimable space. +The concurrent public `stream_archive_sealed_bytes` gauge matched this total. + +## What the format establishes + +Source: Stream `src/internal/storage/segment.zig`, `segment_writer.zig`, and +`segment_footer.zig`, with `docs/jss-format-v1.md` as a readable format guide. +The guide warns of older upstream commit pins; the current writer establishes +that `compressed_size` excludes the 8-byte frame length prefix. + +- The listing gives exact reported sealed file lengths and version checksums. +- Each block index gives compressed frame length, uncompressed body length, + row count, and sequence/time bounds. Reading record bodies is unnecessary. +- The collection index gives collection event counts and a membership bitmask + for each block. It does **not** give a compressed byte contribution for each + collection, or even per-block per-collection row counts. +- A block's compressed frame is shared. Assigning its bytes in proportion to + event counts would be an allocation estimate, not a measured disk size. +- A source's matching-block bytes can be measured as a read footprint, but + footprints overlap. They cannot be summed as independent storage ownership. +- Metadata identifies neither safe deletions nor the space compaction would + reclaim. Changing contents changes compression and may require rewriting. + +## Bounded inspection + +`strata/inspect-storage.py` lists at most ten pages, permits at most +20 requests per run, caps each response at 4 MiB and collection-index expansion +at 32 MiB. It reads headers, block indexes, and collection-index tails for the +first, middle, and final listing entries. No block frames, records, media, +Microcosm APIs, or full-file payloads are read. + +The successful run used 19 requests and 1,833,965 response bytes. An initial +attempt to read whole footers stopped at its cap: the latest footer exceeded +4 MiB. The revised reader skips bloom filters and fetches only the two indexes. +The report records the successful run's budget, not both attempts combined. + +Checks: listing checksum matches header identity; header is reread after each +probe; sizes reconcile exactly: compressed frames + frame prefixes + header + +footer = file length. Collection masks have validated dimensions. This does +not recompute the metadata checksum (which includes skipped bloom filters), +and it does not verify data-frame checksums. Multi-page listing is not atomic. + +| Listing position | File size | Compressed data | Uncompressed blocks | Metadata + framing | Compressed bytes in shared-namespace blocks | +|---|---:|---:|---:|---:|---:| +| First | 269,735,440 | 269,605,093 | 1,263,753,347 | 130,347 | 29,790,149 (11.0%) | +| Middle | 259,045,260 | 258,889,430 | 1,193,461,650 | 155,830 | 48,675,383 (18.8%) | +| Last measured | 273,237,270 | 268,652,030 | 982,730,290 | 4,585,240 | 268,646,139 (>99.9%) | + +These are three concrete examples, not representative samples or archive-wide +compression/ownership estimates. Shared membership includes account-level +sentinels as separate sources. Event-count spheres remain event-count spheres. + +## Product and request boundaries + +The unlisted preview includes a collapsed “Space behind the spheres” explainer: +measured overall file size plus selectable comparisons for the three files. +Measurement date and scope are visible. The report is a static build input, +not refreshed per visit or by a scheduled task. + +Microcosm remains a contextual UFOs link only: **zero API calls from Strata**. +Do not fan out over the namespace inventory, prefetch on pan/zoom, or poll. +A future inline enrichment should be explicitly requested for one selection, +cached, bounded to one in-flight request, and stop on rate limiting rather than +retrying or moving through the inventory. Its time window and observation +coverage must remain distinct from archive storage measurements. + +The existing `my-prefect-server/flows/strata.py` already reads changed segment +collection indexes. If full storage attribution/read-footprint indexing is +wanted, extend that existing checksum-driven pass, reusing its decoded masks +and reading the block index once per changed segment. Persist aggregates keyed +by segment version; do not create a second all-archive crawler. Any shared +block total must deduplicate `(segment, checksum, block index)`. No ingestion, +D1 schema, compaction behavior, or production Stream binary was changed here. -- 2.51.2