discovery before representation #
Research pass, 2026-09-08. This is context for deciding Strata's purpose, not an approved experience or implementation plan. Operational measurements below retain their original dates; this pass did not inspect the production machine or establish current drain completeness.
what archival Jetstream enables #
Upstream's design describes a compressed network archive and live stream for
building AppViews and doing network analysis. Its explicit non-goals include
arbitrary queries and point lookups: the archive is a replay cache for range
scans. The local upstream reference is bluesky-social/jetstream at 289b032
(2026-08-13), docs/README.md sections 1–2, not a claim about latest upstream.
The important change for a builder is being able to choose what to consume after capture began. They can backfill a view, rebuild it, or catch up after being offline, then continue consuming live. The SDK owns the replay-to-live seam; the application owns durable progress, idempotency, and folding markers.
Stream's product context records the architectural motivation as:
origin PDS network -> durable archive -> ephemeral application view
Rebuilding an application need not mean recrawling all origin PDSes. Bobbin's hydration from its Hydrant twin is an attributed example of this architecture, not evidence that Bobbin consumes Stream.
what “history” can honestly mean #
- Bootstrap retrieves repository state available when fetched. It cannot reconstruct every mutation preceding capture.
- Compaction removes superseded and deleted materialization rows. Markers survive. Available replay is not an immutable recording of every old value.
- Witnessed time is the instance's observation time, distinct from record creation. Backfill, catch-up, resync, and timestamp imports matter to its interpretation. Archive order is not a network-wide chronology of creation.
- Coverage belongs to an instance, its enumeration and capture, and what was retrievable. Counts of namespaces are not verified counts of independent apps; counts of event rows are not counts of distinct current records.
- Sequence cursors belong to one instance. Cross-host recovery is approximate and carries different guarantees from replay within one archive.
Some older notes use “full event history” and “whole network” more broadly. The owning compaction and client contracts qualify those phrases.
uses with evidence #
| use | evidence | qualification |
|---|---|---|
| rebuild a filtered view, then follow live | SDK unified subscribe contract | application must fold markers and persist progress |
| replay a particular broadcast's chat | Streamplace example: 17 messages from 633 blocks | retained records only; archive read amplification is visible even in a tiny result |
| analyze or materialize a corpus | network-backfill note records a Bluesky archive-to-ClickHouse drain around 17h | dated, attributed measurement; not a current Stream benchmark |
| inspect archive composition cheaply | Strata's footer reader: about 1.3 KB in two requests per sampled segment | counts reveal composition, not content or historical user activity |
| query the long tail interactively | Strata namespace-partitioned Parquet and DuckDB proof | derived dataset; coverage and reconciliation are separate responsibilities |
The August 17 retrieval comparison is particularly useful: one sparse collection returned 17,537 archive rows for 497 MB versus 13,145 current rows for 3.6 MB through directory plus PDS reads. These answer different questions. The archive's benefit includes one retrieval source and retained events; it does not guarantee the cheapest way to fetch any small collection.
operational lessons relevant to a playground #
- Query cost follows blocks touched, not result size. Sparse scattered rows can be expensive; estimates and bounded reads must precede bulk work.
- Bootstrap layout and mature live layout differ. Per-DID locality degrades as new events arrive interleaved; fresh-archive demos can mislead.
- Compaction changes old files and counts. Strata's footer summaries reconcile checksums, but its drain documentation still leaves selective re-drain open.
- Receiving events does not establish freshness. During catch-up the service can emit old content rapidly. Process uptime is not archive age or coverage.
- Shared serving resources are finite. The August 17 compaction measurement found disk saturation; Strata's drain subsequently gained CPU/memory caps.
- The shelved backup proposal is evidence of a limit, not unfinished scope: restore work grows with the archive while upstream replay retention is finite.
what Strata has, and what it has not established #
Original discussion: replay-use heatmap -> inspect the archive externally -> progressive exploration of the long tail. The explicit constraint was to avoid adding in-process telemetry purely for the playground.
Current Strata summarizes storage slices, and separately queries extracted event metadata. The extracted rows omit record bodies. Neither the heatmap nor the Parquet path currently demonstrates an application's replay-to-live seam. The first idea—visualizing what other consumers replay—requires demand data different from archive composition. Footer counts cannot answer it.
Read-only checks on September 8 found truncated: true on both fm.teal and
sh.tangled partition listings. The Worker returns one listing page and the UI
ignores truncation. Thus “every record an app has ever written” is unsupported
by both the available-history semantics and this concrete retrieval limit.
discovery across the other projects #
The first pass overweighted archive mechanics. Reading Typeahead, pub-search, and plyr.fm changes the working interpretation: the archive is valuable as material from which someone can construct a useful view of the network. Those projects make that concrete in three different ways.
Typeahead: who becomes findable #
Typeahead is community actor search, including identities outside Bluesky.
Its current ingester consumes com.atproto.sync.subscribeRepos directly;
services/src/ingest.zig applies a collection allowlist locally. It is not
currently an archival Jetstream consumer.
The August birds.place incident is a particularly useful case. An account
writing only place.birds.* records was invisible to the previous allowlist,
and identity resolution also failed. The current allowlist includes that
namespace. The broader lexicon-driven discovery proposal remains a design,
not something to count as shipped. Capturing a person's records, recognizing
their identity, extracting a profile, and making them searchable are separate
steps. A person without a Bluesky profile is still a network participant.
The September 5 recovery retrospective reports 377 restored active accounts and checks their search visibility after synchronization. That is a meaningful outcome for recovery; a count of replayed frames alone would not establish it. The report does not establish lossless historical recovery.
Typeahead also separates costly index construction from serving: durable actor data becomes an independently built search artifact, with recent changes in an overlay. Its application index snapshot is different from a Jetstream archive snapshot. Both can be called “snapshot,” but they serve different jobs.
Read: docs/discovery-and-identity.md, docs/collection-discovery.md,
docs/architecture.md, docs/retros/2026-09-05-ingester-recovery-closeout.md,
docs/retros/2026-05-23-standard-site-backfill.md, current ingest allowlist,
README, AGENTS, and recent commit history. Live actor search for zzstoatzz.io
returned nate and two other matching handles during this pass.
pub-search: a corpus across publishing applications #
Pub-search provides the strongest direct example of archival Jetstream
replacing bespoke ingestion infrastructure. The August cutover adopted Stream;
the following unified-client change uses archive replay followed by live
events, and the separate ingester was decommissioned. The current consumer is
backend/src/ingest/jetstream.zig, with nine selected collections and a
persisted sequence. The README still describes the retired ingester, so the
cutover document, implementation, and commit history take precedence.
Its product is one long-form corpus across publishing platforms. Several
platforms share site.standard.document; namespace is not application
identity. Even within that schema, extracting readable content depends on
platform conventions, including structured Leaflet blocks versus flat text.
Deduplication also changes the meaning of volume: the August cleanup removed
27,076 stale republish rows without removing distinct logical documents.
The July corpus audit distinguishes genuinely missing writing, valid dedupe, metadata-only records, excluded sources, and unreachable PDSes. Its 15,174 genuine misses were primarily a historical backfill gap. Those are dated findings, not a current completeness assessment. The lesson for Strata is that “present in the archive,” “indexed,” and “available to discover” are different statements. Pub-search also honors discovery preferences in its discovery surfaces while allowing direct retrieval of a known public URI.
The Atlas already demonstrates progressive exploration of network content: regions, narrower clusters, and individual writing. Its labels derive from semantic neighborhoods, not collection names. The two-dimensional layout is a display projection; clustering uses a separate higher-dimensional embedding. The live map displayed 69,013 documents during this pass. That number describes this displayed dataset, not total network writing.
Its agent surfaces make the same point without a map: curated operations return an author's writing, related posts across platforms, and recommendations with snippets. They package something a reader can understand and act on.
Read: docs/jetstream-cutover.md, docs/reconciliation.md,
docs/content-extraction.md, docs/corpus-audit-2026-07-20.md,
docs/atlas.md, docs/agent-surfaces.md, the architecture portion of
docs/snapshot-pipeline.md, current consumer code, README, and cutover/recent
commit history. Older counts and API claims in overview documents require
care: these documents contain revisions and some contradictory older text.
plyr.fm: records become something you can listen to #
Plyr's public data model includes tracks, likes, timed comments, playlists,
and profiles. It also writes listening records under fm.teal.*. One
application spans multiple namespaces; another client can write records in
plyr's collections. Private application data is outside this public substrate.
The current Python consumer uses the older live Jetstream interface with
wantedCollections and a timestamp cursor. It gates processing on known
artists locally and dispatches durable tasks. It is not evidence of an
already-shipped whole-network archival music rebuild. Own uploads can be
finalized when their PDS write echoes back; records written through another
client can enter the view for a known artist without an upload pending row.
The dead-audioURL retrospective shows why readable records are only part of the experience. Two records referenced unavailable CDN media while their PDS blobs survived. A useful track needs playable media, not merely JSON. Later replay work added revision protection after stale updates walked track metadata backward. The recent player work also distinguishes a published track's identity from an audio content hash shared by multiple uploads.
Quiet collections expose another interpretive trap: no recent events need not mean an unhealthy stream. Plyr detects missing echoes of known writes; it has encountered a connected host silently omitting its collections. These incidents concern the interfaces and hosts used at those dates, not a blanket finding against today's v2 service.
Read: README, STATUS vision and recent work, public lexicon overview,
docs/internal/architecture/jetstream-ingest.md,
docs/internal/retrospectives/2026-06-30-dead-audiourl-ingest.md,
frontend search/collection documentation, and the current Python consumer's
event processing and cursor handling. Inspected the live track listing and
its artist, tag, queue, and sharing context without triggering playback.
the current public Jetstream documentation #
Read the prose and code examples of all seven pages in the Jetstream section on September 8, plus the linked Relay and Consuming the firehose guides:
- Jetstream: filtered decoded events; live, replay, and snapshot as related consumption modes.
- SDK: typed handlers, cancellation, checkpointing, and application indexing examples.
- Network Replay: bounded snapshot or replay followed by live, planning, downloading, metering, and recovery at the seam. Snapshotting is a section here, not another page.
- Webhooks: turn a selected stream into notifications relevant to a watchlist or topic.
- Analytics: materialize a community's posts into DuckDB to ask questions about its activity.
- Streaming moderation: translate documents and authors into reviewable items and moderation workflows.
- Self-hosting: operate an archive, including bootstrap, storage, retention, and serving costs.
- Relay and Consuming the firehose: distinguish raw verified repository events from downstream interpretation.
These are current public documentation, unlike the August local upstream checkout. The overview now recommends v2 for new projects and documents public v2 endpoints; descriptions of the work as merely unannounced or preproduction are stale. The HTTP reference link was located, but its client-rendered schemas were not readable through the web text fetch and are not claimed as reviewed.
The examples repeatedly go from selected records to a concrete outcome: an alert, a queryable dataset, an index, a review item. The defining archival addition is that a consumer can start after data was written, populate its view, and keep it updated through one consumption flow. Tutorials simplify some application concerns; the real projects show the extra work around eligibility, content interpretation, identity, media, and durable processing.
“Point-in-time copy” still requires care: a bounded archive read does not restore superseded values already removed by compaction. Reading and folding available events is distinct from an immutable historical database. Public copy about replaying history must be read alongside the compaction contract.
synthesis before choosing an experience #
The three projects supply concrete evidence for a working hypothesis: Strata has shown the archive's size and organization, but has not yet made the possibility of building new views from shared network data tangible. Its metadata-only drilldown stops before the person, piece of writing, or song that gives those records meaning. This is an interpretation of the research, not a settled brief or an instruction to clone any of these interfaces.
Useful distinctions to preserve in subsequent design work:
- A namespace groups record types; it does not identify one application.
- Events, distinct records, logical works, and people are different units.
- An archive supplies retained data; a consumer supplies interpretation and product-specific inclusion, identity, relationships, and presentation.
- Backfill, application index construction, and live freshness are separate stages even when the end user experiences one coherent product.
- “Can I find/read/play this?” is a stronger demonstration than “rows arrived.”
questions to settle before designing #
- Who should leave with a changed understanding: an app builder, a person exploring the ecosystem, or an archive operator?
- Which capability should they have experienced: rebuilding, replaying, inspecting a corpus, or understanding archive use? These can coexist, but their evidence and success criteria differ.
- Which real example best demonstrates that capability on available data? Verify its completeness, cost, and semantic limits before choosing its form.
- What freshness and coverage can Strata's derived dataset actually support? Check drain completion, missing partitions, pagination, reconciliation, and continued ingestion before promising exhaustive or live results.
source map #
- Stream:
docs/product-copy-context.md,docs/feature-surface.md,docs/upstream-compaction-spec.md,docs/experiment-6-serving-receipt.md,docs/experiment-6-handoff.md,docs/compaction-cost.md,docs/gotchas.md, and the opening status ofdocs/backup-plan.md. - Jetstream SDK:
docs/client.md,docs/failover.md,examples/streamplace_chat.md. - Notes:
protocols/atproto/{firehose,network-backfill,record-retrieval,serving-event-streams}.mdandstorage/columnar-files.md. Measurements are dated evidence, not SLAs. - Strata:
docs/how-it-works.md,src/index.ts,src/ui.js, and original Stream conversation31544ae6-0053-47db-a6a4-d0726cb701d1(August 28–31). - Upstream design: local
~/github.com/bluesky-social/jetstream/docs/README.mdat the revision noted above. The checkout was not refreshed; current public website documentation was read separately as listed above. - Project source roots for the cross-project reading above:
~/tangled.org/zzstoatzz.io/typeahead,~/tangled.org/zzstoatzz.io/pub-search, and~/tangled.org/zzstoatzz.io/plyr.fm.