Storage discovery — 2026-09-09 #
The measured sealed-file listing totals 1,831,276,429,146 bytes (1.83 TB)
across 6,959 files. This is the sum of file lengths, not allocated filesystem
blocks, total host disk use, database size, or a measure of reclaimable space.
The concurrent public stream_archive_sealed_bytes gauge matched this total.
What the format establishes #
Source: Stream src/internal/storage/segment.zig, segment_writer.zig, and
segment_footer.zig, with docs/jss-format-v1.md as a readable format guide.
The guide warns of older upstream commit pins; the current writer establishes
that compressed_size excludes the 8-byte frame length prefix.
- The listing gives exact reported sealed file lengths and version checksums.
- Each block index gives compressed frame length, uncompressed body length, row count, and sequence/time bounds. Reading record bodies is unnecessary.
- The collection index gives collection event counts and a membership bitmask for each block. It does not give a compressed byte contribution for each collection, or even per-block per-collection row counts.
- A block's compressed frame is shared. Assigning its bytes in proportion to event counts would be an allocation estimate, not a measured disk size.
- A source's matching-block bytes can be measured as a read footprint, but footprints overlap. They cannot be summed as independent storage ownership.
- Metadata identifies neither safe deletions nor the space compaction would reclaim. Changing contents changes compression and may require rewriting.
Bounded inspection #
strata/inspect-storage.py lists at most ten pages, permits at most
20 requests per run, caps each response at 4 MiB and collection-index expansion
at 32 MiB. It reads headers, block indexes, and collection-index tails for the
first, middle, and final listing entries. No block frames, records, media,
Microcosm APIs, or full-file payloads are read.
The successful run used 19 requests and 1,833,965 response bytes. An initial attempt to read whole footers stopped at its cap: the latest footer exceeded 4 MiB. The revised reader skips bloom filters and fetches only the two indexes. The report records the successful run's budget, not both attempts combined.
Checks: listing checksum matches header identity; header is reread after each probe; sizes reconcile exactly: compressed frames + frame prefixes + header + footer = file length. Collection masks have validated dimensions. This does not recompute the metadata checksum (which includes skipped bloom filters), and it does not verify data-frame checksums. Multi-page listing is not atomic.
| Listing position | File size | Compressed data | Uncompressed blocks | Metadata + framing | Compressed bytes in shared-namespace blocks |
|---|---|---|---|---|---|
| First | 269,735,440 | 269,605,093 | 1,263,753,347 | 130,347 | 29,790,149 (11.0%) |
| Middle | 259,045,260 | 258,889,430 | 1,193,461,650 | 155,830 | 48,675,383 (18.8%) |
| Last measured | 273,237,270 | 268,652,030 | 982,730,290 | 4,585,240 | 268,646,139 (>99.9%) |
These are three concrete examples, not representative samples or archive-wide compression/ownership estimates. Shared membership includes account-level sentinels as separate sources. Event-count spheres remain event-count spheres.
Product and request boundaries #
The unlisted preview includes a collapsed “Space behind the spheres” explainer: measured overall file size plus selectable comparisons for the three files. Measurement date and scope are visible. The report is a static build input, not refreshed per visit or by a scheduled task.
Microcosm remains a contextual UFOs link only: zero API calls from Strata. Do not fan out over the namespace inventory, prefetch on pan/zoom, or poll. A future inline enrichment should be explicitly requested for one selection, cached, bounded to one in-flight request, and stop on rate limiting rather than retrying or moving through the inventory. Its time window and observation coverage must remain distinct from archive storage measurements.
The existing my-prefect-server/flows/strata.py already reads changed segment
collection indexes. If full storage attribution/read-footprint indexing is
wanted, extend that existing checksum-driven pass, reusing its decoded masks
and reading the block index once per changed segment. Persist aggregates keyed
by segment version; do not create a second all-archive crawler. Any shared
block total must deduplicate (segment, checksum, block index). No ingestion,
D1 schema, compaction behavior, or production Stream binary was changed here.