diff --git a/docs/timestamp-import.md b/docs/timestamp-import.md index 3975931..0ad8e11 100644 --- a/docs/timestamp-import.md +++ b/docs/timestamp-import.md @@ -4,6 +4,33 @@ Stream follows Jetstream V2's operator timestamp-import design. An import changes the timestamp presented to subscribers without changing the immutable witness timestamp used for archive ranges and timestamp cursors. +## What this is for, and the one distinction the whole design turns on + +Records carry **two** timestamps, and an import moves only one of them: + +| | changed by an import? | used for | +|---|---|---| +| **display timestamp** (`time_us` on the wire) | **yes** | what subscribers see | +| **witness timestamp** (`witnessed_at`) | **never** | archive ranges, timestamp cursors, segment header min/max | + +That separation is why an import is safe at all. `witnessed_at` is when *this +instance* observed the event, so it defines segment ordering and is what a +`?cursor=` replay resolves against. If an import could move it, +every sealed segment's header range and every timestamp cursor would become a +lie, and the "segment files sort in creation order, and that order is time +order" invariant would break. So imports are a **presentation** layer over +immutable history, not a rewrite of it. + +The operator use case is backfilled or migrated content whose true creation time +is older than when this instance first saw it — without an import, a repository +imported today would present all its history as today's events. + +Verified 2026-07-30: `--timestamp-import-dir` and `--timestamp-import-token` +both exist in `runtime/cli.zig`, and `/xrpc/network.bsky.jetstream.importTimestamps` +and `/xrpc/network.bsky.jetstream.getImportStatus` are both routed in +`serve/server.zig`. (Those are matched as full literal paths, so grepping the +source for the bare NSID finds nothing.) + ## Source format The source is a plain, seekable RFC 4180 CSV. `uri` and `timestamp` are