--- id: didjson-archive title: The identity layer has a copy that is not the running server status: open crates: [didbot-pds, didbot-serve, didbot-identity, didbot-dns] dependsOn: [deploy] exitCriterion: > One command reports that every account this deployment serves has a current DID document in the archive, names the drift for any that differ, and the archive can be pointed at as a read origin without touching the running server. --- # didjson-archive On `did:plc` the identity layer has a custodian: the directory holds the document, and a resolver asks it rather than the account's server. On `did:web` the document *is* whatever the host serves, which here means it is derived per request — `Registry::did_document` off the `Host` header, through [`DidDocument::for_account`](../crates/didbot-identity/src/document.rs). The consequence is the whole of this epic: the document a resolver reads is computed by the process answering for it, and everything downstream — signature verification, handle resolution, an `at://` URI meaning anything at all — reads it through that computation. This epic gives the identity layer a durable copy of its own, in S3, keyed so that [tombstone-serving](tombstone-serving.md) can serve directly out of it. The archive holds exactly the surface that is already world-readable at each account's own hostname: the document, and the plain-text DID that `/.well-known/atproto-did` answers. Nothing else. A signing key, an agent bearer credential from [`credential`](../crates/didbot-pds/src/credential.rs), the deployment's secret file — none of those belong in a bucket whose whole purpose is that a stranger can read it, and the archive's contents should be defined by that rule rather than by convenience. This is the one design constraint in this epic that is not open for trade-off. ## Granularity, and why there are two shapes One object per hostname is what a static origin needs: a request for `https://mossy-vole.pds.did.bot/.well-known/did.json` has to reach one key without anything running a lookup. One aggregate index per zone is what a completeness check needs: comparing an archive against a live deployment of a hundred thousand accounts should be one download and a set difference, not a hundred thousand `GET`s. The proposal is both, written in the same pass, with the aggregate index naming the per-account keys and their content hashes so the two can be checked against each other rather than trusted separately. Keys are proposed as zone first, then hostname. The zone layout makes that load-bearing rather than cosmetic: [`ZoneRegistry`](../crates/didbot-identity/src/did.rs) admits nested zones and `MultiZoneDns` routes each one to its own provider, so a deployment can be serving `*.pds.did.bot` and `*.agents.pds.did.bot` at once. A zone-first key makes a per-zone cutover a prefix operation, which is what [tombstone-serving](tombstone-serving.md) wants when one zone is frozen and another is still live. - [ ] **Write on the events that change a document, rather than on a timer alone.** The events already exist and already mean this: [`LifecycleEvent`](../crates/didbot-pds/src/lifecycle.rs) covers provisioning and state changes, and `RepoIdentity` on [`RepoEventSink`](../crates/didbot-pds/src/subscribe.rs) is literally "this account's DID document is worth re-resolving". Proposal: an archiver is one more sink on both, so the archive tracks the document for the same reasons a relay does. The sink contract is synchronous with no error channel, so the archiver absorbs its own failures — it queues and retries rather than turning a completed provisioning into a failure, which is what that contract already requires of every sink. - [ ] **A full sweep, at a cadence the process owns.** Event-driven writes are correct until one is dropped, and a queue that absorbs its own failures can drop one by design. The sweep is the answer, and its scheduling is `plan/periodic-backups.md`'s: this epic contributes the job, that one owns when jobs run and what happens when one overruns the next. Proposal for the interval is daily, matching [backup.tf](../infra/pds/backup.tf)'s existing rule so an operator has one cadence to reason about rather than two. - [ ] **Enumeration on the backend, which two epics want.** [`ObjectBackend`](../crates/didbot-pds/src/object_blobs.rs) is `put`/`presigned_get`/`delete` today, and a completeness check needs to list what the archive holds. Adding a list operation also closes the orphan-reclamation item [pds-writes](pds-writes.md) leaves open for exactly the same reason — a crash between a `put` and its log append leaves an object nothing can enumerate — so the operation should be designed once, against both callers, rather than added twice. - [ ] **A real S3 backend, hand-rolled against SigV4.** [pds-writes](pds-writes.md) already names the shape and the reason: `didbot-dns`'s `Route53Dns` signs its own requests rather than pulling in an AWS SDK the build otherwise never uses, and `presigned_get` is a local computation that shares those primitives. This epic is the caller that finally needs the bucket, so the SigV4 work lands here or lands first; which of the two is the owner's call about ordering. - [ ] **Completeness is a reconciliation, and the model already exists in this repository.** `Route53Dns::resync` and [zone reconciliation](../docs/zone-reconciliation.md) do the same job one layer down: enumerate the remote side, compare against local truth, report each disagreement in terms an operator can act on. Proposal: the same shape, with `Registry::accounts` as local truth and the aggregate index as the remote side, reporting three cases separately — an account with no archived document, an archived document that differs from the derived one, and an archived document for a DID the deployment no longer serves. The third is not an error: it is what a frozen account looks like from here, which is the join with [tombstone-serving](tombstone-serving.md). - [ ] **Say what a restore restores.** The document is derived from account state, so the authority for it is the account store and the write-ahead log, and an archived document that differs from the derived one means the store is at a different point in time. Two operations follow and they should be named differently rather than sharing a verb: *reconcile* brings the archive up to what a running server derives, and *serve* points a resolver at the archive with no server involved. The second is [tombstone-serving](tombstone-serving.md)'s; this epic owes it a layout it can serve without transformation. - [ ] **How the hosted-zone ceiling interacts, which depends on the DNS shape.** A Route53 hosted zone defaults to a 10,000-record-set quota — `didbot-dns`'s `route53` module docs say so and say that the API reports it rather than anything local enforcing it. Under the wildcard shape [deployment](../docs/deployment.md) documents, the zone holds a handful of records whatever the account count, so archive size and zone occupancy are independent and this item is a note rather than a constraint. Under a deployment publishing per-agent records they move together: every archived document has a matching record set, and the aggregate index is a cheaper place to count occupancy than paginating `ListResourceRecordSets` a hundred entries at a time. That count is the occupancy [zone-scale](zone-scale.md) asks an operator to be able to watch approaching, so the two items should be closed with one number. - [ ] **Two prefixes with different access, or two buckets.** The documents are public by construction. Anything else this archive grows into — a ledger export, a manifest naming blob keys — has a different audience, and the separation should be structural rather than a convention somebody maintains. Proposal: one bucket, a `public/` prefix with a read policy and everything else without one, and an [adversarial](adversarial.md)-shaped test that asserts an unauthenticated `GET` outside `public/` is refused. - [ ] **Retention and versioning, decided rather than defaulted.** A document changes when a key rotates or a handle moves, and a resolver that cached the old one is the reason [account-types](account-types.md) wants `generation` visible. Proposal: object versioning on, with a lifecycle rule, so the archive can answer "what did this identity look like before the rotation" — and a stated limit on how far back, because an unbounded answer is a storage bill nobody chose. ## Done Nothing closed yet.