id: didjson-archive
title: The identity layer has a copy that is not the running server
status: open
crates: [didbot-pds, didbot-serve, didbot-identity, didbot-dns]
dependsOn: [deploy]
exitCriterion: >
One command reports that every account this deployment serves has a current
DID document in the archive, names the drift for any that differ, and the
archive can be pointed at as a read origin without touching the running
server. #
On did:plc the identity layer has a custodian: the directory holds the
document, and a resolver asks it rather than the account's server. On
did:web the document is whatever the host serves, which here means it is
derived per request — Registry::did_document off the Host header, through
DidDocument::for_account. The
consequence is the whole of this epic: the document a resolver reads is
computed by the process answering for it, and everything downstream —
signature verification, handle resolution, an at:// URI meaning anything at
all — reads it through that computation. This epic gives the identity layer a
durable copy of its own, in S3, keyed so that
tombstone-serving can serve directly out of it.
The archive holds exactly the surface that is already world-readable at each
account's own hostname: the document, and the plain-text DID that
/.well-known/atproto-did answers. Nothing else. A signing key, an agent
bearer credential from
credential, the deployment's
secret file — none of those belong in a bucket whose whole purpose is that a
stranger can read it, and the archive's contents should be defined by that
rule rather than by convenience. This is the one design constraint in this
epic that is not open for trade-off.
One object per hostname is what a static origin needs: a request for
https://mossy-vole.pds.did.bot/.well-known/did.json has to reach one key
without anything running a lookup. One aggregate index per zone is what a
completeness check needs: comparing an archive against a live deployment of a
hundred thousand accounts should be one download and a set difference, not a
hundred thousand GETs. The proposal is both, written in the same pass, with
the aggregate index naming the per-account keys and their content hashes so
the two can be checked against each other rather than trusted separately.
Keys are proposed as zone first, then hostname. The zone layout makes that
load-bearing rather than cosmetic:
ZoneRegistry admits nested zones and
MultiZoneDns routes each one to its own provider, so a deployment can be
serving *.pds.did.bot and *.agents.pds.did.bot at once. A zone-first key
makes a per-zone cutover a prefix operation, which is what
tombstone-serving wants when one zone is frozen and
another is still live.