# Conformance atproto is a protocol other people implement. Every identifier this server mints — a record key, a DID, a handle, an `at://` URI — is a string somebody else's code will parse, and the only way to know it will parse is to check against the same vectors they check against. The official corpus at is CC-0, so it is vendored into this repository under `vendor/atproto-interop-tests/` rather than fetched at test time: a conformance suite that only runs when GitHub is reachable is a suite that stops running. `PROVENANCE.md` in that directory records which upstream commit the copy came from and when. `scripts/refresh-interop-vectors.sh` replaces it and rewrites the provenance, so the copy moves forward deliberately instead of rotting. The `com.atproto.*` lexicon documents are vendored the same way and for the same reason, under `vendor/atproto-lexicons/`, by `scripts/refresh-atproto-lexicons.sh`. Those are not vectors — they are the specification of the wire surface this server claims to implement — and what reads them is described under [the wire](#the-wire-checked-against-the-documents-that-define-it). They are dual MIT/Apache-2.0 rather than CC-0, so the licence files are copied beside them and `PROVENANCE.md` is the attribution. The executable half is `crates/didbot/tests/conformance.rs`. It is an ordinary `cargo test` target and adds no dependency to anything. ## What runs Each syntax suite reads `_syntax_valid.txt` and `_syntax_invalid.txt` and asserts the obvious thing about every line in each. | Suite | Validator | | --- | --- | | `tid` | [`crate::pds::tid::validate`] | | `recordkey` | [`crate::pds::records::validate_record_key`] | | `nsid` | [`crate::nsid_syntax::validate`] | | `handle` | [`crate::identity::validate_handle`] | | `did` | [`crate::identity::validate_did`] | | `atidentifier` | either of the two above | | `aturi` | [`crate::pds::AtUri::parse`] | | `datetime` | [`crate::pds::format::datetime::validate`] | | `language` | [`crate::pds::format::language::validate`] | | `uri` | [`crate::pds::format::uri::validate`] | | `cid` | [`crate::pds::format::cid::validate`] | The `lexicon/` vectors run too, and they are a different shape: rather than a list of strings they are a small catalog of lexicon documents plus two lists of records that must and must not validate against them. They exercise blobs, bytes, unions, refs and closed enumerations — none of which this project's own five lexicons use — so they test the record validator far harder than `lexicons/` can. `crates/didbot/tests/conformance/records.rs` runs them against [`crate::pds::Catalog`]. The `data-model/` vectors are a third shape again, and the only ones that check *bytes* rather than a verdict. Two of the three files are the familiar accept-and-refuse pair, run against [`crate::data::Value::object_from_json`]; the third, `data-model-fixtures.json`, gives a value, the exact DAG-CBOR encoding it must produce, and the CID that encoding must hash to. `crates/didbot/tests/conformance/data_model.rs` checks all three, and the fixtures are why the suite is worth having: an encoder checked against its own decoder round-trips perfectly whatever map ordering or integer width it picked, and would still give every record an identity no other implementation would agree with. Because those files are data rather than code, the runner also checks that it is reading them. A vector file that is present, parses and holds an empty list makes every loop over it run zero times, so each file is asserted non-empty as it is read and the failure names the file. A file nobody opens is the same silence one level up, so the directory listing is compared against the three names the runner reads: a refresh that adds a fourth file fails rather than leaving the new vectors unrun. The `firehose/` vectors are the fifth shape and the misleadingly named ones. They were skipped here for a while on the reading that they tested the protocol's framed DAG-CBOR stream, which this project deliberately does not speak. They do not. `commit-proof-fixtures.json` gives a leaf set, the root over it, a set of adds and deletes, the root after them, and the blocks a subscriber needs to check the operation without holding the repository. `crates/didbot/tests/conformance/firehose.rs` runs all six. That makes it the only upstream vector asserting **a root CID over a set of records**, which is the one claim `mst/` does not make, and the six shapes are chosen to be awkward: a split two layers deep, a leafless split, a split whose neighbour is two layers down, and two commits that merge and split in the same operation. The proofs come from [`crate::data::mst::Tree::covering_proof`], which is the walk to the changed key plus the walks to the keys either side of it — adding or removing a key can split a node or merge two, so a proof of the key alone verifies the lookup and cannot reproduce the root the commit started from. In every fixture the proof this project builds is exactly the size upstream requires, and in four of the six it is smaller than the tree. Upstream's own runner asserts one further thing that is not asserted here: it loads the after-tree from a store holding only the proof, inverts the operations in every order, and requires the before-root back. That needs a tree loadable from partial storage and mutable one operation at a time, and [`crate::data::mst::Tree::build`] is deliberately a pure function of a whole leaf set. What is checked instead is the property the inversion demonstrates — the proof carries every block upstream says it needs. The `mst/` vectors are the fourth shape, and the odd ones: they check two pure functions that a Merkle search tree is built out of, and there is no tree here. A key's height is the count of leading zero bits of its sha-256 divided by two — the fanout is four, so two zero bits buy one layer — and a common prefix length is how many leading bytes two adjacent keys share, which is what lets a node store a suffix instead of a whole key. Both live in [`crate::data::mst`], and `crates/didbot/tests/conformance/mst.rs` runs `key_heights.json` and `common_prefix.json` against them. It also runs `example_keys.txt`, which is a vector file that does not announce itself. Upstream generates it by searching for keys at a wanted height and naming each one after the height it found, so `C4/134079` is an assertion that that key sits at layer four — 156 more height cases, reaching layers three and five, which the nine explicit ones never do. There is now a tree as well, and upstream ships nothing that checks one: a root CID over a set of records is the one claim `mst/` does not make. So it is checked against a repository somebody else's server built. `crates/didbot-data/tests/reference_repository.rs` reads `tests/fixtures/reference-repository.tsv` — 147 records of a real public repository, captured with `goat repo export` and reduced to the two columns a tree is a function of — and requires our root to be the root the reference implementation's commit signed. Two larger repositories were checked the same way while the tree was written, at 1,729 and 10,803 records, and are not vendored because the fixture would be a megabyte. That test is the one that earns its place. The plausible layering rule — the one that collapses a level nothing sits on — passes every self-consistency check a tree can run against itself and reproduces none of the three repositories. Checking the two functions before the tree is the same argument as the data-model fixtures, only sharper. A tree built on a wrong height function is not merely self-consistent, it is *correct*: every lookup succeeds and every proof it makes verifies against itself. The only symptom is a root CID nobody else reproduces. No test a tree runs against itself can find that, and the one thing this suite cannot say — the prefix is counted in bytes, and every candidate upstream ships is ASCII, so a character count passes all of it. That claim is asserted in the crate's own unit tests instead. Three of those validators did not exist before the vectors did. This project minted TIDs without being able to read one, built `at://` URIs without being able to parse one, and accepted a requested handle without checking it was a handle. A format with a writer and no reader is a format nothing can verify, which is the gap a conformance suite exists to find. ## What the vectors cannot tell us A vector says a validator agrees with the specification. It says nothing about whether the strings this server *produces* are ones that validator would accept, and that is the failure that would actually reach a user: an rkey no other implementation can read, a URI that resolves nowhere. So the same test target mints real output and feeds it back — every record key the store hands out is a valid TID, every URI the wire layer builds parses back into the three parts it was built from, every DID and hostname the minter produces is valid under the protocol's own generic rules. That last one is the load-bearing direction. [`AgentDid`] is deliberately much stricter than atproto's DID grammar — `did:web` only, hostname-level only, ports only on loopback — so the generic vectors are *not* run against it; they would report failures that are the entire point of the type. What is asserted instead is the implication that matters: everything we mint is also a DID by the protocol's definition. If that ever stops being true, the strictness has stopped being a subset and has become a dialect. A handle is checked the same way, and then one step further, because a handle is the one identifier here that has to agree with something else. atproto's rule is bidirectional: a handle resolves to a DID, and that DID's document must claim the handle back in `alsoKnownAs`. So the emit tests provision real accounts through a real [`Provisioner`] — with a namer and without, because they are different paths — and assert the loop closes for every one of them. | Property | What is minted, and what it is fed back through | | --- | --- | | every issued handle is a handle | a name composed from a word list and the zone, through [`validate_handle`] | | the well-known answers what the document claims | `Registry::handle_did` against `DidDocument::claims_handle` | | no two accounts answer to one handle | every account in a population of two hundred | | a name nobody holds resolves to nothing | a handle under the zone that was never issued | The second row is the load-bearing one, and it caught a real disagreement: an account with no issued handle claims its own hostname in `alsoKnownAs`, which is the honest `did:web` default, and the well-known used to deny it because nothing was stored. Both answers now derive from one function, so the property is by construction and the test says the construction held. Handle resolution is verified in both directions, including why the DNS TXT resolution method is not implemented and why the reading side checks containment rather than making the round trip. [`AgentDid`]: crate::identity::AgentDid [`Provisioner`]: crate::pds::Provisioner [`validate_handle`]: crate::identity::validate_handle ## Keys and signatures `crypto/` was listed here as blocked on signing, which conflated two things. This server does not sign a *commit* — there is no Merkle search tree to sign over — but it has had signature primitives since the beginning: a keypair per account, the public half published as `publicKeyMultibase` in a DID document, and verification. The vectors ask only about those, so nothing had to be built to run them. | File | What it asserts | | --- | --- | | `w3c_didkey_K256.json` | A private key encodes to the `did:key` the W3C vectors expect, and that string decodes back to the same key | | `signature-fixtures.json` | A signature is accepted only when it is a 64-byte compact pair with a low-S value that verifies | This is the load-bearing direction rather than an incidental one. Every DID document this server serves carries a multibase public key, and a resolver that decodes it differently than we encode it gets a *different key* — which looks like a valid document describing an account whose signatures never verify. Nothing else in the suite would notice. The two rejections are checked for the reason as well as the verdict: the `high-s` fixture must be refused by the low-S rule and the `der-encoded` one by the length. A fixture rejected for the wrong reason keeps passing after the check it is watching is deleted. `ES256` and `w3c_didkey_P256.json` name P-256, and they used to be asserted present and deliberately unrun, because there was nothing to run them against. There is now. [`crate::key::secp256r1`] verifies a P-256 signature and decodes a P-256 `did:key`, and does not sign: the crate's argument for one curve is an argument about *signing*, where `k256` normalizes `s` and `p256` does not, so a P-256 signer here would emit a malleable signature about half the time. Verifying is the other direction — a key from elsewhere is not ours to normalize, it is ours to judge — so the low-`s` rule `k256` applies internally is written out by hand there. Half the network's accounts sign on that curve, so a server that will one day read another server's commit needs to read both. A further test asserts that every fixture in the file is claimed by one of the two curves, because being filtered out by both loops is exactly how P-256 sat unrun. ## The wire, checked against the documents that define it The suites above check strings and bytes. They say nothing about whether a client written against the lexicons can talk to the routes, which is the epic's actual exit criterion. `crates/didbot/tests/conformance/wire.rs` is the automated half. It drives every `com.atproto.*` route the router serves and checks each response body against the `output` schema the *upstream* document declares — `vendor/atproto-lexicons/`, Bluesky's own documents copied unedited by `scripts/refresh-atproto-lexicons.sh`. The requests never leave the process: the router is handed them directly, as in `tests/end_to_end.rs`, so this needs no port, no network and no client written in another language. The distinction between that and an ordinary route test is the whole point. A response type written from a reading of a document is checked by a test written from the same reading, so the two agree with each other and neither agrees with anything else. That is not hypothetical: for months every write route stored the record and answered without a `cid`, which the lexicon has always marked required, and nothing in this repository noticed until the official client refused a response. Reintroducing that bug now fails seven tests with `` `cid` is required and is missing``. | What it checks | How | | --- | --- | | every route's response satisfies its `output` schema | `Catalog::validate_output` against the vendored document | | every request these tests compose satisfies its `input` schema | the same, in the other direction, so a fixture cannot rot and go on proving things | | every route this server serves is a method upstream defines | `serve::ATPROTO_METHODS`, which the route table itself is generated from | | every refusal carries an error name the document declares | plus `InvalidRequest` and `MethodNotImplemented`, which XRPC defines for all methods | | paging conforms on every page, not just the first | `listRecords` is the one response whose shape changes with the request | | a foreign record round-trips and is reported `unknown` | checked against upstream's own `knownValues`, not this project's spelling of them | | a blob upload answers the four-field `blob` shape | the one response whose schema is a lexicon *type* rather than a hand-written object | Two silences are treated as failures rather than passes. A method the vendor directory does not hold, and a method whose document declares no schema for the body in hand, both fail the assertion instead of quietly checking nothing — `com.atproto.sync.getRepo` returns a CAR file and is named as declaring no JSON schema, so that the day it grows one this suite stops skipping it. There is one place this server deliberately says more than its document declares, and it is a named constant with the reason attached rather than a check nobody runs: `com.atproto.repo.getRecord` answers `RepoNotFound` for a repository that was never here, where the document declares only `RecordNotFound`. Those are different things to a caller, and `com.atproto.sync.getRecord` declares the name for the same condition, so it is the protocol's word rather than one invented here. Every other undeclared name fails. ### The same surface, driven by somebody else's client What has been done by hand, and is worth writing down because the result is specific: the official `@atproto/api` was pointed at a development server and run through every `com.atproto.repo.*` route. It validates each response against the lexicon it generated its types from, so it is a conformance check that happens to be somebody else's code. | What it did | Result | | --- | --- | | `describeRepo`, by DID and by handle | accepted | | `getRecord` | accepted | | `getRecord` for a key that is not there | raised its own `RecordNotFoundError` — the typed class the lexicon's declared error name generates | | `getRecord` in a repository that is not here | `XRPCError` with `RepoNotFound` | | `deleteRecord`, twice | accepted both times | | `listRecords` paged with `cursor` past a repository of more than a hundred records | every row arrived exactly once, in key order, and the last page carried no cursor | | `listRecords` with `reverse` | the same rows, oldest first | | `listRecords`, `createRecord`, `applyWrites` | accepted, each response carrying the `cid` the lexicon requires | | `getRecord` naming the `cid` a write returned | accepted, and the same record came back | | `putRecord` carrying the `swapRecord` that is there | accepted, and answered with the new CID | | `putRecord` carrying a `swapRecord` that has moved on | raised its own `InvalidSwapError` — the typed class the lexicon's declared error name generates — and the record was unchanged | | `deleteRecord` carrying a stale `swapRecord` | `InvalidSwapError`, and the record was still there afterwards | | `createRecord` over a `literal:self` record that exists | `XRPCError` with `InvalidRequest`, naming `putRecord` | | `createRecord` carrying the `swapCommit` `getRepo` publishes | accepted | | `createRecord` carrying a `swapCommit` the repository has moved off | `InvalidSwapError` | | `applyWrites` carrying one `swapCommit` for three writes | accepted, three results | Since validation became optimistic the same client also writes a real foreign record — an `app.bsky.feed.post`, composed by `@atproto/api` from its own generated types and carrying a field no version of that lexicon declares — reads it back through `getRecord`, and gets the same object it sent. The write answers `validationStatus: "unknown"`, which is the protocol's word for a record stored under a lexicon the server could not resolve, and a scrobble in the same session answers `"valid"`. That is the epic's exit criterion for the collection set: the client knows nothing about this project, and this project knows nothing about the client's lexicon. Two rows are worth reading twice. `listRecords`, `createRecord` and `applyWrites` used to read "stored the records and then refused the response: `must have the property "cid"`", which was the known gap and the only thing in the surface that failed; naming a record on the way in closed it. And the swap rows used to be one row saying `UnsupportedSwap` — a name the client does not recognise — where the client now raises `InvalidSwapError`, the class generated from the error name the lexicon declares. Seventeen calls, nothing refused that should not have been. The CID a client is handed is checked against something that is not this project. `goat repo ls` over the same repository's CAR export lists each record under exactly the CID `applyWrites` returned for it, so the name a client holds is the name the file uses. The same client drives the three blob methods, and every row passes. It is worth listing separately because a blob is the one thing here whose identity a client can check on its own: the reference `uploadBlob` returns names bytes, and the client recomputes that name from the bytes with its own `multiformats` rather than believing the server's. | What it did | Result | | --- | --- | | `uploadBlob` with a PNG | accepted, and parsed into the client's own `BlobRef` — so the `{"$type": "blob", …}` shape validated against the lexicon that generated it | | the CID, recomputed from the file by the client | identical to the one the upload returned | | `getBlob` | the bytes back, `sha256` identical to the file, `content-type` as uploaded | | `getBlob` after a `SIGKILL` and a restart | identical again, from `--data` | | `getBlob` for a CID nobody uploaded | `XRPCError` with `BlobNotFound` | | `listBlobs` | names the blob; a repository nobody minted is `RepoNotFound` | | `listBlobs` with `since` | `XRPCError` with `InvalidRequest`, refused by name because nothing here counts references from records to blobs | Three sabotage runs go with it, since a store that never refuses anything is not checking anything. One bit flipped in a stored blob makes `getBlob` answer 500 `BlobCorrupt` naming what the bytes actually hash to, having served none of them. An upload cut off half way — the connection reset mid-body — leaves the temporary directory empty and the listing unchanged. And a body over the configured cap is 413 `BlobTooLarge`, a body that would break the account quota is 403 `AccountQuotaExceeded`, and neither leaves a file behind. ## Account state, and where it does not map onto the wire `AgentAccount` carries an `AccountState` (`crates/didbot-pds/src/account.rs`): `provisioning` (document serves, nothing else does — see below), `active`, `frozen` (repository readable, writes refused) or `soft-deleted` (document and verification method still serve, repository emptied, name burned permanently). Hard delete is not a state — it removes the account row entirely, key included, which is this server's only honest way to represent "no evidence it ever existed". Two protocol surfaces are supposed to carry this, and neither carries it cleanly. **`com.atproto.sync.subscribeRepos`'s `#account` event.** The lexicon defines `active` as *"the account has a repository which can be fetched from the host"* — fetchability, not authority to write — and `status` is only meaningful when `active` is false (*"if active=false, this optional field indicates a reason"*). `status` uses `knownValues` rather than a closed `enum`, so a value outside the six listed words (`takendown`, `suspended`, `deleted`, `deactivated`, `desynchronized`, `throttled`) is not rejected — but that open union only lets a deployment name new kinds of *inactive*. It gives no way to say "active but not writable", because `#account` does not model write authority at all. | `AccountState` | `active` | `status` | | --- | --- | --- | | `provisioning` | `false` | `deactivated` — never actually seen: `#account` is not emitted at all while an account is `provisioning`, so this row is what `AccountState::sync_status` would answer if something did read it, not something this server has ever put on the wire. See below. | | `active` | `true` | absent | | `frozen` | `true` | **absent — the lossy case.** A frozen repository is still fetchable, so `active` is honestly `true`, and there is no status word for "readable, not writable". `suspended` was considered and rejected: it implies `active: false`, which would tell a relay to stop fetching a repository that is still there and unchanged — a worse lie than saying nothing. A relay or appview watching `#account` alone cannot learn that an account is frozen. | | `soft-deleted` | `false` | `deleted` — the repository genuinely cannot be fetched, which is exactly what the lexicon's own words mean here. | | *(hard delete)* | *(no event)* | Nothing is emitted for a hard delete beyond whatever `#account` a soft delete or freeze already produced. Once an event is on the wire it is out of this server's control — a relay or appview that cached the repository keeps serving what it has — so "no evidence it ever existed" is true of this server's own surfaces and not of the network. Hard-deleting an account that was never soft-deleted first still emits nothing distinguishing it from an ordinary account simply going quiet; there is no `#account` status for "gone without a trace" either, because the lexicon has no word for an account that was never inactive on the wire before it vanished. | Where a genuinely inactive state this server invents does not fit any of the six known values, the right answer is a namespaced custom status (this project has not needed one yet — every inactive state so far is `deleted`), never reusing `suspended` or another listed word for a meaning the lexicon does not define it as. `takendown`, `suspended` and `deactivated` describe moderation and self-service states this server does not have. They are deferred rather than mapped: nothing here produces them today, and what they would mean is a decision to make alongside the mechanism that creates them, not one to guess at now. **`bot.did.listAgents`.** This server's own surface, and where a frozen (or any other) state is actually visible: `AgentSummary.state` carries the same `AccountState` the `#account` mapping is derived from — one type, read by both surfaces, so they cannot silently drift apart the way a hand-maintained second enumeration would. An operator or dashboard that needs to know an account is frozen reads this, not the firehose. There is no per-agent `get` route yet. When there is, it carries the same field for the same reason. ## Legacy sessions, checked against the lexicon schema `com.atproto.server.{create,refresh,delete,get}Session` are new surface, added by [auth-types](../plan/auth-types.md). `getSession` is not in that epic's original list — it was added because nothing else could check the credential end to end: every `com.atproto.repo.*` write route stays public for now (see that epic's own write-surface note), so `getSession` is the one route that actually requires `Authorization` and answers with the account it names, which is what proves a session works rather than merely that it parses. They join `wire.rs`'s suite the same way every other route does: `procedure` validates the request against `createSession`'s vendored `input` schema before it is sent, and `.conforms` validates the response against each method's vendored `output` schema — `createSession` and `refreshSession` both require `accessJwt`, `refreshJwt`, `handle` and `did`; `getSession` requires `handle` and `did` — and this server's response is checked against exactly that, not against a type written from a separate reading of the same document. There is no official interop corpus for these four methods — `vendor/atproto-interop-tests/` vendors `syntax`, `lexicon`, `crypto`, `data-model`, `mst` and `firehose`, and none of those directories is about a session's wire shape. Rather than invent fixtures and call them vectors, this project validates against the lexicon schema alone, the same as any other route here without a dedicated vector file. If upstream ever publishes session vectors, `scripts/refresh-interop-vectors.sh`'s `VENDORED_DIRS` is where a new directory joins the rest. What is checked beyond schema shape, in `didbot-pds`'s own suite rather than here: app passwords are hashed with Argon2id and never stored reversibly; refresh-token rotation issues a new pair and retires the old one; a refresh token already spent, presented again, revokes its whole family immediately, including the pair a legitimate client is currently holding; and a session minted for one account answers `getSession` with that account's own DID, never another's. ## `com.atproto.server.describeServer`, checked against the lexicon schema Discovery: what this server is, before a caller has an account or even a DID to look up. It joins `wire.rs`'s suite the same way every other route does — `.conforms` validates the response against the vendored `output` schema, which requires `did` and `availableUserDomains` and declares `inviteCodeRequired`, `phoneVerificationRequired`, `blobUploadLimit`, `links` and `contact` all optional. This server answers only what it can answer honestly. `did` is the service DID `com.atproto.server.getSession` and the DID document already agree on. `availableUserDomains` is read out of `Registry::available_user_domains`, which is the deployment's real `didbot_identity::ZoneRegistry` when one is configured and the single configured zone otherwise — never a fixed string — so a several-zone deployment reports every zone it mints handles under. The five optional fields are left off the response entirely rather than filled with an invented value: there is no invite system, no phone verification, no configured blob limit surfaced through `Registry`, and no configured policy or contact document to point at. A test asserts the domain reported is the zone the test server was actually built with, not a value the response type happened to default to. ## The export, checked by something that is not us A CAR file is the one thing this server produces that a stranger can check without trusting it, so it is checked by a stranger. `goat` — the CLI from `bluesky-social/indigo`, a Go implementation sharing no code with this one — reads an export from `com.atproto.sync.getRepo`, parses the commit, walks the Merkle search tree and lists every record. `docs/running-locally.md` has the commands. That check is manual, because vendoring a Go toolchain into `cargo test` is not a trade worth making. What runs unattended is the pair of things it would otherwise be the only evidence for. `didbot-repo/tests/car_framing.rs` walks the CARv1 framing back apart — every section as long as its prefix claims, every block hashing to the CID in front of it, a truncated file refused and a flipped byte breaking exactly one block — without decoding anything. And `didbot-serve`'s `sync` tests write records through the HTTP surface, fetch the CAR through the HTTP surface, decode the commit block with a reader written from the specification rather than from this project's encoder, and verify its signature with a key taken out of `/.well-known/did.json`. That last one is `plan/pds-writes.md`'s exit criterion, run on every commit. The same module does it for a proof. `com.atproto.sync.getRecord` answers with the blocks on one key's path through the tree, and the test walks them the way a stranger would: check the commit's signature with the published key, hash each node block and require it to be the one the previous step named, follow the link, and end at the key or at the node the key would have to be in. Two sabotage cases run alongside it, because a verifier that never refuses anything is not a verifier — a path with a node removed is caught by the missing block, and a node lifted out of a different repository and filed under the CID the path named is caught by the hash. The node reader, like the commit reader, is written from the specification. ## `com.atproto.sync.subscribeRepos` The method is implemented. It is a WebSocket at `/xrpc/com.atproto.sync.subscribeRepos`, and every message is one binary frame holding two concatenated DAG-CBOR values: a header, `{"op": 1, "t": "#commit"}` for a message or `{"op": -1}` for an error, and then the body. `?cursor=` is the bare integer the lexicon declares. This is not `/firehose`. That is this project's own Server-Sent-Events record stream, it carries records as JSON rather than blocks, its cursor is `:` rather than an integer, and the two have separate sequence spaces and separate replay buffers. Neither is derived from the other. Both now announce a removal; `docs/running-locally.md` has `/firehose`'s own shape for one. ### What a `#commit` carries Everything the lexicon marks required. `blocks` is a CARv1 file rooted at the new commit and holding a covering proof of every key the commit touched: the union, over those keys, of the Merkle search tree nodes on each key's own walk from the root and on the walks to its immediate neighbours either side — see `didbot_repo::Repository::covering_proof` and, under it, [`crate::data::mst::Tree::covering_proof`] — plus every block the commit created, which is not always a subset of that proof. `since` is the revision the commit replaced, null for a repository's first. `prevData` is that revision's tree root, so a consumer can check that the operations invert back to the tree it holds — without holding it: a relay is not a consumer that already applied the previous commit, and a diff against blocks it does not have would name nothing it could check. A covering proof is more than the diff `com.atproto.sync.getRepo`'s `since` answers for the same two revisions. That diff, and the plain created-only diff a naive implementation of this frame would carry, both answer "what did this commit add" for a reader who already holds everything before it. A relay holds nothing before it: `com.atproto.sync.subscribeRepos`'s own verification proposal has it invert every operation in the frame, starting from the new tree and reconstructing `prevData`, to check the frame's claim without trusting the frame's own node list. A split or a merge moves a node neither side's own key touched — a node the commit did not create and a created-only diff therefore does not carry — and inversion needs to read it. `crates/didbot/tests/conformance/firehose.rs`'s `a_created_only_diff_is_not_a_covering_proof` measures the gap directly against two of the six upstream commit-proof fixtures. ### Removals A `deleteRecord` is announced exactly as a write is: the same `#commit`, over the same covering proof, with one `op` whose `action` is `delete`, whose `cid` is null as the lexicon requires, and whose `prev` is the CID the key held. `applyWrites` announces its deletions the same way, inside the one commit the batch makes. `prev` is on every op that is not a create — an update carries it too — and it is the field the lexicon calls "required for the inductive firehose". A consumer inverting a commit puts a deleted record back and restores an updated one's previous value, and nothing else in the frame names either. A delete op without it is a frame a relay can parse and cannot check. The covering proof is right for a removal for the same reason it is right for an addition, and by the same code: the proof is taken over the tree *as the commit leaves it*, over every key the commit touched, and a removed key is still a key with a walk and with neighbours either side — its walk now proves absence rather than presence. Upstream's own commit-proof fixtures cover removals (their `dels`), and `crates/didbot/tests/conformance/firehose.rs`'s `each_commit_proof_carries_the_blocks_upstream_requires` runs adds and deletes through the same union. `crates/didbot/tests/subscribe_repos.rs`'s `a_deletions_own_frame_inverts_to_prev_data_from_its_own_blocks_alone` then does what a relay does with a real server's own frames: deletes twenty records one at a time from an out-of-order tree and inverts each removal back to `prevData` from that frame's blocks alone. A consumer that was disconnected across a deletion sees it when it resumes: the deletion is a numbered frame in the one sequence space, so it falls in the gap between the consumer's cursor and now like any other commit, and arrives before the commits after it, once. `a_consumer_disconnected_across_a_deletion_sees_it_on_resume` is that assertion. A consumer whose cursor the buffer can no longer reach is told so — `OutdatedCursor`, above — rather than resumed silently, so a missed removal is never a thing a consumer believes it has and has not. **An account deletion is not a set of record deletions.** Removing an account drops its records without announcing one `#commit` per record; what goes out is the `#account` above, `active: false, status: "deleted"`, which says the whole repository is no longer fetchable. Per-record frames would be worse than useless: they would name commits over a repository that no longer exists for a consumer to reconcile them against. A hard delete announces nothing beyond whatever a soft delete already did — see the account-status table above for why. The lexicon's three deprecated-but-required fields are written at their fixed empty values and never anything else: `rebase` and `tooBig` are always false, and `blobs` is always empty. Nothing here produces a rebase, nothing splits a commit for size, and no record written through this server references a blob the same commit uploaded. `prevData` is absent, rather than wrong, for the first commit a repository makes after a restart that replayed a log written before the tree root was kept. Every later commit carries it. ### Sequence numbers across a restart The lexicon's cursor is an integer with nowhere to name the run of the server that minted it, so the number itself survives a restart: a run reserves a block of numbers in the write-ahead log before handing any of them out, and the next run resumes above the last reservation. A crash therefore *skips* numbers rather than repeating them. The lexicon requires the sequence to increase and not to be gapless, and a consumer whose cursor lands in a skipped range is behind a restart that emptied the replay buffer anyway. A cursor the in-memory replay buffer can no longer reach — because the consumer fell behind, or because the server restarted — gets `#info` naming `OutdatedCursor` and then everything the buffer still holds. Saying so is the point: the consumer has a hole only a re-read of the repositories can fill. A cursor naming a number this server has not assigned is the lexicon's `FutureCursor` error, and the connection closes. ### `#identity` and `#account` `#identity` goes out once, at provisioning, and only once the account is `AccountState::Active` — later than the moment `/.well-known/did.json` starts answering for it, which is when the account is created, `AccountState::Provisioning`. The document resolving and the repository being complete are two different facts, and `#identity` says the second: "this DID resolves *and its repository is fetchable*, go and read it". Sending it any earlier — while the account exists but its initial records are still landing — would be a lie a consumer acts on: it fetches the repository, finds it empty or partial, and caches that. See `AccountState::Provisioning` and `Provisioner::announceable`. It carries the handle this deployment issued, when it issued one. `#account` goes out twice in an account's life: `active: true` beside the `#identity` above, and `active: false, status: "deleted"` once the hostname is withdrawn, the records are dropped and the store no longer answers. Both share one sequence space and one replay buffer with `#commit`, so a consumer resuming from a cursor sees every event of all three kinds after it, interleaved, in order, once each. The ordering at provisioning is fixed: `#identity`, then `#account`, then the commits for the records the repository is born holding — one commit per record, each with a real tree delta and a `prevData` genuinely equal to the tree before it, rather than the two records landing in one uncommitted repository and the second commit's `prevData` claiming a change that never happened. See [agent-accounts](../plan/agent-accounts.md)'s `Provisioning` item for the live-relay failure this ordering fixes. ### The account statuses this deployment does not have The lexicon's `status` has six known values and this server produces one. That is not a gap to be filled later — five of the six describe states nothing here can be in. There is no takedown, no suspension and no deactivation. The two things that can move without an account leaving `active` are its retention and `AccountState::Frozen`, and neither is one of the six. **Retention** is how long an account keeps resolving once it stops being used, and a **pin** is that value set to never; it changes nothing a consumer can observe, because a pinned account and an unpinned one serve the same repository to the same readers. A **freeze** refuses writes and leaves the repository readable, which is not `suspended` (the account is not suspended from anything a relay can see) and not `throttled` (nothing is being slowed); `stateChanged` on `/events` carries it instead. Mapping either onto the closest-looking value would tell a relay something untrue, so neither is announced at all, and this paragraph is where that decision is written down instead of being inferred from an enum with one variant. If this deployment ever grows a real takedown, it gets a real `takendown`. `desynchronized` is likewise never sent, because nothing here produces the condition: a repository is derived from its records on every read, so it cannot be behind a commit it has already announced. ### `excludeIdentity`, `excludeAccount` and `excludeCommit` Read at connect time, alongside `cursor`, and applied as a per-connection delivery filter. A frame of an excluded kind is not sent to that consumer and is not replayed to it either. Sequence numbers are **not** renumbered per consumer. Every consumer shares one global sequence space whatever it excludes, so excluding a kind leaves visible gaps in the numbers a consumer sees — and that is the property that makes the parameter safe: a cursor that skipped past frames the consumer never received still names a point in the one stream, so reconnecting with it is neither an error nor a replay of what was excluded. A value of `false` or `0` turns the exclusion off; the parameter present with any other value, or with none, turns it on. Anything unparseable reads as false, which sends the consumer more than it asked for rather than less. ### What is still not implemented `#sync` is not sent. It exists to hand a consumer a repository's current state without the commits that got it there — for recovering from a broken stream or a data-loss incident — and this server has no path that produces one. `#handle` and `#tombstone` are not in the current lexicon union at all, so there is nothing to implement for them. ## `com.atproto.sync.requestCrawl` and `notifyOfUpdate`, called outward rather than served These are the two `com.atproto.*` methods in this project that are a *client* of, not routes it answers. A relay learns of a new personal data server by opening a `subscribeRepos` connection to it, and nothing makes a relay dial an address it has never heard of — `requestCrawl` is that announcement, and without it the subscribeRepos implementation above is a stream nobody is listening to. `notifyOfUpdate` is upstream's deprecated predecessor for the same request, kept because its own description names a recovery case `requestCrawl`'s does not: resuming after a gap left by a dropped subscription. Neither is in `ATPROTO_METHODS` or `auth::ROUTE_CREDENTIALS`, and `wire.rs`'s harness — which drives routes this server serves — does not exercise either. `didbot_serve::relay` is the client: `RelayClient::request_crawl` and `RelayClient::notify_of_update` post the lexicons' identical `{"hostname": ..}` body to `{relay}/xrpc/com.atproto.sync.requestCrawl` and `{relay}/xrpc/com.atproto.sync.notifyOfUpdate` respectively. Neither is called automatically. The only caller is `didbot_serve::estop_admin`'s `ANNOUNCE`/`NOTIFY` commands, reached over the same local unix socket as the e-stop's own `RELEASE` — an operator asks explicitly, and the socket reply reports what the relay said rather than a log line the operator has to go find. Design decisions, each stated once and checked by a test in `relay.rs`'s and `estop_admin.rs`'s own modules: | Question | Answer | | --- | --- | | When does either fire? | Only when an operator sends `ANNOUNCE` or `NOTIFY` over the e-stop admin socket. Nothing calls either at startup, on a timer, or on any other schedule. | | What credential does it need? | None, the same as every other command on this socket. The socket is `0600` inside a directory the server owns and re-tightens to `0700` on every start, so reaching it at all means host access — which is the whole authorization story, and a stronger one than a shared secret readable from the same filesystem. There is no operator credential anywhere in didbot: a server learns which DID operates it by reading `bot.did.operator` out of the operator's own repository, and is never handed anything to check. | | What does the reply distinguish? | Three outcomes, as `OK `: `{"result":"accepted"}`, a relay that answered but refused (`{"result":"refused","status":..,"body":..}`, covering `HostBanned` and anything else a relay might say), and a relay that could not be reached at all (`{"result":"unreachable","error":..}`). An operator who ran the command sees which one happened. | | Which relay? | `[relay].hostname` in `didbot-config`, or `--relay-hostname` directly — never hardcoded. `None` (the default) leaves `ANNOUNCE`/`NOTIFY` with no relay to reach, reported as `ERR`. | | What about `*.localhost`? | `didbot-dev.rs` never wires the commands up for a loopback zone, whatever `--relay-hostname` says — the same `LoopbackDns::accepts(zone.host())` check that already gates `--tls acme` and the placeholder-secret refusal, because no relay can reach a loopback address. | **What is verified without a real relay:** the request body (`relay::request_crawl_body`) is asserted equal to `{"hostname": ..}`, which is exactly what both vendored `input` schemas require. `RelayClient` is exercised against a real, ephemeral loopback HTTP server standing in for a relay — not a hand-rolled mock of what `reqwest` does — covering a relay that accepts, one that answers a non-2xx status (the `Refused` path), and one that is simply not listening (`Unreachable`, the connection-refused path), for both `request_crawl` and `notify_of_update`. `estop_admin.rs` covers the same three outcomes again through the actual `ANNOUNCE`/`NOTIFY` dispatch, plus the credential checks and a regression test that a relay wired up at `spawn` time hears nothing until a command is actually sent. **What is not verified, and cannot be from here:** everything about a real relay's actual behaviour — whether Bluesky's own relay, or any other, accepts this server's announcement, opens a `subscribeRepos` connection against it, and advances past the initial cursor. That is [federation](../plan/federation.md)'s question, not this suite's, and needs a live deployment and a real relay to answer. There is also no durable, public alerting for either call's failure — an operator sees the outcome in the moment they asked, over the socket, but nothing records it past that; [`plan/alerts.md`](../plan/alerts.md) is where a repository-backed version of this would eventually live, and this branch does not build it. **There is no reverse of either call.** No `com.atproto.*` method un-announces a host; a relay that should stop crawling this server is de-listed out-of-band, by whatever process that relay runs, not by anything this server can send it. This is the same asymmetry the hard-delete note above records one layer down: this server has no way to make a relay that already crawled it forget that it did. ## Everything upstream ships, and the one line that is refused Every directory in the interop corpus is vendored and every file in every one of them is run. There is no suite waiting on anything, and this section used to be a list of four. `mst/` was on that list as blocked on a Merkle search tree, which was true of the directory and true of none of the files in it: everything upstream ships there is a question about a pure function. `data-model/` was blocked on an encoder. `crypto/`'s P-256 half was blocked on a second curve. And `firehose/` was left out on the reading that it tested a frame format — the one reading that was simply wrong, since its only file is a set of tree commit proofs and has nothing to do with how a frame is packed. That mistake cost the strongest available check on the tree for as long as it stood, which is the argument for vendoring a directory before deciding what is in it. `com.atproto.sync.subscribeRepos` is served, and no vendored vector asks about it — see the section below for what it carries and what it does not. One candidate does not pass, and it is upstream's file disagreeing with upstream's specification rather than this server disagreeing with anything. `nsid_syntax_valid.txt` holds an NSID whose domain authority is 283 characters; the NSID specification caps an authority at 253, which is the DNS limit, because the authority is a hostname with its segments reversed. Accepting it would mean accepting a lexicon namespace under a domain that cannot exist. Upstream's own README says the specs are the authority where the two conflict and that such files should be corrected, so the candidate is excused by name with that reason attached, and two things fail if it ever moves: the excusal goes stale if the line leaves the file, and a separate test measures the authority and fails if it comes back within the cap. What the `lexicon/` vectors check is unchanged by the collection set opening. They ask whether a record satisfies a document, and every one of them supplies its own document — so they run against the *explicit* path, which is what a held lexicon still gets. The stance decides which records reach that path, not what it does with them. What no vendored suite asks about is the other branch: a record stored under a lexicon nobody here holds, which is checked instead by the round trip through `getRecord` and by the bytes in the export (`export::tests::a_foreign_record_is_in_the_tree_byte_for_byte`). `lexicon/`, `syntax/cid_*`, `syntax/datetime_*`, `syntax/language_*` and `syntax/uri_*` were on this list until record validation arrived. Each was blocked on the same thing: there was no boundary at which a record's fields were checked, so there was nowhere for a datetime or a language tag to be wrong. `cid` is the odd one out and is worth stating plainly — this server now emits content identifiers, in a CAR export, over a signed commit, but not on the write path: `RecordStore::put` still returns an rkey alone. It validates a client's CIDs because a *client* can write one into a record, which is a different question from whether we mint any — and that is why the suite still runs against the lenient string check in [`crate::pds::format::cid`] rather than against the strict decode in [`crate::data::Cid`], which refuses every multibase but the one an encoder can reproduce byte for byte. `datetime_parse_invalid.txt` and `language_parse_invalid.txt` are a third category upstream, and worth naming because the distinction outlives the two files. They hold candidates whose *syntax* is fine and whose *meaning* is not: `1985-13-12T23:20:50.123Z` is a well-formed timestamp in month thirteen, and `de-DE-1901-1901` is a well-formed language tag that repeats a variant subtag. A parser that only checks shape will accept both, and a suite that ran only the syntax files would look green while it did. `Suite::also_invalid` names the extra file, and both validators run against all three of theirs. ## A second corpus: sixteen thousand tree diffs Everything above is Bluesky's corpus. There is a second one now: David Buchanan's `mst-test-suite`, which asks the question the official vectors do not. Given two Merkle search trees, what changed? That is the question a repository server answers on every commit. A firehose message is a diff — these nodes were created, these were deleted, these records changed, and here are the blocks that prove it — and nothing here checked such an answer against another implementation. The official corpus cannot: it has one file about whole trees, `firehose/`'s commit proofs, and six cases in it. This corpus has 16,384. Upstream mined seven record paths whose heights build a perfect binary tree, deleted every subset of the seven to get 128 trees, and made a case out of every ordered pair. The awkward shapes are not rare in there. A node splits, two nodes merge, the root changes layer because the only key holding it up was deleted, the tree empties, the tree fills. A suite written by hand holds the cases its author thought of. Each case names two CAR files and lists five results. Four are checked in `crates/didbot/tests/conformance/mst_diff.rs`: | Field | What it is | | --- | --- | | `created_nodes` | node blocks the second tree has and the first does not | | `deleted_nodes` | the same the other way round | | `record_ops` | the records that changed, with the value on each side | | `proof_nodes` | the blocks proving those changes against the second root | The fifth, `inductive_proof_nodes`, is checked in `mst_inversion.rs` and is described below. `firehose_cids` is checked by nobody: upstream sets it to the string `TODO` in every case, and the refresh script refuses to vendor a corpus where that has changed without a runner to read it. ### What it took to run one case at all Three things this workspace did not have, because everything before this wrote blocks and never read one back. A DAG-CBOR decoder, [`crate::data::dag_cbor::decode`]. It is a stricter piece of code than the encoder beside it: plain CBOR has five widths for an integer, two length forms for every string and any order for a map's keys, and all of them decode to something. A decoder that accepted them would hand back values whose re-encoding is a different block with a different CID, which is a peer handing this server two spellings of one record and getting two identities for it. Every non-canonical form is an error here, and the tests assert the round trip rather than the reading. A CAR reader, `didbot_repo::car::read`. It hashes every block and checks it against the CID it was filed under, because a CAR file is exactly the claim that these bytes have those names. And a tree that can be loaded rather than built. [`crate::data::mst::Tree`] is this project's own tree: a pure function of a whole leaf set, immutable, sure of its own blocks. [`crate::data::mst::Store`] is the other situation — a pile of blocks that arrived from elsewhere and a root somebody claims. Nothing in it is trusted, so every walk over it can fail: a link may name a block that is not there, a subtree may claim a layer it is not on, a node may not decode. ### The rule that decides how big a proof is `proof_nodes` is one walk per changed path over the tree *after* the change: an inclusion proof for a record that is now there, an exclusion proof for one that is now gone. The walk stops as soon as it is below the layer the path's own height puts it on. That early exit is not an optimisation. A path can only ever be an entry at the one layer its hash puts it on, so a walk that carried on would be reading blocks that cannot hold the answer, and the exclusion proof made of them would be larger while proving the same thing. Removing the rule and running the suite fails 304 cases — all of them proofs that are correct and too big. ### Running a commit backwards `inductive_proof_nodes` is the odd field, because it is not a claim about a diff. It is a claim about what a *verifier* has to read. Bluesky's sync iteration proposal has a firehose consumer check a commit without holding the repository: take the tree after the commit, invert every operation in the message — a create becomes a delete, a delete becomes a create of the previous value — and require the root it ends on to be the root the commit says it came from. The blocks that inversion reads are the blocks the producer had to put in the message, and that set is what the field records. It is not derivable from the diff. Which nodes get read depends on how a tree rewrites itself when a key is inserted into a full node or removed from one that then merges with its neighbour. Checking it needs a tree that can be mutated one operation at a time, which this workspace deliberately does not have. [`crate::data::mst::Tree::build`] is a pure function of a leaf set on purpose: a second way to arrive at a root is a second thing that can be wrong about it, and a repository this size gains nothing from an incremental path. So the rewriter lives in the test that needs it, `mst_inversion.rs`, and not in the crate, where it would be a write path nothing writes through. It is written to match upstream's `atmst` step for step — where it descends, where it splits and merges, and which nodes it reads on the way — because the read set is what is being compared. Two rewriters can reach the same root having read different blocks, and which of those sets is minimal is not a question these vectors answer. The test asserts both things anyway, and in that order: the root first, because that is the property the proposal specifies and a rewriter that lands elsewhere has verified nothing, and the read set second, because it is what catches a rewriter that is correct and reads the whole tree to be so. ### Why this corpus is not vendored Everything else this suite reads is in `vendor/`, and the argument for that is at the top of this page. It does not extend here. The interop vectors are 96kB and the lexicon documents 81kB; this corpus is 16,384 files and 24MB, which is more than the rest of this repository put together. Carrying it would make every clone, every search and every pull request diff pay for a test corpus — and the forge this project lives on refuses a pull request whose diff is larger than it will send, which is how the size stopped being theoretical. The offline argument that justifies vendoring the interop vectors does not carry the weight either. Nothing runs this suite unattended: there is no CI here, so the only thing that runs it is a person at a terminal, and that person has a network. So `scripts/fetch-mst-vectors.sh` clones the suite, and the runner reads it in upstream's own layout — walk `tests/` for `.json` files and assume nothing about the arrangement, which is what upstream's README asks a runner to do. It is cached **per machine**, at `$XDG_CACHE_HOME/didbot/mst-test-suite`, beside the configuration and state this project already keeps under the same directories. Per checkout was the obvious place and the wrong one: the corpus is immutable upstream data that varies with nothing, this project is worked from several worktrees at once, and each of them wanted its own 72MB and failed four tests until somebody fetched one. A clone left in a checkout's own `.cache/` is moved rather than re-downloaded on the first run. `MST_TEST_SUITE` points at a corpus of your own and wins over both. One clone shared by every checkout means a plain run that fast-forwarded it would move the corpus under whatever else is reading it, so a corpus already there is left alone and reported. `--update`, or a ref, changes it on purpose. The CAR files are read from the same clone: 128 of them, 63kB, one tree apiece and no record blocks in any of them, because a diff never looks inside a record. A corpus that is not there **fails** the tests that need it, naming the script to run. It does not skip. A test that passes without having checked anything is the failure this whole page exists to prevent, and it is worse than a red line telling somebody to run one command. ## Adding a suite One line, if a validator exists: `Suite::new("cid").run(cid::validate)`, where the name is the vector files' shared prefix and the argument is anything from `&str` to a `Result`. If a validator does not exist, that is the work — and it belongs in the crate that owns the concept rather than beside the test, or the boundary that actually needs it still has nothing to call. `Suite::run` reports every failure in one panic, naming the file, the line, the candidate and what the validator said. A vector file is a few hundred near-identical strings, so a report that says `handle_syntax_invalid.txt:52` is worth a great deal more than one that says `assertion failed`. ## When a vector fails A failure is a finding about this implementation, not a problem with the harness, and the first move is to fix the implementation. Where a vector genuinely does not apply, `Suite::excused` takes the candidate and a reason, and both appear in the source. Deleting a line from a vendored file would make the copy diverge from upstream silently and turn the next refresh into a conflict nobody can adjudicate. There is one excused candidate today. `nsid_syntax_valid.txt` contains an NSID whose domain authority is 283 characters; the NSID specification caps the authority at 253, which is the DNS name limit. Upstream's own README says the specification wins where the two disagree, and accepting the candidate would mean accepting a lexicon namespace under a domain that cannot exist. ## Whitespace is a test case Several vectors are invalid for no reason other than a leading or trailing space: `one.two.three ` is in `nsid_syntax_invalid.txt` on that basis alone. The whitespace-fixing hooks in `prek.toml` therefore exclude `vendor/`, and the runner never trims a candidate. A fixer reaching those files would edit test cases into passing, and the diff would look like tidying.