--- id: onboarding title: An operator establishes a server before the server can establish anyone else status: open crates: [didbot-pds, didbot-serve, didbot-agentd, didbot-lexicon, didbot-operator] dependsOn: [ownership, deploy] exitCriterion: > An operator takes a freshly applied deployment — instance up, no accounts, nothing configured beyond Terraform's defaults — through generating and installing the real secrets, an operator attestation step, and the first-run operator handshake, then points one agent host's harness at it and gets a working, provisioned account, using only documented steps and no fact this file invents. --- # onboarding Every other epic in this directory describes a system already running: an account exists, a zone resolves, policy is being polled. Nothing describes the hour before any of that is true — Terraform has applied, an EC2 instance is up, and the operator has not yet done anything the server would recognise as them. [dev-setup](dev-setup.md) covers the symmetric gap on the other side, wiring one machine's harness to a stack; this is the half it does not cover, and the half that has to happen first. ## The bootstrap paradox **A provisioning request carries a host's claim.** `bot.did.createAccount` mints only under a claim signed with the node key of a host the operator has vouched for — see "The claim" below. The key is generated on the host and never leaves it, so there is no shared secret to distribute or get wrong. So the paradox is not what it was. It was: the secret that lets an operator act on a fresh deployment must reach the host before the deployment can be acted on. There is no such secret now, and the question that replaces it is narrower — a fresh deployment has to reach a state its operator can claim, using only what the deployment already has. The server mints its own identity through an administrative path that never travels the xrpc write path, so its `did:json` and signing key exist before it serves anything; the operator then claims it from their own machine with `didbot operate`, and the server observes that claim through its own poll. See "The claim" below and `docs/server-lifecycle.md`. What onboarding still owes is the sequence, not the mechanism. The sequence itself is `crates/didbot-onboarding/src/step.rs` — one ordered list, held as data, with each step's checks, the step it waits on, and the remedy beside it. This file points at that list rather than carrying a second copy of it. - [ ] **Walk a first-time operator along it.** `didbot operate --check` runs the four steps a claim waits on, and no command runs the rest of the list. Nothing yet takes somebody through the steps one at a time, against a real deployment, saying what to run on which machine. - [ ] **Say how an operator tells whether the server has seen the claim yet.** The poll is the only thing that moves the server to `claimed`, and its interval is the delay an operator experiences as "nothing happened". - [ ] **Say what a deployment does before it has any operator policy.** [policy](policy.md) puts a pre-flight between provisioning and `claimed`, so an instance is never online and unpolicied. The commands for that half belong here too. ## The claim The top edge of the tree [ownership](ownership.md) describes: the record that says a human operates a server, the command that writes it, how the server decides it has been claimed, and what it does when the claim disappears. [docs/operator-verification.md](../docs/operator-verification.md) is the edge and its timings as a reader meets them. A server nobody has claimed is the honest state of a fresh `did:web` identity, not a lesser version of a claimed one: a reader resolving it sees an empty repository and no operator record. The operator runs `didbot operate ` on their own machine. The command waits for the server to answer over TLS, resolves its `did:web` document, and cross-checks that read against `describeServer`'s own `did`. On its own that cross-check only confirms the server agrees with itself across two routes. What makes the handshake sound is everything beneath it: the operator delegated this hostname's zone, `didbot-tls` proved control of that zone by passing ACME DNS-01, and the certificate that produced is what TLS is now proving belongs to whoever answered. **The operator typing a hostname they themselves delegated is the root of trust; the DID cross-check is a consistency test on top of it.** Having passed both, the command writes `bot.did.operator/` into the operator's own repository. At no point does the operator hand the server a secret or download the server's keys: the server reads the record with no credential, and cannot write there. The claim is not an authorization gate. The operator is configuration, `didbot-pds`'s `--operator`, known at launch; the record is the operator's public statement, for third parties, that they run this server, and what the poll enforces is that they keep it standing. A server nobody has ever claimed is `OperatorTransition::StillUnclaimed`, the ordinary state of a fresh identity: the grace window does not apply, because there is nothing yet to lapse from. **Poll, not watch.** `com.atproto.sync.subscribeRepos` is a whole-host firehose with no per-repository filter, and the operator's own account usually lives on a PDS somebody else runs, so subscribing to all of it to watch one repository is not a default this project can ship. A filtered relay stream is the real "watch", and it depends on infrastructure this deployment does not control. The decision is poll-first: the poll is the fallback when a relay is down, and an emergency-halt path cannot depend on a third party's infrastructure alone. - [ ] **The alarm goes through [alerts](alerts.md).** That epic has no channel yet to write into, so `operator_poll` logs a `tracing::warn` on the transition into a lapse; an operator not watching this process's own logs sees nothing. - [ ] **What happens when the operator's own PDS is down during first run.** The claim is a one-time operation and it blocks on exactly the server the operator has the least control over. Nothing here says what an operator does in that window beyond wait and retry. - [ ] **Operator rotation, and whether more than one operator is possible.** Every record and every mechanism above assumes one operator DID. What replacing it looks like, and whether two simultaneous claims are ever both valid, is undecided. The lost-access failure mode below is the involuntary half of the same question. ## What "first log in" means here There is no human user account on this server, so "log in" is not the right frame and the epic should stop implying one. What exists instead: - `--operator` — a DID string the server writes into every registration record as who is answerable for that account. It is not verified at startup: any string parses. - The `bot.did.operator` handshake — the operator writes a record into their *own* repository and this server reads it (`crates/didbot-serve/src/operator_poll.rs`). Nothing is presented to the server and nothing is compared by it. There is no operator credential, and no HTTP surface is gated on one — see [auth-types](auth-types.md). - The operator's sign-in: atproto OAuth against the operator's own server, asking for `atproto` and nothing else (`crates/didbot-serve/src/operator.rs`). This server compares the DID it signed in against `--operator` and issues its own session, which the dashboard's operator routes take, from a browser or from the `didbot` verbs that drive them. Neither of the first two is the full handshake [ownership](ownership.md) asks for. That epic's own checklist says the agent-side operator claim exists (written by the server at provisioning) and the human's vouch of the server is meant to come from the operator's own repository — but its "three statements" item is still open, including *"the server answers its controlling DID on an unauthenticated status endpoint,"* which is the piece that would let a stranger, or the operator themselves, confirm the server agrees who its operator is. Until that exists, nothing on first boot confirms `--operator` names a DID that has actually agreed to be named — the server just asserts it, on one process argument, forever. - [ ] **A first-run check that `--operator` resolves to something real** before the server starts minting accounts against it — at minimum, that the DID document exists and is fetchable. `--operator` today is `Zone::new`'d alongside the zone but is never itself validated. - [ ] **The unauthenticated status endpoint** ownership.md names, so an operator (or anyone) can ask a running server "who do you say operates you" and compare it against what they meant to configure. This is [ownership](ownership.md)'s item to build; onboarding is where a first-run check would use it. - [ ] **Write down, plainly, that there is no bidirectional claim yet.** Until the status endpoint above exists, "first log in" is `--operator` plus a `bot.did.operator` record this server reads one way, and that is a weaker claim than the rest of `plan/` sometimes assumes. Say so somewhere a new operator reads before they conclude the claim is already checked because the flag is there. ## Required operator attestation The owner's framing names this explicitly, and it is worth being precise about what it is not: [ownership](ownership.md)'s check at creation is about the *account creating* — proving it by its credential records before anything is created beneath it. The operator is a different subject. Nothing in this codebase today asks the operator to prove anything *to the server over HTTP* at all: the handshake runs the other way, with the server reading a record the operator wrote in their own repository. That is a claim about a DID rather than possession of a string, which is the right direction — but this server never checks a signature over a challenge it chose, so it is a claim it reads, not one it verifies interactively. What an operator attestation step would need to establish, before the server mints its first account: - [ ] **That the handshake actually completed** — i.e. this server has read a `bot.did.operator` record naming it, rather than still waiting for one. This is the loud version of the bootstrap-paradox item above, scoped to "before minting begins" rather than "eventually." - [ ] **That the `--operator` DID is one the operator controls**, which is a real proof — sign a nonce with the DID's own key, or an equivalent — not merely a string on the command line. Nothing here does this yet; it is the missing half of "first log in," restated as a checkable step rather than left implicit. - [ ] **Decide whether operator attestation is a one-time ceremony or a standing credential.** A node's attestation is spent once, at provisioning, and the node's ongoing claim to be trusted rests on holding a node credential afterward ([node](node.md)). The operator has no equivalent today, and — since there is deliberately no operator credential at all — no "hold a narrower credential afterward" half to build one out of. [ops-dashboard](ops-dashboard.md)'s operator sign-in is the candidate mechanism. Whether the operator needs a standing credential at all, or the read-only handshake plus host access is the whole story, is this epic's to decide, not to assume. ## The first agent Once a server is up, the handshake has completed, and `--operator` is established, a person still has to get one agent host talking to it. The machine half is the `didbot-claude` plugin and `didbot-agentd`, which takes the server it mints against from `DIDBOT_PDS`. Two things about that are worth knowing: - **The server is an environment variable, set by hand.** A harness client falls back to `http://localhost:3000` when `DIDBOT_PDS_URL` is unset, and the daemon refuses to start without `DIDBOT_PDS`. There is no secret to set alongside either: provisioning takes the host's claim, which `didbot-agentd` signs. - **A session with no usable account provisions one on its first tool call**, so once the variables above are set correctly, that call is what actually creates the account — there is no separate "first login" moment to perform. - [ ] **Documented steps for wiring a harness against a server that is not `didbot-pds`'s own defaults**: the URL, and the host's `become` against that server. - [ ] **A host says which server it is configured for**, so a harness silently pointed at a stale or wrong deployment is caught before the first session rather than after a refusal with no context. ## Failure modes worth naming These are not hypothetical: each is a documented gap or an enforced-but-silent default above, restated as what an operator actually sees when it happens. - [ ] **A server that started but was never given a real secret.** Cannot happen for the attestation secret against a non-`.localhost` zone — `didbot-pds` refuses to start. Can absolutely happen for the operator secret, which only warns, and the warning is the only signal until an e-stop needs releasing. - [ ] **DNS delegated but not propagated.** What it looks like from the operator's seat now has an answer: `didbot operate --check ` reports `zone-delegation` and `dns` separately — which zone above the hostname answers with nameservers, and whether the hostname itself resolves through the operator's own stub resolver — and blocks the checks downstream of a failure rather than reporting three more failures for one cause. What is still open is the second half — how long to wait before it is a bug rather than propagation. Nothing names a number, and a check that has failed for an hour looks exactly like one that has failed for a minute. - [ ] **A certificate that has not issued yet.** `--tls acme` is wired now — `didbot-pds` refuses at startup without `--data`, a real zone and `--route53-zone-id` (`crates/didbot-serve/src/bin/didbot-pds/`), and `infra/pds/templates/user_data.sh.tftpl`'s `ExecStart` passes it. What is still unverified, per [deploy](deploy.md)'s own Done entry, is the protocol exchange itself against a real ACME directory — nobody has run it from an environment with the network access to try. An operator's first boot is where that gets exercised for the first time, and the failure mode worth naming is: a startup that passes every local check above but still cannot complete a DNS-01 challenge against Let's Encrypt, for a reason none of those checks can catch in advance. `didbot operate --check`'s `tls` check is what now catches it — a real HTTPS request from the operator's machine, so a certificate that never issued, was never loaded, or covers the wrong name all fail the same way and say so. It does not make the exchange any more verified; it makes a failed one visible from the seat that matters instead of silent. - [ ] **An operator who has lost access to the DID that operates a server.** There is nothing to rotate — the server holds no credential — but the recovery question is real and unanswered: a server watching a `bot.did.operator` record in a repository its operator can no longer write has no path to being told about a new one, short of restarting it with a different `--operator`. Whether that restart is the answer, or something narrower is, is undecided. ## Done - [x] **A server lifecycle, so "not finished being born" is a state rather than a latch.** `didbot_pds::ServerState` — `booting`, `provisioning`, `unclaimed`, `claimed` — with a capability policy per state and one enforcement point per capability; `docs/server-lifecycle.md` carries the diagram and the table, generated from the types and checked against them by a test. The machinery is `didbot-fsm`, a reusable core (states with a consumer-shaped policy, transitions taken atomically under one lock, one-way capabilities claimed under a permit, persistence as a hook); this lifecycle is its first consumer. The axis is **reads against writes**. `unclaimed` — where a deployment spends most of its early life, and where a server whose claim lapsed returns to — reads everything `claimed` does and writes nothing: no agent accounts, no records, no blobs, no outward announcements. Withdrawing reads would protect nothing and would retroactively break every signature made under a `did:web` document this server has already served. Three things it makes structurally true. A second self-bootstrap keypair cannot be minted, because `ServerLifecycle::begin_bootstrap` is the only source of the `BootstrapPermit` a mint takes by value and it shuts the door in the same critical section. Never-claimed and lapsed are one state with one recovery — run `didbot operate` again — so the bootstrap flow is exercised on every re-claim rather than once per server lifetime. And the emergency stop cannot be thrown before an identity exists: `ServerPolicy::may_halt` is false in `booting` and `provisioning`, so the boot-time self-pause that used to stand in for a lifecycle state is gone. - [x] **The sequence itself, as one list rather than prose in two files.** `didbot-onboarding`: the ordered steps a server passes, each naming the checks that decide it, the step it waits on, and the remedy an operator acts on. A step whose dependency has not passed is **blocked**, so one cause reports as one thing to fix rather than as every consequence of it, and `Run::first_failure` names that cause. A check declares the capabilities it needs, and the environment supplies them. An operator's own machine supplies all of them but a signed-in operator, and a check that needs a capability the environment has not got answers **not checkable here**, naming what is missing. Never a pass and never a silent skip. `didbot operate --check` runs the steps a claim waits on. The policy page checks the document and `describeServer` with `didbot-claim-check` before it writes a claim, and the list's document and description checks call the same crate. - [x] **Claimability checked from the operator's machine, not the server's.** `didbot operate --check ` runs the steps a claim waits on: the zone delegated, DNS resolving, TLS terminating, the `did:web` document serving and naming this host, and `describeServer` agreeing with it. Read-only — no OAuth, no browser, no record — so it is safe to run on a loop while waiting for DNS or an ACME order, and it exits non-zero so a script can wait on it. They live there because a server checking its own DNS confirms only what it wrote, a server checking its own TLS is often checking a loopback path, and the failure worth catching — a *different* deployment answering the same name — is invisible from inside the one that is wrong. The claim flow runs the same checks first, so a check that passes and a claim that then fails on it is a real race rather than two implementations disagreeing. - [x] **`bot.did.pollOperatorClaim`, a nudge with no authority.** `claimed` is entered only by the server observing a standing `bot.did.operator` record through its own poll. This route lets a caller ask it to look sooner: no body, no credential, no operator named, and its whole effect is to wake a task that was going to run anyway. Bounded on the server's **outbound** rate rather than on inbound requests. `MIN_POLL_INTERVAL` is a floor between two polls and every nudge inside one window coalesces into the same single poll, so an unbounded burst costs the operator's PDS one extra request — which the timer was going to spend anyway. `a_burst_of_nudges_produces_at_most_one_extra_poll` fires two hundred and asserts exactly one extra outbound read. - [x] **`bot.did.operator`** (`lexicons/bot/did/operator.json`): `subject`, `creates` and `createdAt`, keyed by the subject's hostname, so one record exists per server and a direct `getRecord` verifies it with no listing. The same record, keyed by an account's hostname, is every other edge of the tree; [ownership](ownership.md) owns its shape. - [x] **The server's own account, before any handshake.** `Provisioner::ensure_server_account` mints the server's keypair and `did:web` on every boot, idempotently, refusing to replace an existing key so a restart never produces a document that no longer matches signatures it already made. Its document is served at `/.well-known/did.json` at once. - [x] **What gates provisioning, and what does not.** A server is launched with its operator already configured; the record is a standing obligation, not an authorization it waits on. What gates provisioning is `didbot_pds::ServerPolicy::provisions_accounts`, false until the readiness gates in `docs/server-lifecycle.md` clear. An emergency stop halts something that is running, so `ServerPolicy::may_halt` is false in every state before the operational one and a throw is refused there. - [x] **The pause a lapse causes.** `didbot_pds::Estop::throw_self(Mode::Pause)` and the existing estop checks refuse provisioning and every token exactly as an operator's own pause would, and `Estop::Cause::OperatorMissing` keeps the two distinguishable. `an_unowned_server_refuses_to_provision_but_still_serves_its_identity` in `crates/didbot-serve/src/tests/mod.rs` proves the refusal over real HTTP, including that `describeServer` keeps answering; `crates/didbot/tests/handshake.rs` proves it is reached only by a claim that lapses, never by a server that was simply never claimed, and that a token issued before the lapse writes again once the claim is back. - [x] **The poll, the grace window, and reversibility.** `didbot_pds::operator::OperatorRelation` plus `didbot-serve::operator_poll`: one `getRecord` at the operator's PDS for the record keyed by this server's hostname, every `DEFAULT_POLL_INTERVAL`; a pause once `DEFAULT_GRACE_WINDOW` has elapsed with no readable record; resumed on its own when the record reappears. The same tick re-reads the record for every root the operator admitted and refreshes the allowances; [ownership](ownership.md) has that half. - [x] **The claim command itself**, `didbot operate `, on the operator's own machine, in a binary that links no server code. It signs in to the operator's own account with the scope `atproto repo:bot.did.operator?action=create&action=update&action=delete`, built from `didbot-scope`'s grammar, and nothing wider (`crates/didbot-operator/src/operate/scope.rs`); waits for TLS; resolves the server's `did:web` document; cross-checks `describeServer`; and writes the record through `com.atproto.repo.putRecord`, keyed at the hostname, so a re-run updates the claim in place. Every seam (`TlsProbe`, `DocumentFetcher`, `ServerDescriber`, `RecordWriter`) is exercised against a fake in the crate's own tests, including the refusal paths, proving nothing is written until every check passes. - [x] **The command holds a session between runs.** `didbot_authstore::FileAuthStore`, `0600` under the operator's config directory; a run resumes what the last one left and opens a browser only when nothing resumes. - [x] **The same shape admits every account.** A key minted where nothing can read it, its public half parked at the server, the record written into the operator's own repository, and the server reading the allowlist from the repository it already polls: `didbot register` and `didbot operate ` are that sequence, and [ownership](ownership.md) owns it. - [x] **An admission asks the operator's own PDS for no `rpc:` scope.** `didbot operate ` asks it for the `repo:bot.did.operator` scope above and writes the record with it. The create at the server carries the operator session `didbot login` keeps, the same identity-only sign-in the dashboard uses; with none kept, the command signs in that way first. The server accepts that session as the caller and still requires the record to name the account. It refuses a service-auth token from the operator's own PDS.