--- id: onboarding title: An operator establishes a server before the server can establish anyone else status: open crates: [didbot-attest, didbot-pds, didbot-serve, didbot-setup, didbot-reconcile] dependsOn: [ownership, deploy] exitCriterion: > An operator takes a freshly applied deployment — instance up, no accounts, nothing configured beyond Terraform's defaults — through generating and installing the real secrets, an operator attestation step, and the first-run ownership handshake, then points one agent host's harness at it and gets a working, provisioned account, using only documented steps and no fact this file invents. --- # onboarding Every other epic in this directory describes a system already running: an account exists, a zone resolves, policy is being polled. Nothing describes the hour before any of that is true — Terraform has applied, an EC2 instance is up, and the operator has not yet done anything the server would recognise as them. [dev-setup](dev-setup.md) covers the symmetric gap on the other side, wiring one machine's harness to a stack; this is the half it does not cover, and the half that has to happen first. ## The bootstrap paradox **A provisioning request carries a host's claim.** `bot.did.provisionAgent` mints only under a claim signed with the node key of a host the operator has vouched for — see [handshake](handshake.md). The key is generated on the host and never leaves it, so there is no shared secret to distribute or get wrong. So the paradox is not what it was. It was: the secret that lets an operator act on a fresh deployment must reach the host before the deployment can be acted on. There is no such secret now, and the question that replaces it is narrower — a fresh deployment has to reach a state its operator can claim, using only what the deployment already has. That question is answered elsewhere and this epic should not re-answer it. The server mints its own identity through an administrative path that never travels the xrpc write path, so its `did:json` and signing key exist before it serves anything; the operator then claims it from their own machine with `didbot-claim`, and the server observes that claim through its own poll. See [handshake](handshake.md) and `docs/server-lifecycle.md`. What onboarding still owes is the sequence, not the mechanism: - [ ] **Write the operator's first commands**, in order, against a real deployment: what to run, on which machine, and what each one is waiting for. `didbot-claim` has the local half and `didbot-claim --check` has the pre-flight; nothing walks a first-time operator through them. - [ ] **Say how an operator tells whether the server has seen the claim yet.** The poll is the only thing that moves the server to `claimed`, and its interval is the delay an operator experiences as "nothing happened". - [ ] **Say what a deployment does before it has any operator policy.** [policy](policy.md) puts a pre-flight between provisioning and `claimed`, so an instance is never online and unpolicied. The commands for that half belong here too. ## What "first log in" means here There is no human user account on this server, so "log in" is not the right frame and the epic should stop implying one. What exists instead: - `--owner` — a DID string the server writes into every registration record as who is answerable for that account. It is not verified at startup: any string parses. - The `bot.did.operator` handshake — the operator writes a record into their *own* repository and this server reads it (`crates/didbot-serve/src/ownership_poll.rs`). Nothing is presented to the server and nothing is compared by it. There is no operator credential, and no HTTP surface is gated on one — see [auth-types](auth-types.md). - Local access to the box, which is what the e-stop admin socket's `0700`/`0600` permissions amount to and the whole of that socket's authorization story. Neither of the first two is the full handshake [ownership](ownership.md) asks for. Ownership's own checklist says the agent-side owner claim exists (written by the server at provisioning) and the human's vouch of the server is meant to come from the owner's own repository — but ownership's "three statements" item is still open, including *"the server answers its controlling DID on an unauthenticated status endpoint,"* which is the piece that would let a stranger, or the operator themselves, confirm the server agrees who its owner is. Until that exists, nothing on first boot confirms `--owner` names a DID that has actually agreed to be named — the server just asserts it, on one process argument, forever. - [ ] **A first-run check that `--owner` resolves to something real** before the server starts minting accounts against it — at minimum, that the DID document exists and is fetchable. `--owner` today is `Zone::new`'d alongside the zone but is never itself validated. - [ ] **The unauthenticated status endpoint** ownership.md names, so an operator (or anyone) can ask a running server "who do you say owns you" and compare it against what they meant to configure. This is ownership's item to build; onboarding is where a first-run check would use it. - [ ] **Write down, plainly, that there is no bidirectional claim yet.** Until the status endpoint above exists, "first log in" is `--owner` plus a `bot.did.operator` record this server reads one way, and that is a weaker claim than the rest of `plan/` sometimes assumes. Say so somewhere a new operator reads before they conclude ownership is already checked because the flag is there. ## Required operator attestation The owner's framing names this explicitly, and it is worth being precise about what it is not: [attestation](attestation.md) is about the *node an agent runs on* — proving a machine before it is trusted with provisioning. The operator is a different subject. Nothing in this codebase today asks the operator to prove anything *to the server over HTTP* at all: the handshake runs the other way, with the server reading a record the operator wrote in their own repository. That is a claim about a DID rather than possession of a string, which is the right direction — but this server never checks a signature over a challenge it chose, so it is a claim it reads, not one it verifies interactively. What an operator attestation step would need to establish, before the server mints its first account: - [ ] **That the handshake actually completed** — i.e. this server has read a `bot.did.operator` record naming it, rather than still waiting for one. This is the loud version of the bootstrap-paradox item above, scoped to "before minting begins" rather than "eventually." - [ ] **That the `--owner` DID is one the operator controls**, which is a real proof — sign a nonce with the DID's own key, or an equivalent — not merely a string on the command line. Nothing here does this yet; it is the missing half of "first log in," restated as a checkable step rather than left implicit. - [ ] **Decide whether operator attestation is a one-time ceremony or a standing credential.** A node's attestation is spent once, at provisioning, and the node's ongoing claim to be trusted rests on holding a node credential afterward ([node](node.md)). The operator has no equivalent today, and — since there is deliberately no operator credential at all — no "hold a narrower credential afterward" half to build one out of. [oauth](oauth.md)'s owner sign-in is the candidate mechanism. Whether the operator needs a standing credential at all, or the read-only handshake plus host access is the whole story, is this epic's to decide, not to assume. ## The first agent Once a server is up, the handshake has completed, and `--owner` is established, a person still has to get one agent host talking to it. [didbot-setup](../crates/didbot-setup/) does the machine half: it describes which services this machine runs and says whether they are answering. What it explicitly does not do, checked against its own source: - **It never sets `DIDBOT_PDS_URL`.** A harness client falls back to `http://localhost:3000` when it is unset, and `didbot-setup` never writes it — its whole flow is built and tested against `didbot-pds`'s own defaults. Pointing a harness at a deployed server is presently "set an environment variable by hand, correctly, on your own," with no `didbot-setup` step that asks for a URL and writes it anywhere. There is no secret to set alongside it: provisioning takes the host's claim, which `didbot-agentd` signs. - **A session with no usable account provisions one on its first tool call**, so once the variables above are set correctly by hand, that call is what actually creates the account — there is no separate "first login" moment to perform. - [ ] **A `didbot-setup` flow — or documented manual steps, if the flow is not worth building yet — for wiring a harness against a server that is not `didbot-pds`'s own defaults**: the URL, landing in the same place `check` already looks, and the host's `become` against that server. - [ ] **`didbot-setup check` should say which server a profile is configured for**, so a harness silently pointed at a stale or wrong deployment is caught before the first session rather than after a refusal with no context. ## Failure modes worth naming These are not hypothetical: each is a documented gap or an enforced-but-silent default above, restated as what an operator actually sees when it happens. - [ ] **A server that started but was never given a real secret.** Cannot happen for the attestation secret against a non-`.localhost` zone — `didbot-pds` refuses to start. Can absolutely happen for the operator secret, which only warns, and the warning is the only signal until an e-stop needs releasing. - [ ] **DNS delegated but not propagated.** What it looks like from the operator's seat now has an answer: `didbot-claim --check ` resolves the name through the operator's own stub resolver and reports `dns FAILED` with the resolver's reason, skipping the checks downstream of it rather than reporting three more failures for one cause. What is still open is the second half — how long to wait before it is a bug rather than propagation. Nothing names a number, and a check that has failed for an hour looks exactly like one that has failed for a minute. - [ ] **A certificate that has not issued yet.** `--tls acme` is wired now — `didbot-pds` refuses at startup without `--data`, a real zone and `--route53-zone-id` (`crates/didbot-serve/src/bin/didbot-pds.rs`), and `infra/pds/templates/user_data.sh.tftpl`'s `ExecStart` passes it. What is still unverified, per [deploy](deploy.md)'s own Done entry, is the protocol exchange itself against a real ACME directory — nobody has run it from an environment with the network access to try. An operator's first boot is where that gets exercised for the first time, and the failure mode worth naming is: a startup that passes every local check above but still cannot complete a DNS-01 challenge against Let's Encrypt, for a reason none of those checks can catch in advance. `didbot-claim --check`'s `tls` check is what now catches it — a real HTTPS request from the operator's machine, so a certificate that never issued, was never loaded, or covers the wrong name all fail the same way and say so. It does not make the exchange any more verified; it makes a failed one visible from the seat that matters instead of silent. - [ ] **Wire the zone reconciler's intent to the account registry.** The machine is built and is `didbot_reconcile::ZoneReconciler` (below, and `docs/zone-reconciliation.md`), but nothing in `didbot-serve` constructs one yet: it takes its intent through a `ZoneIntent` trait, and the implementation that answers "every hostname this deployment published, and what each should point at" from the account store does not exist. Until it does, a deployment still has nobody whose job the zone is. - [ ] **Give the reconciler an authoritative read of a remote zone.** Its other half, and the one that will look done when it is not. `ProviderView` reads through `DnsProvider::target`, which answers from local bookkeeping, so it is a fresh view of the zone only to the extent its refresh hook makes that bookkeeping match. `Route53Dns::resync` is the obvious hook and does not: it inserts what it finds, discards a collision with what is already cached, and removes nothing. Wired that way the reconciler reads the zone every five minutes and reports it in sync no matter what — blind to a record deleted at the AWS console and blind to one repointed there, which are the two drifts it exists to repair. Either `resync` becomes authoritative or the reconciler gets a `ZoneView` that is a real read; the choice belongs with whoever wires the intent above, because a reconciler that cannot see drift is worse than no reconciler — it is a green light over a broken zone. - [ ] **An operator who has lost access to the DID that operates a server.** There is nothing to rotate — the server holds no credential — but the recovery question is real and unanswered: a server watching a `bot.did.operator` record in a repository its operator can no longer write has no path to being told about a new one, short of restarting it with a different `--owner`. Whether that restart is the answer, or something narrower is, is undecided. ## Done - [x] **A zone that drifts after the server is running is somebody's job.** `didbot_reconcile::ZoneReconciler` — the second `didbot-fsm` consumer, and deliberately not folded into the server lifecycle. Its states are about the zone (`in-sync`, `drifted`, `reconciling`, `blocked`) and it runs in every server state including `claimed`, forever, because withdrawing a hostname breaks agents that already exist whether or not an operator's claim currently stands. Unlike `didbot-claim`'s checks it **acts**: it reads the zone, compares it against what this deployment published, and republishes the difference. What "wrong" means is decided per name — a missing record is republished, one pointing elsewhere is replaced, a **wrong TTL is left alone** (structurally: `RecordTarget` carries none, and lowering a TTL before a migration is ordinary practice), `TXT` values belong to the ACME renewal path, and the `NS` delegation above the zone is the operator's. A record it did not publish is **never touched**, and not by a check somebody remembered to write: `Repair` has `Publish` and `Replace` and no third variant, `Drift::Unattributed` yields no repair at all, and a failed read is a `ViewError` rather than an empty zone. It never gates serving either — `ZonePolicy` carries no capability that could, which a test asserts by destructuring it over every state. Bounded at five minutes with backoff doubling to an hour, two confirming reads before anything is written, twenty-five writes per tick, and five attempts per name before it is quarantined and reported. That last budget is the one aimed at people: somebody changing a record by hand gets argued with five times rather than forever. Between the interval and the confirmation count, drift is repaired five to ten minutes after it happens — this epic's answer to the "how long is propagation" question the items above kept deferring. It is also the caller that brought `didbot_fsm::Gates` back. A repair awaiting its confirming read is `Pending`, not `Failed`, and collapsing the two would make a healthy first tick look like an outage — which is exactly the signal the backoff is computed from. - [x] **A server lifecycle, so "not finished being born" is a state rather than a latch.** `didbot_pds::ServerState` — `booting`, `provisioning`, `unclaimed`, `claimed` — with a capability policy per state and one enforcement point per capability; `docs/server-lifecycle.md` carries the diagram and the table, generated from the types and checked against them by a test. The machinery is `didbot-fsm`, a reusable core (states with a consumer-shaped policy, transitions taken atomically under one lock, one-way capabilities claimed under a permit, persistence as a hook); this lifecycle is its first consumer. The axis is **reads against writes**. `unclaimed` — where a deployment spends most of its early life, and where a server whose claim lapsed returns to — reads everything `claimed` does and writes nothing: no agent accounts, no records, no blobs, no outward announcements. Withdrawing reads would protect nothing and would retroactively break every signature made under a `did:web` document this server has already served. Three things it makes structurally true. A second self-bootstrap keypair cannot be minted, because `ServerLifecycle::begin_bootstrap` is the only source of the `BootstrapPermit` a mint takes by value and it shuts the door in the same critical section. Never-claimed and lapsed are one state with one recovery — run `didbot-claim` again — so the bootstrap flow is exercised on every re-claim rather than once per server lifetime. And the emergency stop cannot be thrown before an identity exists: `ServerPolicy::may_halt` is false in `booting` and `provisioning`, so the boot-time self-pause that used to stand in for a lifecycle state is gone. - [x] **Claimability checked from the operator's machine, not the server's.** `didbot_claim::preflight`, exposed as `didbot-claim --check `: DNS resolving, TLS terminating, the `did:web` document serving and naming this host, and `describeServer` agreeing with it. Read-only — no OAuth, no browser, no record — so it is safe to run on a loop while waiting for DNS or an ACME order, and it exits non-zero so a script can wait on it. They live there because a server checking its own DNS confirms only what it wrote, a server checking its own TLS is often checking a loopback path, and the failure worth catching — a *different* deployment answering the same name — is invisible from inside the one that is wrong. The claim flow runs the same function first, so a check that passes and a claim that then fails on it is a real race rather than two implementations disagreeing. - [x] **`bot.did.pollOperatorClaim`, a nudge with no authority.** `claimed` is entered only by the server observing a standing `bot.did.operator` record through its own poll. This route lets a caller ask it to look sooner: no body, no credential, no operator named, and its whole effect is to wake a task that was going to run anyway. Bounded on the server's **outbound** rate rather than on inbound requests. `MIN_POLL_INTERVAL` is a floor between two polls and every nudge inside one window coalesces into the same single poll, so an unbounded burst costs the operator's PDS one extra request — which the timer was going to spend anyway. `a_burst_of_nudges_produces_at_most_one_extra_poll` fires two hundred and asserts exactly one extra outbound read. - [x] **An onboarding surface.** `GET /onboarding` shows which state the server is in, whether it holds its own identity, whether a claim currently stands, which surfaces answer, and what to run next on the operator's own machine. It is a status and guidance page: it accepts no credential, offers nowhere to type one, and exposes no route that changes anything, because nothing exists by which a browser could prove control of the operator DID. Its copy is placeholders; its facts are not.