Identities for entities did.bot
agent llm did
didbot plan onboarding.md
17 kB
Markdown
at main


id: onboarding title: An operator establishes a server before the server can establish anyone else status: open crates: [didbot-pds, didbot-serve, didbot-agentd] dependsOn: [ownership, deploy] exitCriterion: > An operator takes a freshly applied deployment — instance up, no accounts, nothing configured beyond Terraform's defaults — through generating and installing the real secrets, an operator attestation step, and the first-run operator handshake, then points one agent host's harness at it and gets a working, provisioned account, using only documented steps and no fact this file invents. #

onboarding #

Every other epic in this directory describes a system already running: an account exists, a zone resolves, policy is being polled. Nothing describes the hour before any of that is true — Terraform has applied, an EC2 instance is up, and the operator has not yet done anything the server would recognise as them. dev-setup covers the symmetric gap on the other side, wiring one machine's harness to a stack; this is the half it does not cover, and the half that has to happen first.

The bootstrap paradox #

A provisioning request carries a host's claim. bot.did.createAccount mints only under a claim signed with the node key of a host the operator has vouched for — see handshake. The key is generated on the host and never leaves it, so there is no shared secret to distribute or get wrong.

So the paradox is not what it was. It was: the secret that lets an operator act on a fresh deployment must reach the host before the deployment can be acted on. There is no such secret now, and the question that replaces it is narrower — a fresh deployment has to reach a state its operator can claim, using only what the deployment already has.

That question is answered elsewhere and this epic should not re-answer it. The server mints its own identity through an administrative path that never travels the xrpc write path, so its did:json and signing key exist before it serves anything; the operator then claims it from their own machine with didbot operate, and the server observes that claim through its own poll. See handshake and docs/server-lifecycle.md.

What onboarding still owes is the sequence, not the mechanism. The sequence itself is crates/didbot-onboarding/src/step.rs — one ordered list, held as data, with each step's checks, the step it waits on, and the remedy beside it. This file points at that list rather than carrying a second copy of it.

What "first log in" means here #

There is no human user account on this server, so "log in" is not the right frame and the epic should stop implying one. What exists instead:

  • --operator — a DID string the server writes into every registration record as who is answerable for that account. It is not verified at startup: any string parses.
  • The bot.did.operator handshake — the operator writes a record into their own repository and this server reads it (crates/didbot-serve/src/operator_poll.rs). Nothing is presented to the server and nothing is compared by it. There is no operator credential, and no HTTP surface is gated on one — see auth-types.
  • The operator's sign-in: atproto OAuth against the operator's own server, asking for atproto and nothing else (crates/didbot-serve/src/operator.rs). This server compares the DID it signed in against --operator and issues its own session, which the dashboard's operator routes take, from a browser or from the didbot verbs that drive them.

Neither of the first two is the full handshake ownership asks for. That epic's own checklist says the agent-side operator claim exists (written by the server at provisioning) and the human's vouch of the server is meant to come from the operator's own repository — but its "three statements" item is still open, including "the server answers its controlling DID on an unauthenticated status endpoint," which is the piece that would let a stranger, or the operator themselves, confirm the server agrees who its operator is. Until that exists, nothing on first boot confirms --operator names a DID that has actually agreed to be named — the server just asserts it, on one process argument, forever.

Required operator attestation #

The owner's framing names this explicitly, and it is worth being precise about what it is not: ownership's check at creation is about the account creating — proving it by its credential records before anything is created beneath it. The operator is a different subject. Nothing in this codebase today asks the operator to prove anything to the server over HTTP at all: the handshake runs the other way, with the server reading a record the operator wrote in their own repository. That is a claim about a DID rather than possession of a string, which is the right direction — but this server never checks a signature over a challenge it chose, so it is a claim it reads, not one it verifies interactively.

What an operator attestation step would need to establish, before the server mints its first account:

The first agent #

Once a server is up, the handshake has completed, and --operator is established, a person still has to get one agent host talking to it. The machine half is the didbot-claude plugin and didbot-agentd, which takes the server it mints against from DIDBOT_PDS. Two things about that are worth knowing:

  • The server is an environment variable, set by hand. A harness client falls back to http://localhost:3000 when DIDBOT_PDS_URL is unset, and the daemon refuses to start without DIDBOT_PDS. There is no secret to set alongside either: provisioning takes the host's claim, which didbot-agentd signs.

  • A session with no usable account provisions one on its first tool call, so once the variables above are set correctly, that call is what actually creates the account — there is no separate "first login" moment to perform.

Failure modes worth naming #

These are not hypothetical: each is a documented gap or an enforced-but-silent default above, restated as what an operator actually sees when it happens.

Done #

  • The axis is **reads against writes**. `unclaimed` — where a deployment
    spends most of its early life, and where a server whose claim lapsed
    returns to — reads everything `claimed` does and writes nothing: no
    agent accounts, no records, no blobs, no outward announcements.
    Withdrawing reads would protect nothing and would retroactively break
    every signature made under a `did:web` document this server has already
    served.
    
    Three things it makes structurally true. A second self-bootstrap
    keypair cannot be minted, because `ServerLifecycle::begin_bootstrap` is
    the only source of the `BootstrapPermit` a mint takes by value and it
    shuts the door in the same critical section. Never-claimed and lapsed
    are one state with one recovery — run `didbot operate` again — so the
    bootstrap flow is exercised on every re-claim rather than once per
    server lifetime. And the emergency stop cannot be thrown before an
    identity exists: `ServerPolicy::may_halt` is false in `booting` and
    `provisioning`, so the boot-time self-pause that used to stand in for a
    lifecycle state is gone.
    
  • A check declares the capabilities it needs, and the environment
    supplies them. An operator's own machine supplies all of them but a
    signed-in operator, and a check that needs a capability the environment
    has not got answers **not checkable here**, naming what is missing.
    Never a pass and never a silent skip. `didbot operate --check` runs the
    steps a claim waits on. The policy page checks the document and
    `describeServer` with `didbot-claim-check` before it writes a claim,
    and the list's document and description checks call the same crate.
    
  • They live there because a server checking its own DNS confirms only
    what it wrote, a server checking its own TLS is often checking a
    loopback path, and the failure worth catching — a *different*
    deployment answering the same name — is invisible from inside the one
    that is wrong. The claim flow runs the same checks first, so a check
    that passes and a claim that then fails on it is a real race rather
    than two implementations disagreeing.
    
  • Bounded on the server's **outbound** rate rather than on inbound
    requests. `MIN_POLL_INTERVAL` is a floor between two polls and every
    nudge inside one window coalesces into the same single poll, so an
    unbounded burst costs the operator's PDS one extra request — which the
    timer was going to spend anyway. `a_burst_of_nudges_produces_at_most_one_extra_poll`
    fires two hundred and asserts exactly one extra outbound read.