Operations #
Procedures for the deployment infra/pds/templates/user_data.sh.tftpl builds: one
instance, one systemd unit didbot-pds, one container of the release image,
one data volume at /data with the data directory at /data/pds.
| Question | Command |
|---|---|
| Is the unit up | systemctl status didbot-pds |
| What did it say | journalctl -u didbot-pds -f |
| Is a container running | docker ps --filter name=didbot-pds |
| Is the process live and progressing | curl -s https://<zone>/health |
| Are sign-in pushes being dropped | curl -s https://<zone>/health | jq .oauthPushes |
| What does it hold | curl -s https://<zone>/xrpc/bot.did.stats |
Upgrading #
systemctl stop didbot-pds, then confirm withdocker ps --filter name=didbot-pdsthat nothing is running.syncon the host.- Take an on-demand backup of the data volume into the
didbot-pdsvault, or an EBS snapshot of it. Wait for it to report completed. docker pull <image>, then pointExecStartat it in/etc/systemd/system/didbot-pds.service.systemctl cat didbot-pdsshows what is actually running.systemctl daemon-reload, thensystemctl start didbot-pds.- Watch the journal through startup. On a refusal see "Boot refusals"; otherwise run the checks under "Verifying a restore".
Boot refusals #
Each stops the process at startup, before it serves anything, and the unit then restarts every five seconds logging the same message.
| Journal message | What it means | What to do |
|---|---|---|
the data directory at … was written as version … |
The stamp disagrees with the binary | "Rolling back" |
the data directory at … carries a stamp written before this binary's check |
Half the stamp has nothing to compare against | Run the build that wrote it, or delete the directory |
blob bytes are present and pds.wal is not |
Blob files with no log to name them | "Restoring" |
entry <n> did not parse |
Intact bytes this binary does not read as an entry | Run the build that wrote it. Do not delete frames by hand |
| Another process holds the lock | A second writer against one directory | Find and stop it; never mount one data directory twice |
| A permission error writing to the data directory | Ownership, not mode | chown — see "Restoring" |
Rolling back #
A data directory carrying no pds.layout stamp is adopted and stamped with the
opening binary's own layout, so the first run of a new binary restamps it and
the old binary then meets a mismatch on a directory it wrote itself.
The rollback is a restore. Run the old image against a snapshot taken before the upgrade: its stamp is the old layout, or absent, and either is accepted. The cost is everything written since that snapshot — up to twenty-four hours on the daily plan alone, minutes with an on-demand snapshot taken just before the swap.
Backups #
infra/pds/backup.tf puts the one EBS data volume under an AWS Backup plan: vault
didbot-pds, rule daily, schedule cron(0 9 * * ? *) — 09:00 UTC — and
delete_after = 30, thirty daily recovery points. The RPO is twenty-four
hours. aws_backup_selection.pds names the one resource the plan covers:
the EBS data volume mounted at /data.
Snapshot the volume; do not rsync the directory.
Restoring #
- Create a volume from a recovery point in the
didbot-pdsvault, in the target instance's availability zone. - Attach it to a fresh instance, never to the running one, and mount it at
/data, where that instance's boot script chowns/data/pdsto uid 10001. The image runs as that uid and chmods the data directory to0700on every open: the mode a snapshot carries is re-tightened, the owner is the chown. - Start the unit, and run the checks below.
- Detach and delete any test volume.
On a blob bytes are present and pds.wal is not refusal: finish the restore so
the log is present, or move blobs/ aside and start with an empty blob store,
after which getBlob answers NotFound for the blobs the accounts reference.
Verifying a restore #
The startup line is restored the deployment from its write-ahead log, and its
fields are the count of each thing that came back. It may be followed by
compacted the write-ahead log. A the write-ahead log ends in an incomplete write; truncating it warning reports the discarded bytes and the
offset it stopped at; on a snapshot restore that is the torn tail of the
append in flight when the snapshot was taken.
| Check | Confirms |
|---|---|
GET /health returns 200 with ticks advancing between two calls a few seconds apart |
The process is live and its ten-second tick is running |
GET /xrpc/bot.did.stats |
Accounts, records and bytes per collection match what the deployment had, and blobs.held matches its blob inventory |
blobs.missingAtBoot in that response reads {"count": 0, "bytes": 0} |
Every blob the log names had its bytes on the disk this deployment started on. A nonzero count is that many 404s waiting, and bytes is what the restore would have to bring back; the CIDs are in the startup ERROR line |
blobs.discardedAtBoot reads {"count": 0, "bytes": 0} |
The copy carried no blob bytes the log could not account for. A nonzero count is that many files deleted during startup |
GET /xrpc/bot.did.listAgents |
The accounts that should be there are there |
One agent's did.json resolves |
The signing keys came back |
| Fetch one blob that account references | The bytes are on the disk; getBlob re-hashes what it reads and refuses on a mismatch |
Sign-in pushes that leave no decision behind #
POST /oauth/par is unauthenticated, and no answer it gives may depend on
whether an account exists — so a client flooding one agent's pending sign-ins
gets the same 201 an ordinary push gets, and the agent sees only that its
own sign-in never arrives. GET /health carries oauthPushes, three counts
since this process started:
| Count | What moved it |
|---|---|
displaced |
A push past an account's own pending_per_account bound. That account's oldest undecided sign-in made way for it |
unresolvedHint |
A push whose login_hint named no account here |
serverFull |
A push refused because this server holds pending_total undecided records already. This is the one an ordinary caller can see, as a 403 |
Read them twice a minute apart: these are totals, so what says anything is
how fast they move. displaced climbing steadily is a flood aimed at
accounts this deployment holds; unresolvedHint climbing is one probing for
accounts it does not. The journal carries a sampled info line for the
first, one in a hundred, naming the account — sampled because logging every
one would make the log the thing a flood amplifies.
bot.did.stats and bot.did.listAgents are public by default; under
--close-disclosure they answer 403 and the figures are in the startup line.