Identities for entities did.bot
agent llm did
didbot docs zone-reconciliation.md
6.6 kB
Markdown
at commit 18ba4fe0

Zone reconciliation #

didbot_reconcile::ZoneState is what one DNS zone stands at right now, as distinct from didbot_pds::ServerState, which is what this deployment may do. The two are separate machines on purpose: a hostname withdrawn from the zone by hand breaks agents that already exist whether or not the operator's claim currently stands, so repairing it cannot be a step in a bootstrap that finished months ago.

It acts. didbot-claim's checks observe a deployment from the operator's own machine and report; this reads the zone from inside, compares it against what this deployment published, and writes the difference back.

stateDiagram-v2
    [*] --> blocked
    blocked --> in_sync: a read succeeded and nothing is wrong
    blocked --> drifted: a read succeeded and something is
    in_sync --> drifted: a name is missing or points elsewhere
    drifted --> reconciling: a second read confirms the same drift
    drifted --> in_sync: it was propagation, not drift
    reconciling --> in_sync: the repair landed
    reconciling --> blocked: every outstanding repair is quarantined
    in_sync --> blocked: the zone could not be read
state repairs records observation is evidence deployment keeps serving
blocked no no yes
drifted yes yes yes
reconciling yes yes yes
in-sync yes yes yes

The last column is constant, and that is the point: this machine is a repairer, not a gate. serving_is_never_gated_on_this_machines_state asserts it by destructuring ZonePolicy exhaustively over every state, so a capability that could gate a request stops the test compiling.

What counts as wrong #

Decided per name, against the intent — the hostnames this deployment published and what each should point at.

observed verdict what happens
intended name, nothing in the zone missing republished
intended name, a different target pointing elsewhere replaced
a name the intent does not know unattributed reported, never touched
a different TTL not drift ignored
TXT values out of scope ignored
the NS delegation above the zone not ours ignored

A TTL is left alone structurally rather than by choice: RecordTarget carries none, so an intent cannot express one. That is also the right answer — lowering a TTL before a migration is ordinary practice, and a wrong TTL degrades how fresh an answer is, not whether there is one.

TXT is out of scope because one name holds a set of values that the ACME renewal path writes and withdraws within the length of one challenge; a reconciler working from a five-minute snapshot would race it every run.

What it never does #

  • Delete a record it cannot prove stale. Repair has two variants, Publish and Replace, and no third. Drift::Unattributed yields no repair at all, so a name this deployment did not publish has nothing that could act on it. A read that fails produces ViewError, not an empty zone — which is why ZoneView::observe returns a Result where DnsProvider::published returns a bare Vec.
  • Gate serving on its own state. See the table above.
  • Touch a zone it does not manage. ZoneAuthority::ReadOnly blocks before the first read. It never creates a zone under any circumstances, so dns.may_create_zones is not a posture it can reach past: a zone that is not there is ViewError::ZoneAbsent, which blocks and says so.

How often, and how hard #

bound default why
interval 5 minutes Route53Dns's own default record TTL: never re-examine a name faster than a resolver could have seen the last change
backoff double to 1 hour a tick that could not advance waits longer — a gate failed, or every drifted name is quarantined and the zone is blocked. One tick that makes progress clears it
confirmations 2 reads drift must survive a second read before anything is written, so a change in flight is not fought
repairs per tick 25 a wholesale-emptied zone costs a bounded number of writes per tick, not thousands at once
accepted repairs per name 5 a name that keeps drifting back after a repair the provider accepted is quarantined and reported. The counter clears the moment the name reads correct, and a repair the provider refused never adds to it

Between the interval and the confirmation count, drift is repaired five to ten minutes after it happens. That is this epic's answer to plan/onboarding.md's open question of how long drift is propagation and how long it is a fault: long enough that a propagating change is never fought, short enough that a withdrawn hostname is back before an agent's next session.

The attempt budget is aimed at people rather than at providers, and only at people. A human deliberately changing a record by hand gets argued with five times and then reported, rather than forever.

A repair the provider refused is not charged to it, and the two bounds are not interchangeable. Quarantine is permanent for the life of the process — its only exit is observing the name correct — so spending it on a provider fault strands the name: nothing but this machine would ever put the record back, and quarantine is what stopped it. A refused repair fails a gate instead, which is what the backoff is computed from, so a provider outage costs one attempt per tick at a rate falling to hourly and the name is repaired on the first tick after the provider recovers.

Every one of these counters lives in memory, so a restart discards all of them — the backoff included. What keeps that safe rather than expensive is the confirmation count: a repair needs two consecutive reads from one process, so a process that crash-loops faster than its own interval reads the zone and writes nothing at all.

The gates #

didbot_fsm::Gates, declared in causal order, so one failure reports as one cause rather than as every check below it:

  1. zone is managed
  2. zone is readable
  3. intent is known
  4. intended records are present
  5. intended records point here

The last two are the reason this machine brought three-valued gates back to didbot-fsm. A missing record awaiting its confirming read is Pending: it is wrong, and nothing has been attempted, and reporting that as a failure would make the first tick of a healthy process look like an outage — and the backoff is computed from exactly that signal.