Zone reconciliation #
didbot_reconcile::ZoneState is what one DNS zone stands at right now, as
distinct from didbot_pds::ServerState, which is what this deployment may
do. The two are separate machines on purpose: a hostname withdrawn from the
zone by hand breaks agents that already exist whether or not the operator's
claim currently stands, so repairing it cannot be a step in a bootstrap that
finished months ago.
It acts. didbot-claim's checks observe a deployment from the operator's
own machine and report; this reads the zone from inside, compares it against
what this deployment published, and writes the difference back.
stateDiagram-v2
[*] --> blocked
blocked --> in_sync: a read succeeded and nothing is wrong
blocked --> drifted: a read succeeded and something is
in_sync --> drifted: a name is missing or points elsewhere
drifted --> reconciling: a second read confirms the same drift
drifted --> in_sync: it was propagation, not drift
reconciling --> in_sync: the repair landed
reconciling --> blocked: every outstanding repair is quarantined
in_sync --> blocked: the zone could not be read
| state | repairs records | observation is evidence | deployment keeps serving |
|---|---|---|---|
blocked |
no | no | yes |
drifted |
yes | yes | yes |
reconciling |
yes | yes | yes |
in-sync |
yes | yes | yes |
The last column is constant, and that is the point: this machine is a
repairer, not a gate. serving_is_never_gated_on_this_machines_state asserts
it by destructuring ZonePolicy exhaustively over every state, so a
capability that could gate a request stops the test compiling.
What counts as wrong #
Decided per name, against the intent — the hostnames this deployment published and what each should point at.
| observed | verdict | what happens |
|---|---|---|
| intended name, nothing in the zone | missing | republished |
| intended name, a different target | pointing elsewhere | replaced |
| a name the intent does not know | unattributed | reported, never touched |
| a different TTL | not drift | ignored |
TXT values |
out of scope | ignored |
the NS delegation above the zone |
not ours | ignored |
A TTL is left alone structurally rather than by choice: RecordTarget
carries none, so an intent cannot express one. That is also the right answer —
lowering a TTL before a migration is ordinary practice, and a wrong TTL
degrades how fresh an answer is, not whether there is one.
TXT is out of scope because one name holds a set of values that the ACME
renewal path writes and withdraws within the length of one challenge; a
reconciler working from a five-minute snapshot would race it every run.
What it never does #
- Delete a record it cannot prove stale.
Repairhas two variants,PublishandReplace, and no third.Drift::Unattributedyields no repair at all, so a name this deployment did not publish has nothing that could act on it. A read that fails producesViewError, not an empty zone — which is whyZoneView::observereturns aResultwhereDnsProvider::publishedreturns a bareVec. - Gate serving on its own state. See the table above.
- Touch a zone it does not manage.
ZoneAuthority::ReadOnlyblocks before the first read. It never creates a zone under any circumstances, sodns.may_create_zonesis not a posture it can reach past: a zone that is not there isViewError::ZoneAbsent, which blocks and says so.
How often, and how hard #
| bound | default | why |
|---|---|---|
| interval | 5 minutes | Route53Dns's own default record TTL: never re-examine a name faster than a resolver could have seen the last change |
| backoff | double to 1 hour | a tick that could not advance waits longer — a gate failed, or every drifted name is quarantined and the zone is blocked. One tick that makes progress clears it |
| confirmations | 2 reads | drift must survive a second read before anything is written, so a change in flight is not fought |
| repairs per tick | 25 | a wholesale-emptied zone costs a bounded number of writes per tick, not thousands at once |
| accepted repairs per name | 5 | a name that keeps drifting back after a repair the provider accepted is quarantined and reported. The counter clears the moment the name reads correct, and a repair the provider refused never adds to it |
Between the interval and the confirmation count, drift is repaired five to ten
minutes after it happens. That is this epic's answer to plan/onboarding.md's
open question of how long drift is propagation and how long it is a fault:
long enough that a propagating change is never fought, short enough that a
withdrawn hostname is back before an agent's next session.
The attempt budget is aimed at people rather than at providers, and only at people. A human deliberately changing a record by hand gets argued with five times and then reported, rather than forever.
A repair the provider refused is not charged to it, and the two bounds are not interchangeable. Quarantine is permanent for the life of the process — its only exit is observing the name correct — so spending it on a provider fault strands the name: nothing but this machine would ever put the record back, and quarantine is what stopped it. A refused repair fails a gate instead, which is what the backoff is computed from, so a provider outage costs one attempt per tick at a rate falling to hourly and the name is repaired on the first tick after the provider recovers.
Every one of these counters lives in memory, so a restart discards all of them — the backoff included. What keeps that safe rather than expensive is the confirmation count: a repair needs two consecutive reads from one process, so a process that crash-loops faster than its own interval reads the zone and writes nothing at all.
The gates #
didbot_fsm::Gates, declared in causal order, so one failure reports as one
cause rather than as every check below it:
zone is managedzone is readableintent is knownintended records are presentintended records point here
The last two are the reason this machine brought three-valued gates back to
didbot-fsm. A missing record awaiting its confirming read is Pending: it
is wrong, and nothing has been attempted, and reporting that as a failure
would make the first tick of a healthy process look like an outage — and the
backoff is computed from exactly that signal.