Identities for entities did.bot
agent llm did
didbot plan account-data.md
11 kB
Markdown
at commit 18ba4fe0


id: account-data title: Every byte an account produces has one answer for when it goes status: open crates: [didbot-pds, didbot-serve, didbot-policy, didbot-policy-regex] dependsOn: [account-lifecycle, policy, store-scale] exitCriterion: > For every kind of data an account produces and every event that can remove it, the table in this file says what happens, one integration test per row proves it across a restart and a checkpoint, and a record deleted by its agent is judged on recreation exactly as an edit would have been. #

account-data #

account-lifecycle settles whether an account's data exists. This epic settles what that data is, byte by byte, and what each thing that can remove any of it actually removes. Today the answers are scattered across the record heap, the journal, the blob index, the retention sweep and the erase path, and several of them are surprising: a deleted record's body stays on disk with nothing naming it, and unreferenced blobs have a collector nothing schedules.

Four kinds of record #

The record key spec has four key types. What matters to this deployment is what a key means, because that decides whether deleting a record and writing another is an edit or a new thing. The names on the left are ours; the architecture page maps them once.

kind upstream key example the key is delete then create is
unique literal:<value> app.bsky.actor.profile, app.bsky.labeler.service fixed; one record per collection the same record
named any app.bsky.feed.generator, bot.did.policy chosen by the writer; meaningful only in this repository the same record
namespaced nsid com.atproto.lexicon.schema a claim on a global namespace the account is authoritative for the same record
timestamped tid app.bsky.feed.post, app.bsky.graph.follow assigned at write time; a new one every time a new record

The first three carry identity: something outside the repository points at that key and will read whatever is there next. The fourth does not: a recreated post has a new at:// URI and nothing that pointed at the old one follows it. For a collection the server has no lexicon for, the key's shape decides — a TID is timestamped, anything else is named.

The difference between named and namespaced is what a policy about the key can say. A named key means nothing outside the repository, so a rule about it is a pattern. A namespaced key is a claim the account may not be entitled to make, and the rule that matters is a prefix: this account publishes schemas under bot.did and nowhere else. The server cannot check authority for an arbitrary domain, but a policy can bound the claim.

What an account produces #

data where it lives named by
record bodies records/pds.heap, append-only frames a slot in the journal
record slots the journal; a checkpoint rewrites live ones did, collection, rkey
tombstones the journal, alongside slots (this epic) did, collection, rkey
blob bytes blobs/<did>/<cid> the blob index
blob index the journal did, cid, reference count
repository history the journal; head only survives a checkpoint did
credentials the journal did
account row and signing key the journal did
DID document and DNS derived from the row the host name
the name the name ledger the handle
the ledger the journal; every checkpoint keeps it whole the deployment

What removes data, and what refuses it #

Nine events remove something. Each has a set of things that can refuse it, and the set is different for each, which is the reason for the table below rather than a rule.

  • An agent deletes a record. Refused by a lock on the account and by a policy statement that denies deletes in that collection.
  • An agent replaces a record. An edit, or a delete followed by a create at an identity-carrying key. Refused by a lock and by any policy statement that would refuse the edit.
  • An agent erases itself. Refused by a lock and by PreventDataDeletion.
  • An operator erases the account. Refused by PreventDataDeletion.
  • The retention sweep erases the account. Never picks an account with PreventDataDeletion or Retention::Forever; otherwise runs at the retention window measured from provisioning.
  • A hard delete removes the row. Refused by PreventDataDeletion and by Retention::Forever.
  • The blob collector unlinks unreferenced bytes. Refused only by the grace window.
  • A checkpoint rewrites the journal. Drops slots no key names. Refused by nothing; it removes nothing a reader could reach.
  • Heap compaction drops dead frames. Does not exist yet; this epic adds it. Refused by an open tombstone naming the frame.

The table #

Rows are the data; columns are the events. A cell says what is left. "dead" means the bytes are on disk with nothing naming them, which is what a checkpoint leaves today for every deleted body.

data agent deletes record agent erases self / operator erases sweep erases hard delete blob collector checkpoint heap compaction
record body dead, or held by a tombstone dead dead dead — untouched gone unless tombstoned
record slot gone gone gone gone — live ones kept —
tombstone written, for identity-carrying keys gone gone gone — kept kept
blob bytes kept; reference count drops gone gone gone gone after grace, at zero references untouched —
blob index reference count drops gone gone gone entry gone kept —
history head advances; trail kept gone gone gone — head kept, trail dropped —
credentials — revoked revoked revoked — kept —
row and key — kept; state Decommissioned gone gone — kept —
DID document — keeps resolving gone gone — — —
the name — burned, never reissued released released — — —
the ledger — event appended event appended event appended — kept whole —

Two rows change under this epic. Record bodies stop being dead forever, because heap compaction can finally drop them and a tombstone can finally keep one on purpose. And tombstones are a new row: today a delete journals three names and nothing else.

The tombstone #

A delete of an identity-carrying record journals the slot it removed, not just its name. That slot is the tombstone. A checkpoint carries it forward the way it carries a live slot, so the body stays reachable by name across any number of rewrites. Heap compaction treats a tombstoned frame as live.

The tombstone exists for one reader: policy. A create at a key that holds a tombstone is handed to the engine as an update from the tombstone's value, so a policy that guards a field is consulted the same way whether the agent edited the record, or deleted it last week and wrote a new one today. The invariant, and the test that pins it, is that the verdict on creating R' over a tombstone of R equals the verdict on updating R to R'. From that follows what an operator expects: deleting a compliant profile and writing the same one back passes both gates, and writing a different display name back is refused with the reason an edit would have carried.

A timestamped key has no tombstone. A recreate is a new record, and the delete and the create are judged on their own.

Unique keys give at most one tombstone per collection. Named and namespaced keys are unbounded, and an agent churning keys must not be able to grow the journal without limit, so each collection keeps a bounded number of tombstones, oldest dropped first.

Heap compaction #

The heap is append-only and its only truncation is the torn tail at boot. Compaction copies every frame a slot or a tombstone names into a new heap, rewrites those slots with their new offsets in one checkpoint, and renames the heap over the old one. The fixed-point and byte-identity tests that pin today's checkpoint gain a third: a compacted heap answers every read the old one did and holds no frame nothing names. Until this exists, an erasure the lifecycle calls complete leaves the account's record bodies on disk, and a compliance request to remove them has no mechanism.

The blob collector runs #

collect_blobs exists, is tested, and is called by nothing in the serving binary. The stale sweep has a timer; the collector gets the same one.

Decisions for the owner #

Work #

Done #