Identities for entities did.bot
agent llm did
didbot plan store-scale.md
33 kB
Markdown
at commit 18ba4fe0


id: store-scale title: The state is larger than the process holding it status: open crates: [didbot-pds] dependsOn: [] exitCriterion: > The server serves a data set larger than its resident memory, and restarts in a time that does not grow with the history behind it. #

store-scale #

Durable is the memory stores with a write-ahead log underneath. Every record, every account, every signing key, the blob index and the ledger are resident, and the log is the durability rather than the storage. Total capacity is therefore memory, and it is a wall rather than a slope: nothing degrades on the way there.

The second limit is in the same file and is easy to miss. Replay reads the whole log at startup, and compaction happens only at startup — the module says so and calls it a real limit rather than an oversight. So retained state makes restarts slower, slower restarts make restarts rarer, and rarer restarts are fewer compactions. The thing that bounds the log is discouraged by the log growing.

Neither of these bites while accounts are swept aggressively. account-types is a decision to keep more, for longer, so both become load-bearing.

The items below were written one kind at a time and they are one change. Durable storage as a seam is the proposal that treats them that way: one contract every store implements, a second for the kinds whose values are large, a third for the kind that is a function of the others, and a checkpoint the whole set shares. It carries the order the kinds move in and what each move touches.

Durable storage as a seam: a proposal #

Status: awaiting the owner's sign-off. This is the seam the rest of this epic is written against, and the parts of it that are an on-disk format are gathered under What needs signing off. A format is expensive to change once a deployment has written one, which is why it is a decision rather than a review comment.

The change this proposes is general rather than record-shaped. Everything Durable::open holds is a memory store with one log underneath it, and the difference between the kinds is what an entry means and how big the state is, never how durability works. So the proposal is one contract that every kind implements, a second that the large-valued kinds also implement, a third for the one kind whose state is a function of the others, and a checkpoint protocol that all three share.

The three contracts #

Journaled — every store #

pub trait Journaled: Send + Sync {
    /// This store's name in a checkpoint and in a log line.
    const STORE: &'static str;

    /// Applies one replayed fact. Total: a fact was accepted before it was
    /// written, so there is nothing here to refuse.
    fn apply(&self, entry: Entry);

    /// The state, as the shortest sequence of facts that reproduces it.
    fn checkpoint(&self) -> Vec<Entry>;

    /// The journal mark at or before which this store's state is durable
    /// without the journal.
    fn durable_through(&self) -> Mark;
}

apply and checkpoint are the two halves durable::apply and durable::snapshot already are, split per store instead of held in two matches. That split is most of the value on its own: it is what lets each kind be worked on in its own file.

The contract's obligations run in both directions. A store must reconstruct the same state from checkpoint() that it has now, and where its checkpoint may hold less than its history did it must say which observable changes — the agent token store is the standing example, and durable::snapshot already argues it: an expired token is written down rather than dropped, because dropping it would turn an Expired answer into an Unknown one. And the caller must apply facts in journal order and must never apply one twice.

That second obligation is not a convenience. Most of durable::apply is idempotent and Entry::LedgerAppended is not — it appends. So the design takes the strict rule everywhere rather than a per-store one: a journal position is honoured exactly and never crossed. Everything about trimming follows from that sentence.

Bodied — the kinds whose values are large #

pub trait Bodied: Journaled {
    /// Where the bytes live, and under what name.
    fn heap(&self) -> &Heap;
}

A Heap is an append-only file of crate::wal::frame frames — a length and a CRC-32 over both the length and the payload, which is the framing pds.wal uses and which crate::evaluation_log already reuses for a second file for exactly this reason. A value is written to the heap and the journal carries a Slot { offset, len } and the value's content identifier instead of the value.

FileBlobStore is already this, spelled by hand: Entry::BlobUploaded carries a reference and not the bytes, and durable.rs gives the reason in as many words — a log that carried blobs would be a log read into memory in full at every startup to find out what an account is called. The proposal is to make that a contract two kinds share rather than one kind's arrangement.

The ordering the contract requires is the ordering the blob store already keeps, and it is one-directional: bytes reach the medium before the fact that names them does. A crash between them leaves bytes nothing references, which is dead space; the reverse leaves a fact naming bytes nothing has, which is a read that cannot be served. With Durability::Deferred this means the heap syncs before the journal on the same timer, so the values a power cut costs are a suffix of write order in both files — the loss Deferred already documents.

Recovery is decided by the log's own rule rather than by a new one. A slot lying beyond the heap's length is a torn tail: the value is dropped with a line naming it, and every later slot must also lie beyond. A slot lying within the heap's length whose frame fails its CRC is damage in the middle, and the directory is refused — pds.wal's replay says plainly that nothing can tell a torn tail from damage in the middle and that guessing is the failure mode worth avoiding, and that sentence carries over unchanged. Both are decided in one pass, because a slot is either wholly inside the heap's length or it is not.

Derived — the kind that is a function of the others #

pub trait Derived: Journaled {
    /// Rebuilt after every store's checkpoint has loaded, in declared order.
    fn rebuild(&self, from: &Stores);
}

One kind needs this and the reason is worth stating, because it is the place where moving records makes another store's job harder rather than easier. BlobIndex::snapshot deliberately carries no reference count, because durable::snapshot's ordering — blobs before records — reconstructs it as each replayed record's records::blob_refs marks what it points at. Once a record's fact is a slot rather than a body, blob_refs has nothing to read, and the count cannot be reconstructed by ordering any more.

So a derived store either carries what it derives, or declares a rebuild that runs once after the checkpoint has loaded and has the other stores in hand. Carrying the count is the cheaper answer and Derived is the honest shape for the general case, since FileBlobStore::reconcile is already a rebuild of this kind against the disk.

Which kind a type takes #

The three contracts cover the shapes because the shapes differ along exactly two axes, and neither is the one the shapes are usually named by:

  • Is a value large enough that the journal should not carry it? If yes, Bodied. This is the only axis that touches the format.
  • Is the state a function of other stores' state? If yes, Derived.
  • Everything else is Journaled and nothing more, whatever its cardinality.

"Append-only" is not one of the axes, and that is the useful finding. The ledger and the firehose reservation look alike and are opposites: the reservation's apply is a maximum and its checkpoint is one entry, and the ledger's apply appends and its checkpoint is the whole of it. What separates them is idempotence, and the strict position rule above means neither of them needs the contract to know which. Nor is "has expiry" an axis: expiry is something checkpoint() may act on, under the obligation to say what it changes.

The checkpoint, and a journal that can be trimmed #

A write known to be durable need not be replayed again, so the journal ahead of a durable position is the only part a restart reads. The general form of that is a checkpoint.

What a checkpoint holds #

durable::snapshot is already the answer: the whole state, as the shortest log that would reproduce it. What compaction does today is write it as the head of a new log, which is why it can only run when nothing else is writing. A checkpoint is the same bytes in their own file, with the journal position they cover, and that one change is what makes them writable while the journal keeps taking appends.

Keeping "every store contributes" true as stores are added — OAuth grants are the next — should be structural rather than written down. ENTRY_SHAPE is the model: a test that enumerates every Entry variant and requires each to be either carried by some store's checkpoint() or declared as carrying no state, so a store added without contributing fails a test rather than losing data quietly.

How a position is declared, and how the trim point is derived #

durable_through is the per-store declaration, and the global trim point is the minimum over every store. That is the general rule and it is worth having even though today's answer is uniform: every store's only durability is the journal, so every store's position is the last checkpoint's and the minimum is degenerate.

It stops being degenerate exactly when a type moves to durability of its own, which is the whole programme. A Bodied store's bytes are durable at their own sync point; its index is durable at the checkpoint. A store that later keeps its own state file advances its own position without waiting for a global checkpoint, and the minimum is what lets it do that without a flag day. So the method exists so that types can move one at a time, which is the property the owner asked the seam to have.

How the checkpoint survives a crash #

The failure that loses data is a checkpoint claiming more than is durable, and the ordering that forecloses it is one-directional: the checkpoint becomes durable, and only then is anything trimmed. A crash between the two costs nothing — a valid checkpoint over an untrimmed journal is a correct state, because the position says where to start reading and what precedes it is skipped.

The checkpoint is written to a scratch name, framed and CRC-checked the way everything else in this directory is, with the position it covers in its own final frame; then sync_all, a directory sync, a rename over pds.checkpoint, and a directory sync. A checkpoint whose trailer frame is absent or fails its CRC is not a checkpoint, and a boot that finds one falls back to the last position it can prove. The standard is tests/restore.rs's: a copy cut at any frame boundary opens self-consistently or is refused, and the drill grows a second file and then a set of segments to cut.

Trimming, on one EBS volume #

  • Rewrite, which the log did before it was segmented: build the replacement, sync, rename. It copies the whole state, which is why it was a startup operation, and a rename over one file cannot replace several at once — so it is the option a segmented log cannot take.
  • Punching the front out with fallocate. Offsets stay stable, which suits a position, at the cost of a libc dependency, a filesystem-specific syscall, and a file whose length grows forever while its blocks do not. Rejected: a rename reclaims the same space with machinery this repository already has and already tests.
  • Segments — pds.wal.000001, rolled at a size, a whole segment unlinked once the trim point covers its last entry. Unlink is the cheapest reclaim any filesystem offers, it needs nothing new, and it gives a position a stable address, Mark(segment, offset), that a trim cannot renumber. That is what settles it: under a rewrite the offsets restart, so the position has to be re-expressed, and the window in which a checkpoint's position names a file that has been replaced is exactly the window the strict position rule exists to close.

So segments are the trim, for a checkpoint taken at startup and for one taken online alike.

One log, in segments #

One journal, and the reasoning is the checkpoint's rather than a preference. A position is only meaningful as a position in one sequence. Two journals need a position each and a rule for reconciling them, which is a distributed commit inside a single process — and the thing that makes people want two, a slow store holding back a fast one's trim, is what the per-store durable_through minimum already provides. Segments are one order written in pieces, not a second order. A heap is not a journal either: it has no order of its own and is reached only through locators the one journal carries, which is precisely the relationship FileBlobStore already has with it.

The way to make the journal smaller is to keep bytes out of it and to trim what is durable elsewhere, rather than to add more journals.

Records, as the worked example #

Records are the largest kind and the first mover, and the shape of their move is what the per-type agents should read as the pattern.

<data>/records/pds.heap is the Bodied heap. A frame's payload is one record's canonical DAG-CBOR: the bytes records::content_id hashes to name it, and the bytes a CAR export files under that name.

A key does not find a location by name. (did, collection, rkey) finds a Slot in the same three-level BTreeMap MemoryRecordStore::repos is today, with Slot where serde_json::Value is. The key structure stays where it is; the values leave. So the claim is narrower than "not resident", and the measurement above is what it should be judged against: a record costs 1 105 resident bytes today, and under this it costs its three keys, their share of three BTreeMap nodes, and a slot. Capacity stops being a wall at total record bytes and becomes a slope in key count — which is what the exit criterion asks for, and it is not independence from the record count.

FileRecordStore::put keeps check, append, apply, under one lock, with one more append inside it: validate, name and resolve the key outside the lock as now; take the lock; check the precondition against the slot's stored CID; append the body to the heap; append Entry::RecordStored — the three names, the CID and the slot — to the journal; insert the slot.

The seam's guarantees move in one direction only. Written is unchanged, and its cid becomes the hash of bytes on disk as well as of the record handed in, so a read can prove it got what it asked for. Precondition keeps its meaning and gets cheaper: today the check reads the record back and recomputes content_id, and the slot already holds that value. The rkey rules are untouched, and the key still travels in the entry for the reason Entry::RecordPut already gives. RecordError::TooLarge changes which number it is measured against in the generous direction, because the frame limit applies to the body alone rather than to the body plus its envelope, so no record accepted today is refused after. RecordError::StorageFull keeps its shape, with the heap taking the journal's posture of truncating a failed append back and sealing if even that fails.

The cost this moves rather than removes is the read side. RecordStore::snapshot is every record in a repository, and provision::commit_write calls it on every write, because a commit is a rebuild of the whole tree — plan/repo-scale.md's subject, now linear in disk reads as well as in tree nodes. The answer is a buffer in front of the heap, keyed by offset and sized in bytes by a flag, defaulted large enough to hold the repository a write has just rebuilt. That is also what keeps the capacity claim honest: memory stops being the ceiling and becomes a knob.

The order the types move in #

The three foundations come first and only two of them are on each other's critical path.

# Pull Depends on Files
P0 This seam — plan/store-scale.md
P1 The format, all at once P0 wal/mod.rs, layout.rs, tests/upgrade.rs, tests/durability.rs
P2 The contracts, and apply/snapshot split per store P0 new journal.rs, durable.rs, and one arm each into account.rs, records.rs, blobs.rs, ledger.rs, history.rs, names.rs, credential.rs
P3 Heap, and records behind it P1, P2 new heap.rs, records.rs, durable.rs
P4 The checkpoint file P2 new checkpoint.rs, durable.rs, tests/restore.rs
P5 Segments, and trimming P4 wal/mod.rs, checkpoint.rs, tests/restore.rs

P1 and P2 are independent of each other: P1 declares entry variants and filenames and changes no behaviour, and P2 is a refactor that adds no entries. They can be dispatched together.

P2 is the pull that makes the rest dispatchable, and it should be judged on that. Its job is to move each store's replay arm and each store's checkpoint arm out of durable.rs and into that store's own module, leaving durable.rs as wiring. Every per-type pull after it then edits its own file, and the collisions in durable.rs that would otherwise serialise the whole programme do not arise.

After the foundations, the per-type moves are independent of one another with one exception:

Type Kind Depends on Files
Records Bodied P1, P2 records.rs, heap.rs, durable.rs
Blob index Derived P3, P4 blobs.rs
Accounts and signing keys Journaled P2 account.rs
Ledger Journaled, sized P2 ledger.rs
Commit history Journaled P2 history.rs
Name claims and the counter Journaled P2 names.rs
Agent tokens Journaled P2 credential.rs
OAuth grants Journaled P2, PR 538 that pull's own module
Firehose reservations Journaled P2 sequence.rs, durable.rs

The exception is the pair at the top: the blob index's move is caused by the record store's, for the reference-count reason under Derived, and it has to follow P3 rather than run beside it.

Two of these rows are smaller than they look and should be dispatched knowing it. Signing keys are 32 bytes of the 1 167 an account costs and 0.15% of a deployment's memory, as measured above, so the account row is a custody question and the capacity argument for moving it is not there — it should be decided on where a secret ought to live. And the ledger's growth is deliberate: its row is sizing rather than eviction.

The layout stamp moves twice, not nine times #

crate::layout hashes what an entry looks like and what the files are called, and both halves move for this programme. The sequencing that keeps them from moving once per type is P1, and it works because SHAPE hashes the declared ENTRY_SHAPE table rather than the code that writes an entry.

So P1 declares the whole destination in one go: every new variant added to Entry and to ENTRY_SHAPE, every new filename added to NAME_SHAPE — the record directory, the heap, the checkpoint, the segment naming — and LAYOUT to 11. Every older variant stays readable. Behaviour is unchanged, so a directory written before P1 is refused once, by name, saying which half moved, and a directory written after it is read by every later binary in the programme regardless of how many types have moved: a mid-programme directory holds a mixture of old and new variants and every build from P1 onward reads both. A directory is never half-migrated, because the shape it is judged against stopped moving at P1.

The second move is P5's, if segment naming turns out to want a shape P1 could not declare honestly, and the design's aim is that it does not.

the_data_directory_layout_is_pinned in tests/upgrade.rs fails at P1, because files appear in the directory that its pinned path set does not name. That is the tripwire doing the job it was written for, and the pin moves with the change that moved the directory. P1 is therefore a breaking change and says so with a !; the pulls after it are not.

What an operator with an existing data directory does #

Two answers, and this is the first of the decisions below. Under refuse — the policy crate::layout already states — a directory written by an earlier build fails the stamp comparison at P1, says which half moved, and the answer is the module's existing one: check out the build that wrote it, or delete it. Under migrate once, forward only, the older variants stay replayable, layout::check grows an arm accepting exactly the preceding SHAPE, and each per-type pull's first boot rewrites that type's state into its new home and stamps forward. It is one-way — the previous binary refuses the directory afterwards, so the backup taken before the upgrade is the way back — and if it is adopted it has to arrive with a fixture directory written by the preceding build and a test that boots it, which is the standard layout's own argument sets.

The recommendation depends on whether a deployment is holding state worth keeping when P1 lands, which is the owner's to say.

The alternatives, and why not #

An embedded key-value store — redb, fjall, sled, rocksdb — under all of it. This is the strongest alternative and the one a reviewer should press on: it brings range scans that ListParams maps onto directly, someone else's crash recovery, and someone else's fsync discipline, which is worth something in a process that holds every account's signing key. It is rejected for this programme on three grounds. It is a second durability mechanism beside the one already here, with its own log, its own compaction and its own total order, and "the journal says the value is there and the engine does not" is a failure class tests/restore.rs's drill has no shape for — the drill cuts one file, and it would have to cut two engines that recover independently. It is a dependency tree this workspace has not read. And it answers the wrong question first: the contracts above are what make the types movable one at a time, and an engine would still need them to say which type moves when. The place to revisit it is P4, where a store that keeps its own state is doing work the checkpoint would otherwise grow to do.

One file per value — records/<did>/<collection>/<rkey> and its analogues. A key maps to a path, there is no index, and residency really is zero. It puts the whole key space into the filesystem's namespace, which is this file's own open blob fanout question multiplied by records instead of accounts; it turns a list page into a directory scan in the filesystem's order rather than the key order ListParams promises; it adds escapings that layout::names would have to cover, against none for an offset; and a create costs an open, a write and two syncs against one append. It also makes the crash story worse, because a half-written file has no frame to fail.

The journal stays authoritative and the disk is a working set rebuilt at boot. Keep every entry as it is, spill bodies during replay, hold an index. Residency is fixed with no format change, no stamp move and no migration, and the crash window between the two writes vanishes because the working set is discarded and rebuilt. Rejected because it makes this file's second limit worse: a boot goes from replaying into memory to replaying into memory and writing the whole state back out and syncing it, so restarts get slower, which is the feedback loop at the top of this page. It also forecloses the checkpoint, which needs the disk to be an authority.

Shrink the resident value instead. The measured 13.3x is serde_json::Value's overhead over the wire form, and holding Box<serde_json::value::RawValue> or the DAG-CBOR bytes would recover most of it with no crash-safety change at all. It is not this epic — 13x of headroom is a wall further away, and the exit criterion asks for a data set larger than resident memory. It is kept here as the observation that makes an index of slots affordable, and as the thing worth doing anyway if this proposal is not adopted.

What needs signing off #

  1. The frame payload's encoding. Canonical DAG-CBOR — the bytes a content identifier names, verifiable on read, and the same bytes an export files — against the value's JSON, which is what the journal carries today. DAG-CBOR is the recommendation, and its gate is a round-trip test over didbot_data::dag_cbor::decode and Value::to_json for every shape the data model admits.
  2. Whether a Bodied store's entry carries the content identifier. It is recomputable from the bytes. Carrying it makes Precondition free, lets a boot check a slot without decoding it, and costs the journal about sixty bytes per value.
  3. Refuse or migrate, above, which is the one an operator feels.
  4. Whether P1 declares the segment naming, which is what decides whether the stamp moves once or twice.

All four are on-disk format, which is why they are here rather than in a review comment.

Done #

  • What it did before was worse than exhaust. An entry larger than the
    8 MiB a frame's length field is believed for was written, acknowledged,
    and then discarded at the next replay — along with every entry appended
    after it, because a length a reader will not believe is
    indistinguishable from a torn tail. Three acknowledged writes, one
    restart, one survivor. That is now [`WalError::TooLarge`], refused
    before a byte is written, and `RecordTooLarge` (413) on the wire.
    
    The other half is the medium filling. `write_all` reports the first
    error it hit and not how far it got, so a failed append could leave a
    partial frame in the middle of the log and cost every entry after it
    the same way. A failed append is now truncated back to the length the
    log had before it, which leaves the file byte for byte what it was; if
    even that fails the log is sealed and the deployment serves reads only,
    because a fragment at the *end* is the one thing replay knows how to
    drop. Either way the caller gets `StorageFull` (507) — not a 500, which
    would say the deployment is broken when the honest answer is "come
    back". `didbot-pds --log-budget` is the same refusal on a size somebody
    chose rather than on the size the volume turns out to be.
    
    What is not done: nothing asks the filesystem how much room is left, so
    the refusal is at the point of failure rather than ahead of it. That
    needs a `statvfs`, which needs a dependency this workspace does not
    have. The budget is the stand-in and it is off by default.
    
  • Measured, on a record of 83 bytes of DAG-CBOR, by
    `didbot-pds/tests/store_cost.rs` — a counting global allocator, so the
    figure is live heap bytes and not a resident-set high-water mark. Run
    it with `cargo test -p didbot-pds --test store_cost -- --ignored
    --nocapture`; `DIDBOT_COST_SAMPLES` sets the sample count. The numbers
    below are from 2 000; at 200 the record rows are 16% and 22% higher
    and the account rows within 3%. Every JSON object held is an `IndexMap`,
    because `cedar-policy` turns on `serde_json`'s `preserve_order`.
    
    | held | resident bytes each | against its wire form |
    | --- | --- | --- |
    | a record, in the record store | 1 105 | 13.3x |
    | an account and its key, in the account store | 1 167 | — |
    | a record, in a whole deployment | 2 209 | 26.6x |
    | an account, in a whole deployment | 22 008 | — |
    
    The multiplier the item asks about is 13x, and it is not the interesting
    number. The interesting one is the gap between the two account rows: an
    account costs 1 167 bytes in the store that holds accounts and 22 KiB in
    a deployment, so 95% of what an account costs is held somewhere other
    than the account store — its ledger, its commit head and trail, its
    credential, its registration record and its name. A hundred thousand
    accounts is therefore about 2.2 GB before a single record is written.
    
    That reorders the rest of this file. Making records non-resident
    recovers 13x on records and nothing on the 22 KiB, so for an
    account-heavy deployment it is the second fix rather than the first;
    for a record-heavy one it is still the first. Whichever it is, the
    thing to measure again afterwards is this table.