From 45cb22d6464ec3285cad8a2ceca900ffececbf43 Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Sat, 15 Aug 2026 21:43:33 -0500 Subject: [PATCH] docs(data-structures): add three foundational structures --- data-structures/README.md | 3 + data-structures/b-plus-trees.md | 200 ++++++++++++++++++++++++++++ data-structures/ring-buffers.md | 226 ++++++++++++++++++++++++++++++++ data-structures/union-find.md | 200 ++++++++++++++++++++++++++++ site/style.css | 28 ++++ 5 files changed, 657 insertions(+) create mode 100644 data-structures/b-plus-trees.md create mode 100644 data-structures/ring-buffers.md create mode 100644 data-structures/union-find.md diff --git a/data-structures/README.md b/data-structures/README.md index 6e47125..c3ba528 100644 --- a/data-structures/README.md +++ b/data-structures/README.md @@ -3,4 +3,7 @@ implementation notes on the structures themselves: their contracts, tradeoffs, layout, and behavior on real hardware. +- [B+ trees](./b-plus-trees.md) — high-fanout ordered indexes, linked leaves, rebalancing, and bulk loading - [bloom filters](./bloom-filters.md) — probabilistic membership, sizing, hashing, deletion, and locality +- [ring buffers](./ring-buffers.md) — bounded queues and histories, wraparound, ownership, and concurrency +- [union–find](./union-find.md) — disjoint sets, weighted union, path compression, and dynamic limits diff --git a/data-structures/b-plus-trees.md b/data-structures/b-plus-trees.md new file mode 100644 index 0000000..3010549 --- /dev/null +++ b/data-structures/b-plus-trees.md @@ -0,0 +1,200 @@ +# B+ trees + +A B+ tree is a balanced ordered index designed to make each expensive memory or +storage access do a large amount of useful work. Internal nodes contain routing +keys and child pointers; records live only in leaves; leaves form an ordered +linked sequence. + +That separation is the defining move. Internal nodes stay compact and achieve +high fanout, while linked leaves make a range scan continue without walking +back up the tree. + +## shape + +
+ + A three-level B plus tree + A root routes to internal nodes, which route to linked leaf pages containing all records. A highlighted search path reaches key 31, then a range scan follows leaf links to the right. + + routing levels + 30 | 60 + + 10 | 20 + 40 | 50 + 70 | 80 + + + linked leaves — all records live here + 2 6 9 + 10 14 18 + 20 25 29 + 31 36 39 + 40 44 48 + →→→→ + + range scan continues through leaf links + +
A lookup descends through separators. Once it reaches the first matching leaf, a range scan moves horizontally through the record-bearing leaves.
+
+ +A B+ tree of height `h` performs at most one node access per level. If each +internal node has fanout `f`, it indexes on the order of `f^h` leaves. A fanout +in the hundreds keeps even very large trees shallow. + +The word *order* is unfortunately inconsistent across books: it may mean the +maximum number of children or the minimum degree. State capacities directly +when specifying an implementation. + +## invariants + +For a conventional B+ tree: + +- every leaf is at the same depth; +- keys within a node are sorted; +- an internal separator divides the key ranges of adjacent children; +- every record appears in a leaf, never only in an internal node; +- every non-root node obeys an occupancy bound, commonly between half-full and + full; +- linked leaves preserve global key order. + +Separator conventions differ. A parent may store the smallest key in its right +child, the largest key in its left child, or another equivalent fence key. The +choice changes comparison details, not the structure. It must be applied +consistently after splits, merges, and changes to a child's boundary key. + +## lookup and range scan + +Point lookup compares the search key with an internal node's separators, picks +one child, and repeats until reaching a leaf. Searching within each node can use +binary search, interpolation, SIMD comparison, or a short linear scan; the best +choice depends on node size and key representation. + +```text +find(key): + node = root + while node is internal: + node = child selected by node.separators + return search(node.records, key) +``` + +A range scan first performs a lower-bound lookup, then walks records and leaf +links until the upper bound is crossed. For `z` returned records and `B` records +per leaf, its structural cost is approximately + +```text +O(log_f(n) + z/B) node accesses +``` + +This is why a B+ tree supports both selective point queries and ordered scans +without maintaining two unrelated representations. + +## insertion + +Insertion descends to the target leaf and inserts the record in order. If the +leaf overflows: + +1. split its records between a left and right leaf; +2. repair the leaf links; +3. copy a boundary key into the parent with a pointer to the new leaf; +4. recursively split the parent if it overflows; +5. create a new root if the old root splits. + +The tree grows upward, one root split at a time. Copying a separator upward is +important: unlike a classic B-tree split, the corresponding record remains in a +leaf because leaves are the authoritative record layer. + +## deletion + +Deletion removes a record from its leaf. An underfull node can borrow from a +sibling and update the parent separator, or merge with a sibling and remove one +parent entry. Merges may propagate to the root; a root with one child can be +replaced by that child, reducing the height. + +Real systems sometimes relax immediate rebalancing. Delayed merges reduce write +amplification and contention at the cost of lower occupancy. That is a policy +choice only if searches still follow correct separators and all leaves remain +reachable. + +## bulk loading + +Repeated insertion builds a valid tree but does unnecessary searches and +splits when the input is already sorted. Bottom-up bulk loading instead: + +1. packs sorted records into leaf pages at a chosen fill factor; +2. links the leaves; +3. builds each parent level from child boundary keys; +4. repeats until one root remains. + +This is linear in the input size after sorting and produces predictable +occupancy. Leaving deliberate free space is useful for an index that will later +receive random inserts; packing to 100% is appropriate for immutable segments. + +## pages, caches, and fanout + +B+ trees are usually page-shaped rather than pointer-object-shaped. A node is +sized to a disk page, flash page, cache line group, or explicit memory block. +For page size `P`, per-child key bytes `K`, pointer bytes `R`, and fixed overhead +`H`, rough internal fanout is + +```text +f ≈ floor((P - H) / (K + R)) +``` + +More fanout means less height, but node-local search and rewrite cost grow. +Practical layouts therefore use techniques such as prefix compression, +variable-length slots, abbreviated separators, overflow pages for large values, +and separate key and payload storage. + +A page cache adds another invariant: a cursor cannot retain pointers into an +evictable page unless it pins that page or copies the required data. Tree +correctness and memory-lifetime correctness meet at that boundary. + +## concurrency and crash safety + +An in-memory single-writer tree can mutate nodes directly. Concurrent or durable +trees need more machinery: + +- **latch coupling** holds a parent until the child is known safe to modify; +- **B-link trees** add right-sibling links and high keys so readers can recover + from concurrent splits without holding the whole search path; +- write-ahead logging or copy-on-write ordering ensures a crash cannot publish a + parent pointer to an incomplete child; +- page identifiers must remain stable while readers or recovery records refer + to them. + +These mechanisms do not change the abstract map. They make structural changes +observable in a safe order. + +## complexity + +| operation | structural cost | +| --- | --- | +| point lookup | `O(log_f n)` node accesses | +| insert/delete | `O(log_f n)` amortized, including rebalancing | +| lower/upper bound | `O(log_f n)` | +| ordered scan of `z` records | `O(log_f n + z/B)` | +| bottom-up build from sorted input | `O(n)` | +| space | `O(n)` | + +The constants are the point: high fanout turns the logarithm into a very small +number of page accesses. + +## sources + +- Rudolf Bayer and Edward M. McCreight, “Organization and Maintenance of Large + Ordered Indexes,” *Acta Informatica* 1, 1972, + [doi:10.1007/BF00288683](https://doi.org/10.1007/BF00288683). +- Douglas Comer, “The Ubiquitous B-Tree,” *ACM Computing Surveys* 11(2), 1979, + [doi:10.1145/356770.356776](https://doi.org/10.1145/356770.356776). +- Philip L. Lehman and S. Bing Yao, “Efficient Locking for Concurrent Operations + on B-Trees,” *ACM TODS* 6(4), 1981, + [doi:10.1145/319628.319663](https://doi.org/10.1145/319628.319663). +- SQLite, [Database File Format: B-tree Pages](https://www.sqlite.org/fileformat2.html#b_tree_pages). + +### in these projects + +- [burner-redis `ScoreIndex`](https://tangled.org/zzstoatzz.io/burner-redis/blob/main/src/internal/zset/btree.zig) + is a B+-style ordered index for sorted-set members: internal boundary keys, + linked leaves, range cursors, splits, and bottom-up construction. +- `zlite/src/btree.zig` implements point lookup and in-order cursors over SQLite + table B-tree pages; the repository is currently local-only. diff --git a/data-structures/ring-buffers.md b/data-structures/ring-buffers.md new file mode 100644 index 0000000..63be392 --- /dev/null +++ b/data-structures/ring-buffers.md @@ -0,0 +1,226 @@ +# ring buffers + +A ring buffer stores a bounded sequence in a fixed array while treating the end +of the array as adjacent to the beginning. Logical movement is monotonic; +physical positions wrap modulo the capacity. + +It is also called a circular buffer. The structure is simple, but its contract +is not: a full producer may reject, block, or overwrite; readers may consume +items or inspect retained history; and concurrent versions depend on a precise +publication protocol. + +## logical order over physical wraparound + +
+ + A wrapped ring buffer + An eight-slot array contains four logical items in slots six, seven, zero, and one. The read cursor points to slot six and the write cursor points to slot two. Logical order wraps from slot seven to slot zero. + + physical array + + + CD····AB + 01234567 + + + logical order: A → B → C → D + read = 6 + write = 2 + capacity = 8 · length = 4 · next push writes slot 2 · next pop reads slot 6 + +
Array indices wrap, but sequence order does not. Cursors advance monotonically in the logical model and are reduced to a slot only when accessing storage.
+
+ +With capacity `C`, a logical position maps to storage as + +```text +slot(position) = position mod C +``` + +When `C` is a power of two, unsigned positions can use + +```text +slot(position) = position & (C - 1) +``` + +The mask is an implementation optimization, not a different structure. +Monotonically increasing counters are often easier to reason about than cursors +that themselves wrap at `C`: occupancy is `write - read`, while only array +access applies the mask or modulo. + +## distinguishing empty from full + +If only reduced `read` and `write` indices are stored, `read == write` is +ambiguous: it can mean empty or full. Three common representations resolve it: + +1. reserve one slot, so usable capacity is `C - 1` and equality always means + empty; +2. store an explicit length in `[0, C]`; +3. keep monotonic read and write counters, so equality means empty and + `write - read == C` means full. + +The third form composes especially well with sequence numbers and concurrency, +but counter wrap must still be defined. Unsigned subtraction is safe only while +the live distance stays within the chosen half-range assumptions. + +## queue versus rolling history + +A bounded FIFO and a rolling history can use the same representation but have +different full-buffer semantics. + +### bounded queue + +A push into a full queue can: + +- fail immediately, applying backpressure to the caller; +- block until a consumer advances `read`; +- drop the new item. + +Existing unread data remains intact. + +### rolling history + +A push into a full history overwrites the oldest item and advances both ends: + +```text +storage[write mod C] = item +write += 1 +if write - read > C: + read = write - C +``` + +This keeps the newest `C` items. It is appropriate for telemetry, replay +windows, recent events, and diagnostic context—not for work that must be +processed exactly once. + +Naming the policy is part of the type. A generic `push` whose overflow behavior +is implicit invites data loss. + +## operations and complexity + +| operation | cost | +| --- | ---: | +| push one item | `O(1)` | +| pop one item | `O(1)` | +| peek oldest/newest | `O(1)` | +| iterate `n` live items | `O(n)` | +| storage | `O(C)` fixed | + +The fixed bound prevents allocator growth and makes memory use predictable. +Contiguous storage gives good locality, though one logical range may require two +physical slices: + +```text +[first physical run: read .. end] +[second physical run: 0 .. write] +``` + +APIs that expose these two slices can perform batched I/O without copying the +wrapped sequence into a temporary contiguous buffer. + +## ownership and overwrite + +For inline values, overwrite is an assignment. For pointers, strings, or other +owned resources, replacing the oldest slot must first destroy or transfer the +old value. Likewise, `pop` must define whether ownership moves to the caller or +the element is borrowed until the next mutation. + +Useful slot states are therefore not always just “occupied.” A concurrent or +fallible producer may need `empty`, `being written`, and `published` states so a +consumer never observes half-initialized data. + +## replay by sequence number + +A history ring often stores a monotonic sequence number with each item. This +separates identity from physical slot reuse: + +```text +slot = sequence mod C +``` + +A reader asking for everything after cursor `s` must compare `s` with the +oldest retained sequence. If the cursor precedes the retained window, replay is +incomplete and the API should report a gap rather than silently returning a +suffix that looks complete. + +Sequence tags also detect stale reads: slot `i` may exist, but if its stored +sequence is not the requested one, that slot has already been reused. + +## concurrency is a protocol + +A mutex around push and pop is correct and often sufficient. Lock-free rings +need stronger assumptions and are not interchangeable. + +### single producer, single consumer + +An SPSC ring can give each cursor one writer: + +1. producer writes the item into its slot; +2. producer release-stores the new write position; +3. consumer acquire-loads that position before reading the item; +4. consumer release-stores the new read position after consuming it. + +The acquire/release pair publishes slot contents. Relaxing it without a proof +can expose a cursor before the corresponding bytes are visible. + +### multiple producers or consumers + +A shared `fetchAdd` cursor alone is insufficient: it reserves positions but can +publish position `n + 1` before producer `n` has finished writing. Bounded MPMC +rings commonly attach a sequence number to every slot. Producers and consumers +claim a slot only when its sequence denotes the expected generation, then +publish the next generation after writing or reading. + +This solves slot reuse and publication ordering together, at the cost of more +metadata and atomic traffic. False sharing also matters: heavily written producer +and consumer cursors should not occupy the same cache line. + +## failure and shutdown + +A production ring needs behavior for more than full and empty: + +- Can a blocked operation be cancelled? +- Does shutdown drain existing items or discard them? +- How is a producer failure after reservation represented? +- Are dropped items counted and observable? +- Does a snapshot hold the lock, copy items, or tolerate concurrent overwrite? + +A bounded structure turns overload into a decision. That decision should be +visible in metrics and in the API contract. + +## when it fits + +Ring buffers are a strong fit when order matters, capacity is naturally bounded, +and old storage can be reused: + +- producer/consumer handoff; +- network and audio buffers; +- rolling logs and telemetry; +- cursor replay windows; +- scheduler ready queues; +- fixed-window batching. + +They are a poor fit when capacity must grow without a hard policy, arbitrary +middle insertion or removal is common, stable addresses are required, or every +item must survive overload without external backpressure. + +## sources + +- Leslie Lamport, “Specifying Concurrent Program Modules,” *ACM Transactions on + Programming Languages and Systems* 5(2), 1983, + [doi:10.1145/69624.357207](https://doi.org/10.1145/69624.357207). +- Linux kernel documentation, + [Circular Buffers](https://docs.kernel.org/core-api/circular-buffers.html). +- Paul E. McKenney, *Is Parallel Programming Hard, And, If So, What Can You Do + About It?*, + [Circular Buffers](https://mirrors.edge.kernel.org/pub/linux/kernel/people/paulmck/perfbook/perfbook.html). + +### in these projects + +- [pub-search's pending-search buffer](https://tangled.org/zzstoatzz.io/pub-search/blob/main/backend/src/metrics/buffer.zig) + is a mutex-protected rolling queue: when full, it frees and drops the oldest + search before inserting the newest, then drains ownership into a local batch + before performing network I/O. +- [Coral's pulse history](https://tangled.org/zzstoatzz.io/coral/blob/main/backend/src/entity_graph.zig) + uses a monotonic pulse sequence and `sequence mod capacity` to retain a fixed + recent window for frontend replay. diff --git a/data-structures/union-find.md b/data-structures/union-find.md new file mode 100644 index 0000000..6e386b1 --- /dev/null +++ b/data-structures/union-find.md @@ -0,0 +1,200 @@ +# union–find + +Union–find, or the disjoint-set union structure, maintains a partition of a +fixed universe into non-overlapping sets. It answers one question extremely +well: *are these two elements connected by the unions observed so far?* + +Its interface has three operations: + +- `makeSet(x)` creates the singleton set `{x}`; +- `find(x)` returns a representative for the set containing `x`; +- `union(a, b)` merges the two sets when their representatives differ. + +Representatives are implementation identities, not meaningful members chosen by +the caller. Equality of representatives is what proves connectivity. + +## a forest, flattened by use + +The standard representation is a forest of parent pointers. Every set is one +tree; its root is the representative and points to itself, or stores a special +root marker. `find` follows parents to a root. `union` changes one root's parent +to the other root. + +
+ + Path compression in a union-find forest + Before find, nodes seven, six, and four form a chain to root zero. After finding seven, all nodes on that path point directly to zero. + + before find(7)after find(7) + + + 0 + 4 + 6 + 7 + 2 + 5 + + + + 0 + 2 + 4 + 6 + 7 + 5 + + three parent reads reach the rootfuture finds take one parent read + +
Path compression rewrites every parent followed by a find to point at the representative. The partition is unchanged; only its internal shape improves.
+
+ +A compact parent array often stores roots as negative set sizes: + +```text +parent[x] >= 0 → parent index +parent[x] < 0 → x is a root; -parent[x] is the set size +``` + +This combines parent pointers and rank metadata in one machine word per element. +It requires elements to have dense integer identifiers, or a separate map from +external keys to those identifiers. + +## weighted union + +A naive union can repeatedly attach a large tree beneath a singleton and create +a linear chain. **Union by size** attaches the smaller root beneath the larger; +**union by rank** attaches the shallower tree beneath the deeper one and raises +the rank only when equal ranks meet. + +```text +union(a, b): + ra = find(a) + rb = find(b) + if ra == rb: return false + if size[ra] < size[rb]: swap(ra, rb) + parent[rb] = ra + size[ra] += size[rb] + return true +``` + +Without path compression, weighted union keeps tree height logarithmic. With +path compression, each successful or unsuccessful query also makes future +queries cheaper. + +## path compression + +A full compression find first locates the root, then walks the path again and +rewrites every parent: + +```text +find(x): + root = x + while parent[root] is not a root: + root = parent[root] + + while x != root: + next = parent[x] + parent[x] = root + x = next + + return root +``` + +Two one-pass relatives are common: + +- **path splitting** makes every visited node point to its grandparent; +- **path halving** does this for every other node. + +They can have simpler loops and excellent practical behavior. All preserve the +same partition. + +## complexity + +A sequence of `m` operations on `n` elements using weighted union and path +compression costs + +```text +O(m α(n)) +``` + +where `α` is the inverse Ackermann function. It grows so slowly that it is below +5 for any practical input size. The useful interpretation is *amortized almost +constant time*, not literally constant time: one particular `find` may still +walk a path, and the bound applies across the operation sequence. + +| implementation | worst tree height | amortized operation cost | +| --- | ---: | ---: | +| arbitrary linking | `O(n)` | `O(n)` | +| union by size/rank | `O(log n)` | `O(log n)` | +| size/rank + path compression | — | `O(α(n))` | + +Space is `O(n)`. + +## what it does not maintain + +Union–find remembers connectivity, not the edges that established it. It cannot +list a path between two elements unless another graph representation keeps the +edges. It also does not efficiently support splitting a set or deleting an +arbitrary union: one removed edge may or may not disconnect the component, and +the parent forest contains too little information to decide. + +For a changing graph, common choices are: + +- process additions online and rebuild after deletions; +- answer an offline sequence in reverse, turning deletions into unions; +- use rollback union–find with a segment tree over time; +- use a fully dynamic connectivity structure when the added complexity is + justified. + +Path compression conflicts with simple rollback because one `find` mutates many +parents. Rollback implementations usually omit compression, use union by size, +and record each changed word on a stack. + +## applications + +- connected components in an incrementally built undirected graph; +- Kruskal's minimum-spanning-tree algorithm, rejecting edges whose endpoints + already share a representative; +- percolation and image-component labeling; +- equivalence classes in compilers, type inference, and symbolic processing; +- offline connectivity and account/entity merging. + +The fit is strongest when relationships only accumulate during the lifetime of +the structure. + +## representation and correctness + +The fundamental invariant is that every parent chain terminates at exactly one +root. Useful debug checks include: + +- a root's stored size equals the number of elements that reach it; +- the sum of root sizes equals the universe size; +- non-root parents are valid indices; +- `find(find(x)) == find(x)`; +- a successful union decreases the number of roots by exactly one. + +Concurrency is not obtained by merely making parent words atomic. Two unions can +race while choosing roots, lose a size update, or form a cycle. Concurrent +variants need a defined linking order plus compare-and-swap loops, or a lock +around mutations. Coarse locking is often the right first implementation because +ordinary union–find operations are already very cheap. + +## sources + +- Bernard A. Galler and Michael J. Fisher, “An Improved Equivalence Algorithm,” + *Communications of the ACM* 7(5), 1964, + [doi:10.1145/364099.364331](https://doi.org/10.1145/364099.364331). +- Robert Endre Tarjan, “Efficiency of a Good But Not Linear Set Union + Algorithm,” *Journal of the ACM* 22(2), 1975, + [doi:10.1145/321879.321884](https://doi.org/10.1145/321879.321884). +- Robert E. Tarjan and Jan van Leeuwen, “Worst-Case Analysis of Set Union + Algorithms,” *Journal of the ACM* 31(2), 1984, + [doi:10.1145/62.2160](https://doi.org/10.1145/62.2160). + +### in these projects + +- [Coral's lattice](https://tangled.org/zzstoatzz.io/coral/blob/main/backend/src/lattice.zig) + uses a negative-size parent array, weighted union, and full path compression + to track connected occupied sites and detect percolation. Because expiring a + site is a deletion, it periodically rebuilds the structure from live sites. diff --git a/site/style.css b/site/style.css index cde7e44..5ac4e71 100644 --- a/site/style.css +++ b/site/style.css @@ -643,6 +643,34 @@ a:hover { font-size: 13px; font-weight: 600; } +.content .ds-figure .ds-node { + fill: var(--panel); + stroke: var(--muted); + stroke-width: 1.5; +} +.content .ds-figure .ds-node-accent { + fill: var(--accent); + stroke: var(--link); + stroke-width: 2; +} +.content .ds-figure .ds-edge { + fill: none; + stroke: var(--muted); + stroke-width: 1.5; +} +.content .ds-figure .ds-edge-accent { + fill: none; + stroke: var(--link); + stroke-width: 2; +} +.content .ds-figure .ds-muted { + fill: var(--muted); + font-size: 11px; +} +.content .ds-figure .ds-accent { + fill: var(--link); + font-weight: 600; +} .content .ds-figure figcaption { margin-top: 10px; color: var(--muted); -- 2.51.2