From cbbfdc378a2dd27d535d8df714fca3db4286cf75 Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Wed, 8 Jul 2026 23:42:19 -0500 Subject: [PATCH] regroup ai/ by concept; databases/ becomes storage/ MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit the ai groups were retrospectives keyed to projects (pub-search's search, phi's memory, the marvin bot's history) — project notes belong in those repos. regrouped the way the content actually clusters: - ai/retrieval/ (was embeddings/): asymmetric embedding, reciprocal rank fusion, synthesize-before-injecting — query-time mechanics, projects as evidence within - ai/memory/ (dissolves persistent-agents/): message-archives (the transcript layer across marvin's three eras + phi's network-as-archive) and deliberate-and-background-writes (the two write paths, phi/letta convergence, write-time curation, the integration-surface lesson) - ai/cluster-the-2d-projection: loose note until a second cartography note earns a group - agentic-harness -> harness-weight: the finding is the page, Pi is the evidence - databases/ -> storage/ (redis was already stretching 'database'); turbopuffer-in-production moves in as storage/turbopuffer.md — vector store operations are storage operations Co-Authored-By: Claude Fable 5 --- README.md | 4 +- ai/README.md | 16 +++-- ...ection.md => cluster-the-2d-projection.md} | 7 +- ai/embeddings/README.md | 23 ------ ai/local-models/README.md | 3 +- .../{agentic-harness.md => harness-weight.md} | 10 +-- ai/memory/README.md | 18 +++++ ai/memory/deliberate-and-background-writes.md | 56 +++++++++++++++ ai/memory/message-archives.md | 39 +++++++++++ ai/persistent-agents/README.md | 25 ------- ai/persistent-agents/phi-memory-in-parts.md | 59 ---------------- .../slack-bot-memory-rebuilds.md | 70 ------------------- ai/retrieval/README.md | 19 +++++ .../asymmetric-embedding.md | 0 .../reciprocal-rank-fusion.md | 0 .../synthesize-before-injecting.md | 0 architecture/README.md | 6 +- languages/ziglang/ziglua-ffi.md | 2 +- {databases => storage}/README.md | 3 +- {databases => storage}/redis/README.md | 0 {databases => storage}/redis/embedded.md | 0 {databases => storage}/redis/eval-lua.md | 0 {databases => storage}/redis/streams.md | 0 {databases => storage}/sqlite/fts5.md | 0 .../sqlite/sargable-joins.md | 0 .../turbopuffer.md | 0 {databases => storage}/turso.md | 0 27 files changed, 161 insertions(+), 199 deletions(-) rename ai/{embeddings/cluster-the-projection.md => cluster-the-2d-projection.md} (90%) delete mode 100644 ai/embeddings/README.md rename ai/local-models/{agentic-harness.md => harness-weight.md} (87%) create mode 100644 ai/memory/README.md create mode 100644 ai/memory/deliberate-and-background-writes.md create mode 100644 ai/memory/message-archives.md delete mode 100644 ai/persistent-agents/README.md delete mode 100644 ai/persistent-agents/phi-memory-in-parts.md delete mode 100644 ai/persistent-agents/slack-bot-memory-rebuilds.md create mode 100644 ai/retrieval/README.md rename ai/{embeddings => retrieval}/asymmetric-embedding.md (100%) rename ai/{embeddings => retrieval}/reciprocal-rank-fusion.md (100%) rename ai/{embeddings => retrieval}/synthesize-before-injecting.md (100%) rename {databases => storage}/README.md (63%) rename {databases => storage}/redis/README.md (100%) rename {databases => storage}/redis/embedded.md (100%) rename {databases => storage}/redis/eval-lua.md (100%) rename {databases => storage}/redis/streams.md (100%) rename {databases => storage}/sqlite/fts5.md (100%) rename {databases => storage}/sqlite/sargable-joins.md (100%) rename ai/embeddings/turbopuffer-in-production.md => storage/turbopuffer.md (100%) rename {databases => storage}/turso.md (100%) diff --git a/README.md b/README.md index d6cb5d7..ebea1d8 100644 --- a/README.md +++ b/README.md @@ -14,9 +14,9 @@ not a final word. ## organization -- `ai/` — the AI domain: local models, embeddings +- `ai/` — the AI domain: retrieval, memory, local models - `architecture/` — patterns above any specific tool (plus `home-infra/`, the home-box cluster) -- `databases/` — concrete storage systems (redis, sqlite, turso) +- `storage/` — concrete storage systems (redis, sqlite, turso, turbopuffer) - `languages/` — language idioms (python, zig) - `protocols/` — atproto, MCP - `sources/` — raw transcripts and captures (not notes; see its README) diff --git a/ai/README.md b/ai/README.md index 42a4472..312a2ce 100644 --- a/ai/README.md +++ b/ai/README.md @@ -1,14 +1,18 @@ # ai notes on building with and around models, wherever the specific tech lands. +pages are organized by concept; the projects these lessons came from +(pub-search, phi, the marvin slack bot) appear as evidence within them. ## contents -- [local-models/](./local-models/) — running local LLMs agentically on apple silicon: serving, tool-calling, harnesses -- [embeddings/](./embeddings/) — vector search in production: document prep, hybrid fusion, memory synthesis, turbopuffer operations, clustering pipelines -- [persistent-agents/](./persistent-agents/) — what keeps agent state alive between runs: the marvin slack bot's three memory rebuilds, and phi's assembled-from-parts design +- [retrieval/](./retrieval/) — query-time mechanics: asymmetric embedding, rank fusion, synthesis between retrieval and prompt +- [memory/](./memory/) — how agents keep state between runs: message archives, deliberate vs background writes, write-time curation +- [local-models/](./local-models/) — running models on your own hardware: serving, tool-calling, harness weight +- [cluster-the-2d-projection](./cluster-the-2d-projection.md) — reading structure off an embedded corpus: umap, hdbscan, work items from geometry adjacent, filed elsewhere on purpose: [protocols/MCP](../protocols/MCP/) (MCP is -a protocol first; it stays with atproto), and the agent-memory design notes that -live in the [bot repo's docs](https://tangled.org/zzstoatzz.io/bot) beside the -code they describe. +a protocol first; it stays with atproto), [storage/turbopuffer](../storage/turbopuffer.md) +(vector-store operations are storage operations), and the agent-memory design +notes that live in the [bot repo's docs](https://tangled.org/zzstoatzz.io/bot) +beside the code they describe. diff --git a/ai/embeddings/cluster-the-projection.md b/ai/cluster-the-2d-projection.md similarity index 90% rename from ai/embeddings/cluster-the-projection.md rename to ai/cluster-the-2d-projection.md index 23e66db..af10e64 100644 --- a/ai/embeddings/cluster-the-projection.md +++ b/ai/cluster-the-2d-projection.md @@ -1,8 +1,9 @@ # cluster the 2d projection -the phi-atlas flow builds a daily 2d map of everything an agent knows -(~thousands of points: memories, posts, cards, goals) and derives work items -from the map's structure. the pipeline order is the interesting part: +mapping an embedded corpus (thousands of points) into something structure can +be read off — clusters, neighborhoods, work items. the pipeline order is the +part that matters, shown here from the phi-atlas flow (a daily 2d map of +everything an agent knows: memories, posts, cards, goals): ``` 1536-d vectors → umap → 2d coords → hdbscan (twice) → clusters → promotion logic diff --git a/ai/embeddings/README.md b/ai/embeddings/README.md deleted file mode 100644 index 9e94b31..0000000 --- a/ai/embeddings/README.md +++ /dev/null @@ -1,23 +0,0 @@ -# embeddings - -lessons from running vector search in production across three systems: pub-search -(semantic search over atproto publications, zig + voyage + turbopuffer), phi's -memory (an agent's private memory, python + openai + turbopuffer), and the -phi-atlas flow (a daily 2d map of everything the agent knows). - -the recurring work is around the embedding call: preparing documents, fusing -and filtering results, and covering what the vector store won't do for you. - -## notes - -- [asymmetric-embedding](./asymmetric-embedding.md) — set `input_type` at both ends, and the document-prep checklist (truncation, titles, utf-8) -- [reciprocal-rank-fusion](./reciprocal-rank-fusion.md) — fusing keyword and semantic results by rank position, and why the two paths must fail independently -- [synthesize-before-injecting](./synthesize-before-injecting.md) — the stale-memory failure from prompting raw top-k, and the cheap-model pass that fixed it -- [turbopuffer-in-production](./turbopuffer-in-production.md) — operational gotchas: id limits, schema evolution, cold namespaces, stale attributes -- [cluster-the-projection](./cluster-the-projection.md) — the atlas pipeline: umap then hdbscan, and deriving work items from cluster structure - -## sources - -- [pub-search](https://tangled.sh/@zzstoatzz.io/pub-search) — hybrid keyword+semantic search backend -- [bot](https://tangled.org/zzstoatzz.io/bot) — phi's namespace memory -- my-prefect-server — the phi-atlas and docket flows diff --git a/ai/local-models/README.md b/ai/local-models/README.md index cf6f20d..74eb54b 100644 --- a/ai/local-models/README.md +++ b/ai/local-models/README.md @@ -16,8 +16,7 @@ an engine that returns structured tool calls. - [tool-calling](./tool-calling.md) — the real wall for local agents: per-engine tool-call quirks, why does-it-tool is the diagnostic to run first, and the engine that finally returned structured `tool_calls` -- [agentic-harness](./agentic-harness.md) — Pi as a lean headless coding agent over a - local OpenAI-compatible endpoint; lean beats heavy for weak models +- [harness-weight](./harness-weight.md) — a heavy harness breaks tool-calling that the same model handles cleanly under a lean one (observed with Pi) ## stack that worked diff --git a/ai/local-models/agentic-harness.md b/ai/local-models/harness-weight.md similarity index 87% rename from ai/local-models/agentic-harness.md rename to ai/local-models/harness-weight.md index 6f9b555..4ed8f6b 100644 --- a/ai/local-models/agentic-harness.md +++ b/ai/local-models/harness-weight.md @@ -1,8 +1,10 @@ -# agentic harness for local models (Pi) +# harness weight destabilizes weak models -[Pi](https://github.com/badlogic/pi-mono) (`@earendil-works/pi-coding-agent`) is a -lean headless coding agent that points at any OpenAI-compatible endpoint — the right -shape for driving a *local* model, where harness weight directly hurts. +the agent harness around a local model has a measurable cost: a heavy harness +(big system prompt, a dozen built-in tools) breaks tool-calling that the same +model performs cleanly under a lean one. observed with +[Pi](https://github.com/badlogic/pi-mono) (`@earendil-works/pi-coding-agent`), +a lean headless coding agent that points at any OpenAI-compatible endpoint. ## lean beats heavy (for weak models) diff --git a/ai/memory/README.md b/ai/memory/README.md new file mode 100644 index 0000000..3f8965e --- /dev/null +++ b/ai/memory/README.md @@ -0,0 +1,18 @@ +# memory + +how agents keep state between runs. the notes here divide the problem the way +the working systems do: + +- a **transcript layer** — the conversation/thread history itself +- a **long-term layer** — facts and impressions that outlive any one thread, + written either deliberately (a tool call) or in the background (a pipeline) +- **curation** — keeping the long-term layer from rotting, at write time or + read time + +## notes + +- [message-archives](./message-archives.md) — the transcript layer: what has actually stored it across systems and years +- [deliberate-and-background-writes](./deliberate-and-background-writes.md) — the two write paths for long-term memory, and keeping background writes clean + +read-time curation lives with retrieval: +[retrieval/synthesize-before-injecting](../retrieval/synthesize-before-injecting.md). diff --git a/ai/memory/deliberate-and-background-writes.md b/ai/memory/deliberate-and-background-writes.md new file mode 100644 index 0000000..5c53358 --- /dev/null +++ b/ai/memory/deliberate-and-background-writes.md @@ -0,0 +1,56 @@ +# deliberate and background writes + +long-term agent memory gets written two ways, and mature systems end up with +both: + +- **deliberate**: the agent calls a tool — letta's replace-core-memory, + phi's `save_memory`, the marvin slack bot's `store_facts_about_user`. + legible in the transcript, but only fires when the model thinks of it. +- **background**: a pipeline writes without being asked — letta's + asynchronous episodic memory, phi's extraction → reconciliation pass and + its end-of-run residue synthesis. catches what the agent wouldn't have + saved, at the cost of a second system to keep honest. + +phi and letta arrived at this same split independently, which is decent +evidence the split is real rather than one vendor's framing. + +## keeping background writes clean + +a background pipeline that only appends will rot. phi's write path (two +haiku-class calls per batch): + +1. an extractor proposes facts about the user from recent conversation. it + never sees existing memory, so a bad stored fact can't steer new + extraction. +2. each proposal is embedded and matched against the 3 most similar stored + observations; a reconciler returns ADD / UPDATE / DELETE / NOOP. +3. updates supersede in place: old row marked `status: superseded` with a + back-link, reads filter to active rows, history stays queryable. an + UPDATE unions the old row's source references (a refinement keeps its + evidence); a DELETE-and-replace starts fresh (a correction shouldn't + inherit the evidence of the claim it overturns). + +what reaches the prompt carries its pedigree: llm summaries inject as +"trust: low", extracted observations as "trust: medium (N sources, age)", +verbatim transcripts as "trust: high". + +## adopting a background runtime is a reliability decision + +the marvin slack bot spent one day (2025-12-09) on letta's +`agentic-learning` SDK: adopted at 13:32, committed to by 15:15 (*"remove +home-rolled TurboPuffer-based user facts system... automatically extracts +and injects memories without needing explicit tool calls"*), reverted at +23:20 when the SDK's async context manager awaited tasks in `__aexit__` +after the event loop closed, crashing the bot. read narrowly — an SDK +lifecycle bug, no information about the memory model. the durable lesson is +about integration surface: background memory that wraps your request path +makes its bugs your outages, so it has to clear a higher reliability bar +than a store you call explicitly. phi keeps its background passes outside +the request path for the same reason: each prompt block returns empty on +error, and a dead memory store means a thinner prompt instead of a crash. + +## sources + +- [bot](https://tangled.org/zzstoatzz.io/bot) — `src/bot/memory/`, `src/bot/core/residue.py`, `docs/memory.md` +- [marvin](https://github.com/prefecthq/marvin) — commits `30aed2de`, `36da4c1a`, `47da4ce6`, `a593d2b3` (the letta day); `examples/slackbot/src/slackbot/assets.py` +- [letta](https://www.letta.com) — the memgpt lineage; two-memory-type framing via cameron.stream diff --git a/ai/memory/message-archives.md b/ai/memory/message-archives.md new file mode 100644 index 0000000..da62405 --- /dev/null +++ b/ai/memory/message-archives.md @@ -0,0 +1,39 @@ +# message archives + +the transcript layer of agent memory — who said what in which thread — wants +boring, owned storage. two long-running systems, several eras of evidence: + +**marvin's slack bot** rebuilt this layer once per major version: + +- 1.x: a TTL cache in process memory keyed by slack thread + (`TTLCache(maxsize=1000, ttl=86400)` holding `History` objects) — restart + the process, lose everything. +- 2.x: openai-assistants `Thread` objects serialized into prefect JSON + blocks, one entry per slack thread. the only era where the framework + library supplied the persistence primitive (`JSONBlockState`). the same + era's `Application` abstraction — an agent JSON-patching a state object — + was enough to host a working maze game, an early answer to "stateful + agent" that predates the current framework generation. +- 3.x: a bespoke `MessageStore` — each thread's full pydantic-ai + `ModelMessage` list as one JSON object in a `WritableFileSystem` block + (local disk or GCS). by this era the bot runs `pydantic_ai.Agent` directly + and the marvin library is down to a logger and one `cast_async`. + +**phi** (a bluesky agent) doesn't store this layer at all: threads live on +the atproto network, so each run fetches the thread fresh +(`get_thread(uri, depth=100)`, ~200ms). the network is the archive, always +current, shared with every other client. + +the pattern across all of it: the transcript layer ends up in whatever +plain storage is already lying around — RAM, blocks, a bucket, the network — +serialized message lists keyed by thread. the framework-supplied version +(2.x) had the same shape as the hand-rolled ones and didn't outlive the +framework abstraction it depended on. verbatim transcripts also anchor the +trust ladder: phi injects them labeled "trust: high", above extracted facts +and llm summaries, precisely because nothing has paraphrased them. + +## sources + +- [marvin](https://github.com/prefecthq/marvin) — `examples/slackbot/src/slackbot/_internal/message_store.py`; historical `cookbook/slackbot/` +- [bot](https://tangled.org/zzstoatzz.io/bot) — `docs/memory.md` (thread context) +- [marvin 3.x writeup](https://blog.zzstoatzz.io/marvin-3x/) — the replatform context diff --git a/ai/persistent-agents/README.md b/ai/persistent-agents/README.md deleted file mode 100644 index f8f42b9..0000000 --- a/ai/persistent-agents/README.md +++ /dev/null @@ -1,25 +0,0 @@ -# persistent agents - -what actually keeps an agent's state alive between runs, from two codebases with -years of history: marvin's slack bot (three major versions of the same bot, one -brief framework adoption) and phi (a bluesky agent whose whole design is -persistence). - -the reference point throughout is letta (the memgpt lineage): agent-managed -memory as a dedicated runtime, with two kinds of memory — a conscious -replace-core-memory tool call, and an asynchronous background episodic memory. -void, the letta-built bluesky agent, is the most compelling live example of -the approach. these notes record what two other long-running systems actually -run, as comparison material rather than a verdict — the tinkering with letta -itself is ongoing. - -## notes - -- [slack-bot-memory-rebuilds](./slack-bot-memory-rebuilds.md) — the marvin slack bot's memory across three major versions, including a one-day swing through letta's learning-sdk -- [phi-memory-in-parts](./phi-memory-in-parts.md) — phi's persistence: four stores, two cheap model passes, no memory runtime - -## sources - -- [marvin](https://github.com/prefecthq/marvin) — the slack bot's full git history -- [bot](https://tangled.org/zzstoatzz.io/bot) — phi -- [marvin 3.x writeup](https://blog.zzstoatzz.io/marvin-3x/) — the replatform context diff --git a/ai/persistent-agents/phi-memory-in-parts.md b/ai/persistent-agents/phi-memory-in-parts.md deleted file mode 100644 index 6ef82d4..0000000 --- a/ai/persistent-agents/phi-memory-in-parts.md +++ /dev/null @@ -1,59 +0,0 @@ -# phi's memory is assembled from parts - -phi is a bluesky agent that runs a fresh context every ~10 seconds — every run -rebuilds its prompt from scratch, so everything it "remembers" has to live -somewhere concrete between runs. its persistence is four stores and a few -small model passes, each an ordinary component: - -| store | holds | written by | -|---|---|---| -| the network itself | thread history | nobody — fetched live per batch (`get_thread`, ~200ms) | -| turbopuffer | per-user observations + episodic notes | an extraction → reconciliation pipeline (daily) and a `save_memory` tool | -| phi's own PDS (public atproto records) | goals, consent lists, a daily atlas/docket, residue | gated tools and scheduled flows | -| in-process caches | derived prompt blocks | TTL'd reads of the above | - -the parts that do the work a memory framework would claim: - -- **extraction → reconciliation** (two haiku calls): after conversations, an - extractor proposes facts about the user; each proposal is embedded, matched - against the 3 most similar stored observations, and a reconciler returns - ADD / UPDATE / DELETE / NOOP. updates supersede the old row in place - (`status: superseded`, back-link) — reads filter to active rows, history - stays. the extractor never sees existing memory, so bad stored facts can't - steer new extraction. -- **synthesis before injection**: episodic retrieval doesn't paste top-k into - the prompt; a haiku pass dedupes, prefers newer on conflict, and returns - nothing when nothing is relevant (details in - [embeddings/synthesize-before-injecting](../embeddings/synthesize-before-injecting.md)). -- **residue**: a 7-item, 3-day-decaying buffer of what recent runs left - behind, written by a small end-of-run pass into a public PDS record and - injected at the next run's start. capacity and decay are enforced in code. -- **trust labels travel with the content**: llm summaries inject as - "trust: low", extracted observations as "trust: medium (N sources, age)", - verbatim logs as "trust: high" — the downstream model gets told how much to - believe each block. - -two properties fall out of building it this way: - -- **every layer is inspectable separately.** the durable intent is public - atproto records anyone can read; the vector store is queryable directly; - the synthesis prompts are versioned in the repo. when memory misbehaves, - the failing layer is findable. -- **failures degrade instead of crashing.** each prompt block returns empty - on error; a dead vector store means a thinner prompt, and no third-party - runtime wraps the request path (the integration-surface lesson from the - [slack bot's letta episode](./slack-bot-memory-rebuilds.md)). - -phi also lands on the same conscious/background split letta describes, from -different parts: `save_memory` is the conscious tool call, while the -extraction pipeline and residue write memory in the background without the -agent asking. - -none of the pieces is novel — a vector db, some records, two cheap model -passes, a decaying buffer. the design work is in the seams: what gets cleaned -at write time vs read time, what's public vs private, and what enters the -prompt every run vs on demand. - -## sources - -- [bot](https://tangled.org/zzstoatzz.io/bot) — `src/bot/memory/namespace_memory.py`, `src/bot/core/residue.py`, `docs/memory.md` diff --git a/ai/persistent-agents/slack-bot-memory-rebuilds.md b/ai/persistent-agents/slack-bot-memory-rebuilds.md deleted file mode 100644 index e143950..0000000 --- a/ai/persistent-agents/slack-bot-memory-rebuilds.md +++ /dev/null @@ -1,70 +0,0 @@ -# the slack bot rebuilt its memory three times - -marvin's slack bot has run continuously across the library's three major -versions, and its memory has been rebuilt each time. the git history is a -clean record of what "persistent agent" meant in practice at each point. - -## the memory, era by era - -**1.x** (`cookbook/slackbot.py`): a TTL cache in process memory, keyed by slack -thread — `global_cache = TTLCache(maxsize=1000, ttl=86400)` holding marvin -`History` objects. restart the process, lose everything. - -**2.x** (`cookbook/slackbot/`): openai-assistants `Thread` objects serialized -into prefect JSON blocks, one entry per slack thread, plus a hand-rolled -"parent app" tracking per-user notes in another block. this era is the only -one where the marvin *library* supplied the persistence primitive -(`JSONBlockState`, `marvin.beta.assistants`). the same era produced marvin's -`Application` abstraction — an agent holding a state object it maintains by -JSON patch, which was enough to host a working maze game (grid, key, monster, -obstacles) with the board set programmatically and patched by the LLM turn by -turn. an early answer to "stateful agent" that predates the current framework -generation. - -**3.x** (`examples/slackbot/`): the bot stops using marvin's agent machinery -entirely — it runs `pydantic_ai.Agent` directly, and marvin appears only as -`get_logger`, one `cast_async`, and a bare import. memory is a bespoke -`MessageStore`: each thread's full pydantic-ai `ModelMessage` list as one JSON -object in a `WritableFileSystem` block (local disk or GCS), plus a turbopuffer -"user facts" store written by explicit `store_facts_about_user` tool calls. - -across three versions, the durable pattern is a message archive the bot owns, -in whatever storage was already lying around (RAM, prefect blocks, a bucket), -plus an explicit per-user fact store. - -## the letta episode (2025-12-09) - -one day in the history records a swing through letta's `agentic-learning` SDK, -four commits: - -- **13:32** `30aed2de` — wrap the agent run in - `with learning(agent=f"slackbot-{user_id}")` for automatic per-user memory. -- **15:15** `36da4c1a` — commit to it: *"remove home-rolled TurboPuffer-based - user facts system in favor of letta's learning-sdk... automatically extracts - and injects memories without needing explicit tool calls."* -- **15:53** `47da4ce6` — fix: the learning context manager needed `async with`. -- **23:20** `a593d2b3` — revert: the SDK's async context manager awaited - pending tasks in `__aexit__` after the event loop closed, crashing the bot; - the revert *"restores the original TurboPuffer-based user facts system until - the upstream bug is fixed."* - -read narrowly: this says nothing about letta's memory model — the automatic -extract-and-inject pitch was compelling enough to delete a working system the -same afternoon, and what failed was an SDK lifecycle bug in the host process. -the durable lesson is about integration surface: a memory runtime that wraps -your request path makes its bugs your outages, so it has to clear a higher -reliability bar than a store you call explicitly. (letta's model itself — -a conscious replace-core-memory tool call plus an asynchronous background -episodic memory — is the interesting comparison point for -[phi's construction](./phi-memory-in-parts.md), which arrives at a similar -conscious/background split from different parts.) - -## sources - -- [marvin](https://github.com/prefecthq/marvin) — commits `30aed2de`, - `36da4c1a`, `47da4ce6`, `a593d2b3`; `examples/slackbot/src/slackbot/`, - `_internal/message_store.py`; historical `cookbook/slackbot/` -- [marvin 3.x writeup](https://blog.zzstoatzz.io/marvin-3x/) — why 3.x - replatformed on pydantic-ai -- [letta](https://www.letta.com) — the memgpt lineage; the two-memory-type - description comes from letta's own framing (via cameron.stream) diff --git a/ai/retrieval/README.md b/ai/retrieval/README.md new file mode 100644 index 0000000..2afc514 --- /dev/null +++ b/ai/retrieval/README.md @@ -0,0 +1,19 @@ +# retrieval + +query-time mechanics for vector search: how documents and queries get embedded, +how keyword and semantic results combine, and what happens between retrieval +and the prompt. evidence drawn from pub-search (search over atproto +publications) and phi (an agent's memory reads). + +## notes + +- [asymmetric-embedding](./asymmetric-embedding.md) — set `input_type` at both ends, and the document-prep checklist (truncation, titles, utf-8) +- [reciprocal-rank-fusion](./reciprocal-rank-fusion.md) — fusing keyword and semantic results by rank position, and why the two paths must fail independently +- [synthesize-before-injecting](./synthesize-before-injecting.md) — the stale-memory failure from prompting raw top-k, and the cheap-model pass that fixed it + +operational notes on the vector store itself: [storage/turbopuffer](../../storage/turbopuffer.md). + +## sources + +- [pub-search](https://tangled.sh/@zzstoatzz.io/pub-search) — hybrid keyword+semantic search backend +- [bot](https://tangled.org/zzstoatzz.io/bot) — phi's namespace memory diff --git a/ai/embeddings/asymmetric-embedding.md b/ai/retrieval/asymmetric-embedding.md similarity index 100% rename from ai/embeddings/asymmetric-embedding.md rename to ai/retrieval/asymmetric-embedding.md diff --git a/ai/embeddings/reciprocal-rank-fusion.md b/ai/retrieval/reciprocal-rank-fusion.md similarity index 100% rename from ai/embeddings/reciprocal-rank-fusion.md rename to ai/retrieval/reciprocal-rank-fusion.md diff --git a/ai/embeddings/synthesize-before-injecting.md b/ai/retrieval/synthesize-before-injecting.md similarity index 100% rename from ai/embeddings/synthesize-before-injecting.md rename to ai/retrieval/synthesize-before-injecting.md diff --git a/architecture/README.md b/architecture/README.md index fba1b02..7ee6e8b 100644 --- a/architecture/README.md +++ b/architecture/README.md @@ -4,14 +4,14 @@ notes on architectural patterns that span languages, tools, and persistence laye separation rule: -- `databases/` and `protocols/` — concrete systems (redis, sqlite, atproto, MCP) +- `storage/` and `protocols/` — concrete systems (redis, sqlite, turbopuffer, atproto, MCP) - `languages/` — language-specific patterns (zig idioms, python idioms) -- `ai/` — the AI domain (local models, embeddings) +- `ai/` — the AI domain (retrieval, memory, local models) - `architecture/` — the *patterns* that compose those things ## topics -- [background-tasks](./background-tasks.md) — the docket lineage, perpetuals, at-least-once delivery (redis-streams mechanics live in [databases/redis/streams](../databases/redis/streams.md)) +- [background-tasks](./background-tasks.md) — the docket lineage, perpetuals, at-least-once delivery (redis-streams mechanics live in [storage/redis/streams](../storage/redis/streams.md)) - [home-infra/](./home-infra/) — the home-box cluster: cost optimization, public cost declaration, tailscale ## sources diff --git a/languages/ziglang/ziglua-ffi.md b/languages/ziglang/ziglua-ffi.md index 9ea4d1a..f28c7aa 100644 --- a/languages/ziglang/ziglua-ffi.md +++ b/languages/ziglang/ziglua-ffi.md @@ -22,7 +22,7 @@ const zlua = b.dependency("zlua", .{ my_module.addImport("zlua", zlua.module("zlua")); ``` -`lang = .lua51` is the right choice if you're implementing a redis-compatible engine (see [eval-lua](../../databases/redis/eval-lua.md)). other options: `.lua52`, `.lua53`, `.lua54`, `.luau`. +`lang = .lua51` is the right choice if you're implementing a redis-compatible engine (see [eval-lua](../../storage/redis/eval-lua.md)). other options: `.lua52`, `.lua53`, `.lua54`, `.luau`. ## minimal boot diff --git a/databases/README.md b/storage/README.md similarity index 63% rename from databases/README.md rename to storage/README.md index d543bf1..3685566 100644 --- a/databases/README.md +++ b/storage/README.md @@ -1,7 +1,8 @@ -# databases +# storage concrete storage systems, from the building side. - [redis/](./redis/) — streams, eval-lua, embedded - [sqlite/](./sqlite/) — fts5, sargable joins +- [turbopuffer](./turbopuffer.md) — vector store operations: id limits, schema evolution, cold namespaces, stale attributes - [turso](./turso.md) — hosted libSQL: hrana vs export endpoints, auth model diff --git a/databases/redis/README.md b/storage/redis/README.md similarity index 100% rename from databases/redis/README.md rename to storage/redis/README.md diff --git a/databases/redis/embedded.md b/storage/redis/embedded.md similarity index 100% rename from databases/redis/embedded.md rename to storage/redis/embedded.md diff --git a/databases/redis/eval-lua.md b/storage/redis/eval-lua.md similarity index 100% rename from databases/redis/eval-lua.md rename to storage/redis/eval-lua.md diff --git a/databases/redis/streams.md b/storage/redis/streams.md similarity index 100% rename from databases/redis/streams.md rename to storage/redis/streams.md diff --git a/databases/sqlite/fts5.md b/storage/sqlite/fts5.md similarity index 100% rename from databases/sqlite/fts5.md rename to storage/sqlite/fts5.md diff --git a/databases/sqlite/sargable-joins.md b/storage/sqlite/sargable-joins.md similarity index 100% rename from databases/sqlite/sargable-joins.md rename to storage/sqlite/sargable-joins.md diff --git a/ai/embeddings/turbopuffer-in-production.md b/storage/turbopuffer.md similarity index 100% rename from ai/embeddings/turbopuffer-in-production.md rename to storage/turbopuffer.md diff --git a/databases/turso.md b/storage/turso.md similarity index 100% rename from databases/turso.md rename to storage/turso.md -- 2.51.2