diff --git a/routing-memory/SKILL.md b/routing-memory/SKILL.md index a1f9301..60dcc16 100644 --- a/routing-memory/SKILL.md +++ b/routing-memory/SKILL.md @@ -5,13 +5,13 @@ description: Designs and repairs routed agent memory using a lean system layer, # Routing Memory -This helps Letta agents understand how to build a _routing-based memory architecture_. +Design or repair a routing-based memory architecture for a Letta agent or another agent with a Markdown memory repository. -Routing-based memory has lean system/core memory with several indices. When you ask the agent a query, it should use its indices and lean memory to identify which memories to reference. +Routing memory uses a lean always-loaded system layer plus several shallow indices. For each stateful request, the agent follows a retrieval router to the relevant canonical files and falls back to conversation history for exact episodes. -Routing memory is intended to allow agents to have much larger context repositories, at the cost of incurring more tool calls and possibly higher latency as agents sift through their memory. +This allows context repositories to grow beyond what belongs in the prompt, at the cost of more tool calls and higher latency. -Memory design is empirical. There is no "best memory architecture". Routing memory is simply one apporach. Use this skill as a guide to inform your particular architecture. +Memory design remains empirical. Treat this as one practical architecture. Derive it from the agent's actual work and prove it with retrieval tests. ## Core model @@ -40,6 +40,48 @@ What happens when two sources disagree? How will we prove the route works? ``` +## Fit and vocabulary + +Use routing memory when the agent: + +- works across several domains or projects; +- needs durable state across conversations; +- has facts in memory that still fail to appear in answers; +- has duplicated current facts or a growing always-loaded prompt; +- needs exact historical recall without making transcripts canonical current state. + +Skip or defer it when one small, focused memory surface already fits comfortably in context and there is no observed retrieval failure. A router should solve context pressure or activation failures, not manufacture a filing system for its own amusement. + +| Artifact | Role | +| --- | --- | +| Owner | The single file responsible for one class of facts staying correct | +| Index | A shallow, task-shaped view of owners and their load cues | +| Router | An always-loaded decision table from a turn shape to ordered reads | +| Turn shape | The kind of request being made: current status, exact quote, preference, project detail, conflict, or another evidence need | +| Recall | Search over event history for exact episodes and provenance; not canonical current state | +| Activation | The cue and rule that cause stored context to be retrieved for a turn | + +A typical target tree is: + +```text +system/ + persona.md + now.md + retrieval_router.md +indices/ + people.md + projects.md + concepts.md + operations.md +projects/ +concepts/ +reference/ +timeline/ +skills/ +``` + +Adapt the names and omit unused categories. Load [references/templates.md](references/templates.md) when creating owners, indices, router entries, recall briefs, migrations, or probes. + ## Current Letta contract For Letta agents, verify current product details against the official documentation before using CLI commands or changing runtime configuration: @@ -57,7 +99,7 @@ The stable architecture this skill relies on is: - Conversation-history search is separate from MemFS. Compaction can omit exact wording or provenance while the underlying message history remains searchable. - Agent-owned procedural skills live under `$MEMORY_DIR/skills/`. -Use `/init` for ordinary initialization and `/doctor` for a general memory audit. Use this skill when the user specifically wants routed, index-based, out-of-core memory or when those generic workflows produced storage without reliable activation. +Use `/init` to bootstrap or refresh ordinary project memory. Use `/doctor` to audit placement, duplication, and system-prompt usage. Use this skill when the user specifically wants routed, index-based, out-of-core memory or when those generic workflows produced storage without reliable activation. ## Hard boundaries @@ -65,6 +107,7 @@ Use `/init` for ordinary initialization and `/doctor` for a general memory audit - Do not rewrite identity, values, protected user context, or relationship terms as a side effect of organization. - Do not publish private memory, even when the architecture itself is public. - Preserve git history and inspect shared or read-only state before editing. +- Treat externally owned or read-only files as sources, not migration targets. In multi-writer repositories, assign path-level write authority before moving shared state. - Migrate completely. Once ownership is settled, update routes and links in the same change and remove stale duplicate owners. Do not leave a permanent `legacy/` attic. - Keep recall workers read-only. Retrieval authority is not edit, commit, or outbound-message authority. - Do not call a migration successful because files exist. Run retrieval probes. @@ -97,6 +140,16 @@ Unrouted files: Do not edit during discovery unless the user explicitly requested a tiny, obvious repair. +Every durable Markdown memory file should have frontmatter whose `description` states the file's purpose and when it should be retrieved. Describe the role of the file, not merely a summary of its current contents. + +For an empty or nearly empty repository, use a smaller greenfield path: + +1. create identity and hard rules in `system/`; +2. create the retrieval router; +3. create only the first one or two indices the work actually needs; +4. add canonical owners as real state accumulates; +5. seed retrieval probes before expanding the hierarchy. + ## 2. Define canonical ownership Assign one owner to each important or volatile class of information. @@ -143,7 +196,9 @@ Move out of `system/`: - per-project implementation detail not needed on most turns; - raw transcripts and source dumps. -Do not optimize for the smallest possible prompt. Preserve examples, reasons, and semantic cues that help the model notice the route. Cut redundancy and stale detail before cutting identity-bearing or behavior-shaping context. +Do not optimize for the smallest possible prompt. Preserve examples, reasons, and semantic cues that help the model notice the route. + +Measure the actual system-layer context cost. Split only when size, cost, or attention is causing a real problem; there is no universal file-count threshold. Cut redundancy and stale detail before cutting identity-bearing or behavior-shaping context. ## 4. Build several flat indices @@ -176,7 +231,7 @@ Keep common routes one read away when possible. Three layers of disclosure can a ## 5. Write an operational retrieval router -The router is a decision table, not a directory listing. +The router is a decision table, not a directory listing. A turn shape is the evidence need expressed by the request, such as current status, exact prior wording, a stable preference, project state, or a conflict between sources. Each route should include: @@ -217,6 +272,10 @@ Do not begin drafting from latent familiarity and retrieve afterward to decorate If the user explicitly requests a fast response, reduce retrieval depth. Do not silently treat casual phrasing as stateless when it contains a known person, project, place, or recurring decision. +Keyword, semantic, or hybrid search may nominate candidate files when the runtime supports it. Search is a secondary path. It does not replace canonical ownership or the router's responsibility to decide which source governs the answer. + +Use the post-answer update decision tree in [references/templates.md](references/templates.md) before turning every successful retrieval into another memory write. + ## 7. Use recall as evidence recovery Conversation history is an event log, not the primary knowledge base. @@ -238,6 +297,8 @@ Recall procedure: 5. return evidence to the parent agent without editing memory; 6. promote only the durable lesson or retrieval handle into the canonical owner. +Search directly when the episode is narrow and the handles are known. Use an optional recall worker when the search spans many conversations, requires chronological reconstruction, or would crowd the active turn. Give the worker the read-only brief in [references/templates.md](references/templates.md); the parent agent retains all edit, commit, and outbound authority. + Do not copy whole conversations into system memory. Preserve enough handles that future recall can find the source again. ## 8. Separate state from procedure @@ -264,7 +325,7 @@ For a structural change: 5. remove stale duplicate owners in the same pass; 6. search for old paths and old owner language; 7. inspect the final diff for semantic loss; -8. recompile the active context when the runtime requires it. +8. recompile the active context when the runtime requires it. In the Letta CLI, use `/recompile` after changing `system/` so the current conversation receives the new compiled prompt. Do not leave a half-migrated tree with old and new hierarchies both claiming authority. @@ -297,6 +358,19 @@ Repair: Measure routing precision as well as recall. Loading everything prevents misses by recreating prompt bloat. The goal is the smallest sufficient context that still reaches the right evidence. +## Definition of done + +Routing memory is complete only when: + +- every important information class has one canonical owner; +- router and index paths resolve to those owners; +- no stale file still claims duplicate current authority; +- frontmatter descriptions support discovery; +- current-fact, project, exact-quote, conflict, and unknown probes pass; +- shared or read-only authority remains intact; +- the final git state is clean or every remaining change is intentional; +- the report below names unresolved risks. + ## Common failures ### Stored but unrouted