diff --git a/.beads/interactions.jsonl b/.beads/interactions.jsonl index 322a6d2..fe341de 100644 --- a/.beads/interactions.jsonl +++ b/.beads/interactions.jsonl @@ -53,3 +53,4 @@ {"id":"int-fb2d145f","kind":"field_change","created_at":"2026-06-28T16:16:07.11528074Z","actor":"dawn","issue_id":"klbr-oba","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Completed: added reproducible random LongMemEval sampling with question_type stratification."}} {"id":"int-fdca5c0e","kind":"field_change","created_at":"2026-06-28T19:54:14.601069054Z","actor":"dawn","issue_id":"klbr-f0q.1","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Implemented OpPlan/QueryOp classification, serialized op_plan into retrieval traces, preserved lookup defaults, added deterministic planner and pipeline trace tests."}} {"id":"int-9125728a","kind":"field_change","created_at":"2026-06-28T21:21:26.576608546Z","actor":"dawn","issue_id":"klbr-h4l","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Implemented project-local dense/sparse embedding cache, ran 25-sample QA bench, inspected failures, tried bounded planner/high-budget experiments, and wrote report packet."}} +{"id":"int-95ec3c9d","kind":"field_change","created_at":"2026-06-28T22:07:09.303815881Z","actor":"dawn","issue_id":"klbr-jwn","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Completed docs index, current architecture review rewrite, status banners, AGENTS routing, and follow-up implementation issue klbr-u0q."}} diff --git a/.beads/issues.jsonl b/.beads/issues.jsonl index d4c6f25..2b6b202 100644 --- a/.beads/issues.jsonl +++ b/.beads/issues.jsonl @@ -1,3 +1,5 @@ +{"_type":"issue","id":"klbr-u0q","title":"Implement session-first lexical-contained memory retrieval","description":"Build the next retrieval architecture from docs/long-term-memory-arch.md: stage-one session/event candidate generation, lexical containment, and language-agnostic operation planning experiments without adding english cue-word hacks.","design":"Keep EvidencePacket as the rank object. Use exact refs, dense episode/session cards, learned sparse or fts side channels, and graph expansion from strong seeds only. Keep QueryOp shape but replace cue-list policy with a structured classifier plus deterministic fallback.","acceptance_criteria":"A bench profile can retrieve from session/episode cards first and use chunk fts/sparse hits as packet enrichment; lexical-only policy paths are removed or contained behind candidate channels; traces expose candidate session recall and packet/rendered evidence metrics; multilingual fixtures cover non-English cue-free retrieval cases.","status":"open","priority":1,"issue_type":"feature","owner":"90008@klbr.net","created_at":"2026-06-28T22:06:32Z","created_by":"dawn","updated_at":"2026-06-28T22:06:32Z","dependency_count":0,"dependent_count":0,"comment_count":0} +{"_type":"issue","id":"klbr-jwn","title":"Organize evolving memory architecture docs","description":"Make the memory architecture docs distinguish current truth, implementation status, proposals, research inputs, and archived rationale so future agents do not treat older reports as canonical.","acceptance_criteria":"Docs have a clear index and status taxonomy; current memory architecture direction points at docs/long-term-memory-arch.md; older research reports are marked as historical or supporting; AGENTS.md routes future agents through the doc index.","status":"closed","priority":1,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T22:01:20Z","created_by":"dawn","updated_at":"2026-06-28T22:07:09Z","started_at":"2026-06-28T22:01:23Z","closed_at":"2026-06-28T22:07:09Z","close_reason":"Completed docs index, current architecture review rewrite, status banners, AGENTS routing, and follow-up implementation issue klbr-u0q.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-h4l","title":"Cache bench embeddings locally and run 25-sample QA bench","description":"Make LongMemEval bench embeddings use a project-local gitignored sqlite cache so reruns do not pay for the same embeddings twice. Then run a 25-sample stratified QA bench, inspect failures, try bounded non-overfit tweaks if evidence supports them, and prepare a report packet if results remain weak or tuning would be overfit.","design":"Prefer a global project cache path such as benchmarks/cache/*.db over temp/run-local caches. Keep benchmark polling sparse: start the run, wait for artifacts or process completion, then inspect once.","acceptance_criteria":"Embedding cache lives under the project and is ignored by git; cache hits are reused across bench runs; a 25-sample bench artifact exists with failure analysis; any code changes are tested, committed, and pushed.","notes":"Implemented project-local embedding cache at benchmarks/cache/embeddings.db and wired LongMemEval bench LlmClient construction through it. Ran stratified 25-sample QA on seed 4937553249516211047: current-code run benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z scored official accuracy 0.6800. Report packet written to report_packet.md in that run dir. High-budget targeted reruns fixed only dd2973ad and a1cc6108; remaining failures point to fact/timeline synthesis rather than safe cue tuning.","status":"closed","priority":1,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T20:08:32Z","created_by":"dawn","updated_at":"2026-06-28T21:21:27Z","started_at":"2026-06-28T20:08:35Z","closed_at":"2026-06-28T21:21:27Z","close_reason":"Implemented project-local dense/sparse embedding cache, ran 25-sample QA bench, inspected failures, tried bounded planner/high-budget experiments, and wrote report packet.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-f0q.1","title":"Add operation-aware query planning","description":"Introduce a small OpPlan/QueryOp classifier for memory QA queries: lookup, aggregate_count, aggregate_sum, aggregate_avg, order_or_rank, update_resolution, preference_recommendation, and abstain_or_false_premise_check. Use it to switch only high-value classes away from plain lookup behavior.","design":"Keep the label set intentionally small and rule-based first; avoid growing English lexical hacks into the main ranking decision boundary.","acceptance_criteria":"Production memory retrieval traces expose the selected query op and plan; existing lookup behavior remains unchanged by default; aggregate/order/update/preference queries can be detected deterministically in tests.","status":"closed","priority":1,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T19:44:46Z","created_by":"dawn","updated_at":"2026-06-28T19:54:15Z","started_at":"2026-06-28T19:45:04Z","closed_at":"2026-06-28T19:54:15Z","close_reason":"Implemented OpPlan/QueryOp classification, serialized op_plan into retrieval traces, preserved lookup defaults, added deterministic planner and pipeline trace tests.","dependencies":[{"issue_id":"klbr-f0q.1","depends_on_id":"klbr-f0q","type":"parent-child","created_at":"2026-06-28T22:44:46Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":1,"comment_count":0} {"_type":"issue","id":"klbr-f0q.2","title":"Add structured synthesis for planned QA operations","description":"Add a fact-table synthesis path for aggregate count/sum/avg, ordering, previous/latest update resolution, and preference-grounded recommendations. Do not route ordinary conversation lookup through the structured path unless the planner selects it.","design":"Stage A extracts/normalizes fact rows with refs; Stage B computes or renders the answer according to the OpPlan.","acceptance_criteria":"Planned synthesis computes counts/averages/order/update answers from cited fact rows; preference recommendations require personal support and avoid generic distractor answers; lookup questions keep current direct reader path.","status":"open","priority":1,"issue_type":"feature","owner":"90008@klbr.net","created_at":"2026-06-28T19:44:46Z","created_by":"dawn","updated_at":"2026-06-28T19:44:46Z","dependencies":[{"issue_id":"klbr-f0q.2","depends_on_id":"klbr-f0q","type":"parent-child","created_at":"2026-06-28T22:44:46Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-f0q.2","depends_on_id":"klbr-f0q.4","type":"blocks","created_at":"2026-06-28T22:44:56Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} diff --git a/AGENTS.md b/AGENTS.md index 3f37a34..2b3a9d3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -3,10 +3,11 @@ personal ai agent harness in rust. local llm chat daemon with multi-lane long-term memory, tool calling, discord integration, and a ratatui tui / svelte web client. self-hosted, no corporate product feel. (。-`ω´-) memory docs: -- [docs/memory-arch.md](file:///home/mayer/proj/klbr/docs/memory-arch.md) — target architecture and design direction, with current-source status notes. +- [docs/README.md](file:///home/mayer/proj/klbr/docs/README.md) — docs map and status taxonomy; read this first when docs conflict. +- [docs/long-term-memory-arch.md](file:///home/mayer/proj/klbr/docs/long-term-memory-arch.md) — current long-term memory direction after the 2026-06-29 review. - [docs/memory-implementation-status.md](file:///home/mayer/proj/klbr/docs/memory-implementation-status.md) — current implementation, verified commands, and open bead ids. - [docs/memory-benches.md](file:///home/mayer/proj/klbr/docs/memory-benches.md) — benchmark protocol and run commands. -- [docs/evidence-selection.md](file:///home/mayer/proj/klbr/docs/evidence-selection.md) — evidence packet selection/rerank research report. +- [docs/evidence-selection.md](file:///home/mayer/proj/klbr/docs/evidence-selection.md) — supporting evidence packet selection/rerank research report. - [docs/code-intel-tools.md](file:///home/mayer/proj/klbr/docs/code-intel-tools.md) — tree-sitter-first file-local code retrieval/editing tools. --- diff --git a/docs/README.md b/docs/README.md new file mode 100644 index 0000000..be0bc09 --- /dev/null +++ b/docs/README.md @@ -0,0 +1,65 @@ +# docs map + +status: current docs entrypoint +updated: 2026-06-29 + +read this file first when working from docs. the memory docs are evolving: some +files describe current source truth, some describe current target direction, and +some are older research reports kept for rationale. + +## doc states + +- `current`: use as the active source of truth for the named area. +- `direction`: accepted target direction, but not necessarily implemented. +- `supporting`: useful research or design rationale. verify against current + source before treating it as true. +- `historical`: previous report or plan. do not use as canonical when it + conflicts with a `current` or `direction` doc. + +## current memory docs + +- [`long-term-memory-arch.md`](./long-term-memory-arch.md) - current long-term + memory direction after the 2026-06-29 review. use this for retrieval and + architecture decisions. main follow-up: `klbr-u0q`. +- [`memory-implementation-status.md`](./memory-implementation-status.md) - + current implementation map, verification commands, and known open work. +- [`memory-benches.md`](./memory-benches.md) - current benchmark protocol and + run commands. + +## supporting memory reports + +- [`memory-retrieval-research-primer.md`](./memory-retrieval-research-primer.md) + - research brief that led to the evidence-packet work. +- [`evidence-selection.md`](./evidence-selection.md) - packet selection and + rerank research report. mostly useful as rationale for the implemented packet + planner. +- [`planner-grade-evidence-assembly.md`](./planner-grade-evidence-assembly.md) + - operation-aware evidence assembly plan. partially implemented; remaining + work is tracked in beads under `klbr-f0q`. + +## historical memory reports + +- [`memory-arch.md`](./memory-arch.md) - older broad memory architecture report. + keep for lane/ref-substrate rationale, but use `long-term-memory-arch.md` for + current retrieval direction. +- [`principled-evidence-selection-fusion.md`](./principled-evidence-selection-fusion.md) + - earlier evidence-fusion report. +- [`benchmark-issues-report.md`](./benchmark-issues-report.md) - earlier bench + audit. + +## non-memory docs + +- [`code-intel-tools.md`](./code-intel-tools.md) - tree-sitter-first file-local + code retrieval and editing tools. + +## conflict rule + +if docs disagree, trust current source code plus +`memory-implementation-status.md` for what exists today. trust +`long-term-memory-arch.md` for the current target direction. treat old reports +as background unless a current doc explicitly revives a piece of them. + +older supporting and historical reports may contain imported researcher citation +markers such as `filecite` or `academia` handles. do not treat those markers as +valid repo citations; use their prose as rationale and re-check primary sources +when a claim matters. diff --git a/docs/benchmark-issues-report.md b/docs/benchmark-issues-report.md index e4fabdd..2e9f1f0 100644 --- a/docs/benchmark-issues-report.md +++ b/docs/benchmark-issues-report.md @@ -1,5 +1,13 @@ # klbr benchmark issues & gaps report +status: historical +reviewed: 2026-06-29 + +this is an older benchmark audit. keep it as rationale only. use +[`memory-benches.md`](./memory-benches.md) and +[`memory-implementation-status.md`](./memory-implementation-status.md) for +current benchmark protocol and implementation status. + compiled: 2026-06-27 scope: comprehensive breakdown of issues across the entire bench suite (LongMemEval-S/M, Tools Lane, and passive recall). diff --git a/docs/evidence-selection.md b/docs/evidence-selection.md index 7060180..b4f2c71 100644 --- a/docs/evidence-selection.md +++ b/docs/evidence-selection.md @@ -1,5 +1,15 @@ # klbr evidence selection for ref-native memory +status: supporting +reviewed: 2026-06-29 + +this report explains the evidence-packet selection work. much of the packet +builder and packet fusion direction has since landed. use +[`long-term-memory-arch.md`](./long-term-memory-arch.md) for the current +retrieval architecture direction and +[`memory-implementation-status.md`](./memory-implementation-status.md) for live +implementation status. + ## diagnosis the failure pattern you described is not a “retrieval missed the session” problem. it is a **selection granularity** problem: the pipeline is good enough to find the right session, but it sometimes hands the reader the *query-adjacent* chunk instead of the *answer-bearing* evidence packet. in other words, the current stack is already strong at coarse routing and weak at last-mile evidence assembly. that matches the behavior your benchmark note reports: `RecallAny@5 = 1.0` and `RecallAll@5 = 1.0` at session level, while end-to-end QA still fails on cases where the answer lives in a neighboring turn inside the retrieved session. your attached architecture also already frames `MemoryPipeline` as the benchmark-facing production facade and notes that the system is now ref-native enough that deterministic refs, promptable text, edges, and packets are first-class surfaces. fileciteturn0file0 diff --git a/docs/long-term-memory-arch.md b/docs/long-term-memory-arch.md new file mode 100644 index 0000000..df513a8 --- /dev/null +++ b/docs/long-term-memory-arch.md @@ -0,0 +1,256 @@ +# klbr long-term memory architecture direction + +status: direction +reviewed: 2026-06-29 +use with: [`memory-implementation-status.md`](./memory-implementation-status.md), +[`memory-benches.md`](./memory-benches.md) +tracked by: `klbr-u0q` + +## review verdict + +the researcher doc is directionally right: klbr should stop treating lexical +matches and cue lists as hidden retrieval policy. the next durable direction is +session/event-first retrieval, packet-level fusion, and operation-aware +assembly. lexical search should remain a candidate signal, not the thing that +decides what memory means. + +the important correction is scope. klbr already has many pieces from the report: +canonical refs, promptable text, dense and sparse embeddings, fts fallback, +episode cards, evidence packets, packet rrf, packet reranking, and a small +operation planner. the remaining work is not "add packets." it is to make the +first retrieval horizon more session/event based, remove hard-coded english cue +policy, and make aggregate/update/order synthesis less dependent on reader +freestyle. + +## current diagnosis + +recent qa benches show a split between evidence coverage and answer accuracy. +the system can often reach the right session or packet family, but final qa can +still fail because the decisive value is omitted, truncated, mixed with +distractors, or left for the reader to infer without a structured fact table. + +that means the main bottleneck is no longer coarse nearest-neighbor recall. it +is evidence assembly and synthesis: + +- candidate generation can still miss when lexical cues dominate. +- packet construction can still choose the wrong horizon. +- collect-mode questions need fact-group coverage, not just larger top-k. +- update questions need version/timestamp-aware assembly. +- preference questions need user-grounded evidence, not generic topical text. + +## adopted direction + +make memory retrieval packet-first, session-aware, and language-agnostic. + +the default retrieval flow should become: + +```text +query + -> exact ref resolver + -> structured operation planner + -> stage-one session/event candidate generation + dense over episode/event cards + sparse embedding or fts over promptable refs + exact/id/entity side channels + graph expansion only from strong seeds + -> group candidates by evidence group + exact ref | session_id | episode_ref | note_ref + -> build evidence packets + anchor hit + local turn window + episode gist + refs + timestamps + -> packet-level rank fusion + -> optional packet rerank + -> operation-aware packing + -> reader context or deterministic reducer +``` + +chunk hits should act as anchors into packets. they should not be the final +evidence unit. if fts lands on "i went to a play", the packet should include the +neighboring turn or event card where the play is named. + +## lexical containment + +keep lexical infrastructure, remove lexical policy. + +allowed: + +- exact ref and alias resolution. +- sqlite fts5/bm25 as a first-stage candidate channel. +- sparse learned retrieval from the embedding model when available. +- trigram or id/name side indexes for exact-ish entities, refs, code, and + substring-heavy identifiers. +- lexical features as trace/debug evidence. + +not allowed as the main answer: + +- hand-written english morphology such as plural stripping. +- cue-word lists that decide query operation or memory lane by themselves. +- additive lexical bonuses that suppress strong dense/session evidence. +- "fix this miss" keyword patches. + +current code still has a cue-list operation planner in `klbr-core/src/planner.rs`. +that planner is useful scaffolding, but it is technical debt. the replacement +should emit the same small `QueryOp` shape using a language-agnostic classifier +or a structured model call, then keep deterministic fallback behavior for +offline tests. + +## operation planning + +keep the operation set small: + +- `lookup` +- `aggregate_count` +- `aggregate_sum` +- `aggregate_avg` +- `order_or_rank` +- `update_resolution` +- `preference_recommendation` +- `abstain_or_false_premise_check` + +the operation label should change expansion and synthesis, not just metrics. + +```text +lookup: + anchor + small local window + optional episode gist + +aggregate_*: + collect across sessions, track fact groups, stop on coverage not packet count + +order_or_rank: + preserve timestamps and render chronological rows before prose + +update_resolution: + collect competing values, sort by valid/observed time, expose previous/current + +preference_recommendation: + require user-grounded support and demote generic topical evidence + +abstain_or_false_premise_check: + require answer-bearing evidence; otherwise abstain +``` + +for aggregate, order, and update questions, the reader should get a cited fact +table or timeline rows before the final answer. where possible, compute the +answer with a deterministic reducer over extracted facts instead of asking the +reader model to keep all chronology in its head. + +## packet target + +the existing `EvidencePacket` is the right object. future additions should make +it better for grouped evidence and temporal synthesis: + +```rust +struct EvidencePacket { + packet_id: String, + group_key: String, // session_id | episode_ref | exact_ref | note_ref + anchor_refs: Vec, + body_refs: Vec, + timestamps: Vec, + lanes: Vec, + source_scores: SourceScores, + estimated_tokens: usize, +} +``` + +this is a target shape, not a claim about the exact current struct. current +source already has packet kind, session id, anchor ref, refs, bodies, lanes, +signals, and estimated tokens. + +## implementation order + +1. add a session/event-first profile. + dense retrieval over episode cards should seed sessions before chunk hits + decide the evidence horizon. chunk fts/sparse hits should enrich those + session packets. + +2. replace cue-list planner policy. + keep `QueryOp`, but move classification behind a structured classifier with + deterministic tests and a conservative fallback. + +3. finish collect-mode fact coverage. + track fact groups for aggregate/order/update questions. expose sufficiency + and gap state in traces. + +4. render fact rows and timelines. + packets should carry enough timestamp/ref structure for the reader to see + previous/current values and ordered event lists. + +5. add multilingual regression fixtures. + include cross-lingual paraphrase, cjk no-space text, turkish suffix + variation, diacritics, transliteration, and ref/id cases where english cue + words are absent. + +this work is tracked in beads as `klbr-u0q`. + +## ablation matrix + +| profile | stage-one unit | lexical channel | dense channel | graph | rank object | synthesis | +|---|---|---|---|---|---|---| +| `dense-only` | episode/session cards | none | yes | no | packet | free-form | +| `lexical-only` | chunks/notes | fts5 bm25 | no | no | packet | free-form | +| `hybrid-chunk-first` | chunks | fts/sparse | turn/card dense | optional | packet | free-form | +| `session-first` | episode/session cards | fts/sparse side channel | session dense | optional | packet | free-form | +| `fact-card-first` | notes/fact cards | fts/sparse | note-card dense | optional | packet | op-aware | +| `hybrid-packet` | grouped packets | fts/sparse | session+card dense | yes | packet | op-aware | +| `hybrid-packet-no-graph` | grouped packets | fts/sparse | session+card dense | no | packet | op-aware | +| `hybrid-packet-no-planner` | grouped packets | fts/sparse | session+card dense | yes | packet | free-form | + +## metrics + +official qa accuracy stays the public headline, but it is not enough. every run +should also report: + +- candidate session recall. +- packet recall. +- rendered evidence recall. +- answer-value visibility. +- planner/op label distribution. +- fact-group coverage for collect-mode questions. +- final synthesis correctness. +- latency, token cost, and cache stats. + +when diagnosing failures, keep the whenloss-style split: + +```text +tfc = truncated full context +oe = oracle evidence +csm = complete stored memory +rm = retrieved memory +``` + +if `oe` is strong but `csm` is weak, the write path lost answerability. if +`csm` is strong but `rm` is weak, retrieval or packet assembly failed. if `rm` +is strong but qa is weak, synthesis is the bottleneck. + +## benchmark comparability + +do not compare retrieval recall claims to official qa accuracy. + +- LongMemEval v1 is still the current klbr public qa target. +- MemWeave-style LongMemEval-S claims are often correct-session Recall@k, not + official qa accuracy. +- LoCoMo, Mem0, Zep, GraphRAG, and other systems use different readers, judges, + datasets, and budgets unless reproduced locally. +- LongMemEval-V2 is relevant future work because it evaluates memory over web + and enterprise trajectories, but it is not yet the active klbr scoreboard. + +## source notes + +primary sources checked during review: + +- LongMemEval: https://arxiv.org/abs/2410.10813 +- LongMemEval repo: https://github.com/xiaowu0162/longmemeval +- LongMemEval-V2: https://arxiv.org/abs/2605.12493 +- LongMemEval-V2 repo: https://github.com/xiaowu0162/LongMemEval-V2 +- Mem0: https://arxiv.org/abs/2504.19413 +- HippoRAG: https://arxiv.org/abs/2405.14831 +- RAPTOR: https://arxiv.org/abs/2401.18059 +- GraphRAG: https://arxiv.org/abs/2404.16130 +- LightRAG: https://arxiv.org/abs/2410.05779 +- KAG: https://arxiv.org/abs/2409.13731 +- BGE-M3: https://arxiv.org/abs/2402.03216 +- SQLite FTS5: https://www.sqlite.org/fts5.html +- IRCoT: https://arxiv.org/abs/2212.10509 +- WhenLoss-style diagnostics: https://arxiv.org/abs/2605.24579 +- deterministic freshness/conflict assembly: https://arxiv.org/abs/2606.01435 +- Zep/Graphiti: https://arxiv.org/abs/2501.13956 +- MemWeave README: https://github.com/sachinsharma9780/memweave/blob/main/README.md diff --git a/docs/memory-arch.md b/docs/memory-arch.md index 03f3d32..2e05096 100644 --- a/docs/memory-arch.md +++ b/docs/memory-arch.md @@ -1,5 +1,14 @@ # a stronger memory architecture for your llm harness +status: historical +reviewed: 2026-06-29 + +this is an older broad memory architecture report. keep it for lane/ref-substrate +rationale, but do not treat it as the current retrieval roadmap. use +[`long-term-memory-arch.md`](./long-term-memory-arch.md) for the current +direction and [`memory-implementation-status.md`](./memory-implementation-status.md) +for what is implemented today. + ## bottom line the best move is not to bolt on yet another memory subsystem. it is to make one clear memory stack with four lanes and one address space: a **live context lane** for immediate conversational continuity, an **episodic lane** for concrete events and scenes, a **semantic lane** for durable notes and distilled knowledge, and a **profile/procedural lane** for stable preferences, policies, and habits. all of them should sit on top of a single reference substrate so every turn chunk, note paragraph, summary, asset, and derived memory is addressable with a stable id and exact provenance. that gives you the thing your current harness is closest to but not fully doing yet: deterministic recall when exact references exist, plus associative retrieval when they do not. fileciteturn0file0 citeturn2academia1turn4academia1turn13academia5turn3academia3 @@ -8,7 +17,8 @@ i would not replace your sqlite-plus-files direction with a graph database. your ## current-source map -this doc is the target architecture, not a point-in-time implementation report. +this historical doc is target architecture from an earlier phase, not a +point-in-time implementation report. for current verification and open beads, use [`docs/memory-implementation-status.md`](./memory-implementation-status.md). diff --git a/docs/memory-benches.md b/docs/memory-benches.md index 31f1527..cf41195 100644 --- a/docs/memory-benches.md +++ b/docs/memory-benches.md @@ -1,4 +1,13 @@ -the current benchmark situation feels weird because it is benchmarking *lanes* instead of the memory system: retrieval dev/test, passive recall, router, tools-lane, sweeps, older longmem wrappers, and newer longmemeval commands all coexist, while the report itself says the full runtime really has three retrieval channels: passive semantic recall, active memory tools, and deterministic reflink resolution. it also says the benchmarks mostly evaluate retrieval policy rather than the full continuous agent loop. that is the core smell. +# memory benchmark protocol + +status: current +updated: 2026-06-29 + +this is the active benchmark protocol doc. if benchmark status conflicts with +[`memory-implementation-status.md`](./memory-implementation-status.md), verify +against current source and update both docs. + +the current benchmark situation feels weird because it is benchmarking *lanes* instead of the memory system: retrieval dev/test, passive recall, router, tools-lane, sweeps, older longmem wrappers, and newer longmemeval commands all coexist, while the report itself says the full runtime really has three retrieval channels: passive semantic recall, active memory tools, and deterministic reflink resolution. it also says the benchmarks mostly evaluate retrieval policy rather than the full continuous agent loop. that is the core smell. my recommendation: make **longmemeval the main public scoreboard**, but not the only eval in the repo. diff --git a/docs/memory-implementation-status.md b/docs/memory-implementation-status.md index e81286c..f54f4c7 100644 --- a/docs/memory-implementation-status.md +++ b/docs/memory-implementation-status.md @@ -1,6 +1,11 @@ # memory architecture implementation status -updated: 2026-06-27 +status: current +updated: 2026-06-29 + +read this with [`long-term-memory-arch.md`](./long-term-memory-arch.md). this +file says what exists today; the architecture doc says where the memory stack is +going next. ## implemented @@ -55,6 +60,20 @@ updated: 2026-06-27 - `--official-eval-cmd` - manifest/report/store stats output +## open direction + +- `klbr-u0q`: implement session/event-first retrieval and lexical containment. +- replace the cue-list operation planner in `klbr-core/src/planner.rs` with a + language-agnostic structured classifier while keeping deterministic fallback + tests. +- add a session/event-first retrieval profile where dense episode cards seed + sessions and chunk hits enrich evidence packets. +- finish collect-mode fact-group sufficiency and gap diagnosis for aggregate, + order, and update-resolution questions. +- render fact rows/timeline rows for operation-aware synthesis. +- add multilingual retrieval fixtures so english lexical shortcuts cannot become + hidden policy. + ## verification ```bash diff --git a/docs/memory-retrieval-research-primer.md b/docs/memory-retrieval-research-primer.md index 95193d1..575d187 100644 --- a/docs/memory-retrieval-research-primer.md +++ b/docs/memory-retrieval-research-primer.md @@ -1,5 +1,14 @@ # memory retrieval research primer +status: supporting +reviewed: 2026-06-29 + +this is a research brief from the evidence-packet phase. some current-state +claims are historical. use [`long-term-memory-arch.md`](./long-term-memory-arch.md) +for current direction and +[`memory-implementation-status.md`](./memory-implementation-status.md) for live +implementation status. + ## why this report exists klbr's new memory stack is ref-native and lane-aware enough that the next hard diff --git a/docs/planner-grade-evidence-assembly.md b/docs/planner-grade-evidence-assembly.md index f62ff7c..1c8eafa 100644 --- a/docs/planner-grade-evidence-assembly.md +++ b/docs/planner-grade-evidence-assembly.md @@ -1,5 +1,13 @@ # klbr’s next fix is planner-grade evidence assembly +status: supporting +reviewed: 2026-06-29 + +this plan is partially implemented under beads epic `klbr-f0q`. it remains +useful for collect-mode, fact-row, and structured-synthesis work, but +[`long-term-memory-arch.md`](./long-term-memory-arch.md) is now the higher-level +current architecture direction. + ## diagnosis the good news is that the earlier “right session, wrong chunk” problem looks substantially less central now. your internal benchmark gap report already described the original failure as a selection-granularity bug, where session recall was high but qa accuracy lagged because answer-bearing neighboring turns were getting dropped; that report also notes the move to `EvidencePacket` planning as the production fix. your memory system report, meanwhile, describes `MemoryPipeline` as the shared facade for observing sessions, retrieving evidence, assembling context, and answering through the real stack rather than a benchmark-only shortcut. in other words, klbr is no longer mostly failing because it cannot *find* the right area; it is now more often failing because it does not always *plan the right evidence set and reasoning mode* for the question type. fileciteturn0file0 fileciteturn0file1 @@ -311,4 +319,4 @@ fourth, keep lexical retrieval on a short leash. sqlite fts5 and tokenizer-level the biggest risks are pretty predictable. collect mode can bloat context if you let it turn into “retrieve everything,” so the stopping condition has to be fact-group sufficiency, not packet count. timeline extraction can introduce its own errors, so every structured row must keep its supporting ref. preference mode can become overly conservative and say “i don’t know” too often if it requires too much personal support; but that is still better than regressing into generic recommendations. and multilingual behavior will get worse, not better, if you keep growing english cue lists; tokenizer-level lexical retrieval and multilingual sparse/dense channels are the safer bet there. cross-domain leakage is also a real risk when extra budget adds semantically adjacent but task-irrelevant memories, which is exactly why preference and temporal questions need stricter planner rules than plain lookup. citeturn12academia3turn12academia0turn6academia1turn8view0turn8view1 -if i had to compress all of this into one sentence: **the next strong version of klbr should stop being a very good packet retriever with a hopeful reader, and become a query-planned evidence system that knows when it is doing lookup, collection, chronology, update resolution, or preference grounding.** that is the cleanest way to turn your current “budget helps coverage but not answers” plateau into actual qa gains. fileciteturn0file0 fileciteturn0file1 citeturn12academia2turn5academia2turn10academia3turn10academia0 \ No newline at end of file +if i had to compress all of this into one sentence: **the next strong version of klbr should stop being a very good packet retriever with a hopeful reader, and become a query-planned evidence system that knows when it is doing lookup, collection, chronology, update resolution, or preference grounding.** that is the cleanest way to turn your current “budget helps coverage but not answers” plateau into actual qa gains. fileciteturn0file0 fileciteturn0file1 citeturn12academia2turn5academia2turn10academia3turn10academia0 diff --git a/docs/principled-evidence-selection-fusion.md b/docs/principled-evidence-selection-fusion.md index 094b69f..e6cc444 100644 --- a/docs/principled-evidence-selection-fusion.md +++ b/docs/principled-evidence-selection-fusion.md @@ -1,5 +1,11 @@ # principled evidence selection and fusion for klbr +status: historical +reviewed: 2026-06-29 + +this is an earlier evidence-fusion report. keep it as rationale only. use +[`long-term-memory-arch.md`](./long-term-memory-arch.md) for current direction. + ## executive summary klbr’s next bottleneck is exactly what your bench result suggests: the system is often retrieving the right **session** but assembling the wrong **evidence object**. the current stack already has the right raw ingredients for a better design — canonical refs, aliases, edges, promptable text, FTS, embeddings, and a benchmark-facing `MemoryPipeline` facade with `observe_session`, `retrieve_evidence`, `assemble_context`, `answer`, and `complete_stored_context` — but the remaining gap is selection policy, not storage substrate. the uploaded architecture report also already frames the system as three channels today: passive semantic recall, active memory tools, and deterministic reflink resolution, with the caution that runtime and benchmark behavior are not yet fully identical. fileciteturn0file1 @@ -648,4 +654,4 @@ the safest rollout plan is therefore: - then wire runtime passive recall to it behind a config flag; - finally remove the old same-session dedupe path once the deterministic fixtures and LongMemEval smoke stay green. fileciteturn0file0turn0file1 -the short version is: **klbr should stop fusing raw hits and start selecting budgeted evidence packets**. that one move contains the same-session dedupe bug, gives exact refs a principled home, keeps SQLite and canonical refs, constrains lexical policy, and makes both benchmarking and runtime behavior meaningfully more truthful to what the agent actually needs to answer. \ No newline at end of file +the short version is: **klbr should stop fusing raw hits and start selecting budgeted evidence packets**. that one move contains the same-session dedupe bug, gives exact refs a principled home, keeps SQLite and canonical refs, constrains lexical policy, and makes both benchmarking and runtime behavior meaningfully more truthful to what the agent actually needs to answer.