From e5c89ba61ceb48ad16d09de3cdaafeab7cecca11 Mon Sep 17 00:00:00 2001 From: dawn <90008@klbr.net> Date: Tue, 30 Jun 2026 14:14:45 +0300 Subject: [PATCH] track passive memory followups --- .beads/issues.jsonl | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/.beads/issues.jsonl b/.beads/issues.jsonl index 86a0bcc..8b38f30 100644 --- a/.beads/issues.jsonl +++ b/.beads/issues.jsonl @@ -1,10 +1,13 @@ +{"_type":"issue","id":"klbr-54l.3","title":"Stabilize op planner labels for LongMemEval update and count cases","description":"The same-seed regression changed planner behavior heavily: old op_plan_counts had update_resolution=10 and lookup=3, while current has update_resolution=0 and lookup=15. That can disable collect/update behavior and shrink packet budgets even when retrieval finds the right area. Harden the structured planner path and fallback behavior for aggregate_count, aggregate_sum, aggregate_avg, order_or_rank, update_resolution, and preference_recommendation without reintroducing hidden English cue lists as the main policy.","acceptance_criteria":"Known klbr-54l rows get stable op labels across reruns or model-parser fallback; update_resolution no longer collapses to lookup on the same 25-row sample; tests cover malformed structured planner output and paraphrased update/count questions; implementation remains language-agnostic rather than relying on hardcoded English cue lists.","status":"open","priority":1,"issue_type":"bug","owner":"90008@klbr.net","created_at":"2026-06-30T11:13:17Z","created_by":"dawn","updated_at":"2026-06-30T11:13:17Z","labels":["bench","longmemeval","memory","planner"],"dependencies":[{"issue_id":"klbr-54l.3","depends_on_id":"klbr-54l","type":"parent-child","created_at":"2026-06-30T14:13:17Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l.3","depends_on_id":"klbr-54l.1","type":"blocks","created_at":"2026-06-30T14:14:16Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} +{"_type":"issue","id":"klbr-54l.2","title":"Expand session-first episode anchors into supporting turn-window packets","description":"The klbr-54l trace research found a concrete failure class: stage_one finds the right session, but session-first emits thin episode_bridge packets from episode/card candidates and the final context drops the actual answer-bearing turn span. For 6b168ec8, the right session is found and a late episode bridge contains the three-bikes evidence, but the rendered context spends budget on irrelevant fts turn windows. In session-first mode, strong episode/session anchors should enrich or produce supporting turn_window packets from the same session's best seed refs instead of relying only on expand_edges from the episode card.","acceptance_criteria":"A production-pipeline fixture covers the candidate-session-hit / thin-episode-bridge-dropped failure; when the answer session is a strong stage_one group, packet planning includes a non-thin supporting turn_window from that session; 6b168ec8 no longer loses the three-bikes evidence before reader time without increasing global read budget.","status":"open","priority":1,"issue_type":"bug","owner":"90008@klbr.net","created_at":"2026-06-30T11:13:10Z","created_by":"dawn","updated_at":"2026-06-30T11:13:10Z","labels":["bench","evidence-packets","longmemeval","memory"],"dependencies":[{"issue_id":"klbr-54l.2","depends_on_id":"klbr-54l","type":"parent-child","created_at":"2026-06-30T14:13:10Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l.2","depends_on_id":"klbr-54l.1","type":"blocks","created_at":"2026-06-30T14:14:16Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} +{"_type":"issue","id":"klbr-54l.1","title":"Trace-diff old-pass/current-fail LongMemEval passive QA rows","description":"Build a reusable trace-diff for the same-seed LongMemEval-S 25-row regression in klbr-54l before changing ranking or planner code. Compare baseline benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z against current benchmarks/runs/longmemeval-s/klbr-full/2026-06-30_qa_sample25_seed4937553249516211047_current for old-pass/current-fail ids: 6b168ec8, c14c00dd, b5ef892d, 46a3abf7, 720133ac, gpt4_385a5000, dad224aa. Report per question: op_plan, stage_one answer-session rank, candidate vs packet recall, packet order/kind/session, answer-bearing selected/rendered/in-context, answer_value_visible, critical_fact_row_visible, packet/context tokens, omissions, and reader hypothesis.","acceptance_criteria":"A command, test, or bench report emits a per-question diff for the listed ids; 6b168ec8 clearly shows the candidate-session-hit but answer-packet-not-rendered class; output is durable enough to rerun during klbr-54l without hand-inspecting trace.jsonl.","status":"open","priority":1,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-30T11:13:02Z","created_by":"dawn","updated_at":"2026-06-30T11:13:02Z","labels":["bench","longmemeval","memory","passive-recall"],"dependencies":[{"issue_id":"klbr-54l.1","depends_on_id":"klbr-54l","type":"parent-child","created_at":"2026-06-30T14:13:02Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":6,"comment_count":0} {"_type":"issue","id":"klbr-7yo.5","title":"Document folgezettel memory semantics and add regression coverage","description":"Document the new memory ontology for agents: titles are APIs, bodies are attractors, source refs are required for grounded notes, and old memory CRUD tools are not the agent-facing surface. Add focused tests and a small smoke/regression path so the new trail traversal does not silently regress.","acceptance_criteria":"Docs describe mk_* tools and note-writing policy; tests cover exposed tool names, zettel write/follow/recall/revise/link/fleeting flows, and trail packet expansion. If LongMemEval regression klbr-54l is still open, docs should state how the new topology relates to that failure mode without claiming it is solved unless verified.","status":"closed","priority":1,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-30T01:21:30Z","created_by":"dawn","updated_at":"2026-06-30T01:39:41Z","closed_at":"2026-06-30T01:39:41Z","close_reason":"Added folgezettel docs, AGENTS/status updates, exposed-tool tests, mk flow tests, and zettel trail packet regression coverage; klbr-54l remains open for bench validation.","labels":["benchmarks","docs","folgezettel","memory","tests","tools"],"dependencies":[{"issue_id":"klbr-7yo.5","depends_on_id":"klbr-7yo","type":"parent-child","created_at":"2026-06-30T04:21:29Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-7yo.5","depends_on_id":"klbr-7yo.2","type":"blocks","created_at":"2026-06-30T04:21:31Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-7yo.1","title":"Add folgezettel note topology to the memory substrate","description":"Represent zettel trails explicitly on top of existing refs/edges/markdown notes. A zettel title is the API handle; the body is a sourced attractor. Need parent/child continuation, branch/cross-link, supersession/revision, and fleeting/unplaced capture without replacing immutable turn/source refs.","acceptance_criteria":"MemoryStore can write zettel/fleeting markdown notes with frontmatter for title, parent/trail metadata, source refs, and status; edges encode follows/branch/source/supersedes relations; index/follow/recall queries can recover roots, children, siblings/path, full note bodies, and source refs.","status":"closed","priority":1,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-30T01:21:19Z","created_by":"dawn","updated_at":"2026-06-30T01:39:23Z","started_at":"2026-06-30T01:21:44Z","closed_at":"2026-06-30T01:39:23Z","close_reason":"Completed folgezettel note/fleeting storage paths, continuation/source/supersession edges, and trail lookup substrate.","labels":["folgezettel","memory","tools"],"dependencies":[{"issue_id":"klbr-7yo.1","depends_on_id":"klbr-7yo","type":"parent-child","created_at":"2026-06-30T04:21:19Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":2,"comment_count":0} {"_type":"issue","id":"klbr-7yo.2","title":"Implement mk_* folgezettel memory tools","description":"Add the agent-facing mk_* tool family: mk_index, mk_search, mk_follow, mk_recall, mk_remember, mk_revise, mk_link, and mk_fleeting. Tools should make the note move explicit and hide old memory CRUD semantics from the model.","acceptance_criteria":"Each mk_* tool has a focused JSON schema, concise success/error output, source-ref handling where required, and tests. mk_follow returns structure/titles only; mk_recall reads one note fully; mk_remember requires a claim-like title and parent/root placement; mk_fleeting captures unplaced scratch.","status":"closed","priority":1,"issue_type":"feature","owner":"90008@klbr.net","created_at":"2026-06-30T01:21:19Z","created_by":"dawn","updated_at":"2026-06-30T01:39:27Z","closed_at":"2026-06-30T01:39:27Z","close_reason":"Implemented mk_index/mk_search/mk_follow/mk_recall/mk_remember/mk_revise/mk_link/mk_fleeting with focused tool-flow coverage.","labels":["folgezettel","memory","tools"],"dependencies":[{"issue_id":"klbr-7yo.2","depends_on_id":"klbr-7yo","type":"parent-child","created_at":"2026-06-30T04:21:19Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-7yo.2","depends_on_id":"klbr-7yo.1","type":"blocks","created_at":"2026-06-30T04:21:29Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":2,"comment_count":0} {"_type":"issue","id":"klbr-7yo.3","title":"Expose only mk_* memory tools to agents","description":"Swap memory_tools() and normal all_tools() exposure away from remember/recall/context_for/fetch_memories/write_memory_note/edit_memory/list_memories toward the mk_* ontology so agents do not see two competing memory models. Compatibility with old exposed tool names is explicitly not required.","acceptance_criteria":"Runtime agent/reflection registries expose mk_* memory tools only, with old tools either internal-only or removed from the normal registry; tests assert the exposed tool names; docs explain migration and current tool surface.","status":"closed","priority":1,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-30T01:21:19Z","created_by":"dawn","updated_at":"2026-06-30T01:39:32Z","closed_at":"2026-06-30T01:39:32Z","close_reason":"memory_tools/all_tools now expose mk_* only for memory; old public memory tool names are removed from the normal registry and asserted in tests.","labels":["agent","folgezettel","memory","tools"],"dependencies":[{"issue_id":"klbr-7yo.3","depends_on_id":"klbr-7yo","type":"parent-child","created_at":"2026-06-30T04:21:19Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-7yo.3","depends_on_id":"klbr-7yo.2","type":"blocks","created_at":"2026-06-30T04:21:30Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-7yo.4","title":"Teach retrieval and context assembly to walk zettel trails","description":"After retrieval finds a zettel/note entrypoint, packet assembly should expand local folgezettel context: parent/path, previous/next/siblings when available, children/branches under budget, and source refs. This should improve answer-ready context without replacing dense/fts entrypoint search.","acceptance_criteria":"Evidence planner or pipeline can produce trail-aware packets from zettel refs; trace fields expose trail expansion; packets include source bodies near zettel attractor bodies; focused tests cover matched note -\u003e parent/child/source expansion under budget.","status":"closed","priority":1,"issue_type":"feature","owner":"90008@klbr.net","created_at":"2026-06-30T01:21:19Z","created_by":"dawn","updated_at":"2026-06-30T01:39:36Z","closed_at":"2026-06-30T01:39:36Z","close_reason":"EvidencePlanner adds zettel trail_support packet bodies from parent/sibling/child/source refs, including chunk-to-note normalization and regression coverage.","labels":["folgezettel","memory","retrieval","tools"],"dependencies":[{"issue_id":"klbr-7yo.4","depends_on_id":"klbr-7yo","type":"parent-child","created_at":"2026-06-30T04:21:19Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-7yo.4","depends_on_id":"klbr-7yo.1","type":"blocks","created_at":"2026-06-30T04:21:30Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-7yo","title":"Make memory folgezettel-native for agents","description":"Replace the agent-facing memory ontology with folgezettel-style tools and note semantics. Titles are APIs: compact reusable claims. Bodies are attractors: sourced reasoning/provenance that helps the model reconstruct why the title is true. The implementation should keep immutable turns/refs as substrate, but expose mk_* tools instead of the old memory CRUD/search tool set so agents learn to edit trails rather than buckets.","design":"Root idea: retrieval finds entrypoints, folgezettel topology decides local traversal. Compatibility with old agent-facing memory tools is not required; keep old internals only if useful during migration. Expose one ontology: mk_index, mk_search, mk_follow, mk_recall, mk_remember, mk_revise, mk_link, mk_fleeting. Titles are APIs; bodies are attractors grounded in source refs.","acceptance_criteria":"Agent-facing memory tools are mk_* only; zettel notes have claim-like API titles and attractor bodies with source refs; agents can index/search/follow/recall/write/revise/link/fleeting-capture trails; retrieval/context assembly can include local trail packets with source refs; docs and focused tests cover the new ontology and old-tool exposure is removed from normal agent tool registries.","status":"closed","priority":1,"issue_type":"epic","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-30T01:20:58Z","created_by":"dawn","updated_at":"2026-06-30T01:39:45Z","started_at":"2026-06-30T01:21:44Z","closed_at":"2026-06-30T01:39:45Z","close_reason":"Completed folgezettel-native agent memory surface, zettel/fleeting storage, trail-aware packet support, docs, and focused regression tests.","labels":["folgezettel","memory","tools"],"dependency_count":0,"dependent_count":0,"comment_count":0} -{"_type":"issue","id":"klbr-54l","title":"Investigate LongMemEval QA regression after memory evidence changes","description":"Same-seed LongMemEval-S 25-sample QA dropped after the recent memory evidence/planner/session-first changes. Baseline run benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z scored 0.6800. Current run benchmarks/runs/longmemeval-s/klbr-full/2026-06-30_qa_sample25_seed4937553249516211047_current scored 0.4400 with the same sample seed 4937553249516211047 and local judge model. Session recall stayed similar, but answer-bearing/rendered evidence and packet/context budget dropped: answer_bearing_ref_in_context 0.80 -\u003e 0.36, answer_bearing_ref_selected 0.84 -\u003e 0.44, packet_tokens_total_mean 4412.64 -\u003e 1811.28. Op plans also changed heavily: old update_resolution=10, current update_resolution=0 and lookup=15. Regressions old-pass/current-fail: 6b168ec8, c14c00dd, b5ef892d, 46a3abf7, 720133ac, gpt4_385a5000, dad224aa. Fix old-fail/current-pass: gpt4_59149c77.","acceptance_criteria":"Identify whether the loss comes from structured model planning, packet selection/rendering budget, answer-bearing ref propagation, or reader prompting; add a focused regression/smoke that prevents the same same-seed sample from losing old passing rows without an explicit expected-metric update.","status":"open","priority":1,"issue_type":"bug","owner":"90008@klbr.net","created_at":"2026-06-30T01:00:03Z","created_by":"dawn","updated_at":"2026-06-30T01:00:03Z","dependency_count":0,"dependent_count":0,"comment_count":0} +{"_type":"issue","id":"klbr-54l","title":"Investigate LongMemEval QA regression after memory evidence changes","description":"Same-seed LongMemEval-S 25-sample QA dropped after the recent memory evidence/planner/session-first changes. Baseline run benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z scored 0.6800. Current run benchmarks/runs/longmemeval-s/klbr-full/2026-06-30_qa_sample25_seed4937553249516211047_current scored 0.4400 with the same sample seed 4937553249516211047 and local judge model. Session recall stayed similar, but answer-bearing/rendered evidence and packet/context budget dropped: answer_bearing_ref_in_context 0.80 -\u003e 0.36, answer_bearing_ref_selected 0.84 -\u003e 0.44, packet_tokens_total_mean 4412.64 -\u003e 1811.28. Op plans also changed heavily: old update_resolution=10, current update_resolution=0 and lookup=15. Regressions old-pass/current-fail: 6b168ec8, c14c00dd, b5ef892d, 46a3abf7, 720133ac, gpt4_385a5000, dad224aa. Fix old-fail/current-pass: gpt4_59149c77.","acceptance_criteria":"Identify whether the loss comes from structured model planning, packet selection/rendering budget, answer-bearing ref propagation, or reader prompting; add a focused regression/smoke that prevents the same same-seed sample from losing old passing rows without an explicit expected-metric update.","status":"open","priority":1,"issue_type":"bug","owner":"90008@klbr.net","created_at":"2026-06-30T01:00:03Z","created_by":"dawn","updated_at":"2026-06-30T01:00:03Z","dependencies":[{"issue_id":"klbr-54l","depends_on_id":"klbr-54l.1","type":"blocks","created_at":"2026-06-30T14:14:18Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l","depends_on_id":"klbr-54l.2","type":"blocks","created_at":"2026-06-30T14:14:19Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l","depends_on_id":"klbr-54l.3","type":"blocks","created_at":"2026-06-30T14:14:19Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l","depends_on_id":"klbr-54l.4","type":"blocks","created_at":"2026-06-30T14:14:20Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l","depends_on_id":"klbr-54l.5","type":"blocks","created_at":"2026-06-30T14:14:20Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l","depends_on_id":"klbr-54l.6","type":"blocks","created_at":"2026-06-30T14:14:20Z","created_by":"dawn","metadata":"{}"}],"dependency_count":6,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-u0q","title":"Implement session-first lexical-contained memory retrieval","description":"Build the next retrieval architecture from docs/long-term-memory-arch.md: stage-one session/event candidate generation, lexical containment, and language-agnostic operation planning experiments without adding english cue-word hacks.","design":"Keep EvidencePacket as the rank object. Use exact refs, dense episode/session cards, learned sparse or fts side channels, and graph expansion from strong seeds only. Keep QueryOp shape but replace cue-list policy with a structured classifier plus deterministic fallback.","acceptance_criteria":"A bench profile can retrieve from session/episode cards first and use chunk fts/sparse hits as packet enrichment; lexical-only policy paths are removed or contained behind candidate channels; traces expose candidate session recall and packet/rendered evidence metrics; multilingual fixtures cover non-English cue-free retrieval cases.","notes":"2026-06-29 partial: removed hardcoded english cue-list planner, english lane routing, english negation/entity-boundary packet filters, and wh/pronoun neighbor-expansion heuristics. OpPlan now comes from structured model JSON when an llm endpoint is configured, otherwise conservative lookup. Core and bench tests pass.\n2026-06-29 partial: landed initial session-first retrieval profile for klbr-full/dense-only/session-first profiles. Retrieval now groups archival candidates by session_id, prefers episodic/event-card anchors, adds chunk fts/sparse/dense hits as enrichment, serializes stage_one.session_candidates, and reports CandidateSessionRecall metrics in klbr-bench. fts-only now writes episode notes but skips embedded episode-memory rows so lexical ablations do not require the dense embedder at ingest. Verified core/bench tests, continuous-loop, and a one-row session-first/fts-only retrieval-only trace smoke.","status":"closed","priority":1,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T22:06:32Z","created_by":"dawn","updated_at":"2026-06-29T12:04:10Z","started_at":"2026-06-28T22:10:12Z","closed_at":"2026-06-29T12:04:10Z","close_reason":"Completed session/event-first lexical-contained slice: no english cue-list planner/filters, fts remains a candidate channel, trigram fts covers no-space multilingual exact-ish retrieval, traces expose session candidates/packet metrics, and docs clarify fts vs semantic multilingual retrieval.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-jwn","title":"Organize evolving memory architecture docs","description":"Make the memory architecture docs distinguish current truth, implementation status, proposals, research inputs, and archived rationale so future agents do not treat older reports as canonical.","acceptance_criteria":"Docs have a clear index and status taxonomy; current memory architecture direction points at docs/long-term-memory-arch.md; older research reports are marked as historical or supporting; AGENTS.md routes future agents through the doc index.","status":"closed","priority":1,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T22:01:20Z","created_by":"dawn","updated_at":"2026-06-28T22:07:09Z","started_at":"2026-06-28T22:01:23Z","closed_at":"2026-06-28T22:07:09Z","close_reason":"Completed docs index, current architecture review rewrite, status banners, AGENTS routing, and follow-up implementation issue klbr-u0q.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-h4l","title":"Cache bench embeddings locally and run 25-sample QA bench","description":"Make LongMemEval bench embeddings use a project-local gitignored sqlite cache so reruns do not pay for the same embeddings twice. Then run a 25-sample stratified QA bench, inspect failures, try bounded non-overfit tweaks if evidence supports them, and prepare a report packet if results remain weak or tuning would be overfit.","design":"Prefer a global project cache path such as benchmarks/cache/*.db over temp/run-local caches. Keep benchmark polling sparse: start the run, wait for artifacts or process completion, then inspect once.","acceptance_criteria":"Embedding cache lives under the project and is ignored by git; cache hits are reused across bench runs; a 25-sample bench artifact exists with failure analysis; any code changes are tested, committed, and pushed.","notes":"Implemented project-local embedding cache at benchmarks/cache/embeddings.db and wired LongMemEval bench LlmClient construction through it. Ran stratified 25-sample QA on seed 4937553249516211047: current-code run benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z scored official accuracy 0.6800. Report packet written to report_packet.md in that run dir. High-budget targeted reruns fixed only dd2973ad and a1cc6108; remaining failures point to fact/timeline synthesis rather than safe cue tuning.","status":"closed","priority":1,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T20:08:32Z","created_by":"dawn","updated_at":"2026-06-28T21:21:27Z","started_at":"2026-06-28T20:08:35Z","closed_at":"2026-06-28T21:21:27Z","close_reason":"Implemented project-local dense/sparse embedding cache, ran 25-sample QA bench, inspected failures, tried bounded planner/high-budget experiments, and wrote report packet.","dependency_count":0,"dependent_count":0,"comment_count":0} @@ -37,6 +40,9 @@ {"_type":"issue","id":"klbr-wmz.2","title":"Route runtime passive recall through the canonical memory pipeline","description":"Bench retrieval now goes through MemoryPipeline, but runtime passive recall in klbr-core/src/agent.rs still embeds the prompt, calls MemoryStore::get_searchable, runs retrieval::retrieve_exact over legacy memories, then injects Context memory packets. That bypasses fts, exact refs, markdown notes, graph expansion, and the canonical lane/lifecycle path the docs describe.","design":"Avoid duplicating retrieval logic in agent.rs. Either make MemoryPipeline usable by AgentRuntime or extract a shared retrieval facade that both MemoryPipeline and runtime passive recall call.","acceptance_criteria":"Runtime passive recall uses the same lane-aware canonical retrieval and context packet assembly policy as the production pipeline; recalled packets can include fts/exact/dense/graph candidates from refs and markdown notes; archived/tombstoned/suppressed refs do not leak; klbr-core/src/instructions.md matches the actual memory packet format; tests or a focused integration fixture cover passive recall from a markdown note and from an explicit ref.","notes":"Runtime passive recall now calls MemoryPipeline::retrieve_evidence and injects shared EvidencePacket XML via Context::inject_evidence_packets; no-model runtime packet fixture covers turn-window expansion. Remaining acceptance is blocked on klbr-wmz.1 because dense search still starts from legacy memory rows before ref mapping.","status":"closed","priority":1,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-26T17:53:19Z","created_by":"dawn","updated_at":"2026-06-26T19:44:55Z","started_at":"2026-06-26T19:37:32Z","closed_at":"2026-06-26T19:44:55Z","close_reason":"Completed: runtime passive recall now uses MemoryPipeline/EvidencePlanner packets, instructions document current packet XML, and fixtures cover turn-window recall plus explicit markdown refs; dense canonical dependency completed in klbr-wmz.1.","labels":["architecture","memory","retrieval","runtime"],"dependencies":[{"issue_id":"klbr-wmz.2","depends_on_id":"klbr-wmz","type":"parent-child","created_at":"2026-06-26T20:53:19Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-wmz.2","depends_on_id":"klbr-wmz.1","type":"blocks","created_at":"2026-06-26T22:37:53Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-wmz.1","title":"Move dense retrieval onto canonical refs and embedding_items","description":"Current source still has dense retrieval on the legacy memory-row surface: klbr-core/src/pipeline.rs search_dense calls MemoryStore::get_searchable and retrieval::retrieve_exact, while fts/exact retrieval uses canonical refs and promptable_text. The schema already has refs, promptable_text, markdown_note_chunks, and embedding_items, so dense retrieval should not be the odd path out.","design":"Prefer a ref-native embedding index backed by embedding_items. Backfill embeddings from promptable_text, keep memory-id aliases as compatibility aliases, and make klbr/full versus dense-only profiles exercise the same canonical identity layer as fts and exact retrieval.","acceptance_criteria":"Dense candidate generation works over canonical ref ids for memories, turn chunks, markdown note chunks, episode notes, profile notes, and procedural notes; lane and lifecycle filtering come from refs/ref_metadata instead of memory tags alone; benchmark traces return canonical refs for dense hits; regression tests cover a markdown-note-only hit and a tombstoned/suppressed ref not leaking through dense search.","status":"closed","priority":1,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-26T17:53:11Z","created_by":"dawn","updated_at":"2026-06-26T19:43:27Z","started_at":"2026-06-26T19:38:02Z","closed_at":"2026-06-26T19:43:27Z","close_reason":"Completed: dense candidate generation now lazily backfills embedding_items from active promptable refs, scores canonical ref embeddings directly, and tests markdown-note dense hits plus suppressed-ref filtering.","labels":["architecture","memory","refs","retrieval"],"dependencies":[{"issue_id":"klbr-wmz.1","depends_on_id":"klbr-wmz","type":"parent-child","created_at":"2026-06-26T20:53:11Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":1,"comment_count":0} {"_type":"issue","id":"klbr-wmz","title":"Finish memory architecture follow-through","description":"Tracks the remaining memory architecture work identified from docs/memory-arch.md, docs/memory-benches.md, docs/memory-implementation-status.md, and current klbr-core/klbr-bench source. Current status says pipeline, typed memory packets, markdown notes, edge mirroring, lifecycle projection, and benchmark runner integration exist; this epic is for gaps still present in source/docs.","acceptance_criteria":"Close when the child issues are complete, docs/memory-implementation-status.md is updated from current verification, and the architecture docs no longer point at missing or stale follow-up work.","status":"closed","priority":1,"issue_type":"epic","owner":"90008@klbr.net","created_at":"2026-06-26T17:52:54Z","created_by":"dawn","updated_at":"2026-06-26T23:45:35Z","closed_at":"2026-06-26T23:45:35Z","close_reason":"All 19 memory architecture child issues are closed; docs/status were updated from current verification; remaining official evaluator run is tracked separately as external blocked klbr-1yn.","labels":["architecture","memory"],"dependency_count":0,"dependent_count":0,"comment_count":0} +{"_type":"issue","id":"klbr-54l.6","title":"Probe passive writer retention for atomic facts and events","description":"Research on memory systems suggests passive QA can fail because write-time compression loses entity/slot/value/time facts before retrieval ever has a chance. Klbr now writes episode_event_card artifacts and generic fact_rows, but the next passive-recall improvement pass should add writer-side probes for counts, updates, preference constraints, and temporal facts so we can tell whether the passive writer retained the answer-bearing atom before tuning retrieval.","acceptance_criteria":"Fixtures or diagnostics check that answer-bearing entity/slot/value/time atoms exist in stored promptable artifacts before retrieval; failures are reported separately from packet/ranking misses; at least count, update_resolution, temporal order, and preference-constraint cases are covered.","status":"open","priority":2,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-30T11:13:57Z","created_by":"dawn","updated_at":"2026-06-30T11:13:57Z","labels":["bench","longmemeval","memory","writer"],"dependencies":[{"issue_id":"klbr-54l.6","depends_on_id":"klbr-54l","type":"parent-child","created_at":"2026-06-30T14:13:56Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l.6","depends_on_id":"klbr-54l.1","type":"blocks","created_at":"2026-06-30T14:14:18Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} +{"_type":"issue","id":"klbr-54l.5","title":"Make WhenLoss-style passive QA diagnostics first-class","description":"The passive QA bench should be treated as best-effort passive recall, not the whole memory-system score. The klbr-bench --diagnostic whenloss mode exists, but regression work needs a first-class joined report that separates write-side loss from retrieval/packet/rendering loss per question. Add a report path for tfc, oracle-evidence, complete-stored-memory, and retrieved-memory scores joined with packet metrics and local/official QA outcomes.","acceptance_criteria":"A normal regression run can emit per-question tfc/oe/csm/rm scores plus write_gap and retrieval_gap; report rows join those scores with op_plan, answer_bearing_ref_selected/rendered/in_context, answer_value_visible, and official/local QA labels; docs explicitly frame LongMemEval passive QA as best-effort passive recall rather than agentic folgezettel memory quality.","status":"open","priority":2,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-30T11:13:41Z","created_by":"dawn","updated_at":"2026-06-30T11:13:51Z","labels":["bench","diagnostics","longmemeval","memory"],"dependencies":[{"issue_id":"klbr-54l.5","depends_on_id":"klbr-54l","type":"parent-child","created_at":"2026-06-30T14:13:41Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l.5","depends_on_id":"klbr-54l.1","type":"blocks","created_at":"2026-06-30T14:14:17Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} +{"_type":"issue","id":"klbr-54l.4","title":"Add answer-session packet survival guardrail to context packing","description":"Current passive QA can spend the final read budget on unrelated high-scoring fts turn windows while a packet from a strongly supported answer session exists but is too late or too thin to render. Add a generic runtime-safe packing/ranking guardrail: when a session group has strong multi-signal support in stage_one, preserve at least one useful non-thin support packet from that session before unrelated single-signal packets, without using gold labels or answer ids outside benchmark metrics.","acceptance_criteria":"A fixture where stage_one finds the relevant session but context packing omits its useful packet fails before the change and passes after; traces expose first relevant/answer-session packet rank for diagnostics; answer_bearing_ref_rendered improves on old-pass/current-fail rows without packet_tokens_total_mean exploding back into unbounded context dumping.","status":"open","priority":2,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-30T11:13:26Z","created_by":"dawn","updated_at":"2026-06-30T11:13:26Z","labels":["bench","context-packing","longmemeval","memory"],"dependencies":[{"issue_id":"klbr-54l.4","depends_on_id":"klbr-54l","type":"parent-child","created_at":"2026-06-30T14:13:25Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-54l.4","depends_on_id":"klbr-54l.1","type":"blocks","created_at":"2026-06-30T14:14:17Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} {"_type":"issue","id":"klbr-f0q.5","title":"Add planner-grade benchmark metrics and fixtures","description":"Add rendered-evidence metrics and deterministic production-pipeline fixtures for count, average truncation, ordered temporal sessions, previous-vs-latest, and preference distractor regressions.","design":"Metrics should separate retrieval selection, packet rendering, fact-row visibility, and final answer correctness rather than collapsing them into one recall score.","acceptance_criteria":"Bench reports split selected/rendered/visible evidence metrics; tests cover the concrete failure families from the 20-sample LongMemEval brief; docs list smoke commands for planner/fact-table runs.","notes":"Partial metrics support landed early: LongMemEval reports op_plan_counts plus answer_bearing_ref_selected, answer_bearing_ref_rendered, and answer_value_visible. Remaining: critical_fact_row_visible, collect_mode_fact_group_recall, preference/update-specific correctness metrics, and deterministic fixtures for all failure families.","status":"closed","priority":2,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T19:44:47Z","created_by":"dawn","updated_at":"2026-06-29T12:04:05Z","started_at":"2026-06-29T12:04:00Z","closed_at":"2026-06-29T12:04:05Z","close_reason":"Completed split bench metrics for selected/rendered/value/fact-row visibility and collect-mode fact-group recall, plus deterministic tests for count, sum/avg, order, update, preference distractor, lookup fallback, and cjk no-space trigram retrieval.","dependencies":[{"issue_id":"klbr-f0q.5","depends_on_id":"klbr-f0q","type":"parent-child","created_at":"2026-06-28T22:44:46Z","created_by":"dawn","metadata":"{}"},{"issue_id":"klbr-f0q.5","depends_on_id":"klbr-f0q.2","type":"blocks","created_at":"2026-06-28T22:44:56Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-oba","title":"add stratified random longmemeval sampling","description":"support reproducible random subset runs for the LongMemEval pipeline, stratified by question_type so small incremental qa runs still cover categories.","status":"closed","priority":2,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T16:09:37Z","created_by":"dawn","updated_at":"2026-06-28T16:16:07Z","started_at":"2026-06-28T16:10:01Z","closed_at":"2026-06-28T16:16:07Z","close_reason":"Completed: added reproducible random LongMemEval sampling with question_type stratification.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-8dq","title":"default longmemeval qa output under benchmarks/runs","description":"make the qa bench less annoying to run by defaulting its output directory under benchmarks/runs when --out is not supplied.","status":"closed","priority":2,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T15:55:48Z","created_by":"dawn","updated_at":"2026-06-28T16:16:07Z","started_at":"2026-06-28T15:56:03Z","closed_at":"2026-06-28T16:16:07Z","close_reason":"Completed: LongMemEval pipeline run now defaults --out under benchmarks/runs.","dependency_count":0,"dependent_count":0,"comment_count":0} @@ -59,6 +65,7 @@ {"_type":"issue","id":"klbr-wmz.8","title":"Validate the LoCoMo adapter and comparison protocol","description":"docs/memory-benches.md recommends LoCoMo as the secondary public suite for memweaver-style comparison, and docs/memory-implementation-status.md says a flexible locomo adapter exists. Current source parses locomo-ish shapes inside klbr-bench/src/longmemeval.rs, but there is no verified dataset fixture, loader test, scoring protocol, or documented run result proving the adapter matches the intended benchmark semantics.","design":"Keep this as a comparison harness task, not a claim about beating another system. The output should make dataset version, reader model, token budget, and scoring method explicit.","acceptance_criteria":"A small LoCoMo-shaped fixture or documented local dataset path exercises the adapter; loader tests cover supported input shapes and session ordering; the benchmark manifest/report clearly labels LoCoMo runs and avoids claiming comparability without matched reader/budget/scoring; docs include the exact command and current verified result or blocker.","status":"closed","priority":2,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-26T17:54:09Z","created_by":"dawn","updated_at":"2026-06-26T23:44:57Z","started_at":"2026-06-26T23:42:41Z","closed_at":"2026-06-26T23:44:57Z","close_reason":"Added LoCoMo smoke fixture, loader tests for session/haystack/conversation shapes, protocol notes in report/manifest, docs with exact command/result, and verified locomo retrieval-only smoke.","labels":["architecture","benchmarks","locomo","memory"],"dependencies":[{"issue_id":"klbr-wmz.8","depends_on_id":"klbr-wmz","type":"parent-child","created_at":"2026-06-26T20:54:09Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-wmz.5","title":"Upgrade episodic notes from transcript cards to source-grounded event cards","description":"klbr-core/src/pipeline.rs render_episode_card creates an episode note from timestamp, source refs, and truncated role lines. docs/memory-arch.md asks for source-grounded scene memory that preserves who, when, where, what changed, what was said, available attachments, and supporting raw refs. The current implementation is addressable, but still too transcript-shaped for temporal/update reasoning.","design":"Do not synthesize unsupported vivid details. The event card should summarize only what source turns or attachments support, with raw refs retained as the escape hatch.","acceptance_criteria":"Episode artifacts have a stable structured representation for time anchors, participants/entities, changes/decisions, salient quotes or snippets, attachments when present, and source refs; the markdown rendering remains human-editable; retrieval/context packets can expose the structured gist without losing provenance; tests cover an episode with multiple turns and verify source refs, session id resolution, and useful promptable chunks.","status":"closed","priority":2,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-26T17:53:44Z","created_by":"dawn","updated_at":"2026-06-26T23:36:56Z","started_at":"2026-06-26T23:33:59Z","closed_at":"2026-06-26T23:36:56Z","close_reason":"Episode notes now render as structured episode_event_card artifacts with event metadata, source-grounded timelines, stated facts/decisions, attachment markers, source refs, and promptable retrieval coverage; added core regression.","labels":["architecture","episodic","memory","provenance"],"dependencies":[{"issue_id":"klbr-wmz.5","depends_on_id":"klbr-wmz","type":"parent-child","created_at":"2026-06-26T20:53:44Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-wmz.4","title":"Materialize profile and procedural lanes as first-class notes","description":"The schema and MemoryGarden know about profile and procedural lanes, but the production pipeline observe_session path currently writes raw turns plus episodic notes/memories. Stable preferences, standing instructions, and workflows still depend mostly on tags or model-authored remember calls instead of a first-class markdown-note flow with source refs and update policy.","design":"Build on MemoryGarden and upsert_markdown_note rather than adding a new store. Treat legacy memory tags as routing hints, not the durable source of truth for profile/procedural knowledge.","acceptance_criteria":"There is an explicit writer/reflection path for profile_note and procedural_note artifacts; new notes include source refs, frontmatter, stable paths under profile/ or procedural/, and canonical refs/chunks; updates use supersession or versioning instead of silent overwrite; user-confirmation or policy gates are documented for stable profile changes; tests cover creating and updating one profile note and one procedural note from source turns.","status":"closed","priority":2,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-26T17:53:36Z","created_by":"dawn","updated_at":"2026-06-26T23:42:13Z","started_at":"2026-06-26T23:37:20Z","closed_at":"2026-06-26T23:42:13Z","close_reason":"Added write_memory_note reflection tool for source-grounded profile/procedural markdown notes, ref supersession updates, source/policy frontmatter, stable lane paths, and tests for profile/procedural create/update paths.","labels":["architecture","markdown","memory","profile"],"dependencies":[{"issue_id":"klbr-wmz.4","depends_on_id":"klbr-wmz","type":"parent-child","created_at":"2026-06-26T20:53:36Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0} +{"_type":"issue","id":"klbr-d76","title":"Design a separate folgezettel agentic memory eval","description":"LongMemEval passive QA is useful as best-effort passive recall, but it does not test the mk tool surface, titles-as-apis, sourced note writing, revisions, links, or trail-aware recall. Design a small agentic memory eval where the model must use mk_index, mk_search, mk_follow, mk_recall, mk_remember, mk_revise, mk_link, and mk_fleeting over a folgezettel tree, then measure sourced note quality, title usefulness, revision behavior, and later retrieval through trail_support.","acceptance_criteria":"A design doc or bench skeleton defines tasks, scoring, and traces for agentic folgezettel memory separately from LongMemEval passive QA; it includes at least note creation, revision/supersession, linking, search/follow/recall, and later answer-from-zettel scenarios; docs make clear this is not a replacement for passive LongMemEval recall.","status":"open","priority":3,"issue_type":"feature","owner":"90008@klbr.net","created_at":"2026-06-30T11:14:04Z","created_by":"dawn","updated_at":"2026-06-30T11:14:04Z","labels":["agentic","bench","folgezettel","memory"],"dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-u3t.4","title":"Integrate BGE-M3 sparse retrieval integration path","description":"Research and implement a second sparse retrieval channel from BGE-M3 sparse embeddings (if the embedder service supports exposing sparse weights) as a multilingual fallback to FTS BM25.","status":"closed","priority":3,"issue_type":"task","owner":"90008@klbr.net","created_at":"2026-06-27T17:26:43Z","created_by":"dawn","updated_at":"2026-06-27T17:37:28Z","closed_at":"2026-06-27T17:37:28Z","close_reason":"Implemented BGE-M3 sparse embeddings retrieval path via API fallback, customized SQLite LIKE-based sparse match scoring in memory.rs, and integrated into pipeline.rs. Verified all tests pass.","dependencies":[{"issue_id":"klbr-u3t.4","depends_on_id":"klbr-u3t","type":"parent-child","created_at":"2026-06-27T20:26:43Z","created_by":"dawn","metadata":"{}"}],"dependency_count":0,"dependent_count":1,"comment_count":0} {"_type":"issue","id":"klbr-mpy","title":"Improve tool-required router separability","description":"Fresh router artifact benchmarks/models/router/linear/out-router-linear-iter-15 reports test tool-required false-memory rate 0.6913 and tool-required recall 0.3087 after adding the metric/calibration surface. Improve training data, features, or model shape so tool-required queries stop looking like memory queries without regressing memory false-abstain.","status":"closed","priority":3,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-27T15:08:07Z","created_by":"dawn","updated_at":"2026-06-27T15:17:20Z","started_at":"2026-06-27T15:12:36Z","closed_at":"2026-06-27T15:17:20Z","close_reason":"Added tool-required training weighting for the linear router and regenerated out-router-linear-iter-16. Test tool→memory improved 0.6913→0.0940, tool-required recall 0.3087→0.9060, and memory false-abstain remained 0.0000 on test/holdout.","dependencies":[{"issue_id":"klbr-mpy","depends_on_id":"klbr-b4z","type":"blocks","created_at":"2026-06-27T18:08:16Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-b4z","title":"Regenerate router calibration artifacts","description":"After klbr-9yo, rerun the router bench with the calibrated tool-required metrics and commit fresh benchmarks/models/router model/report outputs so future sweeps track tool→memory and tool-required recall from generated artifacts.","status":"closed","priority":3,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-27T14:56:42Z","created_by":"dawn","updated_at":"2026-06-27T15:08:17Z","started_at":"2026-06-27T15:01:12Z","closed_at":"2026-06-27T15:08:17Z","close_reason":"Regenerated router linear artifact out-router-linear-iter-15 with tool-required metrics, updated benchmark helper default to the fresh model, and filed klbr-mpy for residual tool→memory quality work.","dependencies":[{"issue_id":"klbr-b4z","depends_on_id":"klbr-9yo","type":"blocks","created_at":"2026-06-27T17:56:51Z","created_by":"dawn","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0} -- 2.51.2