diff --git a/.beads/issues.jsonl b/.beads/issues.jsonl index 1632d28..ce7f70e 100644 --- a/.beads/issues.jsonl +++ b/.beads/issues.jsonl @@ -1,3 +1,4 @@ +{"_type":"issue","id":"klbr-54l","title":"Investigate LongMemEval QA regression after memory evidence changes","description":"Same-seed LongMemEval-S 25-sample QA dropped after the recent memory evidence/planner/session-first changes. Baseline run benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z scored 0.6800. Current run benchmarks/runs/longmemeval-s/klbr-full/2026-06-30_qa_sample25_seed4937553249516211047_current scored 0.4400 with the same sample seed 4937553249516211047 and local judge model. Session recall stayed similar, but answer-bearing/rendered evidence and packet/context budget dropped: answer_bearing_ref_in_context 0.80 -\u003e 0.36, answer_bearing_ref_selected 0.84 -\u003e 0.44, packet_tokens_total_mean 4412.64 -\u003e 1811.28. Op plans also changed heavily: old update_resolution=10, current update_resolution=0 and lookup=15. Regressions old-pass/current-fail: 6b168ec8, c14c00dd, b5ef892d, 46a3abf7, 720133ac, gpt4_385a5000, dad224aa. Fix old-fail/current-pass: gpt4_59149c77.","acceptance_criteria":"Identify whether the loss comes from structured model planning, packet selection/rendering budget, answer-bearing ref propagation, or reader prompting; add a focused regression/smoke that prevents the same same-seed sample from losing old passing rows without an explicit expected-metric update.","status":"open","priority":1,"issue_type":"bug","owner":"90008@klbr.net","created_at":"2026-06-30T01:00:03Z","created_by":"dawn","updated_at":"2026-06-30T01:00:03Z","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-u0q","title":"Implement session-first lexical-contained memory retrieval","description":"Build the next retrieval architecture from docs/long-term-memory-arch.md: stage-one session/event candidate generation, lexical containment, and language-agnostic operation planning experiments without adding english cue-word hacks.","design":"Keep EvidencePacket as the rank object. Use exact refs, dense episode/session cards, learned sparse or fts side channels, and graph expansion from strong seeds only. Keep QueryOp shape but replace cue-list policy with a structured classifier plus deterministic fallback.","acceptance_criteria":"A bench profile can retrieve from session/episode cards first and use chunk fts/sparse hits as packet enrichment; lexical-only policy paths are removed or contained behind candidate channels; traces expose candidate session recall and packet/rendered evidence metrics; multilingual fixtures cover non-English cue-free retrieval cases.","notes":"2026-06-29 partial: removed hardcoded english cue-list planner, english lane routing, english negation/entity-boundary packet filters, and wh/pronoun neighbor-expansion heuristics. OpPlan now comes from structured model JSON when an llm endpoint is configured, otherwise conservative lookup. Core and bench tests pass.\n2026-06-29 partial: landed initial session-first retrieval profile for klbr-full/dense-only/session-first profiles. Retrieval now groups archival candidates by session_id, prefers episodic/event-card anchors, adds chunk fts/sparse/dense hits as enrichment, serializes stage_one.session_candidates, and reports CandidateSessionRecall metrics in klbr-bench. fts-only now writes episode notes but skips embedded episode-memory rows so lexical ablations do not require the dense embedder at ingest. Verified core/bench tests, continuous-loop, and a one-row session-first/fts-only retrieval-only trace smoke.","status":"closed","priority":1,"issue_type":"feature","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T22:06:32Z","created_by":"dawn","updated_at":"2026-06-29T12:04:10Z","started_at":"2026-06-28T22:10:12Z","closed_at":"2026-06-29T12:04:10Z","close_reason":"Completed session/event-first lexical-contained slice: no english cue-list planner/filters, fts remains a candidate channel, trigram fts covers no-space multilingual exact-ish retrieval, traces expose session candidates/packet metrics, and docs clarify fts vs semantic multilingual retrieval.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-jwn","title":"Organize evolving memory architecture docs","description":"Make the memory architecture docs distinguish current truth, implementation status, proposals, research inputs, and archived rationale so future agents do not treat older reports as canonical.","acceptance_criteria":"Docs have a clear index and status taxonomy; current memory architecture direction points at docs/long-term-memory-arch.md; older research reports are marked as historical or supporting; AGENTS.md routes future agents through the doc index.","status":"closed","priority":1,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T22:01:20Z","created_by":"dawn","updated_at":"2026-06-28T22:07:09Z","started_at":"2026-06-28T22:01:23Z","closed_at":"2026-06-28T22:07:09Z","close_reason":"Completed docs index, current architecture review rewrite, status banners, AGENTS routing, and follow-up implementation issue klbr-u0q.","dependency_count":0,"dependent_count":0,"comment_count":0} {"_type":"issue","id":"klbr-h4l","title":"Cache bench embeddings locally and run 25-sample QA bench","description":"Make LongMemEval bench embeddings use a project-local gitignored sqlite cache so reruns do not pay for the same embeddings twice. Then run a 25-sample stratified QA bench, inspect failures, try bounded non-overfit tweaks if evidence supports them, and prepare a report packet if results remain weak or tuning would be overfit.","design":"Prefer a global project cache path such as benchmarks/cache/*.db over temp/run-local caches. Keep benchmark polling sparse: start the run, wait for artifacts or process completion, then inspect once.","acceptance_criteria":"Embedding cache lives under the project and is ignored by git; cache hits are reused across bench runs; a 25-sample bench artifact exists with failure analysis; any code changes are tested, committed, and pushed.","notes":"Implemented project-local embedding cache at benchmarks/cache/embeddings.db and wired LongMemEval bench LlmClient construction through it. Ran stratified 25-sample QA on seed 4937553249516211047: current-code run benchmarks/runs/longmemeval-s/klbr-full/2026-06-28_210327.693783Z scored official accuracy 0.6800. Report packet written to report_packet.md in that run dir. High-budget targeted reruns fixed only dd2973ad and a1cc6108; remaining failures point to fact/timeline synthesis rather than safe cue tuning.","status":"closed","priority":1,"issue_type":"task","assignee":"dawn","owner":"90008@klbr.net","created_at":"2026-06-28T20:08:32Z","created_by":"dawn","updated_at":"2026-06-28T21:21:27Z","started_at":"2026-06-28T20:08:35Z","closed_at":"2026-06-28T21:21:27Z","close_reason":"Implemented project-local dense/sparse embedding cache, ran 25-sample QA bench, inspected failures, tried bounded planner/high-budget experiments, and wrote report packet.","dependency_count":0,"dependent_count":0,"comment_count":0}