diff --git a/.agents/skills/beads/SKILL.md b/.agents/skills/beads/SKILL.md new file mode 100644 index 0000000..a5a3344 --- /dev/null +++ b/.agents/skills/beads/SKILL.md @@ -0,0 +1,80 @@ +--- +name: beads +description: Use when working in a repository that uses bd or Beads for durable project task tracking, issue dependencies, blocker management, multi-session handoff, or shared work memory. Trigger when the user asks to find ready work, claim or close tasks, create follow-up work, inspect blockers, recover project context, or choose between local planning and persistent project tracking. +--- + +# Beads + +Use Beads as the shared project task system. Local plans, scratch files, and personal memories are useful, but they are not the durable source of truth for project work. + +## First Step + +Run: + +```bash +bd prime +``` + +If that prints nothing, check whether the repository has an active Beads workspace: + +```bash +bd where +``` + +## Preferred Route + +Use the `bd` CLI when shell access is available. It is the most compact and direct Beads interface. + +## Core CLI Workflow + +1. Find work: + +```bash +bd ready +bd list --status=open +bd list --status=in_progress +``` + +2. Inspect before editing: + +```bash +bd show +``` + +3. Claim work atomically: + +```bash +bd update --claim +``` + +4. Create durable follow-up work when implementation reveals new tasks: + +```bash +bd create "Short title" --description="Why this exists and what needs to be done" --type=task --priority=2 +``` + +5. Close completed work: + +```bash +bd close --reason="Completed" +``` + +## What Belongs In Beads + +Use Beads for: + +- shared project tasks +- blockers and dependencies +- discovered follow-up work +- work that must survive thread reset, compaction, or handoff +- status that another person or agent should be able to resume + +Use agent-local planning tools only for the current turn's execution checklist. Do not treat them as shared project state. + +## Rules + +- Do not create markdown TODO files as the source of truth when Beads is available. +- Do not use `bd edit`; it opens an interactive editor. Use `bd update` flags instead. +- Prefer `--json` when parsing `bd` output programmatically. +- If hooks are installed, `bd prime` may already be injected. Run it manually when context is missing. +- Do not auto-close or mutate tasks unless the work is actually complete. diff --git a/.agents/skills/beads/agents/openai.yaml b/.agents/skills/beads/agents/openai.yaml new file mode 100644 index 0000000..09c3b8f --- /dev/null +++ b/.agents/skills/beads/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "Beads" + short_description: "Project task tracking with bd" + default_prompt: "Use $beads to inspect ready work and manage durable project tasks." diff --git a/AGENTS.md b/AGENTS.md index c0e1152..1b956a6 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -521,3 +521,25 @@ bd close # Complete work - NEVER say "ready to push when you are" - YOU must push - If push fails, resolve and retry until it succeeds + + +## Beads Issue Tracker + +Use Beads (`bd`) for durable task tracking in repositories that include it. Use the `beads` skill at `.agents/skills/beads/SKILL.md` (project install) or `~/.agents/skills/beads/SKILL.md` (global install) for Beads workflow guidance, then use the `bd` CLI for issue operations. + +### Quick Reference + +```bash +bd ready # Find available work +bd show # View issue details +bd update --claim # Claim work +bd close # Complete work +bd prime # Refresh Beads context +``` + +### Rules + +- Use `bd` for all task tracking; do not create markdown TODO lists. +- Run `bd prime` when Beads context is missing or stale. +- Keep persistent project memory in Beads via `bd remember`; do not create ad hoc memory files. + diff --git a/docs/evidence-selection.md b/docs/evidence-selection.md new file mode 100644 index 0000000..e520d42 --- /dev/null +++ b/docs/evidence-selection.md @@ -0,0 +1,404 @@ +# klbr evidence selection for ref-native memory + +## diagnosis + +the failure pattern you described is not a “retrieval missed the session” problem. it is a **selection granularity** problem: the pipeline is good enough to find the right session, but it sometimes hands the reader the *query-adjacent* chunk instead of the *answer-bearing* evidence packet. in other words, the current stack is already strong at coarse routing and weak at last-mile evidence assembly. that matches the behavior your benchmark note reports: `RecallAny@5 = 1.0` and `RecallAll@5 = 1.0` at session level, while end-to-end QA still fails on cases where the answer lives in a neighboring turn inside the retrieved session. your attached architecture also already frames `MemoryPipeline` as the benchmark-facing production facade and notes that the system is now ref-native enough that deterministic refs, promptable text, edges, and packets are first-class surfaces. fileciteturn0file0 + +the core architectural implication is pretty clean: **stop treating retrieved objects as final context units**. exact refs, FTS hits, dense hits, and graph hits should be treated as *seeds* for packet construction, not as the packet itself. long-term memory benchmarks make this distinction matter. LongMemEval explicitly decomposes memory design into indexing, retrieval, and reading, and reports that value granularity, key expansion, and time-aware query expansion all materially affect retrieval and QA. recent diagnostic work goes further and shows that end-to-end scores alone cannot tell you whether write-time compression lost evidence or retrieval-time assembly failed to surface it, which is exactly why a “right session, wrong chunk” failure can hide under perfect session recall. citeturn18view0turn19view0 + +there is a second smell here too: the current lexical helpers in `pipeline.rs` are clearly useful as scaffolding, but they are doing too much policy work. if `lexical_relevance`, `term_variants`, cue-list lane routing, and type bonuses are making ranking decisions, then klbr is effectively learning a retrieval policy from ad hoc english-ish heuristics. that is risky because SQLite FTS5 already gives you a much more principled lexical surface: phrase queries, prefix queries, `NEAR`, boolean operators, `bm25()`, built-in tokenizers, custom tokenizers, locale support, and synonym support. you do not need to invent your own little half-stemmer in rust unless you enjoy multilingual pain. citeturn5view0turn6view2turn17view1turn17view2 + +## recommended fusion architecture + +my recommendation is to make **evidence packets** the only rankable object after first-stage retrieval. not chunks, not sessions, not raw refs. + +a packet should look like this conceptually: + +```rust +struct EvidencePacket { + packet_id: String, + session_id: Option, + anchor_ref: String, + packet_kind: PacketKind, // exact_ref | turn_window | episode_bridge | graph_bridge + refs: Vec, // canonical refs included in packet + bodies: Vec, // promptable bodies, already ordered + signals: PacketSignals, // per-source ranks/scores/counts + lanes: Vec, + estimated_tokens: usize, +} +``` + +the retrieval flow should then become: + +```text +exact refs / fts / dense turn / dense episode / graph seeds + -> packet builder + -> packet fusion + -> packet rerank + -> budgeted packet packer + -> xml memory_packets +``` + +that sounds slightly more elaborate than the current merge path, but it is actually the simplest way to preserve the invariants you care about: exact refs stay exact, provenance stays attached to canonical refs, packets remain evidence-not-instructions, and you still do not have to dump whole sessions. it also lines up with existing retrieval research. LongMemEval’s own analysis says granularity matters and that decomposed units plus better keys and time-aware query expansion improve retrieval and downstream QA. late-interaction work such as ColBERT shows why fine-grained matching matters in the first place: pooled representations often blur the local signal that actually answers the query. recent conversational-memory retrieval work makes the same point even more directly: turn-level late interaction beats pooled session representations, and lexical–dense fusion can help, but what helps depends on the benchmark regime. citeturn18view0turn15academia0turn13view0turn14academia0 + +the concrete packet types i would use are these: + +- **exact-ref packet**: if the query contains `[d...]`, `[m...]`, or any canonical/alias ref, resolve it first and always include it. one-hop graph expansion is allowed here, but only as packet enrichment, not as a competing candidate list. +- **turn-window packet**: if FTS or dense retrieval hits a turn chunk, anchor the packet on that chunk and pull a small local window, usually `[-1, +1]` turns relative to the matched turn. +- **episode-bridge packet**: if a session-level episode card is hit, include the episode gist plus the best one or two supporting turn chunks from that session. +- **graph-bridge packet**: if a linked note or derived memory is hit, include the anchor plus a capped number of directly linked supporting refs. + +that design gives you the thing your failures wanted: when FTS lands on “i went to a play…” the packet builder turns that into “matched turn + next turn + tiny episode gist,” which is exactly where “the glass menagerie” tends to live. hierarchical retrieval work points in the same direction: retrieve small child units, then reconstruct slightly larger parent context for the reader. HRR, H-RAG, and HiKEY all report better retrieval or QA behavior when the system separates precise retrieval units from larger generation-time context assembly, instead of making the reader guess from a single isolated chunk or an undifferentiated whole document. citeturn8academia0turn8academia1turn8academia2 + +for fusion, i would split the problem in two. + +**first-stage fusion** should be rank-based and conservative. the channels you have are heterogeneous: exact refs are effectively deterministic, BM25/FTS scores live on one scale, dense scores on another, and graph expansion is mostly structural. because of that, raw score fusion is brittle unless you calibrate it carefully. the hybrid retrieval literature is pretty explicit that fusion behavior depends on how you normalize and tune scores; Bruch, Gai, and Ingber show that convex score combination can outperform RRF when calibrated, but also that RRF has different robustness properties and that fusion behavior is not automatic. with the current amount of supervision you have, a weighted rank fusion at the **packet** level is the safer first move. citeturn10academia0turn13view0 + +**second-stage fusion** should happen after packet construction, not before. once you have packets, you can safely use query-normalized features and a reranker or learned blend because the object you are scoring is finally the thing you actually care about: a budgeted evidence set that might answer the question. + +my default scoring recipe would be: + +```rust +packet_rrf = + w_fts * rrf(rank_fts) + + w_dense * rrf(rank_dense) + + w_graph * rrf(rank_graph) + + w_exact * exact_bonus + + w_agree * agreement_bonus + + w_local * local_answerability_bonus; +``` + +where: + +- `exact_bonus` is huge and effectively pins explicit refs. +- `agreement_bonus` fires when two or more channels support the same packet. +- `local_answerability_bonus` is small and only rewards packets that include more than one local ref around an anchor, so “matched chunk + neighbor” beats “matched chunk alone.” + +because recent conversational-memory work found that dense late interaction helps most on multi-hop and temporal questions while BM25 stays very strong on lexical or adversarial ones, this kind of packet-level fusion gives you the right division of labor without letting one brittle signal erase the others. it is also very compatible with your constraint set: sqlite stays, canonical refs stay, benchmark and runtime can share the same planner, and nothing here depends on english-only lexical hacks. citeturn13view0turn18view0 + +## replacing same-session dedupe + +the current same-session dedupe should be removed from `merge_candidates` and replaced with **same-session grouping plus capped packet emission**. + +the current behavior loses because “same session” is being used as a proxy for “duplicate evidence.” those are not the same thing. two seeds from the same session may be complementary, not redundant: one can be the question-adjacent setup and the other the answer-bearing neighbor or episode card. your failure examples are exactly this pattern. + +the replacement rule should be: + +```text +dedupe by: + canonical ref identity + identical window hash + identical packet hash + +never dedupe by: + session id alone +``` + +more concretely: + +```rust +fn merge_candidates_v2( + seeds: Vec, + cfg: MergeConfig, +) -> Vec { + let mut by_session: HashMap, Vec> = bucket_by_session(seeds); + let mut packets = Vec::new(); + + for (session_id, session_seeds) in by_session.iter_mut() { + session_seeds.sort_by(source_then_rank); + + let anchors = cluster_by_anchor(session_seeds, cfg.max_turn_gap); + + for cluster in anchors { + let packet = build_packet(cluster, cfg); + if !is_duplicate_packet(&packet, &packets) { + packets.push(packet); + } + } + + let extras = emit_non_overlapping_packets(session_seeds, cfg.max_packets_per_session); + for packet in extras { + if !is_duplicate_packet(&packet, &packets) { + packets.push(packet); + } + } + } + + rank_packets(packets, cfg) +} +``` + +i would define `cluster_by_anchor` this way: + +- if two seeds resolve to the same canonical ref, merge them. +- if two turn-chunk seeds are in the same session and within `±1` logical turn step, merge them. +- if one seed is a turn chunk and another is the session episode card, merge the episode gist into the packet instead of choosing one. +- if seeds are in the same session but non-overlapping and far apart, allow more than one packet from that session, up to a cap like `max_packets_per_session = 2`. + +that gives you a very specific replacement for the lossy session dedupe: + +```rust +struct MergeConfig { + max_turn_gap: usize, // usually 1 + max_packets_per_session: usize, // usually 2 + max_graph_neighbors: usize, // usually 2 or 3 + episode_gist_tokens: usize, // usually 60..120 +} +``` + +the round-robin diversity guard should happen **after** packet construction, not before. you still want to stop one session from flooding the final context, but the right mechanism is “at most two packets per session in the final ranked set,” not “first packet wins, all other same-session evidence dies.” + +this is also where you should fold graph expansion. graph hits should no longer arrive as free-floating candidates that compete with FTS and dense at the same layer. they should enrich a packet that already has a good seed, unless the query contains an explicit ref. that keeps the graph useful without letting it act like glitter thrown over the ranking list. + +## lexical scoring policy + +my position is: **keep lexical retrieval, remove lexical heuristics from the scoring-critical path, and constrain lane routing.** + +specifically: + +- keep SQLite FTS5 as a first-stage candidate generator; +- stop using `lexical_relevance`, `lexical_entry_score`, `lexical_terms`, and `term_variants` as production ranking policy; +- stop using english cue lists as the main lane boundary; +- if you keep those helpers at all, keep them only behind a debug profile or as trace features. + +SQLite FTS5 is already good enough to be your lexical substrate. it supports phrase search, prefix search, `NEAR`, boolean combinations, `bm25()`, custom tokenizers, locale-aware tokenization hooks, and synonym support. its `unicode61` tokenizer is the default and handles unicode token classes; its trigram tokenizer supports general substring matching and can backstop entity fragments, mixed scripts, or languages where word boundaries are not trivial. by contrast, the porter tokenizer is explicitly english-oriented, and SQLite’s own docs warn that porter stemming is designed for english and may or may not help in other languages. that makes porter-like custom plural stripping in rust exactly the wrong direction for klbr’s production path. citeturn5view0turn6view0turn6view1turn6view2turn17view1 + +so the lexical strategy i would actually ship is: + +```text +lexical retrieval: + primary index: FTS5 unicode61 + bm25 + fallback index: FTS5 trigram for entity/code/subword recall + query form: parameterized MATCH built from escaped user text + optional features: quoted phrase probes, NEAR probes, alias probes from ref_aliases + locale/synonym support: tokenizer-level, not rust string hacks +``` + +in practice that means: + +```rust +fn search_lexical(query: &str, lanes: &[MemoryLane], limit: usize) -> Result> { + let q = build_fts_query(query)?; // escaping, phrase extraction, aliases + let hits_primary = search_unicode61(q, lanes, limit * 4)?; + let hits_trigram = search_trigram(q, lanes, limit * 2)?; + Ok(fuse_fts_hits(hits_primary, hits_trigram, limit)) +} +``` + +`build_fts_query` should be boring and safe. escaping, maybe a quoted-phrase probe for longer noun phrases, maybe alias expansion from `ref_aliases`, maybe a second `NEAR` probe for multi-term queries. what it should **not** do is homebrew stemming, english plural rules, or hand-coded intent classification. + +there is also a clean future path here because you are already using BGE-M3. BGE-M3 is explicitly designed to support **dense retrieval, sparse retrieval, and multi-vector retrieval** in more than 100 languages. so if you later want a multilingual lexical-ish signal that is less brittle than BM25 on paraphrase and less hacky than your current helpers, the principled add-on is “BGE-M3 sparse as another retrieval channel,” not “more english regexes in `pipeline.rs`.” i still would not make that the first fix; packetization and dedupe are higher priority. but it is a much better future direction than growing the cue-list forest. citeturn4academia0 + +lane routing should get the same treatment. exact refs should still override everything. after that, lane priors should come from **metadata and ref type**, not lexical vibes. a turn chunk from the current or recent context should bias toward `Live`/`Episodic`. a stable note or profile ref should bias toward `Semantic`/`Profile`. if you want a learned router later, small and explicit is fine. but “contains ‘when’ -> episodic” should not remain the main boundary for a system that is otherwise this careful about refs and provenance. + +## reranking and neighbor expansion + +the reranker should rank **packets**, not chunks and not whole sessions. + +chunk-level reranking is what got you into this mess: it can happily pick the semantically adjacent mention and miss the local answer turn. whole-session reranking is the opposite failure mode: too much filler, not enough answer density. reranking a packet that contains “matched chunk + local window + tiny episode gist + source refs” is the sweet spot because it keeps local answer-bearing detail while preserving enough context for the model to understand why the chunk matters. hierarchical retrieval papers keep rediscovering this exact tradeoff: small retrieval units help precision, larger reconstructed context helps the reader. citeturn8academia0turn8academia1turn8academia2turn9academia0 + +my recommendation for neighbor expansion is also simple: + +```text +if seed is exact ref: + include anchor + include one-hop linked refs, capped by relation + include local siblings if anchor is a turn chunk + +if seed is turn chunk: + include matched chunk + include previous turn + include next turn + include tiny episode gist if session exists + +if seed is episode card: + include episode gist + include top 2 supporting chunks from same session + +if seed is graph note: + include anchor + include direct supporting refs only +``` + +that `[−1, +1]` window should be the default because your observed failures are literally “right turn plus next turn names the answer.” it is also language-agnostic. you do not need to solve answer typing up front to get most of the win. later, if you want a slightly smarter policy, you can make the window adaptive: + +```rust +fn expand_turn_window(anchor: TurnRef, query: &str, cfg: WindowConfig) -> Vec { + let mut refs = vec![anchor.ref_id.clone()]; + refs.extend(prev_turn(anchor.session_id, anchor.turn_ord, 1)); + refs.extend(next_turn(anchor.session_id, anchor.turn_ord, 1)); + + if looks_incomplete(anchor.body) || query_is_wh_like(query) { + refs.extend(best_extra_neighbor(anchor.session_id, anchor.turn_ord, query)); + } + + refs +} +``` + +but i would treat `looks_incomplete` and `query_is_wh_like` as optional later passes, because the fixed `±1` window solves the class of failures you have right now without another heuristic thicket. + +on the reranker itself, i would be careful. BERT-style passage rerankers are powerful in standard IR; passage reranking with BERT was a huge win on MS MARCO-style passage retrieval. but recent conversational-memory retrieval work found that an **off-the-shelf web-search cross-encoder reranker hurt** after lexical–dense fusion on conversational-memory benchmarks, and specifically by a non-trivial amount in the tested setup. so my recommendation is not “remove reranking,” it is “demote reranking from judge to tie-breaker until it proves itself on your fixtures.” citeturn9academia0turn13view0 + +that means: + +- rerank only the top `M` packets, not raw seeds; +- never let reranker score alone suppress an exact-ref packet; +- do not use reranker thresholds that drop all packets from a session if packet fusion already says the session is strong; +- log both pre-rerank and post-rerank packet order so you can see whether rerank helped or killed answer-bearing packets. + +a practical packet rerank input template would be: + +```text +query: +{query} + +packet: +[session={session_id} kind={packet_kind} source_signals={signals}] +{packet_body} + +task: +rank how likely this packet contains the minimal evidence needed to answer the query. +``` + +note the wording there. you do not want a generic “relevance” reranker. you want an **answer-bearing evidence** reranker. + +## evaluation and regression plan + +the key permanent guardrail is to stop treating `RecallAny@5` as enough. LongMemEval gives you gold evidence and task categories, and WhenLoss gives you the right diagnostic framing for write-side vs retrieval-side degradation. so the eval stack should include both public smoke runs and deterministic local fixtures that directly target the failure mode you found. citeturn18view0turn19view0 + +the metrics i would add immediately are: + +```text +answer_bearing_ref_in_context +gold_session_in_context +gold_session_plus_neighbor_in_context +packet_recall_any@k +packet_recall_all@k +same_session_dedupe_suppression_count +fts_only / dense_only / graph_only / full deltas +packet_tokens_mean +false_recall_rate +``` + +`answer_bearing_ref_in_context` should be the star. if the right session is present but the answer-bearing ref is absent, that should show up as a first-class regression, not as a vague QA miss. + +the deterministic fixture set should include at least these six cases: + +- **same-session next-turn answer**. FTS hits the setup turn, answer is named in the next turn. expected behavior: turn-window packet brings the answer-bearing neighbor into context. +- **same-session previous-turn answer**. dense or FTS hits a follow-up turn, but the answer was stated one turn earlier. expected behavior: backward window works too. +- **same-session dual-signal merge**. FTS hits one turn chunk, dense hits the episode card or another turn in the same session. expected behavior: signals merge into one packet or two non-overlapping packets, but none are dropped just because the session id matches. +- **explicit-ref plus graph edge**. query contains a direct ref. expected behavior: exact ref is pinned into context, graph neighbors enrich but do not replace it. +- **multilingual lexical stress**. same fact expressed with inflection, diacritics, substring, or non-space-separated script. expected behavior: unicode/trigram FTS plus dense retrieval succeeds without english stemming hacks. +- **false recall adversary**. multiple sessions share surface tokens, but only one contains the answer. expected behavior: packet fusion plus rerank prefers the answer-bearing packet, and neighbor expansion does not explode context. + +if you want two more, i would add **supersession/update** and **archived-provenance visibility**, because those tend to rot quietly in memory systems and then show up as “why did it recall the stale thing.” LongMemEval explicitly includes knowledge updates and abstention, so it is worth keeping those visible in your local fixture layer too. citeturn18view0 + +a minimal regression harness shape would be: + +```rust +#[test] +fn next_turn_answer_is_not_lost_after_fts_hit() { + let trace = run_fixture("same_session_next_turn_answer"); + assert!(trace.metrics.gold_session_in_context); + assert!(trace.metrics.answer_bearing_ref_in_context); + assert_eq!(trace.answer, "target"); +} +``` + +for LongMemEval smoke, i would keep one command that exercises the real production pipeline and is small enough to run constantly: + +```bash +rtk cargo run -p klbr-bench -- run \ + --suite longmemeval-s \ + --data benchmarks/inputs/datasets/longmemeval_s_cleaned_subset30.json \ + --profile klbr-full \ + --top-k 5 \ + --budget-read 5000 \ + --graph-depth 1 \ + --limit 5 \ + --out /tmp/klbr-lme-s-smoke +``` + +and then add two targeted one-question regressions for the exact known failures once `--question-id` is available in your normal loop. the point is not just to keep a public benchmark number green; it is to make “right session, wrong chunk” impossible to reintroduce without tripping a test. + +i would also wire in a WhenLoss-style diagnostic mode for klbr itself: + +```text +oe -> oracle gold turns only +csm -> complete stored context, budgeted deterministically +rm -> actual retrieved packets +``` + +then track: + +```text +delta_write = score(oe) - score(csm) +delta_retr = score(csm) - score(rm) +``` + +that tells you whether a new regression came from write-time memory construction or read-time evidence selection, which matters a lot once packetization, reflections, and compaction are all in play. citeturn19view0 + +## implementation sketch and rollout order + +this is the rollout order i would actually do, because it is the shortest path to fixing the observed failures without turning the codebase into soup. + +first, extract a shared `EvidencePlanner` and make both `MemoryPipeline::retrieve_evidence` and runtime passive recall call it. your attached report already notes that runtime passive recall in `agent.rs` still has legacy retrieval paths, so this refactor is not optional if you want benchmark wins to survive contact with production. fileciteturn0file0 + +second, replace `merge_candidates` with packet grouping and per-session caps. do **not** touch embeddings, schema, or graph modeling first. the failure is already explainable without any of that. + +third, delete the hand-written lexical score from the production ranking path and replace it with FTS-native BM25 plus packet fusion. if you want to keep the old helpers, keep them behind a `debug-lexical-heuristics` profile so you can compare traces. + +fourth, rerank packets, not chunks, and make the reranker non-authoritarian until it proves itself on the deterministic fixtures. + +this is the skeleton i would start from: + +```rust +async fn retrieve_evidence(query: BenchQuery, budget: ContextBudget) -> Result { + memory.sync_reference_indexes()?; + + let exact = resolve_exact_refs(&query)?; + let fts_turns = search_fts_turn_chunks(&query, budget.seed_k)?; + let fts_notes = search_fts_notes(&query, budget.seed_k / 2)?; + let dense_turns = search_dense_turns(&query, budget.seed_k).await?; + let dense_episodes = search_dense_episode_cards(&query, budget.seed_k / 2).await?; + + let packets = build_packets( + exact, + fts_turns, + fts_notes, + dense_turns, + dense_episodes, + PacketConfig { + turn_window: 1, + max_packets_per_session: 2, + max_graph_neighbors: 2, + episode_gist_tokens: 96, + }, + )?; + + let mut ranked = fuse_packets_weighted_rrf(packets); + ranked = rerank_top_packets(&query, ranked, budget.rerank_k).await?; + ranked = budget_packets(ranked, budget.max_tokens); + + Ok(trace(query, ranked)) +} +``` + +and this is the one line i would want taped to the monitor while implementing it: + +```text +seeds are not context. packets are context. +``` + +## risks and failure modes + +the biggest multilingual risk is obvious: if you replace the current english-ish heuristics with nothing, but you never upgrade lexical retrieval to use FTS5’s proper tokenizer features, you can still fail badly on non-english or non-space-delimited text. `unicode61` is a much better default than hand-rolled rust token splitting, but it is not magic for every language; trigram helps for substring/entity recall, and SQLite’s locale-aware custom tokenizers exist for a reason. porter-style stemming should stay out of the critical path unless the locale is explicitly english. citeturn6view0turn6view1turn17view1 + +the biggest false-recall risk is packet expansion. `±1` local windows are cheap and usually right for conversational QA, but they can still pull in distractor turns. that is why neighbor expansion should stay **small, local, and packet-scoped**, and why graph expansion should enrich packets rather than spray free-floating candidates into the ranked list. the mitigation is not “never expand”; it is “expand predictably, cap aggressively, and score the packet as a packet.” hierarchical retrieval literature is pretty consistent that context reconstruction helps, but uncontrolled parent expansion hurts token efficiency and can dilute answer-bearing evidence. citeturn8academia0turn8academia1turn8academia2 + +the biggest ranking risk is pretending that raw channel scores are commensurate when they are not. BM25, cosine distance, binary exactness, and graph support are different species. that is why first-stage rank fusion is the safer default here, and why a learned or tuned score combiner belongs after packet construction, not before. the hybrid retrieval literature shows that score fusion can beat RRF when calibrated, but it also shows that fusion behavior is sensitive to normalization and setup. klbr should not bet correctness on score comparability it does not yet have. citeturn10academia0turn7academia1 + +the final risk is organizational, not algorithmic: if you fix this only in `klbr-bench`, you will get a beautiful benchmark curve and a runtime that still uses the old passive-recall path. your own architecture notes already warn that benchmark behavior and runtime behavior are not yet identical. so the durable fix is a shared planner module, shared packet type, shared packet renderer, and shared trace schema across bench and runtime. otherwise the next person will “optimize” one path and quietly resurrect `klbr-wmz.10` in the other. fileciteturn0file0 + +the short version is: keep sqlite, keep refs, keep exact resolution, keep evidence packets, and stop letting same-session dedupe or english-ish lexical scaffolding act like truth. the next hard problem in klbr is not more memory storage. it is **packetizing the right evidence and ranking that packet like it might actually answer the question**. \ No newline at end of file