# plan.md — M3 (OpenAI server) execution plan > Continuation doc for a fresh session. Covers: current state, the M2.6 > prefix cache that landed, the M3.3 chat-template research that's done, > and the detailed next steps for M3.2 onward. **Read ROADMAP.md §M3 first** > for the milestone spec; this file is the execution-level plan. ## Where things stand (HEAD c3148d1; M3 uncommitted) M1 (real prefill, 749 tok/s) DONE. M2 (serving engine: API split, sampling, stop conditions, detokenizer, multi-slot scheduler) DONE — last M2 commit `709da80`. M2.6 (prefix cache) DONE. **M3 (OpenAI server) DONE (2026-07-20)** — M3.1 skeleton, M3.3 chat-template renderer, M3.2 /v1/chat/completions, M3.4 tool-call parsing, M3.5 lifecycle (disconnect cancel + per-request log), M3.6 smoke (14/14). All uncommitted; awaiting commit decision. Next: M4 (long-context perf) or the remaining MVP done-criterion #7 (Docker image). Uncommitted work this session: - M3.3: `src/qxmx_chat_template.{h,cpp}` (ChatML renderer), `tools/marker_probe.cpp`, `tests/chat_template_test.cpp` + `tests/tpl_crosscheck.py` (byte- + token-id- exact gate vs jinja2 + llama-tokenize, all 4 fixtures green). - M3.2: `src/qxmx_json.h` (extracted jval parser), `src/qxmx_request.h` (OpenAI request parser), `tools/qxmx_serve.cpp` rewritten (driver-thread model, SSE + JSON + usage, prefix cache, 400 validation, cancellation hook), `scripts/smoke.sh` (10/10 green). `meson.build` wires the new sources. All gates green: `qxmx_diff` 0.108, `chat_template_test` + `tpl_crosscheck.py` (byte- and id-exact), `prefix_cache_engine_test`, `smoke.sh` 10/10, engine binaries build clean. Known pre-existing (NOT M3.2): `scheduler_test` solo- vs-sequential last-token divergence (the M2.5 isolation gate still passes). --- ## M3.1 — DONE (qxmx_serve skeleton) Vendored cpp-httplib v0.18.3 into `third_party/httplib.h`. New `tools/qxmx_serve.cpp` + `qxmx_serve` meson target (SYCL, `engine_sources_naive`). Flags: `--host` (127.0.0.1), `--port` (8080), `--slots` (2), `--ctx-per-slot` (0=default→64K), `--ctk` (q8_0). Endpoints live: `GET /health` → `ok`; `GET /v1/models` → real model id (reads `general.name` from GGUF = `"Bonsai-27B"`, NOT guessed); `POST /v1/chat/completions` → 501 stub (M3.2 fills it). Per-request logger. **Bug fixed in 5b65a96:** `engine::on_open_slot` allocates `max_ctx` of KV *per slot* (`qxmx_gpu.cpp:1126`). The M3.1 skeleton multiplied by `--slots`, over-allocating. Fix: `max_ctx = ctx_per_slot` (each slot gets exactly its budget). Default ctx bumped 16K→64K (2×64K = 11.9 GB on 32 GB B70, comfortable). --- ## M2.6 — prefix cache (LANDED core: M2.6.1, M2.6.2, M2.6.3) **Full design record: memory `qxmx_m2_6_prefix_cache_design.md`.** Read it before touching the cache. Key points: ### What landed - `src/qxmx_prefix_cache.{h,cpp}` — pure-host cache module. Block size 512 (= `QXMX_CHUNK`). Hash = parent_hash mixed with FNV-1a over 512 token ids (chained, 64-bit, **verify token ids on hit — collisions are corruption**). Two stores: KV blocks (host-RAM quantized K+V copy, 16 attn layers) + SSM checkpoints (host-RAM 48 ssm_state + 48 ssm_conv, ~149 MB each, at block boundaries). Two LRU lists (KV leaf-first TODO — currently evict-anywhere, correct but suboptimal; SSM evict-anywhere). Budget 16 GB default each, evict-while-over-budget in `kv_put`/`ssm_put`. - `engine::prefill_cached(slot, ids, n, cache)` on `engine_i` (default no-op pass-through to `prefill` for the CPU ref). GPU impl does RESTORE (walk blocks, DMA host→device KV+SSM on hit, jump `n_tokens`) + PREFILL remainder with INLINE snapshotting. - Tests: `prefix_cache_test` (pure-host gate) + `prefix_cache_engine_test` (in-engine gate, SYCL). Both green. ### The critical gotcha (cost ~1h, MUST NOT relearn) **SSM state is mutable (updated in-place by `run_chunk_b`).** A checkpoint AT boundary `(b+1)*512` must be captured the MOMENT that boundary is crossed. Snapshotting all blocks at prefill-end captures the state at position `n` (it's advanced through all of P), corrupting every checkpoint. Symptom: cached vs fresh prefill diverged WILDLY (not drift — completely wrong state: "Paris is a city in France" vs "I'm not sure what you mean by Paris"). Fix: drive `prefill_step` one call at a time, snapshot after each call that leaves `i` on a block boundary. The chunk granularity (PC_BLOCK=512, chunk=512 default) means the chunk completing a boundary leaves `i == (b+1)*512` exactly; the tail (<16) only runs after all full blocks. KV blocks are immutable once written (fixed offset) but snapshotted inline alongside their boundary's SSM checkpoint to share one `q.wait()`. ### -ctk interaction (verified) `-ctk` affects ONLY the KV block size (K row: fp16=2048 B, q8_0/fp8=1088 B; V row always 640 B — TurboQuant 4-bit, `-ctk`-independent). SSM checkpoint is ALWAYS ~149 MB (48 layers × fp32 state+conv, never quantized — `malloc_shared` at `qxmx_gpu.cpp:1120`). The cache's `init(n_attn, n_ssm, k_row, v_row, ssm_state, ssm_conv)` takes the engine's actual row strides, so per-block + per-ckpt sizes are computed from whatever `-ctk` produced. 16 GB budget ≈ 104 SSM checkpoints regardless of `-ctk`; 780 (fp16) or 1213 (q8_0/fp8) KV blocks. ### Remaining M2.6 work (DEFER to after M3 core, or as noted) - **M2.6.4** — real KV leaf-first eviction (parent→child hash map) + ref-count wired into the restore path (`kv_ref` on hit, `kv_deref` on slot close). Currently evict-anywhere; correct (never evicts a ref>0 block) but suboptimal for prefix-chain preservation. NOT gate-bearing. - **M2.6.5** — system+tools prompt boundary checkpoint placement (Marconi pattern, for cross-request sharing across maki subagent tasks). Needs M3.3's chat-template boundary detection first. Every-boundary placement already covers the session-resume pattern (dominant case). --- ## M3.3 — chat template: DONE (renderer landed 2026-07-19; gated) **Read memory `qxmx_m3_3_chat_template.md` for the full record (corrected marker facts + gotchas).** The 7764-byte Jinja `tokenizer.chat_template` (Qwen3-style ChatML) is hardcoded in `src/qxmx_chat_template.{h,cpp}` — a full Jinja interpreter is post-MVP. Gate: `tests/tpl_crosscheck.py` renders 4 fixtures with jinja2 + tokenizes with `llama-tokenize`; all match BYTE-EXACT on text AND token-id-for-token-id vs llama.cpp. The `jval` JSON parser/serializer used by the renderer was extracted into `src/qxmx_json.h` (shared header) for M3.2's request parser. Original research notes (some marker descriptions were WRONG — see the memory file's "gotcha #1") retained below for history: - **Basic turn:** `<|im_start|>{role}\n{content}<|im_end|>\n` - **Generation prompt:** `<|im_start|>assistant\n` + thinking prelude (`enable_thinking` controls whether the think-open marker is emitted). - **Tool definitions (when `request.tools` non-empty):** a synthetic system message is prepended with `# Tools` header, `` block (one JSON per tool, `tojson`), a format example, and an `` reminder. The model is told to wrap function calls in a tool-call XML wrapper tag (~10/11-char open/close tags). - **Tool call (model GENERATES, M3.4 parses):** tool-call-open, then ``, `` blocks, ``, tool-call-close. Multiple `` blocks for parallel calls. - **Tool result (`message.role == "tool"`):** wrapped as a user turn with result-open/result-close markers (~6/7-char). `<|im_start|>user` only before the FIRST tool message in a run. - **multi_step_tool / last_query_index:** edge case for multi-turn tool + thinking; DEFER for MVP (treat `preserve_thinking` as false). - **Special token ids:** bos=248044, eos=248046 (eos IS `<|im_end|>`), pad=248044, add_bos=0 (encode adds NO specials by default). `<|im_start|>` and `<|im_end|>` are CONTROL tokens (likely single-token). The tool/function/parameter/thinking markers are MULTI-TOKEN text strings, parsed at the text layer in M3.4, not via token ids. **ONE THING STILL TO VERIFY (the session got cut short here):** the exact tokenization of `<|im_start|>` / `<|im_end|>` (confirm single-token) and the multi-token breakdown of the tool markers. **Redo the tokenizer probe CLEANLY using the `write` tool, NOT a heredoc** — a heredoc + the template's literal markers blew up the shell formatting in this session. The probe: load the tokenizer, call `find_by_text` for each marker, print ids. Build the marker strings at runtime from split parts (e.g. `std::string(""`) to avoid putting the literal markers in the tool-call payload. --- ## Next steps (detailed) ### M3.3 — chat template renderer (DO FIRST, before M3.2) M3.2 needs messages→tokens, which needs the chat template. Per ROADMAP §M3, hardcode the template (a full Jinja interpreter is post-MVP). **Plan:** 1. Redo the special-token-id probe cleanly (see above) — confirms `<|im_start|>`/`<|im_end|>` are single tokens. 2. New `src/qxmx_chat_template.{h,cpp}` (pure host, no SYCL): a C++ `render(messages, tools, add_generation_prompt, enable_thinking) -> std::string` that emits the exact bytes the Jinja would. Parse the OpenAI request's messages (role/content, optional reasoning_content, optional tool_calls for assistant turns, role=="tool" for results) + tools array (name/description/parameters → JSON). 3. Use a small JSON writer (shared with M3.2's request parsing — start a `src/qxmx_json.{h,cpp}` or inline a minimal one; tools need `tojson` with proper escaping). 4. Validate: render a few fixtures (plain chat, system+tools, multi-turn with tool results), tokenize, eyeball the token stream. Cross-check against the special-token probe. **Scope discipline:** MVP handles role ∈ {system, user, assistant, tool}, content as string (defer multipart content / vision), no preserve_thinking (drop prior reasoning_content), enable_thinking=false default (skip the reasoning prelude — keeps outputs clean for tool use). ### M3.2 — `/v1/chat/completions` endpoint (SSE + JSON + usage) — DONE (2026-07-20) **Shipped:** - `src/qxmx_json.h` — extracted the `jval` recursive JSON parser/serializer from M3.3 into a shared header (header-only); serializer emits HF default spacing (", "/": "). Used by the chat template (tool tojson) AND the request parser. - `src/qxmx_request.h` — `parse_chat_request(body, eos_id, ctx_per_slot, &req)` fills `chat_request{messages, tools, sp, stopp, stream, max_tokens, seed, has_seed, enable_thinking, error}`. Fields: messages (required), tools, stream, temperature, top_p, top_k, max_tokens (or max_completion_tokens), stop (string|array), seed, enable_thinking. Content may be a string OR array of `{type:"text",text:...}` (text parts concatenated, non-text dropped). Assistant tool_calls parsed into `chat_tool_call{name, arguments_json}` (arguments normalized to a string). Defaults: temperature=0 (greedy), max_tokens=256, enable_thinking=false. - `tools/qxmx_serve.cpp` rewritten — the real endpoint (replaced the M3.1 501 stub). See the **Driver thread** design below. - `scripts/smoke.sh` — 10-check smoke gate (models, non-stream, stream, 400s, concurrency). All green. **Driver-thread design (the crux — engine is single-threaded):** A `server_core` struct owns the engine + tokenizer + prefix cache + a fixed pool of N pre-opened slots (free-stack) + a request queue (mutex+condvar) + ONE driver thread. cpp-httplib runs one worker per HTTP request; HTTP workers `core.submit(dreq)` and block draining a per-request `chunk_channel` (mutex+condvar+deque). The driver loop: count active slots, pull pending requests (block on condvar ONLY when fully idle — critical: must not block while a slot is mid-request, or the slot never advances), assign to free slots, drive every active slot ONE step, reap finished slots (push finish sentinel, reset slot, return to free pool). Each `step` mirrors `scheduler::step` (M2.5): prefill phase -> `prefill_cached` (one blocking call: restore + remainder + inline snapshot) then sample first token; decode phase -> stream prev token via `incremental_decoder` + push_text to channel, stop-check, forward+sample next. **Key decisions / deviations:** - **Prefill is NOT chunk-interleaved across slots** (regression from M2.5's scheduler): `prefill_cached` is monolithic (restore+remainder+snapshot in one blocking call), so a long prefill on slot A stalls slot B's decode. The prefix cache keeps turn N+1 prefills short (just the delta) — the common path after turn 1 — so this is acceptable for MVP. Chunk- interleaving-during-prefill is post-MVP (would need `restore_cached` split out of `prefill_cached` + stepwise snapshotting). ROADMAP #5 (prefix reuse) is the hard gate and IS met. - **Prefix cache** is a single global `prefix_cache` shared across all slots; `prefill_cached(slot, ids, n, &cache)` does restore + inline snapshotting. Verified: identical >512-tok prompt sent twice -> 2nd request ~2x faster (restored 1/512 block, bit-identical output). - **Streaming lifetime**: `set_chunked_content_provider` with a resource releaser that `delete`s the `driver_request` after the provider completes (the driver has already nulled `st.dreq` via reap before `ch.finish()` wakes the worker — happens-before via the channel mutex). Non-stream: worker `dreq.reset()` after the channel drains. - **Cancellation hook** (M3.5): `chunk_channel::cancelled` atomic; the driver checks it each step and finishes the slot (finish_reason "length" — no "cancelled" in the OpenAI spec). The SSE provider sets it on `sink.write` failure (client disconnect). Full disconnect polling is M3.5. - **400 validation**: missing/empty messages, prompt+max_tokens > ctx_per_slot. **Endpoint gate (green):** `curl /v1/chat/completions` stream:true emits SSE deltas + finish + [DONE] + usage; stream:false returns valid OpenAI JSON with choices[]+usage. Two parallel requests complete; qxmx_diff 0.108 unchanged; prefix_cache_engine_test green. (MVP done-criteria #1.) **Pre-existing (NOT M3.2):** `scheduler_test` "solo scheduler ids != sequential" fails on the last token — independent of M3.2 (M3.2 touches neither qxmx_scheduler nor qxmx_gpu); the M2.5 isolation gate (interleaved==sequential) still PASSES. Logged for later; out of M3.2 scope. ### M3.4 — tool-call parsing — DONE (2026-07-20) **Shipped:** - `src/qxmx_tool_call.h` — header-only `parse_tool_call_block(body, &calls)`. Parses the text between the TC_OPEN/TC_CLOSE specials: one or more `` blocks, each with `\nVALUE\n` entries. Values that jparse as JSON keep that type (numbers, bools, objects round-trip; the template emits non-strings via tojson); everything else is a string (strings are emitted bare). Structural errors (no `push` would leak the literal marker text into content). Text between them accumulates in a per-slot buffer; at TC_CLOSE the block is parsed. - Well-formed block -> one `stream_event` per ` fall back to plain content (markers reconstructed from `tok.decode(id, skip_special=false)`) + stderr log. - Whitespace-only content after a completed tool call is dropped (the "\n" separator between parallel TC pairs would otherwise dirty message.content). - Reap flushes the decoder tail + any unterminated block (plain content). - `chunk_channel` now carries `stream_event` (content OR one tool call). SSE worker: tool call -> 2 chunks (id/type/name + arguments delta); role rides the first chunk of either kind. finish_reason = "tool_calls" iff >=1 structured call AND the stop mapped to "stop" (length/ctx stays honest). - `scripts/smoke.sh`: +2 checks (tool-call fixture, stream + non-stream; python3 validates structured shape + arguments JSON). **12/12 green.** **Gates:** tool_call_test green; smoke.sh 12/12; qxmx_diff 0.108 unchanged. Live fixture (get_weather, Paris): non-stream -> `tool_calls[0].function. {name:get_weather, arguments:{"city": "Paris"}}`, content "", finish_reason tool_calls; stream -> name chunk + arguments chunk + finish chunk + [DONE]. ### M3.5 — lifecycle — DONE (2026-07-20) **Shipped (tools/qxmx_serve.cpp):** - **Disconnect cancellation, polled per token.** SSE provider: every event write failure sets `ch.cancelled` and returns (decode writes once per token -> per-token poll). Idle stretches (long prefill, no events) are covered by `chunk_channel::pop_for(ev, 1000)` returning `tick`; the provider writes an SSE comment heartbeat (`: hb`) whose failure means the client is gone. Driver checks `is_cancelled()` at each step top and reaps with finish_reason "length" (the spec has no "cancelled"). A disconnect during the monolithic `prefill_cached` is caught within ~1 s but only honored when the prefill call returns (chunk-interleaved prefill is post-MVP). Cancelled-while-pending requests are dropped at assign time. - **UAF fix (found by inspection before it fired).** The M3.2 design had the SSE resource-releaser `delete` the driver_request while the driver still held a raw `st.dreq` — on the disconnect path the provider returns BEFORE reap, so the driver's next step dereferenced freed memory. `driver_request` is now `shared_ptr`: both the provider lambda capture and the driver (pending queue + slot) hold refs; no releaser at all. - **Per-request log line** at reap: slot, prompt tok, prefill tok/s, decode tok/s (wall-clock as the client experienced it — slots share the GPU), gen tok, finish_reason (+" (cancelled)" flag). - **400 validation** (prompt+max_tokens > ctx_per_slot, missing messages): landed with M3.2. - **MVP deviation:** non-stream requests can't detect disconnect (nothing is written until completion); they run to completion bounded by max_tokens. Documented in the handler comment. **Live validation:** mid-decode abort -> "finish=length (cancelled)" after 45 tok / 1.6 s; mid-prefill abort (3213-tok prompt, curl --max-time 1) -> heartbeat caught it, slot reaped right after prefill returned; server healthy after both. ### M3.6 — smoke tests — DONE (2026-07-20) `scripts/smoke.sh` (14/14 green): models, non-stream, stream, 400 x2, two-parallel concurrency, tool-call fixture x2 (stream + non-stream, structured tool_calls + finish_reason), disconnect/cancel + post-cancel health. Standing gates: qxmx_diff 0.108, tool_call_test, chat_template test + tpl_crosscheck (M3.3), prefix_cache_engine_test (M2.6). ### M3.5 — lifecycle (original spec, superseded by the DONE record above) Cancellation on client disconnect (poll per token — essential), HTTP 400 when prompt+max_tokens > ctx-per-slot, per-request log line (slot, prompt tok, prefill tok/s, decode tok/s, finish_reason). ### M3.6 — smoke tests (original spec, superseded by the DONE record above) `scripts/smoke.sh`: non-stream, stream, tool-call fixture, two parallel requests (concurrency), disconnect/cancel. Plus the standing `qxmx_diff` regression gate. --- ## Reminders (from AGENTS.md, don't relearn) - One subtask per build, validate, measure, log. If a subtask can't be done as specified, STOP and report. - The engine is the only honest perf oracle. `QXMX_PROFILE=1` on `qxmx_run` for phase breakdown. - Never `.wait()` between device kernels (drains the pipeline ~1 ms each); only the final prefill sync, `phase()` profiling waits, and the `forward()`/`run_chunk_b()` start-of-call drain (M2.5 USM host-write fix) are legitimate. - USM: `malloc_shared` for GPU RMW, `malloc_device` for GPU-only. `malloc_host` for RMW = 2× correct output from stale L2 (never use). - `setvars.sh` must be sourced in the SAME shell command as the GPU build/run, else the GPU isn't visible. - icpx ad-hoc compiles need `--gcc-install-dir=/usr/lib/gcc/x86_64-linux-gnu/14` (handled inside meson). - Don't modify `qxmx_ref.{h,cpp}` (CPU oracle), the tokenizer, or historical spike harnesses. - Don't push unless asked; never force-push or amend commits you didn't create. ## Quick commands ```bash source /opt/intel/oneapi/setvars.sh >/dev/null 2>&1 # every shell M=~/models/bonsai/Ternary-Bonsai-27B-Q2_g64.gguf meson setup --reconfigure build >/dev/null 2>&1 && meson compile -C build ./build/qxmx_diff "$M" "The capital of France is" # decode gate (0.108) ./build/qxmx_run "$M" -f bench/session_4mb.txt -n 1 # prefill tok/s ./build/prefix_cache_test # M2.6.1 gate ./build/prefix_cache_engine_test "$M" bench/session_4mb.txt /tmp/delta.txt -n 12 # M2.6.3 gate (cached==fresh); needs a delta file (printf 'Now tell me about Paris.\n' > /tmp/delta.txt) ./build/scheduler_test "$M" "prompt A" "prompt B" -n 12 # M2.5 gate ./build/qxmx_serve --port 18080 "$M" # M3 server (M3.1 skeleton) curl -s http://127.0.0.1:18080/health # -> ok curl -s http://127.0.0.1:18080/v1/models # -> {"object":"list",...} ``` ## Memory files to read before touching each area - `qxmx_m2_6_prefix_cache_design.md` — the cache (before any M2.6.4/5 work or wiring it into M3.2). - `qxmx_m3_3_chat_template.md` — the chat template (before M3.3/M3.2). - `qxmx_m2_5_scheduler.md` — the scheduler + the USM host-write ordering bug (before the M3.2 threading design). - `qxmx_sycl_gotchas.md` — in-order queue, USM kinds, the malloc_host RMW bug (before any GPU code). - `shell_gotchas.md` — the icpx `--gcc-install-dir` gotcha + the no-filter-pipe rule.