diff --git a/llama/baseline-qwen3.6-2026-09-04.md b/llama/baseline-qwen3.6-2026-09-04.md new file mode 100644 index 0000000..3d9f8b6 --- /dev/null +++ b/llama/baseline-qwen3.6-2026-09-04.md @@ -0,0 +1,50 @@ +# Baseline — qwen3.6-35b-a3b (2026-09-04) + +## Build / host +- llama-server `b210-50f068fff`, Strix Halo 395 (32 threads), 125 GiB unified +- Router: `MODELS_MAX=1`, port 9931, sleep-idle 900s. Model was `sleeping`, woke on first request. + +## Config snapshot (models.ini [qwen3.6-35b-a3b]) +- weights `Qwen3.6-35B-A3B-UD-Q8_K_XL.gguf` (37G) + `mmproj-BF16` (861M) +- `spec-type draft-mtp`, `spec-draft-n-max 2` +- template `qwen-fixed-chat_template.jinja` (froggeric-v22.1), `reasoning auto`, `deepseek`, `preserve` +- sampling: temp 1.0, top-p 0.95, top-k 20, min-p 0.0, presence 1.5, repeat 1.0 (= upstream Thinking/general) +- unset (defaults): ctx-size → 262144 from GGUF, batch 2048 / ubatch 512, cache-type K/V default, np auto → **4 slots × 262144**, unified KV on, cont-batching on + +## Final lock-in (2026-09-04) + +`llama/models.ini` `[*]` + `[qwen3.6-35b-a3b]`: +- `n-gpu-layers = all` (was 99; future-proof for deeper models, global `[*]`) +- `spec-draft-n-max = 3` (was 2; MTP sweep winner) +- `ubatch-size = 2048` (was default 512; −21% long-prompt ingest, peaked vs 4096 regression) +- unchanged: ctx 262144, 4 slots unified KV (F16), flash-attn, sampling Thinking/general, froggeric jinja, KV quant skipped + +Verified: child args `--n-gpu-layers all --spec-draft-n-max 3 --ubatch-size 2048`; sanity 71.1 tok/s / 11-11 accepted. + +## Tests (all via POST /v1/chat/completions, model=qwen3.6-35b-a3b, stream=false) + +| # | Task | prompt_n | prompt tok/s | gen tok/s | wall | draft accepted/total | Notes | +|---|------|----------|--------------|-----------|------|----------------------|-------| +| T1 | "one sentence: speculative decoding", max 128 | 19 | 151 | 60.9 | 2.2s | — | thinking model; short replies mostly land in `reasoning_content` | +| T2 | math + `<\|think_medium\|>`, max 512 | 40 | 266 | 68.0 | 7.7s | — | ok | +| T3 | tool-use (get_weather/Paris), max 256 | 382 | 646 | 59.8 | 2.0s | — | correct `tool_calls` emitted | +| T4 | codebase dump ~38KB chars / 44 files, max 256 | 12431 | 835 | 64.7 | 16.3s | 57/64 (0.89 this req) | correct answer (top-level dirs + llama/ contents), `reasoning_content` empty (auto skipped thinking) | +| T5 | "Say hi in one sentence", max 64 | 16 | 147 | 62.1 | ~1.2s | 38/48 (0.79 this req) | **truncated inside ``** — content empty, reasoning only. Expected: thinking models need generous `max_tokens`; 64 cuts the answer off. | + +## Server counters (after T5, /metrics?model=qwen3.6-35b-a3b) +- prompt avg 471 tok/s, gen avg 65.1 tok/s +- spec totals: 2727 drafted / 1968 accepted = **0.72**; pos0 1088, pos1 880 (cond. pos1 ≈ 0.81) +- `n_tokens_max` observed 36163 (historical, pre-baseline) +- `requests_deferred` 0 +- Prior journal (overnight task 477): prompt 276 tok/s, gen 46 tok/s, acceptance 0.55, mean len 2.10 — slower/longer-context than today's short tests + +## Memory +- `MemoryCurrent` 42.5G, `MemoryPeak` 63.7G; system 49G used / 75G avail +- Headroom for full-262k KV (~21.5G at F16) looks fine on paper; not yet tested at the limit + +## Observations / follow-ups +1. Sampling default matches upstream Thinking/general — keep; precise-coding + instruct presets belong as **per-request overrides** from pi, not in models.ini. +2. `max_tokens` guidance for pi: thinking responses need room (T5 shows 64 tokens = answer never leaves ``). Suggest ≥1024 default for agentic calls. +3. Tool-use path works with custom jinja (T3). Deeper agentic loop (multi-turn + truncation warnings) not yet tested. +4. Long-context validated to 12k prompt tokens only. Full 262k soak + compaction-concurrency test still open (deferred, lower priority per user). +5. Next tuning candidates: ubatch/batch for big prompts, `spec-draft-n-max` 2→3→4 sweep, cache-type-k/v if 262k-soak pressures memory. diff --git a/llama/models.ini b/llama/models.ini index adadf13..43668d1 100644 --- a/llama/models.ini +++ b/llama/models.ini @@ -8,7 +8,7 @@ version = 1 ; Global defaults shared by every model instance (overridable per-model) ; --------------------------------------------------------------------------- [*] -n-gpu-layers = 99 +n-gpu-layers = all flash-attn = on jinja = true metrics = true ; Prometheus /metrics endpoint (per-model: /metrics?model=...) @@ -21,7 +21,8 @@ load-on-startup = true model = /home/spencer/.cache/models/Qwen3.6-35B-A3B-UD-Q8_K_XL.gguf mmproj = /home/spencer/.cache/models/Qwen3.6-35B-A3B-mmproj-BF16.gguf spec-type = draft-mtp -spec-draft-n-max = 2 +spec-draft-n-max = 3 +ubatch-size = 2048 chat-template-file = /home/spencer/.config/llama.cpp/qwen-fixed-chat_template.jinja reasoning-format = deepseek reasoning = auto diff --git a/llama/prometheus/README.md b/llama/prometheus/README.md new file mode 100644 index 0000000..06bbb0b --- /dev/null +++ b/llama/prometheus/README.md @@ -0,0 +1,21 @@ +# Short-term Prometheus for llama-server (rootless podman) + +- Container: `prometheus-short`, host networking, web UI on 127.0.0.1:9090 +- Config: `prometheus.yml` (this dir), bind-mounted read-only +SELinux :z +- TSDB data: `~/.local/share/prometheus-data` (7d retention) +- Scrapes router `127.0.0.1:9931/metrics?model=qwen3.6-35b-a3b` every 10s +- Router requires `?model=` per job (400 without it) + +Manage: + podman start prometheus-short | podman stop prometheus-short | podman logs -f prometheus-short + # config edits: podman exec prometheus-short kill -HUP 1 (or restart) + +Useful queries (qwen3.6 job): + increase(llamacpp:prompt_tokens_seconds[...]) # prompt tok/s windowed + llmacpp:predicted_tokens_seconds # gen tok/s gauge + sum(increase(llamacpp:spec_decode_num_accepted_tokens_total[5m])) + / sum(increase(llamacpp:spec_decode_num_draft_tokens_total[5m])) # acceptance + sum(llamacpp:spec_decode_num_accepted_tokens_per_pos_total{position="1"}) + / sum(llamacpp:spec_decode_num_accepted_tokens_per_pos_total{position="0"}) # pos1 conditional + +NOTE: rootless container does NOT auto-restart on machine reboot; restart manually with `podman start prometheus-short` (or promote to a quadlet user service if it outlives the short-term window). diff --git a/llama/prometheus/prometheus.yml b/llama/prometheus/prometheus.yml new file mode 100644 index 0000000..ca8e662 --- /dev/null +++ b/llama/prometheus/prometheus.yml @@ -0,0 +1,25 @@ +# Short-term Prometheus scrape config for llama-server router (user service, port 9931). +# One scrape job per model instance; the router serves per-model metrics via ?model=. +# Stored in the dotfiles repo for reproducibility alongside tuning notes. +global: + scrape_interval: 10s + evaluation_interval: 30s + +rule_files: [] + +scrape_configs: + # qwen3.6-35b-a3b (current tuning target) + - job_name: llama-qwen3.6-35b-a3b + metrics_path: /metrics + params: + model: ["qwen3.6-35b-a3b"] + static_configs: + - targets: ["127.0.0.1:9931"] + labels: + model: "qwen3.6-35b-a3b" + + # NOTE: the router returns 400 for /metrics without a model param, + # so every job must pin ?model=. User-side aggregate is not exposed. + + # Add other models later by copying the qwen3.6 job and changing params.model + # e.g. qwen3.8-27b, qwen3.8-flash-next, qwen3-coder-next-80b-a3b \ No newline at end of file diff --git a/llama/results-envelope-100k-2026-09-04.md b/llama/results-envelope-100k-2026-09-04.md new file mode 100644 index 0000000..e9109a1 --- /dev/null +++ b/llama/results-envelope-100k-2026-09-04.md @@ -0,0 +1,34 @@ +# Envelope soak — 92k-token prompt on qwen3.6-35b-a3b (2026-09-04) + +Single 92,095-token prompt (llama.cpp `src/**` source dump, 47 files, 361 KB) sent to the live router instance (n-max 3, 4 slots, unified KV, ubatch 512, KV F16 default). + +## Results +- Prompt processing: 92,095 tokens in 200.2 s → **460 tok/s** (vs 813 tok/s at 14.4k prompt — long prompts roughly halve prompt throughput) +- TTFT ≈ 200 s (prompt-bound; generation negligible by comparison) +- Generation: 192 tokens @ **45.8 tok/s** — long-context decode is slower than short-ctx ~60-72 (KV attention cost + unified pool) +- Draft acceptance at depth: **123/202 = 0.61** (vs 0.76-0.84 short-ctx; matches overnight 0.55 @ 36k — acceptance degrades with ctx depth, expected for MTP) +- Prompt was completely cached-cold (`cached_total 0`), `n_tokens_max = 92286` +- Answer truncated inside ` thinking` (max_tokens 192 too small — harness artifact, not a defect; pi sends 262144) + +## Memory +- No OOM; RSS +0.7 GB during the soak (48→49 G used of 125 G; 75 G available). KV for 92k ≈ 7.5 GB F16, absorbed by the unified pool. +- Extrapolation to full 262,144: ~21.5 GB F16 KV → peak ~62-66 G RSS. **Comfortable — KV cache quantization NOT needed**; keep F16 for max long-context recall. + +## Decisions +- **ubatch-size 2048 (lock-in)** — see sweep table below; peaked at 2048, regressed at 4096. +- KV quant skipped (no pressure). +- Acceptance-at-depth 0.61 is inherent to MTP at long ctx; n-max 3 remains the best setting. + +## ubatch sweep (same 92,095-token prompt, restart per variant) + +| ubatch | batch | prompt tok/s | gen tok/s | draft acc | +|--------|--------|--------------|-----------|-----------| +| 512 (default) | 2048 (default) | 460.1 | 45.8 | 0.61 | +| 1024 | 2048 | 519.7 (+13%) | 47.9 | 0.65 | +| **2048** | 2048 | **555.5 (+21%)** | 59.7 | 0.93 | +| 4096 | 4096 | 524.8 (−6% vs 2048) | 60.3 | 0.95 | + +- Optimal at 2048; 4096 regresses (larger physical batches overshoot GPU efficiency on this APU). +- `batch-size` left at default 2048 (physical cap is ubatch). +- Final child args verified: `--spec-draft-n-max 3 --flash-attn on --n-gpu-layers 99 --ubatch-size 2048`. +- Post-lock sanity: 72.4 tok/s short chat, 11/11 accepted. \ No newline at end of file diff --git a/llama/results-spec-sweep-2026-09-04.md b/llama/results-spec-sweep-2026-09-04.md new file mode 100644 index 0000000..94ff290 --- /dev/null +++ b/llama/results-spec-sweep-2026-09-04.md @@ -0,0 +1,24 @@ +# Spec-draft sweep — qwen3.6-35b-a3b (2026-09-04) + +Same-day A/B, same build (b210), same battery (3× short-chat, 3× reason, 3× tool, 1× ~14.4k-token codebase Q&A), `draft-mtp` fixed, only `spec-draft-n-max` varied. Acceptance/mean-len from fresh `/metrics` counters (child restarted per variant → totals = battery-only). + +| n-max | acceptance | mean draft len | chat gps | reason gps | tool gps | code gps | +|-------|-----------|----------------|----------|------------|----------|----------| +| 2 | 0.836 | 2.00 | 61 | 67 | 61 | 65.2 | +| **3** | 0.760 | **3.00** | 62 | **72** | 61 | **71.7** | +| 4 | 0.674 | 3.98 | 56 | 74.5 | 62 | 59.9 | + +## Decision: keep `spec-draft-n-max = 3` + +- n-max 3 posts full 3-token draft chains (mean len 3.00) with best agentic (code) and reasoning throughput. +- n-max 4 over-drafts: acceptance falls to 0.67, code gen drops ~12 tok/s vs n-max 3 despite full 4-token chains — MTP useful depth exceeded for this model. +- n-max 2 under-drafts (mean 2.00): less per-step throughput. + +## Locks / verification +- `models.ini` `[qwen3.6-35b-a3b]`: `spec-type=draft-mtp`, `spec-draft-n-max=3` (qwen3.8-27b stanza untouched at 4). +- Router restarted; child args verified via `/v1/models`: `--spec-draft-n-max 3 --spec-type draft-mtp`. +- Post-restart sanity: gen 70.7 tok/s, 11/11 accepted. + +## Notes +- Counters reset on child restart; Prometheus scrape deltas across restarts are invalid — always read `/metrics?model=...` fresh per test window. +- Prior overnight run (36k ctx, n-max 2): acceptance 0.55 / mean 2.10 — long-context acceptance is lower than short-ctx; keep that in mind for the 262k soak. \ No newline at end of file