Something went wrong. Try again.
User-specific application configuration is traditionally stored in so called dotfiles, these are my own.
Something went wrong. Try again.
123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263version = 1
; ============================================================================; llama.cpp router presets — Strix Halo (Radeon 8060S, 126 GiB unified VRAM); ============================================================================
; ---------------------------------------------------------------------------; Global defaults shared by every model instance (overridable per-model); ---------------------------------------------------------------------------[*]n-gpu-layers = allflash-attn = onjinja = truemetrics = true ; Prometheus /metrics endpoint (per-model: /metrics?model=...)
; ---------------------------------------------------------------------------; Qwen3.6-35B-A3B — Alibaba, MoE 3B-active + multimodal. Fast all-rounder.; ---------------------------------------------------------------------------[qwen3.6-35b-a3b]load-on-startup = truemodel = /home/spencer/.cache/models/Qwen3.6-35B-A3B-UD-Q8_K_XL.ggufmmproj = /home/spencer/.cache/models/Qwen3.6-35B-A3B-mmproj-BF16.ggufspec-type = draft-mtpspec-draft-n-max = 3ubatch-size = 2048chat-template-file = /home/spencer/.config/llama.cpp/qwen-fixed-chat_template.jinjareasoning-format = deepseekreasoning = autoreasoning-preserve = truetemperature = 1.0top-p = 0.95top-k = 20min-p = 0.0presence-penalty = 1.5repeat-penalty = 1.0
; ---------------------------------------------------------------------------; Qwen3.8-Flash-Next — Qwen4-preview arch (qwen4exp), 125B total / 6B active.;; Speculative decoding: DRAFT-MTP (unsloth qwen4exp/mtp fork build provides it; upstream PR; ggml-org/llama.cpp#27836 still can't load head-only MTP GGUFs). Head = unsloth shared-Q4_K_M.; n-gram/PLE table is lazy-from-disk in the fork loader (TENSOR_READ_LAZY, no flag needed).; ---------------------------------------------------------------------------[qwen3.8-flash-next]; unsloth UD-Q4_K_XL + MTP head (their qwen4exp/mtp fork build — REQUIRED for MTP).; n-gram/PLE table is served from disk lazily by the fork loader by default (TENSOR_READ_LAZY).; The shared head excludes embed_tokens (shares the main model's).model = /home/spencer/.cache/models/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.ggufmmproj = /home/spencer/.cache/models/Qwen3.8-Flash-Next-mmproj-BF16.ggufspec-draft-model = /home/spencer/.cache/models/mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.ggufchat-template-file = /home/spencer/.config/llama.cpp/qwen-fixed-chat_template.jinjareasoning-format = deepseekreasoning = autoreasoning-preserve = truespec-type = draft-mtpspec-draft-n-max = 5temperature = 1.0top-p = 0.95top-k = 20min-p = 0.0presence-penalty = 0.0repeat-penalty = 1.0