# design: a self-improvement loop for zigman this document explains what we're building and why, so a reviewer can poke holes in it. feedback especially wanted on the **open questions** at the end. ## what zigman is a single static binary that reads the [zig language reference](https://ziglang.org/documentation/master/) in the terminal. the langref is one big static HTML page with a stable anchor per section; `zigman ` fetches it once, caches it, and prints the matching section as clean markdown. it exists so humans (and agents) can read precise zig language semantics without a browser. it also ships as a cross-client `SKILL.md` (agentskills.io format) under `skills/zigman/`. scope is deliberate: zigman covers the **language**, not the standard library. ## the actual goal zigman is a means to an end. the end is a **virtuous feedback loop**: use a small, local, open model (gemma3 via ollama on the author's machine) to *improve zigman*, where the model uses zigman itself to learn the zig it needs. concretely: > give a weak model a real zig coding task and the zigman tool. measure whether > having zigman makes it write better zig. when it doesn't, that failure tells us > how to make zigman (and its skill) better. repeat. the bet: if a weak model writes correct zig *more often with zigman than without*, zigman is delivering real value — and the gap between "with" and "without" is a number we can optimize. the same loop surfaces zigman's own quality problems (see axes below). framing: this is partly a yak-shave to prove you can stand up a *loop that achieves an objective* with local models. dumb models are fine — even desirable — as long as they're not *too* dumb to use a tool at all; the interesting question then becomes "how do you enable a not-very-smart model to use a tool effectively." ## why zig is a good domain for this unlike most agent evals, zig has an **objective oracle: the compiler.** we do not need an LLM to grade whether code is correct — `zig test` is ground truth: it compiles *and runs* the agent's function against a hidden test (oracle inputs are forced to runtime, so it executes, not just type-checks). that removes the "who judges the judge" problem from the core signal. ## architecture of the eval the stack is all local, no GUI, and each layer is a CLI we orchestrate from a small stdlib-only uv package (`evals/`, the `zigman-eval` command) — no agent framework of our own: - **agent harness: [Pi](https://github.com/badlogic/pi-mono)** (`@earendil-works/pi-coding-agent`), run headless (`pi -p --mode json`). lightweight built-in `read`/`bash`/`edit`/`write` tools and native agentskills.io `--skill` loading. we parse its JSON event stream for the tool-call trace. - **model: gemma-4-12B** served OpenAI-compatible by `mlx_vlm.server` on `:1234` (LM Studio drops in on the same port). model-agnostic: any Pi provider/model works, so we can sweep dumb→smart. gemma-4 is the generation that finally tool-calls *and* writes competent zig locally (gemma3/qwen2.5 via ollama did not — see findings). - **objective oracle:** after the run we read the agent's `solution.zig`, append the hidden oracle `test`, and `zig test`. pass/fail, no opinion. `selftest` validates every oracle against a reference solution before any model runs. each task is a real one-shot coding ask (e.g. "implement `pub fn cstrLen(...)`") plus a hidden oracle test and a reference solution. **ablation:** the zigman arm loads `skills/zigman` via `--skill` and has the `zigman` CLI on PATH (the agent runs it via `bash`); the bare arm doesn't. lift = pass-rate delta. ## current state (what's proven) - the pipeline runs end-to-end, fully local: Pi + gemma-4 (MLX) writes `solution.zig`, runs `zig test` itself via `bash`, and the hidden oracle passes. a committed sample trace lives at [`evals/sample_report.json`](evals/sample_report.json); live runs write to `evals/reports/` (gitignored). - the corpus is seven intentionally reference-dependent tasks (`@intCast` single-arg form, `for (xs, 0..) |x, i|` index capture, exhaustive enum switch, labeled-block break, error-union `try`, sentinel array slicing, sentinel pointer len), all selftest-validated. ### the two-day detour, and what it taught us (serving layer, per does-it-tool's thesis) getting a *local* model to use tools at all was the whole fight, and the lesson is that it's a **serving-layer / model-generation** problem, never "local can't": 1. **ollama + gemma3/qwen2.5 was a dead end for code-bearing tool calls.** gemma3-tools is a prompt *hack* (no native tool-calling; a Modelfile coerces bare JSON that ollama string-parses) and it fences code-bearing calls; stock qwen2.5-coder emits bare JSON ollama won't parse; llama3.1 narrates. each a different per-model quirk. 2. **opencode and `SkillsToolset` made it worse** — their large prompt/tool overhead destabilized weak models' tool formatting. lean harness wins. 3. **the fix was the engine+model, not the harness** (this is Vicki Boykis's "local models are good now" point): **gemma-4 served by mlx_vlm returns proper structured `tool_calls`** — immune to fencing because the call isn't text — and writes competent zig. swap the serving layer, don't hand-tune Modelfiles. so "how do you enable a not-too-dumb model to use a tool" is overwhelmingly a plumbing question (engine, template, tool-call format), not an IQ one. ## the two loops - **capability loop (objective):** does zigman lift a weak model's zig correctness? ablation (with-zigman vs bare) scored by the compiler. this is the measurement. - **quality loop (multi-axis):** the model, *using* zigman, reveals zigman's own weaknesses across explicit axes — API sensibility, discoverability, output quality, ergonomics, portability, performance, robustness, maintainability, aesthetics, distribution (see `evals/AXES.md`). its friction is the signal; fixes stay within the existing command surface. ## open questions (where we want feedback) 1. **eval content is the crux.** we now ship a 7-task reference-dependent corpus (per Codex's task-mining suggestion) and a bare-vs-zigman ablation that flags which tasks are reference-dependent (bare not perfect). open: is "bare fails, zigman passes" a stable enough signal on a noisy weak model to *promote* a task, or do we need N trials + a significance bar? and what's the principled way to keep sourcing gap tasks as models improve (version-sensitive builtins age out)? 2. **judge with only a weak local model.** prefect uses Opus as judge; we have no cloud keys in the loop's environment, so today the judge is also gemma. how much should we trust a weak judge, and where does the objective oracle suffice without it? 3. **gaming risk.** the agent writes to a file and we append a hidden oracle. if a task lets the agent write its own tests, it can pass vacuously. is appending a hidden oracle to the agent's final file robust enough, or do we need stronger isolation between the agent's code and the grading harness? 4. **does the loop actually close?** we can measure lift and collect critiques, but the "act" step (improving zigman) is still human. how much of that can/should be automated, and how do we avoid overfitting zigman to one weak model's quirks? 5. **is ablation the right top-line metric**, or should we weight by whether the model used the tool, latency, token cost, etc.? ## non-goals - not trying to make gemma a good zig programmer; gemma is an instrument. - not expanding zigman's command surface to chase eval numbers — improvements stay within the existing API (`zigman `, `-l`, `-V`, `-o`, completions). - not a general code-eval framework; the compiler oracle is what makes this clean, and that's specific to compiled languages.