zig langref cli nate.tngl.io/zigman
zigman DESIGN.md
7.9 kB
Markdown
at main

design: a self-improvement loop for zigman #

this document explains what we're building and why, so a reviewer can poke holes in it. feedback especially wanted on the open questions at the end.

what zigman is #

a single static binary that reads the zig language reference in the terminal. the langref is one big static HTML page with a stable anchor per section; zigman <query> fetches it once, caches it, and prints the matching section as clean markdown. it exists so humans (and agents) can read precise zig language semantics without a browser. it also ships as a cross-client SKILL.md (agentskills.io format) under skills/zigman/.

scope is deliberate: zigman covers the language, not the standard library.

the actual goal #

zigman is a means to an end. the end is a virtuous feedback loop: use a small, local, open model (gemma3 via ollama on the author's machine) to improve zigman, where the model uses zigman itself to learn the zig it needs. concretely:

give a weak model a real zig coding task and the zigman tool. measure whether having zigman makes it write better zig. when it doesn't, that failure tells us how to make zigman (and its skill) better. repeat.

the bet: if a weak model writes correct zig more often with zigman than without, zigman is delivering real value — and the gap between "with" and "without" is a number we can optimize. the same loop surfaces zigman's own quality problems (see axes below).

framing: this is partly a yak-shave to prove you can stand up a loop that achieves an objective with local models. dumb models are fine — even desirable — as long as they're not too dumb to use a tool at all; the interesting question then becomes "how do you enable a not-very-smart model to use a tool effectively."

why zig is a good domain for this #

unlike most agent evals, zig has an objective oracle: the compiler. we do not need an LLM to grade whether code is correct — zig test is ground truth: it compiles and runs the agent's function against a hidden test (oracle inputs are forced to runtime, so it executes, not just type-checks). that removes the "who judges the judge" problem from the core signal.

architecture of the eval #

the stack is all local, no GUI, and each layer is a CLI we orchestrate from a small stdlib-only uv package (evals/, the zigman-eval command) — no agent framework of our own:

  • agent harness: Pi (@earendil-works/pi-coding-agent), run headless (pi -p --mode json). lightweight built-in read/bash/edit/write tools and native agentskills.io --skill loading. we parse its JSON event stream for the tool-call trace.
  • model: gemma-4-12B served OpenAI-compatible by mlx_vlm.server on :1234 (LM Studio drops in on the same port). model-agnostic: any Pi provider/model works, so we can sweep dumb→smart. gemma-4 is the generation that finally tool-calls and writes competent zig locally (gemma3/qwen2.5 via ollama did not — see findings).
  • objective oracle: after the run we read the agent's solution.zig, append the hidden oracle test, and zig test. pass/fail, no opinion. selftest validates every oracle against a reference solution before any model runs.

each task is a real one-shot coding ask (e.g. "implement pub fn cstrLen(...)") plus a hidden oracle test and a reference solution. ablation: the zigman arm loads skills/zigman via --skill and has the zigman CLI on PATH (the agent runs it via bash); the bare arm doesn't. lift = pass-rate delta.

current state (what's proven) #

  • the pipeline runs end-to-end, fully local: Pi + gemma-4 (MLX) writes solution.zig, runs zig test itself via bash, and the hidden oracle passes. a committed sample trace lives at evals/sample_report.json; live runs write to evals/reports/ (gitignored).
  • the corpus is seven intentionally reference-dependent tasks (@intCast single-arg form, for (xs, 0..) |x, i| index capture, exhaustive enum switch, labeled-block break, error-union try, sentinel array slicing, sentinel pointer len), all selftest-validated.

the two-day detour, and what it taught us (serving layer, per does-it-tool's thesis) #

getting a local model to use tools at all was the whole fight, and the lesson is that it's a serving-layer / model-generation problem, never "local can't":

  1. ollama + gemma3/qwen2.5 was a dead end for code-bearing tool calls. gemma3-tools is a prompt hack (no native tool-calling; a Modelfile coerces bare JSON that ollama string-parses) and it fences code-bearing calls; stock qwen2.5-coder emits bare JSON ollama won't parse; llama3.1 narrates. each a different per-model quirk.
  2. opencode and SkillsToolset made it worse — their large prompt/tool overhead destabilized weak models' tool formatting. lean harness wins.
  3. the fix was the engine+model, not the harness (this is Vicki Boykis's "local models are good now" point): gemma-4 served by mlx_vlm returns proper structured tool_calls — immune to fencing because the call isn't text — and writes competent zig. swap the serving layer, don't hand-tune Modelfiles.

so "how do you enable a not-too-dumb model to use a tool" is overwhelmingly a plumbing question (engine, template, tool-call format), not an IQ one.

the two loops #

  • capability loop (objective): does zigman lift a weak model's zig correctness? ablation (with-zigman vs bare) scored by the compiler. this is the measurement.
  • quality loop (multi-axis): the model, using zigman, reveals zigman's own weaknesses across explicit axes — API sensibility, discoverability, output quality, ergonomics, portability, performance, robustness, maintainability, aesthetics, distribution (see evals/AXES.md). its friction is the signal; fixes stay within the existing command surface.

open questions (where we want feedback) #

  1. eval content is the crux. we now ship a 7-task reference-dependent corpus (per Codex's task-mining suggestion) and a bare-vs-zigman ablation that flags which tasks are reference-dependent (bare not perfect). open: is "bare fails, zigman passes" a stable enough signal on a noisy weak model to promote a task, or do we need N trials + a significance bar? and what's the principled way to keep sourcing gap tasks as models improve (version-sensitive builtins age out)?
  2. judge with only a weak local model. prefect uses Opus as judge; we have no cloud keys in the loop's environment, so today the judge is also gemma. how much should we trust a weak judge, and where does the objective oracle suffice without it?
  3. gaming risk. the agent writes to a file and we append a hidden oracle. if a task lets the agent write its own tests, it can pass vacuously. is appending a hidden oracle to the agent's final file robust enough, or do we need stronger isolation between the agent's code and the grading harness?
  4. does the loop actually close? we can measure lift and collect critiques, but the "act" step (improving zigman) is still human. how much of that can/should be automated, and how do we avoid overfitting zigman to one weak model's quirks?
  5. is ablation the right top-line metric, or should we weight by whether the model used the tool, latency, token cost, etc.?

non-goals #

  • not trying to make gemma a good zig programmer; gemma is an instrument.
  • not expanding zigman's command surface to chase eval numbers — improvements stay within the existing API (zigman <query>, -l, -V, -o, completions).
  • not a general code-eval framework; the compiler oracle is what makes this clean, and that's specific to compiled languages.