zig langref cli nate.tngl.io/zigman
README.md

zigman-eval #

does zigman help a local coding agent write correct zig? an ablation (bare vs. zigman) over a corpus of reference-dependent zig tasks, run through Pi (the agent) and scored by zig test (the objective oracle — it runs the code, not just compiles it).

pure stdlib; it orchestrates the pi and zig CLIs as subprocesses.

layout #

src/zigman_eval/
  tasks.py    # the corpus + the compiler oracle (zig_test_source, grade, validate_oracles)
  config.py   # Config + resolution (model via arg/env, zigman via PATH, skill by walking up)
  pi.py       # run one task through Pi headless; parse its JSON events -> ArmResult
  loop.py     # ablate() -> Report; render() the table; save() the json
  cli.py      # `zigman-eval selftest | run`
tests/        # the oracles must accept their reference solutions

use #

first serve a tool-calling-capable local model (see the project's local-models/ notes — gemma-4 via mlx_vlm.server works):

mlx_vlm.server --model ~/models/gemma-4-12B-it-8bit --port 1234

then, from this directory:

uv run zigman-eval selftest                       # validate oracles, no model
uv run zigman-eval run --model <provider-model-id>   # ablate the corpus
uv run zigman-eval run --model <id> -k intcast --arms zigman
ZIGMAN_EVAL_MODEL=<id> uv run zigman-eval run     # model via env
uv run pytest                                     # the oracle tests

the --model is whatever your Pi provider calls the model (for the local provider configured in ~/.pi/agent/models.json, that's typically the model path).