OCaml library to run Deepseek v4 within a program on macOS
Cuda 33%
Metal 28%
C 19%
OCaml 9%
POV-Ray SDL 6%
C++ 4%
<1%
Perl <1%
Python <1%
Shell <1%
Dune <1%

README.md

ds4 -- local agents with the DS4 inference engine #

OCaml bindings for the DS4 inference engine, which runs the DeepSeek V4, DeepSeek V4.1 Flash, GLM 5.3 Flash and Qwen3.8 Flash Next language models locally, plus ds4-agent, a plain command over them.

Requires OCaml 5. Metal needs macOS on Apple Silicon and the Xcode Metal toolchain; CUDA needs Linux and the CUDA toolkit. Downloading a model needs uv or hf on your PATH.

Commands #

ds4-agent is built once per backend:

Command Backend Availability
ds4-agent-metal Metal macOS on Apple Silicon
ds4-agent-cuda CUDA Linux, with DS4_CUDA=yes
ds4-agent-cpu CPU everywhere

The examples below use ds4-agent-metal. Substitute ds4-agent-cpu to run without a GPU.

list shows which models exist, download fetches one, chat sends one prompt and prints the reply, and agent runs the model in a tool-using loop. A --model naming no file is refused before anything is loaded.

dune build builds Metal on macOS and CPU everywhere. CUDA is off unless you ask for it:

DS4_CUDA=yes dune build

That compiles for the GPU doing the build. Set DS4_CUDA_ARCH (sm_89 for an L4, sm_90 for an H100) to build for another one, and DS4_CUDA_HOME if the toolkit is not on the PATH. A GPU must hold the whole model in memory to run at full speed, so size the card to the model or use ds4-agent-cpu.

dune build
dune exec -- ds4-agent-metal list                  # what models exist
dune exec -- ds4-agent-metal download q4           # fetch one
dune exec -- ds4-agent-metal chat "Explain monads in one sentence."
dune exec -- ds4-agent-metal agent -d ./workspace "Summarise README.md"

Run ds4-agent-metal <command> --help for the options of each.

Choosing a model #

For DeepSeek, use the 0731 build of Flash, DeepSeek's 2026-07-31 retrain. Pick the largest quantisation your machine has memory for.

Memory Alias Target Size
256 GB+ q4 q4-imatrix-0731 153 GB
128 GB q2q4 q2-q4-imatrix-0731 98 GB
96 GB q2 q2-imatrix-0731 81 GB

DeepSeek V4.1 Flash is newer and larger. Its Engram tables stay on disk and are read as needed, so it uses less memory than its size on disk. It runs on Metal only.

Memory Alias Target On disk Resident
512 GB v41-q4 ds41-q4 483 GB 294 GB
256 GB v41-q2 ds41-q2 341 GB 152 GB

The Q4 is published in two parts, which download fetches and joins into one file. Joining needs another 37 GB free.

GLM 5.3 Flash and Qwen3.8 Flash Next run on the same engine. Qwen's n-gram tables stay on disk as V4.1's Engram tables do, and it runs on Metal only.

Memory Alias Target On disk Resident
256 GB+ glm-q4 glm53-q4 178 GB 178 GB
96 GB+ glm-q2 glm53-q2 90 GB 90 GB
128 GB+ qwen-q4 qwen38-q4 165 GB 70 GB
64 GB+ qwen-q2 qwen38-q2 137 GB 42 GB

Each of GLM 5.3, DeepSeek V4.1 Flash and Qwen3.8 Flash Next also has a small vision sidecar (glm-vision, v41-vision, qwen-vision, all under 1.2 GB), which agent --vision FILE loads beside the language model.

DeepSeek V4 Flash Vision Experimental is a checkpoint of its own, frozen before the 0731 retrain, rather than the regular Flash weights with a sidecar added. All three of its quants share one sidecar, ds4f-vision-encoder (alias vision-exp-encoder).

Memory Alias Target Size
96 to 128 vision-q2 ds4f-vision-q2 81 GB
128 GB vision-q2q4 ds4f-vision-q2-q4 91 GB
256 GB+ vision-mxfp4 ds4f-vision-mxfp4 145 GB

Models are stored under $XDG_DATA_HOME/ds4. With no --model, the best DeepSeek build you have downloaded is used, V4.1 ahead of V4 at the same quantisation, so chat needs no arguments after download. A GLM or Qwen model is chosen by asking for it, as in chat -m qwen-q4.

A target is named for the build it carries, and the short aliases follow the newest one. Asking for q4 gets whichever build is current, while q4-imatrix-preview still names the first Flash release. list dims the builds a newer one has superseded.

Two things affect the choice. The quantisations above are built with an imatrix, which weights the quantisation by measured activation statistics. mxfp4-0731 is not, so despite being larger it is not necessarily better. Memory use is also more than the model's size, because the KV cache grows with the context, and on Apple Silicon the GPU can address only about 75% of memory by default. sysctl iogpu.wired_limit_mb reports that limit.

The much larger PRO models are also listed, and need 512 GB.

Chatting #

chat sends one prompt and prints the reply.

ds4-agent-metal chat "Hello?"
ds4-agent-metal chat -m q2 "Hello?"                # a particular model
ds4-agent-metal chat -m /path/to/model.gguf "Hello?"
ds4-agent-metal chat --think high "Prove it."      # reason before answering
ds4-agent-metal chat --ctx 16384 "..."             # a larger context

The model keeps nothing between runs. A --model that names no file is refused before anything is loaded.

GLM 5.3 and Qwen3.8 Flash Next carry a draft head, which chat and agent arm on a GPU backend so that each step can commit tokens it drafted as well as the one it sampled. A Qwen agent runs about a quarter faster for it. Greedy decoding gives the same text either way, and --no-mtp turns it off.

--think opens the conversation with the model's own reasoning-effort instruction, so a reasoning model reasons in the form it was trained for.

The agent #

agent puts the same model in a loop with tools, so it can work across several turns.

ds4-agent-metal agent -d ./workspace "Rename Foo to Bar in lib/"
ds4-agent-metal agent -d ./workspace < prompts.txt

If the workspace holds an AGENTS.md, it joins the system prompt, so a project states its conventions once. Ctrl-C interrupts a reply and withdraws its prompt, and a second Ctrl-C at the prompt leaves.

With a prompt argument it runs that one exchange and exits. Without one it reads prompts from standard input, one per line, into one conversation, until the end of input. At a terminal it shows a ? prompt once the model has loaded, and Ctrl-D ends the session. The reply goes to standard output. The prompt, each tool call, the first line of its result, and any change to the context go to standard error, so a reply can be captured on its own.

list, tree, read, read_lines, find, grep and stat read the filesystem, edit changes part of a file, write creates one or replaces the whole of one and append adds to the end of one. There is no shell.

--vision FILE loads the sidecar that matches the model and adds view_image, confined to the same capability as the other file tools. GLM 5.3, DeepSeek V4.1 Flash and Qwen3.8 Flash Next each publish one, as the glm53-vision, ds41-vision and qwen38-vision download targets.

grep quotes each matching line as it is in the file, so it can be quoted back to edit. write and edit replace a file by renaming a new one over it, so an interrupted change leaves the old file whole, and edit answers with the changed lines numbered and how far the lines after them moved.

The filesystem tools reach only what they are granted. The agent starts with a capability for the workspace given by -d, and any directory outside it must be requested with open_dir. The grant is reported on standard error. A capability rejects .. and symlinks that lead out of it, so no path the model supplies can escape.

A reply is bounded at 16384 tokens, which --max-tokens sets, and a tool call the model does not finish inside that is discarded rather than half made. The agent tells the model and gives it another turn, which is what append is for. A call that is malformed, rather than unfinished, is reported and retried. Work can be cancelled cooperatively.

The agent starts with a context of 32768 tokens, which --ctx sets. A conversation that outgrows it moves to a larger one, up to --max-ctx, without reloading the model. Once it cannot grow further, the model is asked to summarise its own conversation, which is then replaced by that summary and a verbatim tail of the most recent exchanges, so a long-running task keeps going rather than stopping.

Library #

The ds4 library is the same functionality without the command line. V4.generate sends one prompt, V4.Session keeps a conversation, Agent adds the tool loop, and Toolbox provides the tools. The backend is chosen by depending on ds4.metal, ds4.cuda or ds4.cpu. ds4.cli holds the model catalogue, the list, download and chat subcommands, and in Coder the parts of the agent subcommand, and ds4.dsml the prompt encoding and reply parsing. See the interface files in lib/, cli/ and dsml/.

Maintenance #

See ARCH.md for the repository layout, the vendored engine and how to update it, the FFI, and how to quantise a DeepSeek release yourself.