# qxmx This project contains an experimental single-purpose inference engine with one goal: running the [Bonsai-27B ternary-quantized model](https://prismml.com/news/bonsai-27b) as fast as possible on Intel Arc GPU hardware. I have an Arc Pro B70. Once I discovered that Intel's GPU hardware natively supports a 2-bit integer format (and AMD's Radeon Pro R9700 does not!), one which happens to map perfectly in-hardware to the ternary values `[-1, 0, +1]`... I wanted to see how hard it would be to bootstrap a new inference engine from scratch using modern LLMs (mainly GLM-5.2 and Kimi-K3). KV cache has user-selectable K quantization (`fp16`, `fp8`, and `q8_0`), but V quantization is currently hardcoded to 4-bit TurboQuant. There's an OpenAI-chat-completions-compatible endpoint built into `qxmx_serve`. DSpark speculative decoding is supported by the model but not yet by this software. Nor is any vision functionality currently implemented. One thing I've learned is that despite being compatible with a wide variety of open models, llama.cpp is *really* well-optimized and hard to beat! I was inspired by https://github.com/antirez/ds4/ and https://github.com/JustVugg/colibri/ but don't blame them! None of this is their fault. ## download the model You'll need either of the ternary 27B `Q2_0` files from Prism-ML's huggingface; both work, but I mainly developed and tested with [Ternary-Bonsai-27B-Q2_g64.gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf/blob/main/Ternary-Bonsai-27B-Q2_g64.gguf) specifically. In the future it may make more sense to use `hf download` to grab the appropriate files once we support speculative decoding or vision features, but for now, just this one GGUF file is all we need. ## the easy path You can build the Dockerfile and then run `docker --device /dev/dri ...` to pass your Intel GPU through (right now only one GPU is supported but multiple Arc GPU support is planned for a future release). That's probably the easiest way to get it running locally. An example command: ```bash docker run --rm --name qxmx_serve --device /dev/dri -p 8080:8080 -v ~/models:/models:ro qxmx -- ./build/qxmx_serve --host 0.0.0.0 --port 8080 --slots 2 --ctx-per-slot 131072 --ctk q8_0 /models/Ternary-Bonsai-27B-Q2_g64.gguf ``` This will get you a local web server that speaks something similar to the OpenAI `/chat/completions` endpoint at `http://localhost:8080/v1/chat/completions`, and you can see the single model at `http://localhost:8080/v1/models`. Tool-calling in [maki](https://maki.sh) seems to work. Performance optimizations are ongoing but speed is now competitive with llama.cpp on B70. ## building from source The author has only tested this on Debian Trixie with some packages from testing/unstable. You'll need to ensure you have run `apt install meson ninja-build pkg-config libstdc++-14-dev`, and then you'll have to set up the oneapi apt repo and install the `intel-oneapi-dpcpp-cpp-2026.1` package from it. Once that's all done, you should be able to run: ```bash source /opt/intel/oneapi/setvars.sh meson setup --native-file meson/icpx.ini build meson compile -C build ``` Steal from the Dockerfile. Or if you're on Fedora or Arch or whatever, send me a PR to update these instructions once you figure out the right steps to get it working for you.