diff --git a/AGENTS.md b/AGENTS.md index 878216b..c03aa23 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -69,11 +69,11 @@ without `--gcc-install-dir=/usr/lib/gcc/x86_64-linux-gnu/14`. Inside meson it's handled; for any ad-hoc icpx compile (a device-info probe, etc.) you **must pass `--gcc-install-dir` explicitly** or you get cryptic link errors. -Targets: `qxmx_run`, `qxmx_run_naive` (A/B), `qxmx_diff`, `qxmx_tok`, +Targets: `qxmx_run`, `qxmx_diff`, `qxmx_tok`, `qxmx_info`, `qxmx_ref_run`, plus the `tests/` harnesses (`fa_kernel_test`, `attn_forward_batch_test`, `ssm_batch_test`, `ffn_batch_test`, `elementwise_batch_test`, `gemm_engine_test`, `gemm_tune`, `fa_fused_bench`, -`jm_probe`, …). +…). ## Commands @@ -149,8 +149,7 @@ visible. Runtime knobs: `QXMX_CHUNK` (chunk size, default 512), 32-wide tiles are listed but throw "undefined builtin" / "not supported" — instantiating them poisons the whole program build at submit time. Always validate a tile by running it inside a try/catch; never trust - `matrix_combinations` alone. (See memory `qxmx_joint_matrix_b70.md`, - `tests/jm_probe.cpp`.) + `matrix_combinations` alone. (See memory `qxmx_joint_matrix_b70.md`.) - **Inline-asm DPAS in plain SYCL is BLOCKED.** vISA JIT rejects the asm dialect ("parsing vISA inline assembly failed: syntax error, unexpected IDENT"). Obsoleted by `sycl::joint_matrix` — do not revive. diff --git a/README.md b/README.md new file mode 100644 index 0000000..a6ab000 --- /dev/null +++ b/README.md @@ -0,0 +1,41 @@ +# qxmx + +This project contains an experimental single-purpose inference engine with one goal: running the [Bonsai-27B ternary-quantized model](https://prismml.com/news/bonsai-27b) as fast as possible on Intel Arc GPU hardware. + +I have an Arc Pro B70. Once I discovered that Intel's GPU hardware natively supports a 2-bit integer format (and AMD's Radeon Pro R9700 does not!), one which happens to map perfectly in-hardware to the ternary values `[-1, 0, +1]`... I wanted to see how hard it would be to bootstrap a new inference engine from scratch using modern LLMs (mainly GLM-5.2 and Kimi-K3). KV cache has user-selectable K quantization (`fp16`, `fp8`, and `q8_0`), but V quantization is currently hardcoded to 4-bit TurboQuant. There's an OpenAI-chat-completions-compatible endpoint built into `qxmx_serve`. + +DSpark speculative decoding is supported by the model but not yet by this software. Nor is any vision functionality currently implemented. + +One thing I've learned is that despite being compatible with a wide variety of open models, llama.cpp is *really* well-optimized and hard to beat! + +I was inspired by https://github.com/antirez/ds4/ and https://github.com/JustVugg/colibri/ but don't blame them! None of this is their fault. + +## download the model + +You'll need [Ternary-Bonsai-27B-Q2_g64.gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf/blob/main/Ternary-Bonsai-27B-Q2_g64.gguf) specifically. In the future it may make more sense to use `hf download` to grab the appropriate files once we support speculative decoding or vision features, but for now, just this one GGUF file is all we need. + +## the easy path + +You can build the Dockerfile and then run `docker --device /dev/dri ...` to pass your Intel GPU through (right now only one GPU is supported but multiple Arc GPU support is planned for a future release). That's probably the easiest way to get it running locally. An example command: + +```bash +docker run --rm --name qxmx_serve --device /dev/dri -p 8080:8080 -v ~/models:/models:ro qxmx -- ./build/qxmx_serve --host 0.0.0.0 --port 8080 --slots 2 --ctx-per-slot 131072 --ctk q8_0 /models/Ternary-Bonsai-27B-Q2_g64.gguf +``` + +This will get you a local web server that speaks something similar to the OpenAI `/chat/completions` endpoint at `http://localhost:8080/v1/chat/completions`, and you can see the single model at `http://localhost:8080/v1/models`. + +Tool-calling in [maki](https://maki.sh) seems to work. Speed is acceptable on a B70, but there's definitely room for further performance optimizations. + +## building from source + +The author has only tested this on Debian Trixie with some packages from testing/unstable. + +You'll need to ensure you have run `apt install meson ninja-build pkg-config libstdc++-14-dev`, and then you'll have to set up the oneapi apt repo and install the `intel-oneapi-dpcpp-cpp-2026.1` package from it. Once that's all done, you should be able to run: + +```bash +source /opt/intel/oneapi/setvars.sh +meson setup --native-file meson/icpx.ini build +meson compile -C build +``` + +Steal from the Dockerfile. Or if you're on Fedora or Arch or whatever, send me a PR to update these instructions once you figure out the right steps to get it working for you.