A single-model (Bonsai-27B) inference engine for Intel Arc GPUs

M1.7b.12: naive register-recurrence DeltaNet -- 591 tok/s (+18% over WY/UT) master

Port of llama.cpp's gated_delta_net algorithm (ggml/src/ggml-sycl/ gated_delta_net.cpp): the HD-row state column lives in REGISTERS (per-thread private array) instead of SLM, and the per-token recurrence is a scalar loop -- no matmuls, no SLM, no WY/UT decomposition, no UT solve. src/qxmx_deltanet_naive.cpp: gpu_deltanet_batched now dispatches to the naive register recurrence. Layout: one WG per v-head (NVH=48), LWS=HD=128 threads (one thread per state column); each thread holds its column's HD=128 state rows in a private float Sd_col[HD]. Same fp op order as the decode oracle (gpu_deltanet_rmsnorm): L2-normalize q,k per key head via group reduce, per-token decay+kv_mem+delta+rank-1-update+attn (per-thread, no reduce since each thread owns its full column), folded group_rmsnorm over the HD output values. State loaded from global once, written once. tests/deltanet_naive_test.cpp: validates gpu_deltanet_batched vs the decode per-token oracle across 8 configs (chunk 1/13/32/64/100/128/256/512, zero/ rand s0, slow/fast/mixed decay). Bit-exact: abs ~3e-6 (tighter than WY/UT's ~0.01 -- naive uses the same scalar op order as the oracle, no tf32 reassociation). Built against the naive engine source set. meson.build: deltanet_naive_test target added (naive source set). Result (bench/session_4mb.txt, 5289 tok, 5 runs each): wyut: 10.57 s = 500 tok/s (SSM wyut phase 2.9 ms/layer) naive: 8.95 s = 591 tok/s (SSM naive phase 1.2 ms/layer, -59%) +18% prefill throughput. Output unchanged (' Paris. The'). Per-chunk breakdown (M=256): ssm 49%->34%, ffn 37%->43%, attn 12%->22% (rel shares shift as SSM shrank). All tests pass on both targets. Next levers: bump chunk to 512+ (naive scales; WY/UT can't -- the UT solve is O(chunk^2)), and tune the FFN/attention GEMMs which are now the larger shares. llama-bench ub=512 hit 810 tok/s -- chunk size is the next win.