A single-model (Bonsai-27B) inference engine for Intel Arc GPUs

ssm_timing_test: add naive (gpu_deltanet_batched) timing line master

Notes: the naive impl dispatches via gpu_deltanet_batched (the meson-selected entry), so ssm_timing_test -- which links the wyut source set -- reports the wyut time for that line, not the naive time. Kept as a placeholder; a true naive isolated measurement needs a naive-set timing target. The in-pipeline SSM phase time (QXMX_PROFILE) is the authoritative number: 1.1 ms/layer. Tried + reverted M1.7b.14 (hoisting q,k L2-norm to a separate pre-kernel to eliminate the 3x GQA redundancy): net neutral -- the pre-kernel's launch + 16-key-head reduce work costs as much as the saved 32 redundant in-kernel L2-norms. Not worth the complexity. The in-kernel L2-norm stays.


+1
1 changed file