A single-model (Bonsai-27B) inference engine for Intel Arc GPUs
Something went wrong. Try again.
ssm_timing_test: add naive (gpu_deltanet_batched) timing line master
Notes: the naive impl dispatches via gpu_deltanet_batched (the meson-selected entry), so ssm_timing_test -- which links the wyut source set -- reports the wyut time for that line, not the naive time. Kept as a placeholder; a true naive isolated measurement needs a naive-set timing target. The in-pipeline SSM phase time (QXMX_PROFILE) is the authoritative number: 1.1 ms/layer. Tried + reverted M1.7b.14 (hoisting q,k L2-norm to a separate pre-kernel to eliminate the 3x GQA redundancy): net neutral -- the pre-kernel's launch + 16-key-head reduce work costs as much as the saved 32 redundant in-kernel L2-norms. Not worth the complexity. The in-kernel L2-norm stays.
Author Chris Lee Date (Jul 19, 2026, 6:02 PM UTC) Commit 47dee175 47dee17526ea6d02c5d3b3fdb0e44bcdb871a616 Parent 918d4d0a 918d4d0a4363726c5700a76bcc64534de7c8102b