A single-model (Bonsai-27B) inference engine for Intel Arc GPUs

attn: native::exp in decode softmax (6.7% attn at 33k ctx) master

sycl::exp -> sycl::native::exp for the two exp() calls per token in the FlashDecoding scan's online-softmax recurrence. Inputs are bounded softmax diffs (score - m_new, max_s - m_new) -- the same numerical class the M1.7c fused FA kernel (qxmx_fa_fused.cpp:284) already uses native::exp for. Measured on B50 at 33k context (QXMX_PROFILE=2): attn 122.8 -> 114.6 ms (-6.7%), total decode step 186.7 -> 178.7 ms. ~12.7M exp calls/step at that depth, so the cheaper intrinsic matters on the scan's critical path. Correctness: qxmx_diff mean|diff|=0.1064 (was 0.1111 on this B50; B70 reference 0.108) -- slightly better, gate green, greedy argmax unchanged.


+9 -2
1 changed file