attn: v3 FD-in-WG decode scan (16.4 tok/s at 63k ctx, beats llama.cpp scan) master
Zero-barrier scan: each of the WG's 8 SGs scans its own strided token stream with SG-local online-softmax partials (fold kernel: 2 barriers/ token = 49.7% XVE idle from barrier convergence at 1 WG/EU). 8-way SG merge through SLM once per WG (same math as attn_fd_merge). 3-head register split (acc[3][8]=24 regs): one WG per (kv-head, head-group, split) -> 512 WGs = 2/EU. V dequant 1x per token (v1 vec: 8x redundant). Supersedes the v1 vec attempt (register spill 13KB -> 320B; dynamic ALU instructions now 46% of fold's; VTune-verified). Measured (B70, q8_0 K, QXMX_PROFILE=2): 63.7k ctx: attn 41.8 -> 28.3 ms, TG 74.3 -> 61.0 ms (13.4 -> 16.4 tok/s) ~3.4k ctx: attn 5.9 -> 5.4 ms llama.cpp SYCL fattn-vec reference point: ~30 ms attn at 65k with fp16 KV (2.3x more KV bytes than our q8+TurboQuant4) -- our scan is now ahead. Gates: fa_kernel_test 40/40 PASS; qxmx_diff mean|d| = 0.1048 (gate 0.108, best yet). fp16-vs-q8_0 K A/B at 65k: 0.7% (null -- K format is not the scan bottleneck). Co-Authored-By: Kimi-K3 <kimi-k3@moonshot.ai>