Example client for the highly experimental divepool embedding firehose

match M1 Metal compute: bf16 weights, float32 GEMM master

MLX on M1 stores weights in bf16 but computes all matmul in float32 — Metal's simdgroup_matrix uses the FP32 ALU pipeline and MLX kernels hardcode AccumType=float. Load weights in bf16 for identical rounding, then upcast entire model to float32 for inference. Cosine similarity improves from ~0.99 to ~0.995+ at 768d. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>


+18 -8
1 changed file