Something went wrong. Try again.
Example client for the highly experimental divepool embedding firehose
Something went wrong. Try again.
match M1 Metal compute: bf16 weights, float32 GEMM master
MLX on M1 stores weights in bf16 but computes all matmul in float32 — Metal's simdgroup_matrix uses the FP32 ALU pipeline and MLX kernels hardcode AccumType=float. Load weights in bf16 for identical rounding, then upcast entire model to float32 for inference. Cosine similarity improves from ~0.99 to ~0.995+ at 768d. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Author jraedisch Co-author Claude Opus 4.6 (1M context) Date (Apr 3, 2026, 1:13 AM UTC) Commit cab5478d cab5478d1c5016d87914f174fee9214394ac4242 Parent fae2c151 fae2c1515f76c12d77b69c163db5e0fe11654075