MLX mixed fp16 and bf16 type promotion in arithmetic
Parent: Mac local LLMs: MLX kernels, numerics and internals · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
MLX's promotion table gives float32 for float16 combined with bfloat16, in either operand order.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- MLX's promotion table gives float32 for float16 combined with bfloat16, in either operand order. [source]
- The promotion table in `mlx/dtype.cpp` is documented in a source comment as following JAX type promotion rules. [source]
- Float16 combined with float16 gives float16, and bfloat16 combined with bfloat16 gives bfloat16. [source]
- Float16 or bfloat16 combined with any bool, uint8 to uint64, or int8 to int64 type keeps the half type. [source]
- Float16 or bfloat16 combined with float32 gives float32. [source]
- `add` computes `promote_types(a.dtype(), b.dtype())` and applies `astype` to both operands before broadcasting, so the primitive runs at the promoted type. [source]
- `subtract` uses the same `promote_types` pattern as `add`. [source]
- `result_type` over a list of arrays folds `promote_types` starting from bool, so n-ary ops promote the same way. [source]
- `TypeToDtype<double>` converts to float32 in `mlx/dtype.cpp`. [source]
- MLX's NumPy interop page says NumPy has no bfloat16, so bfloat16 arrays must be cast to float16 or float32 before `np.array`. [source]
- MLX maintainers declined adding fp8 dtypes in a December 2024 reply, citing that bf16 is already emulated on older machines and fp8 would be slower than fp32 and add library footprint. [source]
- In the Gemma 4 float16 guard, keeping `std_bias` in bf16 against fp16 hidden states makes the standardize subtraction run in float32, which avoids the fp16 overflow in that op but changes the output dtype to float32. [source]
Children
- No children recorded.