macOS 27 TensorOps FP8 FP4 MX scale planes
Parent: Mac local LLMs: Speed, bandwidth and prefill · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Data plane: an `MTLTensorDescriptor` with a quantized `dataType`; the session's example is `MTLTensorDataTypeMetalFloat8E4M3` (FP8 E4M3) with `usage = MTLTensorUsageCompute`.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Data plane: an `MTLTensorDescriptor` with a quantized `dataType`; the session's example is `MTLTensorDataTypeMetalFloat8E4M3` (FP8 E4M3) with `usage = MTLTensorUsageCompute`. [source]
- Scales plane: an `MTLTensorAuxiliaryPlaneDescriptor` with `dataType = MTLTensorDataTypeMetalFloat8UE8M0` and `blockFactors = {32, 1}`, registered in an `MTLTensorAuxiliaryPlaneDescriptorMap` under `MTLTensorPlaneTypeScales`, then attached as `tensorDesc.auxiliaryPlanes`. The data, scales and metadata are packed into one tensor object. Each scale element covers a 32x1 block of the data plane (32 elements along the first listed dimension, which the example orders as {NumCols, NumRows}). [source]
- Shader side: `tensor<device metal_fp8_e4m3_format, dextents<int, 2>, tensor_handle, scales_plane>` where `scales_plane = tensor_blockwise<tensor_plane_scales, device metal_fp8_ue8m0_format, 32, 1>`. Swapping the tag `tensor_handle` for `tensor_inline` builds a temporary tensor on the shader stack from buffer pointers. [source]
- Tiling: `slice` on a quantized tensor slices the data and scales planes together according to the block size; the `matmul2d_descriptor` and op setup are identical to non-quantized tensors and TensorOps handles dequantization. Apple's guidance is to feed quantized data straight into TensorOps so it can use any available acceleration. [source]
- Custom formats: if a format is not native, dequantize into a cooperative tensor (registers) and pass it as a matmul input (supported since 26.3), avoiding the threadgroup-memory round trip of dequantizing to f16 in threadgroup memory. [source]
- Flash-attention pieces in 27: `reduce_rows` with a max reduction and initial -INFINITY into a row-reduction destination cooperative tensor; `map_iterator` to map between the 2-D tensor and the reduction tensor; `get_left_input_cooperative_tensor<float, half, float>(ctQK)` to reuse the QK result as the left input of the second matmul; `is_compatible_as_left_input` / `is_compatible_as_right_input` must be checked because cooperative tensor layouts differ by data type. [source]
- Alignment: the new data types have additional alignment requirements compared with 16-bit types; Apple tells developers to check the Metal documentation. [source]
- TensorOps data types by release (tech talk 111432): bfloat in 26.1, cooperative tensors as matmul input in 26.3 ("custom dequantization routines inside your kernel"), 4/8-bit integer tensors in 26.4, fp4/fp8/int2 plus MX scale planes in 27. The 26.4 SDK also added the "M or N a multiple of 16" static_assert in the existing dossier. [source]
- Format match: the documented scale plane is E8M0 with 32x1 blocks, which matches MLX `mxfp4` and `mxfp8` (group size 32, E8M0 scales) but not MLX `nvfp4` (group size 16, per the group-size-16 note in MLX split-K code); NVFP4's E4M3 scale and 16-element block are not shown in the session. Whether macOS 27 accepts an E4M3 scale plane or a block factor of 16 is not stated in the sources read. [source]
- No `metal_fp8`, `tensor_blockwise` or `ue8m0` reference appears in the fetched MLX device, matmul, quantized or SDPA sources, so MLX NAX kernels today dequantize in-kernel rather than using native MX tensors. [source]
- Affine (scale plus bias, group sizes 32/64/128) quantization, the dominant MLX and GGUF-K style format, has no described scale-plane equivalent; it would use the custom-dequant cooperative-tensor path. [source]
- Pre-release status: the session is a WWDC26 preview of 27; names such as `MTLTensorDataTypeMetalFloat8UE8M0` may change before release. [source]
- None found. Apple says llama.cpp, MLX and PyTorch already use the Neural Accelerators; no source shows any of them using the 27 native MX tensors yet. [source]
- Hardware acceleration scope: the session says TensorOps uses "any available hardware acceleration" but does not say whether FP8/FP4/MX run natively on M5 Neural Accelerators or only on a future chip; no rate figures given. [source]
- Which fp4 variant (E2M1) naming and alignment constants the API uses; not shown. [source]
- Whether MLX or llama.cpp will move quantized matmul to native MX planes and what that does to decode bandwidth; none in sources read. [source]
- macOS and iOS 27 extend TensorOps data types to 4- and 8-bit floating point and 2-bit integer, after 4- and 8-bit integer support in 26.x. [source]
- In macOS and iOS 27 a single MTLTensor can hold quantized data plus a scales plane in FP8 E8M0 block-wise scale-factor format. [source]
- Each element of the scale plane applies to a block of elements in the data plane, set by `blockFactors`. [source]
- The session's host example uses `MTLTensorDataTypeMetalFloat8E4M3` for the data plane and `MTLTensorDataTypeMetalFloat8UE8M0` with blockFactors {32, 1} for the scales plane. [source]
- The scales plane is registered with `MTLTensorPlaneTypeScales` in an `MTLTensorAuxiliaryPlaneDescriptorMap` assigned to `tensorDesc.auxiliaryPlanes`. [source]
- In MSL the MXFP8 tensor type is `tensor<device metal_fp8_e4m3_format, dextents<int, 2>, tensor_handle, scales_plane>` with `tensor_blockwise<tensor_plane_scales, device metal_fp8_ue8m0_format, 32, 1>`. [source]
- Replacing `tensor_handle` with `tensor_inline` creates the tensor on the shader stack from buffer pointers. [source]
- Calling `slice` on quantized tensors slices the data and scales planes together according to the block size. [source]
- Setting up `matmul2d` with quantized tensors is identical to normal tensors and TensorOps handles dequantization. [source]
- For a custom format, dequantizing into a cooperative tensor and passing it to `matmul2d` avoids a threadgroup-memory round trip. [source]
- The new MX and E8M0 types have additional alignment requirements versus larger data types. [source]
- macOS 27 allows a cooperative tensor as a direct matmul input via `get_left_input_cooperative_tensor`, with `is_compatible_as_left_input` to check layout compatibility. [source]
- TensorOps provides `reduce_rows` (e.g. max with initial -INFINITY) and `map_iterator` for flash-attention softmax on cooperative tensors. [source]
- In macOS 26.3 cooperative tensors became valid matmul inputs, enabling custom dequantization inside a kernel. [source]
- The existing MLX NAX code paths contain no reference to `metal_fp8`, `tensor_blockwise` or `ue8m0`. [source]
- MLX nvfp4 uses group size 16, which does not match the documented 32x1 E8M0 scale block. [source]
Children
- No children recorded.