<!-- llms-explorer concept facts · https://llms-explorer.com/tree/mlx-0-32-2-tensor-unit-expert-sorted-quantized-m/ · pack 2026-10-05 · ~214 tokens -->

# MLX 0.32.2 tensor-unit expert-sorted quantized matmul row-limit bug on M5

> Is there an upstream ml-explore/mlx issue or fix PR for the row-limit bug? Searches returned none.

Parent: [Mac local LLMs: Speculative decoding and MTP](https://llms-explorer.com/tree/mac-local-llms-speculative-decoding-and-mtp/) · 1 facets · 2 facts · page: https://llms-explorer.com/tree/mlx-0-32-2-tensor-unit-expert-sorted-quantized-m/

## Facts

- Is there an upstream ml-explore/mlx issue or fix PR for the row-limit bug? Searches returned none. — source: `asserted`
- The concept is already covered by mtplx-flash-decoding-verify-attention-kernel-sdp.md (32,767-row limit, 19.4% of cold prompt lengths above 3,277 tokens, 2.12.1 padding workaround, MLX 0.32.2 pin); a search of the cache and the web found no source that adds to or contradicts it. — source: `asserted`
