MLX 0.32.2 tensor-unit expert-sorted quantized matmul row-limit bug on M5
Parent: Mac local LLMs: Speculative decoding and MTP · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Is there an upstream ml-explore/mlx issue or fix PR for the row-limit bug? Searches returned none.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Is there an upstream ml-explore/mlx issue or fix PR for the row-limit bug? Searches returned none. [source]
- The concept is already covered by mtplx-flash-decoding-verify-attention-kernel-sdp.md (32,767-row limit, 19.4% of cold prompt lengths above 3,277 tokens, 2.12.1 padding workaround, MLX 0.32.2 pin); a search of the cache and the web found no source that adds to or contradicts it. [source]
Children
- No children recorded.