Continuous batching and aggregate throughput on MoE
Parent: Mac local LLMs: Serving ops and multi-model · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Aggregate-vs-concurrency curves for MoE under mlx-lm BatchGenerator or llama-server slots remain unpublished in sources read; the existing open question stands.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Aggregate-vs-concurrency curves for MoE under mlx-lm BatchGenerator or llama-server slots remain unpublished in sources read; the existing open question stands. [source]
- No source read in this batch adds an MoE aggregate-throughput measurement beyond those already held in continuous-batching-on-mlx.md and moe-active-parameter-decode-on-unified-memory.md. [source]
- A third-party guide lists vllm-mlx for multi-user serving at "2-3.4x scaling" without model or hardware, so it adds no usable MoE data. [source]
Children
- No children recorded.