<!-- llms-explorer concept facts · https://llms-explorer.com/tree/continuous-batching-and-aggregate-throughput-on/ · pack 2026-10-05 · ~273 tokens -->

# Continuous batching and aggregate throughput on MoE

> Aggregate-vs-concurrency curves for MoE under mlx-lm BatchGenerator or llama-server slots remain unpublished in sources read; the existing open question stands.

Parent: [Mac local LLMs: Serving ops and multi-model](https://llms-explorer.com/tree/mac-local-llms-serving-ops-and-multi-model/) · 1 facets · 3 facts · page: https://llms-explorer.com/tree/continuous-batching-and-aggregate-throughput-on/

## Facts

- Aggregate-vs-concurrency curves for MoE under mlx-lm BatchGenerator or llama-server slots remain unpublished in sources read; the existing open question stands. — source: `asserted`
- No source read in this batch adds an MoE aggregate-throughput measurement beyond those already held in continuous-batching-on-mlx.md and moe-active-parameter-decode-on-unified-memory.md. — source: `asserted`
- A third-party guide lists vllm-mlx for multi-user serving at "2-3.4x scaling" without model or hardware, so it adds no usable MoE data. — [source](https://blog.starmorph.com/blog/apple-silicon-llm-inference-optimization-guide)
