LLM Model Routing, Cascades & Mixture-of-Agents

Parent: AI Agent Ecosystems · researched 2026-06-01T00:04:23.191Z· 18 sources · 11 concepts · skill llm-routing-cascades

Reference under the ai-agent-engineering hub. The multi-model serving-decision layer — choosing/orchestrating WHICH model(s) answer each request to ride the cost/quality/latency Pareto frontier — dist

LLM Model Routing, Cascades & Mixture-of-Agents

Children

Frontier under this node: Cost/quality/latency Pareto modeling, Failure modes (routing collapse, tail miscalibration), Mixture-of-Agents (MoA layered proposers + aggregator), Model cascades + deferral/abstention (FrugalGPT, threshold calibration), Output ensembling & fusion (LLM-Blender PairRanker + GenFuser), Predictive routing (RouteLLM router taxonomy, PGR/CPT), Route-by-difficulty / complexity (Route-to-Reason, RADAR, think-vs-non-think), Router evaluation & benchmarks (RouterBench, AIQ, RouterArena), Routing tooling landscape (OpenRouter, LiteLLM, NotDiamond, Martian, vLLM Semantic Router), Semantic / prompt caching as a routing layer (GPTCache), Speculative cascades (token-level deferral, vs speculative decoding)

← the whole tree · 3D view· how to read this page