# Model routing
Model routing is the practice of directing different tasks to different AI models based on the task's requirements. Instead of using one model for everything, a routing layer selects the most appropriate model considering capability, cost, speed, and context needs.
This is the [[Receptionist AI Design Pattern]] applied at the model level. A simple classification (or even a smaller, faster model) decides whether a task needs a powerful reasoning model like [[GPT-5.6|GPT-5.6 Sol]] or [[Claude Fable 5]] or can be handled by a faster, cheaper model like Claude Haiku. [[AI Subagents]] naturally implement this: the parent agent uses a premium model while subagents can run on lighter models for routine tasks like code review or file exploration.
Routing criteria:
- **Task complexity**: simple lookups vs. multi-step reasoning
- **Latency requirements**: real-time responses vs. background processing
- **Cost sensitivity**: high-volume tasks benefit from cheaper models
- **Context length**: some tasks need large [[Context Window|context windows]], others don't
- **Specialization**: some models excel at code, others at creative writing or analysis
[[OpenRouter]] and [[AI Gateway]] solutions provide infrastructure for model routing, offering a unified API across multiple providers with automatic fallback, load balancing, and cost optimization.
Dedicated routers take this further by deciding per request rather than per application: [[Not Diamond]] predicts the right model for each input without sitting in the request path, [[Ramp Router]] reads coding-agent conversation signals to escalate only on hard turns, and [[Requesty]] and [[Cheaper Inference]] rank routes by cost. [[Martian]] attacked the underlying question directly, treating model prediction as an interpretability problem.
The trade-off is routing accuracy. Misrouting a complex task to a cheap model produces bad output. Misrouting a simple task to an expensive model wastes money. Getting this right requires [[AI Observability]] to track quality per model per task type.
## Routing before vs checking after
Routing has a sibling that's easy to confuse with it: the [[AI Model Cascades|model cascade]]. A router decides *before* any model answers ("this looks hard, send it to the big model"). A cascade sends everything to the cheap model first and escalates *after* checking the answer. Routers bet on predicting difficulty; cascades bet on detecting mistakes. RouteLLM (LMSYS, 2024) is the reference router paper: trained on Chatbot Arena preferences, it cut costs by over 85% on MT Bench against GPT-4 alone while keeping 95% of its quality. FrugalGPT (Stanford, 2023) is the reference cascade. You can combine both.
## Routing with a decision model
The router itself doesn't have to be an LLM. A [[Decision Models (DMs)|decision model]] like [[Jev]] answers "which handler?" in ~100 ms for a fraction of a cent, and returns a probability you can gate on. TypeSafe's intent routing pattern asks two questions in one call (a Choice over the intent, a Score for complexity) and routes in code:
- intent confidence < 0.5 → human
- `order_status` → a plain database lookup, no LLM at all
- `product_question` / `return_exchange` → two different specialist LLMs, each with its own context
- `complaint` → human if complexity is high *or* the model isn't sure about the complexity; otherwise a complaint-resolution LLM
Two details I find worth copying. First, the router doesn't only choose between models; it also picks plain code or a person when that's the right handler. Second, confidence is checked on *every* question the routing depends on, not only the main one. See [[Confidence-Gated Routing]] for the threshold logic and [[Speculative Fan-Out]] for why asking both questions at once costs almost nothing.
## References
- [Intent routing (TypeSafe docs)](https://docs.typesafe.ai/patterns/intent-routing)
- [Ong et al., RouteLLM (2024)](https://arxiv.org/abs/2406.18665)
- [Chen, Zaharia, Zou, FrugalGPT (2023)](https://arxiv.org/abs/2305.05176)
## Related
- [[Receptionist AI Design Pattern]]
- [[Not Diamond]]
- [[Martian]]
- [[Ramp Router]]
- [[Requesty]]
- [[Cheaper Inference]]
- [[LiteLLM]]
- [[Portkey]]
- [[AI Subagents]]
- [[AI Agent Orchestration]]
- [[AI Gateway]]
- [[OpenRouter]]
- [[Large Language Models (LLMs)]]
- [[Context Window]]
- [[AI Observability]]
- [[AI Model Cascades]]
- [[Confidence-Gated Routing]]
- [[Speculative Fan-Out]]
- [[Decision Models (DMs)]]
- [[Jev]]
- [[AI Agent Routing]]