# Model routing Model routing is the practice of directing different tasks to different AI models based on the task's requirements. Instead of using one model for everything, a routing layer selects the most appropriate model considering capability, cost, speed, and context needs. This is the [[Receptionist AI Design Pattern]] applied at the model level. A simple classification (or even a smaller, faster model) decides whether a task needs a powerful reasoning model like [[GPT-5.6|GPT-5.6 Sol]] or [[Claude Fable 5]] or can be handled by a faster, cheaper model like Claude Haiku. [[AI Subagents]] naturally implement this: the parent agent uses a premium model while subagents can run on lighter models for routine tasks like code review or file exploration. Routing criteria: - **Task complexity**: simple lookups vs. multi-step reasoning - **Latency requirements**: real-time responses vs. background processing - **Cost sensitivity**: high-volume tasks benefit from cheaper models - **Context length**: some tasks need large [[Context Window|context windows]], others don't - **Specialization**: some models excel at code, others at creative writing or analysis [[OpenRouter]] and [[AI Gateway]] solutions provide infrastructure for model routing, offering a unified API across multiple providers with automatic fallback, load balancing, and cost optimization. Dedicated routers take this further by deciding per request rather than per application: [[Not Diamond]] predicts the right model for each input without sitting in the request path, [[Ramp Router]] reads coding-agent conversation signals to escalate only on hard turns, and [[Requesty]] and [[Cheaper Inference]] rank routes by cost. [[Martian]] attacked the underlying question directly, treating model prediction as an interpretability problem. The trade-off is routing accuracy. Misrouting a complex task to a cheap model produces bad output. Misrouting a simple task to an expensive model wastes money. Getting this right requires [[AI Observability]] to track quality per model per task type. ## Routing before vs checking after Routing has a sibling that's easy to confuse with it: the [[AI Model Cascades|model cascade]]. A router decides *before* any model answers ("this looks hard, send it to the big model"). A cascade sends everything to the cheap model first and escalates *after* checking the answer. Routers bet on predicting difficulty; cascades bet on detecting mistakes. RouteLLM (LMSYS, 2024) is the reference router paper: trained on Chatbot Arena preferences, it cut costs by over 85% on MT Bench against GPT-4 alone while keeping 95% of its quality. FrugalGPT (Stanford, 2023) is the reference cascade. You can combine both. ## Routing with a decision model The router itself doesn't have to be an LLM. A [[Decision Models (DMs)|decision model]] like [[Jev]] answers "which handler?" in ~100 ms for a fraction of a cent, and returns a probability you can gate on. TypeSafe's intent routing pattern asks two questions in one call (a Choice over the intent, a Score for complexity) and routes in code: - intent confidence < 0.5 → human - `order_status` → a plain database lookup, no LLM at all - `product_question` / `return_exchange` → two different specialist LLMs, each with its own context - `complaint` → human if complexity is high *or* the model isn't sure about the complexity; otherwise a complaint-resolution LLM Two details I find worth copying. First, the router doesn't only choose between models; it also picks plain code or a person when that's the right handler. Second, confidence is checked on *every* question the routing depends on, not only the main one. See [[Confidence-Gated Routing]] for the threshold logic and [[Speculative Fan-Out]] for why asking both questions at once costs almost nothing. ## References - [Intent routing (TypeSafe docs)](https://docs.typesafe.ai/patterns/intent-routing) - [Ong et al., RouteLLM (2024)](https://arxiv.org/abs/2406.18665) - [Chen, Zaharia, Zou, FrugalGPT (2023)](https://arxiv.org/abs/2305.05176) ## Related - [[Receptionist AI Design Pattern]] - [[Not Diamond]] - [[Martian]] - [[Ramp Router]] - [[Requesty]] - [[Cheaper Inference]] - [[LiteLLM]] - [[Portkey]] - [[AI Subagents]] - [[AI Agent Orchestration]] - [[AI Gateway]] - [[OpenRouter]] - [[Large Language Models (LLMs)]] - [[Context Window]] - [[AI Observability]] - [[AI Model Cascades]] - [[Confidence-Gated Routing]] - [[Speculative Fan-Out]] - [[Decision Models (DMs)]] - [[Jev]] - [[AI Agent Routing]]