Routing Across the Model Frontier
Why routing matters and why it is a runtime decision
The model landscape has changed quickly. Production systems can now choose across providers, model families and price points, each with different strengths in reasoning, coding, context length, latency, reliability and cost.
That makes a single default model increasingly restrictive. Always use the strongest model and you overpay for routine work. Always use a cheaper one and you lose quality on harder tasks. A frontier model may be worth it for a difficult coding problem and wasted on classification or extraction.
Model routing resolves this by choosing a model for each request. A decision layer weighs expected performance against cost, latency and availability. Simple requests stay on efficient models, hard ones are escalated, and other providers act as fallbacks.
| Candidate model | Expected performance | Relative cost | Typical use |
|---|---|---|---|
| Efficient model | 76% | Low | Routine, high-volume work |
| Balanced model | 84% | Medium | General-purpose requests |
| Frontier model | 89% | High | Difficult or quality-sensitive tasks |
The highest-scoring model doesn't always need to win. If a cheap model is already very likely to succeed, paying much more for a small improvement adds little. A leaderboard asks which model is strongest on average. A router asks which model best fits this request.
The economic case
The cost gap between models can be substantial. For a request with 2,000 input tokens and 500 output tokens, at OpenAI's list prices:
| Model | Input price / 1M tokens | Output price / 1M tokens | Approx. request cost |
|---|---|---|---|
| GPT-5 nano | $0.05 | $0.40 | ~$0.0003 |
| GPT-5 | $1.25 | $10.00 | ~$0.0075 |
That's a 25× difference for the same request within one model family.
For a question like "What is the capital of Peru?", both models will get it right. Performance is equal, so the cheaper model is the better choice. On a hard reasoning task, the performance gap may be large enough to justify the extra cost. That is the basic economics of routing: spend more only when the extra capability is likely to change the outcome.
The core technical problem: predicting before generation
The router has to make its choice before it knows how any candidate model would answer. If we could run all candidates, score every response and pick the best, routing would be easy, but we would already have paid the cost and latency of running them all.
Routing is therefore a prediction problem. For a prompt and a candidate model, we want to estimate performance before generation takes place. A useful decomposition is:
Predicted performance ≈ Prompt difficulty + Global model quality + Prompt-specific model fit
The prompt difficulty and global model quality terms account for a large amount of variation, but we must also include prompt-specific model fit. Model A may be stronger overall while Model B is unusually well suited to one coding task or query pattern. Catching those variations is what makes routing truly prompt-aware.
At Telluvian, we split the problem into two estimates: XPerf, how likely each model is to succeed on the next turn, and XCost, what it would cost to serve that turn.
XPerf
XPerf estimates the probability that each candidate model will successfully answer the next request in a conversation.
| Candidate model | Predicted success |
|---|---|
| Model A | 82% |
| Model B | 71% |
| Model C | 46% |
An 82% XPerf score means that, before generation, we estimate roughly an 82% chance the model will answer the next request successfully under the evaluation definition being used.
XPerf looks across the entire messages array: earlier user messages, previous answers and the new request. A follow-up such as "Can you rewrite the second option?" may be trivial in isolation but impossible to interpret without the earlier conversation, so the prediction has to be based on the conversation state, not just the final message.
XPerf is independent of cost. Pricing, cache state and other operating constraints are handled elsewhere in the routing policy.
Why activations, not embeddings
Embeddings are the obvious baseline. They are cheap to compute and easy to store. Moreover, they naturally capture a request's topic, vocabulary, and similarity to other requests. If an embedding could predict which model will perform best, there would be little reason to run a proxy language model just to extract its hidden states.
But semantic similarity is not the same as routing relevance. Two requests on the same subject can place very different demands on a model. One may need straightforward extraction, another multi-step reasoning, tool use, careful instruction following or long-context retrieval.
Activations are produced while a language model processes the request itself. Those internal representations may reflect properties that predict model behaviour, such as task difficulty, uncertainty, novelty and reasoning demand, as well as topic.
We compared activation features with a long-context embedding model that processed the same prompts without truncation. Activations performed better on almost all benchmarks (see the research log). Our hypothesis is that activations preserve information specifically relevant to how language models will perform, which is exactly the signal a router needs.
XCost
In a multi-turn conversation, cost depends on more than list price. It also depends on which models have already processed which parts of the session.
Consider a conversation that has moved between models. Model A handled the first turn. Model B then received the conversation so far and generated the second answer. When the third message arrives, the candidates are in very different cache states.

| Candidate model | Cached context by Turn 3 | Additional context to process | Predicted success | Expected serving cost |
|---|---|---|---|---|
| Model A | Q1 + A1 | Q2 + A2 + Q3 | 74% | Medium |
| Model B | Q1 + A1 + Q2 + A2 | Q3 | 78% | Low |
| Model C | None | Full conversation | 95% | High |
XCost estimates:
Expected cost = uncached input cost + cached input cost + expected output cost
Almost all of this is known at routing time: prices are fixed, conversation length is known, and we track what each model has cached. The only real unknown is the length of the response.
Here Model B is cheapest because it has most of the conversation cached. Model C is far more likely to succeed but has to process everything from scratch. A warm cache makes staying put cheaper, but it shouldn't lock a conversation in if another model is much more likely to succeed.
Turning predictions into a decision
XPerf and XCost are predictions. Choosing between them is a policy, and the right one depends on the product.
| Model | XPerf | XCost |
|---|---|---|
| A | 87% | Low |
| B | 89% | Medium |
| C | 91% | High |
A high-volume consumer product may pick A. A quality-sensitive research tool may pick C. A latency-sensitive app may rule one out entirely. Conceptually:
Utility(model) = XPerf − Cost penalty − Latency penalty − Reliability penalty + Product preferences
The weights are product decisions. Keeping them out of the predictor means it can be improved without baking in commercial preferences.
The separation also makes decisions legible. If a request goes to a cheap model, we can see whether that's because the predicted gap was small, the cache was warm, or latency ruled out the alternatives. When a result is poor, we can tell whether the model, the prediction or the policy was at fault.
Where we go next
Workload-specific routing. A legal assistant, a coding agent and a support tool send very different requests and value different things: careful reasoning, speed, tool use, consistent formatting. Given examples from one application, a router can learn where particular models are unusually strong or cheap for that workload, instead of relying on a global ranking.
Latency as part of the request. Some requests need an answer now: live support, interactive assistants, in-product coding. Others can wait: batch processing, background research, long-running agents. Requests that can wait open up batching, scheduling and cheaper capacity. This is the idea behind services like Sail Research, which trade latency tolerance for lower inference cost. Routing can apply the same principle at the decision layer.
The goal is a router that asks, for every turn:
- How likely is each model to handle this well?
- What will this turn cost, given the session's cache state?
- How quickly is the answer needed?
- Which trade-off fits this workload?
Routing as a layer in the AI stack
The model ecosystem keeps getting more varied, and that variety only helps if applications can use it. Routing makes model choice a runtime decision, so new models and providers become options without rebuilding around them.
At Telluvian, our aim is to improve performance while reducing cost by sending each request to the most appropriate model: efficient models where they're enough, stronger ones where they add value, and the cache and latency taken into account throughout.