Overhead view of a dark rail yard where several tracks cross and diverge at a set of points
All posts
Yutong Wang

Routing Across the Model Frontier

Why routing matters and why it is a runtime decision

The model landscape has changed quickly. Production systems can now choose across providers, model families and price points, each with different strengths in reasoning, coding, context length, latency, reliability and cost.

That makes a single default model increasingly restrictive. Always use the strongest model and you overpay for routine work. Always use a cheaper one and you lose quality on harder tasks. A frontier model may be worth it for a difficult coding problem and wasted on classification or extraction.

Model routing resolves this by choosing a model for each request. A decision layer weighs expected performance against cost, latency and availability. Simple requests stay on efficient models, hard ones are escalated, and other providers act as fallbacks.

Candidate modelExpected performanceRelative costTypical use
Efficient model76%LowRoutine, high-volume work
Balanced model84%MediumGeneral-purpose requests
Frontier model89%HighDifficult or quality-sensitive tasks

The highest-scoring model doesn't always need to win. If a cheap model is already very likely to succeed, paying much more for a small improvement adds little. A leaderboard asks which model is strongest on average. A router asks which model best fits this request.

The economic case

The cost gap between models can be substantial. For a request with 2,000 input tokens and 500 output tokens, at OpenAI's list prices:

ModelInput price / 1M tokensOutput price / 1M tokensApprox. request cost
GPT-5 nano$0.05$0.40~$0.0003
GPT-5$1.25$10.00~$0.0075

That's a 25× difference for the same request within one model family.

For a question like "What is the capital of Peru?", both models will get it right. Performance is equal, so the cheaper model is the better choice. On a hard reasoning task, the performance gap may be large enough to justify the extra cost. That is the basic economics of routing: spend more only when the extra capability is likely to change the outcome.

The core technical problem: predicting before generation

The router has to make its choice before it knows how any candidate model would answer. If we could run all candidates, score every response and pick the best, routing would be easy, but we would already have paid the cost and latency of running them all.

Routing is therefore a prediction problem. For a prompt and a candidate model, we want to estimate performance before generation takes place. A useful decomposition is:

Predicted performance ≈ Prompt difficulty + Global model quality + Prompt-specific model fit

The prompt difficulty and global model quality terms account for a large amount of variation, but we must also include prompt-specific model fit. Model A may be stronger overall while Model B is unusually well suited to one coding task or query pattern. Catching those variations is what makes routing truly prompt-aware.

At Telluvian, we split the problem into two estimates: XPerf, how likely each model is to succeed on the next turn, and XCost, what it would cost to serve that turn.

XPerf

XPerf estimates the probability that each candidate model will successfully answer the next request in a conversation.

Candidate modelPredicted success
Model A82%
Model B71%
Model C46%

An 82% XPerf score means that, before generation, we estimate roughly an 82% chance the model will answer the next request successfully under the evaluation definition being used.

XPerf looks across the entire messages array: earlier user messages, previous answers and the new request. A follow-up such as "Can you rewrite the second option?" may be trivial in isolation but impossible to interpret without the earlier conversation, so the prediction has to be based on the conversation state, not just the final message.

XPerf is independent of cost. Pricing, cache state and other operating constraints are handled elsewhere in the routing policy.

Why activations, not embeddings

Embeddings are the obvious baseline. They are cheap to compute and easy to store. Moreover, they naturally capture a request's topic, vocabulary, and similarity to other requests. If an embedding could predict which model will perform best, there would be little reason to run a proxy language model just to extract its hidden states.

But semantic similarity is not the same as routing relevance. Two requests on the same subject can place very different demands on a model. One may need straightforward extraction, another multi-step reasoning, tool use, careful instruction following or long-context retrieval.

Activations are produced while a language model processes the request itself. Those internal representations may reflect properties that predict model behaviour, such as task difficulty, uncertainty, novelty and reasoning demand, as well as topic.

We compared activation features with a long-context embedding model that processed the same prompts without truncation. Activations performed better on almost all benchmarks (see the research log). Our hypothesis is that activations preserve information specifically relevant to how language models will perform, which is exactly the signal a router needs.

XCost

In a multi-turn conversation, cost depends on more than list price. It also depends on which models have already processed which parts of the session.

Consider a conversation that has moved between models. Model A handled the first turn. Model B then received the conversation so far and generated the second answer. When the third message arrives, the candidates are in very different cache states.

Diagram titled 'Why Session State Matters for Routing'. Turn 1: the user asks Question 1 and Model A generates Answer 1, so Model A's cache holds Question 1 and Answer 1. Turn 2: Question 2 is routed to Model B, which receives the conversation so far and generates Answer 2, so Model B's cache holds Question 1, Answer 1, Question 2 and Answer 2. Turn 3: the user asks Question 3 and the router must decide which model answers. By Turn 3, candidate models can have very different cache states.
Candidate modelCached context by Turn 3Additional context to processPredicted successExpected serving cost
Model AQ1 + A1Q2 + A2 + Q374%Medium
Model BQ1 + A1 + Q2 + A2Q378%Low
Model CNoneFull conversation95%High

XCost estimates:

Expected cost = uncached input cost + cached input cost + expected output cost

Almost all of this is known at routing time: prices are fixed, conversation length is known, and we track what each model has cached. The only real unknown is the length of the response.

Here Model B is cheapest because it has most of the conversation cached. Model C is far more likely to succeed but has to process everything from scratch. A warm cache makes staying put cheaper, but it shouldn't lock a conversation in if another model is much more likely to succeed.

Turning predictions into a decision

XPerf and XCost are predictions. Choosing between them is a policy, and the right one depends on the product.

ModelXPerfXCost
A87%Low
B89%Medium
C91%High

A high-volume consumer product may pick A. A quality-sensitive research tool may pick C. A latency-sensitive app may rule one out entirely. Conceptually:

Utility(model) = XPerf − Cost penalty − Latency penalty − Reliability penalty + Product preferences

The weights are product decisions. Keeping them out of the predictor means it can be improved without baking in commercial preferences.

The separation also makes decisions legible. If a request goes to a cheap model, we can see whether that's because the predicted gap was small, the cache was warm, or latency ruled out the alternatives. When a result is poor, we can tell whether the model, the prediction or the policy was at fault.

Where we go next

Workload-specific routing. A legal assistant, a coding agent and a support tool send very different requests and value different things: careful reasoning, speed, tool use, consistent formatting. Given examples from one application, a router can learn where particular models are unusually strong or cheap for that workload, instead of relying on a global ranking.

Latency as part of the request. Some requests need an answer now: live support, interactive assistants, in-product coding. Others can wait: batch processing, background research, long-running agents. Requests that can wait open up batching, scheduling and cheaper capacity. This is the idea behind services like Sail Research, which trade latency tolerance for lower inference cost. Routing can apply the same principle at the decision layer.

The goal is a router that asks, for every turn:

  • How likely is each model to handle this well?
  • What will this turn cost, given the session's cache state?
  • How quickly is the answer needed?
  • Which trade-off fits this workload?

Routing as a layer in the AI stack

The model ecosystem keeps getting more varied, and that variety only helps if applications can use it. Routing makes model choice a runtime decision, so new models and providers become options without rebuilding around them.

At Telluvian, our aim is to improve performance while reducing cost by sending each request to the most appropriate model: efficient models where they're enough, stronger ones where they add value, and the cache and latency taken into account throughout.