XPerf:让模型路由从“选最强模型”走向“预测谁最适合当前问题”XPerf: From Predicting Model Performance to Better Model Routing
作者:王宇彤 Yutong Wang
研究时间:2026 年 9 月起
状态:持续迭代中
当一个产品可以同时调用多个大语言模型时,“哪个模型最强”其实不是最有用的问题。不同模型在代码、数学、知识问答、长文本理解和 Agent 任务上的优势并不完全相同,而调用成本、延迟和输出长度也各有差异。
因此,更实际的问题是:
面对当前这个具体请求,哪个模型最适合回答它?
XPerf 是 Telluvian 路由系统中的模型表现预测组件。它接收用户的 prompt,并为每个候选模型预测预期表现,例如:
| 候选模型 | 预期表现 |
|---|---|
| Model A | 0.80 |
| Model B | 0.60 |
| Model C | 0.47 |
XPerf 并不直接负责整个调用决策。成本、延迟、缓存、输出长度和产品策略仍由外层系统处理;它首先试图把一个更基础的问题做好:同一个 prompt 对不同模型而言,成功的可能性分别有多大?
整个项目也是围绕这个问题逐渐展开的。最初,我们先建立最简单的 router baseline;随后加入 embedding 和语言模型内部 activation;再逐步审计、扩展数据并改进评测方式。随着实验推进,研究问题也从“能否预测模型表现”逐渐变成了更困难的一件事:预测得准之后,为什么仍然不一定选得对?
一个 Router 怎样才算真的有用?
模型路由很容易被一个看起来不错的总体准确率误导。假如一个 router 能达到 85% accuracy,但始终调用同一个最强模型就能达到 87%,那么这个 router 虽然“会预测”,却没有创造真正的选模价值。
因此,项目最开始并没有急着训练复杂模型,而是先建立一套之后一直沿用的评测框架。
| 评测方式 | 它回答的问题 |
|---|---|
| Best single model | 如果永远固定选择一个模型,最好能做到什么程度? |
| Oracle | 如果事后知道所有模型的真实表现,理论上还能提升多少? |
| Prompt-hash clean split | 面对同一类任务中的新问题,是否还能泛化? |
| Benchmark holdout | 面对训练时完全没见过的任务类型,是否还能迁移? |
| Per-benchmark / headroom | 路由收益来自哪些任务?数据中还剩多少选模空间? |
其中,Best single model 是最重要的现实基线。如果 router 无法稳定超过它,就意味着动态选模暂时还没有比“永远调用一个强模型”更有价值。
Oracle 则是另一个极端。它无法部署,因为它假设我们事先知道每个模型在每道题上的真实得分,但它可以告诉我们数据本身究竟有没有值得学习的模型选择空间。后续几乎所有 XPerf 实验,都建立在这套评测逻辑上。
从 Embedding Router 到 Activation Probe
第一轮实验采用的是相对简单、容易复现的路由方法,包括 embedding kNN、Ridge,以及 MiniLM、BGE、MPNet 等通用文本 embedding。
这类方法的逻辑很直接:先把 prompt 映射成一个向量,再利用历史上语义相近问题的模型表现推测当前应该选择哪个模型。它们是很重要的 baseline,因为如果一个简单的 frozen embedding 已经足够完成路由,那么引入更昂贵的内部 activation 就没有太大必要。
在此基础上,我进一步尝试了 activation probe。这里不再只使用模型最终输出的文本 embedding,而是直接读取一个 proxy language model 在处理 prompt 时产生的 hidden states。背后的假设是:模型内部状态可能不只编码“这道题在说什么”,还包含题目难度、推理结构和其他与最终模型表现相关的信息。
LLMRouterBench:第一次看到明显的 Router 信号
在 LLMRouterBench 的完整 activation dump 上,我训练了一个 shared-trunk MLP。不同候选模型共用一部分表示网络,再分别预测各自的表现。
验证集最终选择出的配置为 Layer 52,使用最后 32 个 token 的 mean pooling,hidden size 为 64,dropout 为 0.5,并对三个随机 seed 的结果进行集成。
| Clean test 指标 | 结果 |
|---|---|
| Shared-trunk router accuracy | 64.54% |
| Best single model | 58.63% |
| 绝对提升 | +5.91 个百分点 |
这是一个重要的早期结果。至少在相同 benchmark 集合中的未见 prompt 上,activation probe 已经能够超过固定最佳模型,说明内部表征中确实存在可用于模型选择的信号。
但如果只看 64.54% 这个数字,很容易对结果过于乐观。真正的问题出现在更严格的 benchmark holdout 中:
| Benchmark holdout 检查 | 结果 |
|---|---|
| Leave-one-benchmark-out router accuracy | 56.34% |
| 同一设置下强固定模型 | 58.44% |
| Within-headroom 相对表现 | −24.17% |
| Validation BCE 最低点 | 约第 3 epoch |
训练过程中,training loss 持续下降,但 validation loss 大约在第三个 epoch 后开始反弹。与此同时,一旦把整个 benchmark 留到测试阶段,router 的表现便跌到固定模型以下。
这说明模型在熟悉的任务分布中确实学到了有用规律,但其中相当一部分规律可能与 benchmark 本身绑定得太紧。
换句话说,模型可能已经很会判断:
“这种 benchmark 通常哪个模型比较强。”
但距离真正理解:
“这一个具体 prompt 为什么更适合某个模型。”
还有明显差距。
数据比想象中更重要:SPROUT 与 Epoch
这一结果让研究重点逐渐从“继续调模型结构”转向了数据本身。
Routing dataset 有一个很容易被忽视的问题:记录数很多,并不意味着问题足够多样。 十万次模型回答可能只是几千个 prompt 的重复采样,而一个 router 真正需要学习的是不同任务、不同能力和不同模型之间的关系。
因此,在继续训练之前,我先对 SPROUT 和 Epoch 两个数据源做了系统审计。
SPROUT:先确认这里有没有值得路由的空间
SPROUT 共包含 43,288 条记录和 14 个候选模型。完成数据检查并基于 prompt text 重建 clean hash split 后,clean test 上出现了一个非常重要的结果:
| 项目 | 结果 |
|---|---|
| 总记录数 | 43,288 |
| 候选模型数 | 14 |
| 官方 split 跨集合重复 prompt | 108 |
| Best single model:o3-mini | 80.91% |
| Oracle | 95.00% |
| 可路由 headroom | 14.10 个百分点 |
Oracle 比最佳固定模型高出 14.10 个百分点,说明这个数据集里确实存在明显的模型选择空间。与此同时,官方 split 中有 108 个 prompt 跨集合重复,因此后续采用了基于 prompt text 的 clean hash split,避免相同问题同时出现在训练和测试中。
目前 SPROUT 已经完成数据审计与训练准备,但尚未重新进行 activation extraction 和新的 probe 训练。
Epoch:64 万次回答,其实只有 3,340 个问题
Epoch 数据库的规模看起来更大:
| 审计维度 | 结果 |
|---|---|
| Attempts | 642,752 |
| 独立 prompt | 3,340 |
| Runs | 1,067 |
| 原始模型标识 | 277 |
| Prompt × model 聚合记录 | 189,506 |
| 可匹配 frontier 模型 | 25 |
| Ungraded attempts | 361 |
| 仅一次 scored attempt 的 prompt × model 对 | 66.42% |
这里最值得注意的是第一行和第二行之间的差异。642,752 次 attempts 实际只对应 3,340 个独立 prompt,因此不能简单把几十万条回答当成几十万种不同的训练问题。
重复 attempts 仍然很有价值,它们可以帮助估计同一个模型在同一个问题上的表现分布和不确定性;但它们不能替代 benchmark 和任务类型上的真正多样性。
审计中还确认了几个数据处理细节:N grade 代表 non-credit,而不是缺失值;361 条 ungraded attempts 需要排除;66.42% 的 prompt × model 组合只有一次评分;此外,SWE-Bench 中还有 Claude Code 与 Codex 的 run-level benchmark provenance 混淆问题。Epoch 与 SPROUT 则没有发现 exact prompt overlap。
这部分工作虽然没有直接产生新的 router accuracy,却为后续训练建立了更可靠的数据基础。
升级 Routing Dataset:从预测表现到跨模型 Calibration
新的 routing dataset 到位后,我完成了 GPU activation extraction,并首先选择拥有完整 10-model coverage 的 prompt 进行训练。
| 数据项目 | 结果 |
|---|---|
| 原始 prompt × model 记录 | 45,409 |
| 独立 prompt | 6,027 |
| 完整 10-model coverage pool | 3,904 |
| 内部 3-model subset | 2,123 |
| 提取的 proxy layers | 23 |
只使用 3,904 个完整覆盖的 prompt,是为了确保所有候选模型拥有相同的训练机会,避免缺失标签本身影响模型比较。训练、选层与最终测试也严格分离:训练集用于拟合,validation 用于选择 layer,held-out test 只在最后报告。
最终验证选择 Layer 44:
| Held-out test 指标 | 结果 |
|---|---|
| XPerf router | 86.22% |
| Best fixed model:GPT-6 Astra | 86.87% |
| Oracle | 90.64% |
| 平均 per-model AP | ≈ 0.949 |
| Dataset-prior AP | ≈ 0.822 |
这一轮实验出现了一个非常关键的矛盾。
平均 per-model AP 已经达到约 0.949,远高于 dataset prior。这说明 probe 对“某个模型会不会成功”拥有很强的预测能力。
但最终 router accuracy 只有 86.22%,仍略低于固定使用 GPT-6 Astra 的 86.87%。
能够分别预测每个模型的表现,并不意味着这些预测已经能够被直接横向比较。
这也第一次把 XPerf 的核心问题明确地指向 cross-model calibration:不同模型的预测概率可能各自很准确,但它们的尺度未必完全一致,因此简单取最大值并不一定会选出最合适的模型。
Shared Trunk:不同模型之间确实存在共享结构
随后,我在 V2 interim 数据上比较了独立 heads 和 shared-trunk architecture。
| 方法 | Held-out accuracy |
|---|---|
| Independent V2 heads | 89.71% |
| Shared-trunk V2 | 91.43% |
| Best fixed:Gemini 3.8 Flash | 92.57% |
| Oracle | 94.86% |
Shared trunk 比独立 heads 提高了 1.71 个百分点,说明不同候选模型的表现预测之间存在可以共享的结构,而不是 17 个完全互不相关的问题。
不过,91.43% 仍没有超过 92.57% 的最佳固定模型。也就是说,共享表征让 prediction 变得更好,但“更好的 prediction”依然没有完全转化成“更好的 routing”。
这一阶段同时完成了 V2 router checkpoint、deployment manifest、ordered output-model mapping,以及 layer、pooling、prompt formatting 和 normalisation 等部署说明,并保存了 routing with / without layer swap、score histogram、top-1 frequency 和 correlation matrix 等诊断结果。这样,XPerf 不再只是一个 notebook 中的实验,而开始具备被产品侧加载、复现和检查的完整部署接口。
把“表现”和“成本”拆开:XPerf、XLength 与 XCost
随着研究继续推进,系统边界也变得更加清晰。XPerf 最适合专注于预测模型表现,而成本和输出长度应该作为独立信号处理,而不是全部塞进一个 router 中。
因此,后续增加了 XCost 与 XLength。XLength 预测模型完成当前请求时可能产生多少 completion tokens,为外层成本估计提供输入;XCost 则用于比较不同路由策略与简单 prompt-length heuristic 的质量—成本关系。
整个架构可以简单表示为:
XPerf → 预期回答质量
XLength → 预期输出长度
↓
外层策略 → 综合质量、成本、延迟与缓存做最终选择这种拆分的好处是职责清晰。XPerf 不需要自己预测最终账单,XLength 也不需要判断答案质量;产品层可以根据不同使用场景决定如何权衡这些信号。
Activation 的优势,只是因为它“看得更长”吗?
早期 embedding baseline 有一个很明显的潜在不公平因素:MPNet 的输入长度大约只有 384 tokens,而 activation probe 可以看到更长的 prompt。
因此,activation 表现更好可能有两种解释:一种是 hidden states 本身真的包含更多与模型表现有关的信息;另一种只是 activation 看到了更多文本。
V6 专门用来区分这两个解释。
数据集包含 7,313 个 prompt、17 个模型目标,以及固定的 train / validation / test buckets。Embedding baseline 换成了 Qwen3-Embedding-8B,它拥有 32K context window,并能够在不截断的情况下编码全部 7,313 个 prompt。除此以外,masked-BCE objective、width-64 shared trunk 和三 seed validation selection 等训练设置保持一致。
最终对比如下:
| Held-out test BCE ↓ | Qwen activation | Qwen3-Embedding-8B |
|---|---|---|
| Pooled BCE | 0.2497 | 0.3278 |
| Macro-by-benchmark BCE | 0.2973 | 0.3295 |
Activation probe 在 11 个 benchmark 中的 10 个上取得更低 BCE。
这说明此前看到的优势并不能简单归因于 context window。即便 embedding model 能完整看到长 prompt,Qwen3.8-27B Layer 50 activation 仍然提供了更强的模型表现预测信号。
从目前的实验看,普通 embedding 很擅长描述“prompt 是什么”,而语言模型内部 activation 似乎还能保留一些与“模型面对这个 prompt 会表现怎样”有关的额外结构。
Prediction 已经很强,Router 到底还缺什么?
走到这里,XPerf 已经可以相当不错地预测不同模型的表现,但研究中开始出现一个越来越明显的可能性:模型主要学到了两类信号。
第一类是 Prompt Difficulty——这道题整体有多难;第二类是 Global Model Quality——某个模型平均而言有多强。
这两种信息足以产生非常好的 per-model prediction。例如,当一个非常困难的问题出现时,所有模型的成功概率可能同时降低;简单问题则可能让所有模型的预测一起升高。
问题在于,router 真正需要的不只是这些共同变化。
它还需要知道:
在这一个具体 prompt 上,哪个模型会相对于自己的平均水平表现得更好?
这就是 prompt-specific residual。
一个模型可能整体比另一个模型强,但在代码、数学、事实检索或特定推理结构上,相对优势可能发生变化。真正值得路由的部分,恰恰来自这种 prompt-specific 的相对差异。
下一步:把“预测得准”变成“选得对”
接下来的实验会直接检验上面的假设。
首先是 Residual Target。与其让模型直接预测每个候选模型的原始成功率,可以改为预测它相对于当前 prompt 平均表现的偏差。这样做的目的,是先削弱所有模型共同拥有的 difficulty signal,让模型把更多容量用于学习候选模型之间真正不同的部分。
另一个重要实验是建立一个非常简单、可解释的 Additive Baseline:
它表达的是一个非常强的 null hypothesis:也许根本不需要学习复杂的“模型—问题匹配关系”,只要知道题目有多难,再知道每个模型平均有多强,就已经可以解释大部分预测结果。
如果复杂的 activation probe 与这个简单 baseline 表现接近,就说明当前模型主要学习到的仍然是共享 difficulty 和 global model quality。只有当 residual model 能稳定超过这个 baseline,才更有证据说明模型开始捕捉真正的 prompt-specific routing signal。
后续还会继续测试 per-model calibration、pairwise ranking、regret minimisation 和带不确定性的 expected utility,并通过 leave-one-benchmark-out 评测观察这些信号在新任务上的迁移能力。这里的目标不再只是继续把 BCE 降低一点,而是让训练目标更直接地服务于最终路由决策。
为什么数据多样性可能比模型结构更重要?
目前几个阶段的实验共同指向了同一个问题:router 的瓶颈很可能不只是模型结构,而是训练数据本身。
如果训练数据只覆盖少数 benchmark,一个很强的 probe 完全可能学会识别 benchmark,然后记住“这个 benchmark 通常哪个模型表现最好”。这种策略在随机 clean split 中可以获得不错结果,却很容易在 benchmark holdout 中失效。
因此,下一阶段的数据扩展不会单纯追求更多 attempts,而会更关注 benchmark 和任务类型的多样性、frontier model coverage、run-level provenance、重复测量以及能够支持 leave-one-benchmark-out 的数据结构。Epoch、SPROUT 和后续 routing datasets 也会逐步标准化到统一的 prompt / model / score 格式,同时保留来源和置信度等信息。
这也是为什么前面的数据审计并不是项目中的辅助工作,而是 router 能否真正泛化的核心部分。
从 Research Probe 到真正可用的 Router
回顾整个项目,XPerf 最初的问题其实很简单:
我们能不能根据一个 prompt 预测不同模型的表现?
目前的实验已经说明,这件事是可以做到的,而且 activation probe 能够提供相当强的预测信号。与此同时,我们也逐渐建立起了 embedding baseline、activation extraction、shared-trunk training、clean split、benchmark holdout、SPROUT 与 Epoch audit、V2/V6 router、XLength、deployment manifest 和更稳定的评测流程。
但真正接近产品的问题比最初的 prediction task 更难:
如何把对模型表现的预测,转化成稳定、可泛化、真正优于固定模型的选模能力?
这意味着下一阶段的重点不会只是继续优化一个单独的 loss,也不会只是让 per-model AP 再高一点。更重要的是确认 XPerf 是否真正捕捉到了不同模型面对同一个 prompt 时的相对优势,并且这种优势能否迁移到训练时没有见过的新任务上。
最终,XPerf 想回答的不是“哪个模型总体最强”。而是:
面对这一次具体请求,哪个模型最适合回答它?
Author: Yutong Wang
Research period: September 2026 onwards
Status: Ongoing
When a product can call several large language models, asking which model is "the strongest" is not necessarily the most useful question. Different models have different strengths across coding, mathematics, knowledge retrieval, long-context tasks and agentic workloads, while cost, latency and output length vary as well.
A more useful question is:
For this particular request, which model is most likely to give the best answer?
XPerf is Telluvian's model-performance prediction component. Given a user prompt, it estimates the expected performance of each candidate model. For example:
| Candidate model | Expected performance |
|---|---|
| Model A | 0.80 |
| Model B | 0.60 |
| Model C | 0.47 |
XPerf does not make the entire serving decision on its own. Cost, latency, caching, output length and product-level preferences are handled separately by the surrounding system. Its job is narrower: for the same prompt, how likely is each candidate model to perform well?
The project has gradually evolved around that question. We began with simple routing baselines, then introduced text embeddings and internal model activations, before moving on to dataset auditing, larger routing datasets and stricter evaluation. As the experiments progressed, the problem became more specific: if we can already predict model performance reasonably well, why does that not always translate into better model selection?
Establishing a Reliable Routing Baseline
Routing results can look impressive without necessarily being useful. A router that achieves 85% accuracy, for example, is not particularly valuable if always choosing the same strong model already achieves 87%.
For that reason, the project began by establishing a set of evaluation baselines that have remained useful throughout the work.
| Evaluation | What it tells us |
|---|---|
| Best single model | How well can we do by always selecting the same model? |
| Oracle | How much routing headroom exists if we knew the true result in advance? |
| Prompt-hash clean split | Does the router generalise to unseen prompts from familiar task distributions? |
| Benchmark holdout | Does it transfer to task types that were absent from training? |
| Per-benchmark / headroom analysis | Where does the routing gain come from, and how much room for improvement remains? |
The best single model is the most important practical baseline. If a router cannot consistently beat it, dynamic model selection has not yet demonstrated an advantage over simply calling one strong model every time.
The oracle sits at the other extreme. It cannot be deployed because it assumes we already know how every model performed on each prompt, but it tells us whether the dataset contains meaningful routing headroom in the first place. This evaluation framework became the basis for nearly all subsequent XPerf experiments.
From Embedding Routers to Activation Probes
The first round of experiments used relatively simple and reproducible routing methods: embedding kNN, Ridge regression, and general-purpose text embeddings such as MiniLM, BGE and MPNet.
The idea is straightforward. A prompt is converted into a vector representation, and model performance on semantically similar historical prompts is then used to estimate which model is likely to work well on the current request. These methods provide an important baseline: if a frozen text embedding is already sufficient for routing, there may be little reason to introduce a more expensive activation-based system.
The next step was to test an activation probe. Instead of relying only on a final text embedding, this approach reads the hidden states produced by a proxy language model while processing the prompt. The working hypothesis is that these internal representations may encode not just what the prompt is about, but also aspects of difficulty, reasoning structure and other information related to downstream model performance.
LLMRouterBench: The First Clear Routing Signal
On the full LLMRouterBench activation dump, I trained a shared-trunk MLP in which the candidate models share part of the representation network before producing their own performance predictions.
The validation-selected configuration used Layer 52, mean pooling over the final 32 tokens, a hidden size of 64, dropout of 0.5, and an ensemble across three random seeds.
| Clean test metric | Result |
|---|---|
| Shared-trunk router accuracy | 64.54% |
| Best single model | 58.63% |
| Absolute improvement | +5.91 percentage points |
This was an important early result. On unseen prompts drawn from the same set of benchmarks, the activation-based router outperformed the best fixed model, suggesting that the internal representation contained useful model-selection signal.
The picture changed under stricter benchmark-level evaluation.
| Benchmark holdout check | Result |
|---|---|
| Leave-one-benchmark-out router accuracy | 56.34% |
| Strong fixed model under the same setting | 58.44% |
| Within-headroom relative performance | −24.17% |
| Lowest validation BCE | Around epoch 3 |
Training loss continued to fall, while validation loss started to rise after roughly the third epoch. At the same time, once an entire benchmark was held out from training, router performance fell below the fixed-model baseline.
This suggests that the model had learnt useful structure within familiar task distributions, but that some of this structure was too tightly tied to the benchmarks themselves.
In practical terms, the model appeared to be learning something closer to:
"Models of this type usually perform well on this benchmark."
rather than fully learning:
"This particular prompt is better suited to this particular model."
That distinction became increasingly important in the later stages of the work.
Data Matters More Than It First Appears: SPROUT and Epoch
These results shifted the focus away from simply tuning model architecture and towards the training data itself.
Routing datasets have an easy-to-miss limitation: a large number of records does not necessarily mean a large amount of task diversity. Hundreds of thousands of model responses may still come from only a few thousand distinct prompts, whereas a useful router needs to learn differences across tasks, capabilities and models.
Before doing more training, I therefore audited two important data sources: SPROUT and Epoch.
SPROUT: Establishing Whether There Is Enough Routing Headroom
SPROUT contains 43,288 records across 14 candidate models. After checking data quality and rebuilding a clean split based on prompt text, the clean test set showed a substantial gap between the best fixed model and the oracle.
| Item | Result |
|---|---|
| Total records | 43,288 |
| Candidate models | 14 |
| Duplicate prompts crossing the official splits | 108 |
| Best single model: o3-mini | 80.91% |
| Oracle | 95.00% |
| Routing headroom | 14.10 percentage points |
The oracle outperforming the best fixed model by 14.10 percentage points is important because it shows that the dataset contains genuine model-selection opportunity.
The audit also found 108 prompts duplicated across the official splits. Subsequent preparation therefore uses a prompt-text-based clean hash split to prevent the same prompt from appearing in both training and evaluation.
SPROUT has now been audited and prepared for training, although a new activation extraction and probe-training pass has not yet been run.
Epoch: 642,752 Attempts, but Only 3,340 Distinct Prompts
The Epoch database looks considerably larger at first glance:
| Audit dimension | Result |
|---|---|
| Attempts | 642,752 |
| Distinct prompts | 3,340 |
| Runs | 1,067 |
| Raw model identifiers | 277 |
| Prompt × model aggregates | 189,506 |
| Matched frontier models | 25 |
| Ungraded attempts | 361 |
| Prompt × model pairs with only one scored attempt | 66.42% |
The most important comparison is between the first two rows. The 642,752 attempts correspond to only 3,340 distinct prompts, so the number of responses cannot be treated as equivalent to the number of distinct training problems.
Repeated attempts are still useful because they provide information about variation and uncertainty in a model's performance on the same prompt. They are not, however, a substitute for broader benchmark and task diversity.
The audit also clarified several labelling issues. An N grade means non-credit rather than a missing label; 361 ungraded attempts should be excluded; 66.42% of prompt × model pairs have only a single scored attempt; and the SWE-Bench data contained a run-level provenance issue that mixed Claude Code and Codex benchmark labels. No exact prompt overlap was found between Epoch and SPROUT.
This stage did not directly produce a higher router score, but it gave the later training work a much more reliable data foundation.
Expanding the Routing Dataset: From Performance Prediction to Cross-Model Calibration
Once the newer routing dataset was available, I completed GPU activation extraction and began by training only on prompts with complete coverage across all 10 candidate models.
| Data item | Result |
|---|---|
| Raw prompt × model records | 45,409 |
| Distinct prompts | 6,027 |
| Complete 10-model coverage pool | 3,904 |
| Internal 3-model subset | 2,123 |
| Extracted proxy layers | 23 |
Restricting the first experiment to the 3,904 fully covered prompts ensured that every candidate model had the same training opportunities and prevented missing labels from distorting the comparison. Training, model selection and final evaluation were also kept separate: the training set was used for fitting, validation for selecting the layer, and the held-out test set only for the final report.
Layer 44 was selected on validation.
| Held-out test metric | Result |
|---|---|
| XPerf router | 86.22% |
| Best fixed model: GPT-6 Astra | 86.87% |
| Oracle | 90.64% |
| Mean per-model AP | ≈ 0.949 |
| Dataset-prior AP | ≈ 0.822 |
This experiment exposed one of the most important tensions in the project.
Mean per-model AP was already around 0.949, substantially above the dataset prior. In other words, the probe had become very good at predicting whether an individual model would succeed.
Yet the final router accuracy was 86.22%, slightly below the 86.87% achieved by always using GPT-6 Astra.
Accurately predicting each model in isolation does not necessarily mean that those predictions are directly comparable across models.
This brought cross-model calibration into focus. Individual probability estimates may each be accurate while still operating on slightly different scales, meaning that simply taking the largest prediction does not always select the best model.
Shared Trunk: Learning Structure Across Candidate Models
The next experiment used the V2 interim dataset to compare independent prediction heads with a shared-trunk architecture.
| Method | Held-out accuracy |
|---|---|
| Independent V2 heads | 89.71% |
| Shared-trunk V2 | 91.43% |
| Best fixed: Gemini 3.8 Flash | 92.57% |
| Oracle | 94.86% |
The shared trunk improved accuracy by 1.71 percentage points over the independent heads. This suggests that performance prediction across candidate models contains shared structure rather than being 17 entirely separate prediction problems.
However, 91.43% still remained below the 92.57% fixed-model baseline. Better prediction had again not fully translated into better routing.
This stage also moved the work closer to deployment. I prepared the V2 router checkpoint, deployment manifest, ordered output-model mapping, and documentation covering layer selection, pooling, prompt formatting and normalisation. Diagnostic outputs were also saved for routing with and without layer swapping, along with score histograms, top-1 frequency and correlation matrices.
At this point, XPerf was no longer just a notebook experiment: it had a reproducible model package that could be loaded and inspected by the product side.
Separating Performance from Cost: XPerf, XLength and XCost
As the system developed, its component boundaries became clearer. XPerf is best kept focused on model performance, while cost and output length can be treated as separate signals rather than folded into a single router.
This led to the addition of XCost and XLength. XLength predicts how many completion tokens a model is likely to generate for a given request, providing an input to the product layer's cost estimate. XCost evaluates the quality–cost trade-off of different routing strategies and compares them with simpler heuristics such as prompt length.
The overall design can be summarised as:
XPerf → Expected answer quality
XLength → Expected output length
↓
Outer policy → Final decision across quality, cost,
latency and cachingThe advantage of this separation is that each component has a clear responsibility. XPerf does not need to estimate the final API bill, while XLength does not need to judge answer quality. The surrounding product logic can decide how those signals should be weighted for a particular use case.
V6: Testing the Value of Activations with Long-Context Embeddings
The earlier embedding baseline had a potential disadvantage: MPNet could only see roughly 384 tokens, whereas the activation probe had access to much longer prompts.
That left two possible explanations for the activation advantage. The hidden states might genuinely contain richer information about model performance, or they might simply perform better because they had access to more of the prompt.
V6 was designed to separate these two possibilities.
The dataset contains 7,313 prompts, 17 model targets and fixed train, validation and test buckets. For a fairer representation comparison, the embedding baseline was replaced with Qwen3-Embedding-8B, which supports a 32K context window and can encode all 7,313 prompts without truncation. The masked-BCE objective, width-64 shared trunk and three-seed validation selection were otherwise kept consistent.
The final comparison was:
| Held-out test BCE ↓ | Qwen activation | Qwen3-Embedding-8B |
|---|---|---|
| Pooled BCE | 0.2497 | 0.3278 |
| Macro-by-benchmark BCE | 0.2973 | 0.3295 |
The activation probe achieved lower BCE on 10 of the 11 benchmarks.
This makes it difficult to explain the earlier advantage purely in terms of context length. Even when the embedding model can see the full long prompt, Layer 50 activations from Qwen3.8-27B still provide a stronger signal for predicting downstream model performance.
At least on the current experiments, general-purpose embeddings appear to be good at representing what a prompt is about, while the proxy model's internal activations appear to retain additional structure related to how models are likely to perform on that prompt.
Current Diagnosis: Difficulty, Model Quality and Prompt-Specific Residuals
By this stage, XPerf had become reasonably strong at predicting individual model performance. The remaining evidence, however, suggests that a large part of this predictive power may come from two relatively simple signals.
The first is Prompt Difficulty: how difficult the request is overall.
The second is Global Model Quality: how strong a particular model is on average.
These two signals can already explain strong per-model prediction. On a difficult prompt, the predicted success probabilities for all models may fall together; on an easy prompt, they may all rise together.
A useful router needs more than that. It also needs to capture:
whether a particular model is likely to perform better or worse than its usual level on this specific prompt.
This is the prompt-specific residual.
One model may be stronger overall, but relative performance can still vary across coding, mathematics, factual retrieval or particular reasoning structures. Those prompt-specific changes are where much of the real routing value should come from.
Next Step: Turning Predictive Performance into Routing Advantage
The next round of experiments is designed to test this interpretation directly.
The first direction is a Residual Target. Instead of predicting each candidate model's raw success probability, the model can predict its deviation from the average performance on that prompt. This removes some of the difficulty signal that is shared across all candidate models and forces the predictor to focus more directly on where the models differ.
The second experiment is a deliberately simple and interpretable Additive Baseline:
It represents a strong null hypothesis: perhaps there is no need to learn a complicated prompt–model matching function at all. Knowing how difficult the prompt is, together with the average quality of each model, may already explain most of the predictive performance.
If a sophisticated activation probe performs similarly to this simple baseline, that would suggest that most of the current signal comes from shared difficulty and global model quality. If a residual model can consistently outperform it, there would be stronger evidence that the model is beginning to capture genuine prompt-specific routing signal.
Future work will also continue to examine per-model calibration, pairwise ranking, regret minimisation and uncertainty-aware expected utility, alongside leave-one-benchmark-out evaluation for measuring transfer to new task distributions. The aim is no longer simply to reduce BCE a little further; it is to make the training objective more closely reflect the final routing decision.
Data Diversity and Cross-Task Generalisation
Several stages of the project now point towards the same conclusion: the main bottleneck may not be model architecture alone, but the diversity of the training data.
When a training set covers only a small number of benchmarks, a sufficiently powerful probe can learn to recognise the benchmark and associate it with whichever model tends to perform best there. That can work well under a random clean split while failing badly once an entirely new benchmark is held out.
The next stage of data expansion will therefore focus less on simply collecting more attempts and more on broadening benchmark and task diversity, improving frontier-model coverage, preserving run-level provenance, retaining repeated measurements and constructing datasets that support meaningful leave-one-benchmark-out evaluation.
Epoch, SPROUT and future routing datasets will gradually be standardised into a common prompt / model / score format while retaining source and confidence information. Data auditing is therefore not a side task in this project; it is central to whether the router can ultimately generalise.
From Research Probe to Deployable Router
Looking back over the project, the original XPerf question was relatively simple:
Can we predict how different models will perform from the prompt alone?
The product-level problem is harder:
How do we turn model-performance predictions into model choices that are stable, generalisable and genuinely better than a fixed-model baseline?
The next stage is therefore not just about optimising another loss function or pushing per-model AP slightly higher. The more important question is whether XPerf can capture the relative advantage between candidate models on the same prompt, and whether that advantage survives on tasks that were absent from training.
XPerf is not ultimately trying to answer:
"Which model is strongest overall?"
It is trying to answer:
"For this request, which model is the right one to call?"