An open book held up in a dark room, with a thin band of warm light falling across its pages
All posts
Lysander Mawby

Interpretability and Efficiency

TL;DR: Through interpretability, we can understand how large models think. By reading their thoughts, we can extract useful work from the prefill stage of their inference rather than just the decode phase. As prefill is far cheaper, faster, and more optimised for current hardware, this lets us achieve astonishing efficiency across a whole new domain of LLM applications.

Introduction

Mechanistic interpretability, like any young scientific field, is poorly defined. It is less a bag of techniques than a goal: to understand how large pre-trained models work. Why do they generalise so well? What are their failure modes, and why? How can we prove they are aligned to the goals they claim?

Early mechanistic interpretability (mech interp) work focused on image models, the most visually intuitive models of their day. The convolutional maps of deep image classifiers show a clear pattern. Early layers do basic extraction like edge detection. Middle layers start to look like maps of faces or parts of whatever animal is being classified. Late layers do a relatively simple job of combining those learned representations. Looking at these maps turns an intractable problem into something intuitive.

Since the LLM era was kicked off circa 2022, model scale and complexity have exploded. Despite this, most of the large reasoning models we are most eager to understand still take the form of large transformer models.

Transformers have some unique properties which make them amenable to interpretation. Most notably, they are entirely built around an information bottleneck referred to as the residual stream, where many useful concepts can be read.

Transformers also have this curious property that they read far more quickly and cheaply than they write. The prefill stage, where a model reads text, is highly parallelisable and efficient on modern GPUs. The decode stage, on the other hand, where the model iteratively generates text, is notoriously inefficient.

A modern LLM might read 10,000 tokens/s, but only be able to write 50 tokens/s at economically viable batch sizes. Decode only gets remotely near prefill's speed on heavily specialised purpose-built chips, while prefill works on general-purpose hardware readily available today.

Ordinarily we can only use the written tokens, so we wait for slow, expensive decode before getting anything useful. But if we can understand a model's thoughts, we can do useful work without waiting for more tokens.

Transformers Overview

The transformer architecture is a very well covered topic. We encourage the reader to refer to any one of these excellent materials to learn more:

Overview

History

Inference

To see why interpretability brings such large efficiency gains, you only need two concepts: the residual stream and attention maps.

Residual Stream

Consider some text tt being processed by a language model. It is first tokenised into a list of NN tokens T=(T1,T2,⋯ ,TN)\mathbf{T} = \left( T_1, T_2, \cdots, T_N \right), where Ti∈{1,…,V}T_{i} \in \{1, \dots, V\} and VV is the vocabulary size.

A language model is then a function f(T)=logits∈RVf(\mathbf{T}) = \text{logits} \in \mathbb{R}^V, turning a list of tokens into logits giving the probability of each possible next token.

The tokens are first embedded by an embedding matrix WE∈RV×dmodelW_{E} \in \mathbb{R}^{V \times d_{\text{model}}}, where dmodeld_{\text{model}} is the model dimension. Looking up each token's row gives us the first residual stream state, h0=WE[T]∈RN×dmodelh_{0} = W_{E}[\mathbf{T}] \in \mathbb{R}^{N \times d_{\text{model}}}, with one row per token.

Each of the LL layers then looks much the same. It takes the previous residual stream state hℓ−1h_{\ell - 1} and applies two operations: multi-head attention (MHA), which moves information between tokens, and a standard multi-layer perceptron (MLP), which processes each token on its own. The result is the next state hℓh_{\ell}.

hℓ−1′=hℓ−1+MHA(hℓ−1)hℓ=hℓ−1′+MLP(hℓ−1′)\begin{aligned} h_{\ell - 1}' &= h_{\ell - 1} + \text{MHA}(h_{\ell - 1}) \\ h_{\ell} &= h_{\ell - 1}' + \text{MLP}(h_{\ell - 1}') \end{aligned}

So each token has a fixed-length vector that every layer adds to. Layer ℓ\ell only sees the state from layer ℓ−1\ell - 1. If a concept needs to be available across several layers, such as the language being spoken, the type of query, or the emotion being depicted, it has to be written to the residual stream and left there.

The residual stream is a global information bottleneck, sharing information across the model. By understanding how information is represented in it (and there has been real progress here, such as the linear representation hypothesis, sparse autoencoders and the Jacobian Lens), we should be able to read a model's thoughts just by looking at this tight bottleneck!

Attention Maps

Let us turn back to the multi-head attention (MHA) operation mentioned earlier. This is the only way in which token representations can interact with one another. In a transformer, without the attention operation, every residual stream state would be evolving entirely independently.

The exact mathematical description of the MHA operation isn't too important for us, but is stated below for clarity.

MHA(hℓ−1)=concati(softmax(QiKiTdhead)Vi)WO\text{MHA}(h_{\ell-1}) = \text{concat}_{i} \left( \text{softmax} \left( \frac{Q_{i}K_{i}^{T}}{\sqrt{d_{\text{head}}}} \right) V_{i} \right) W_{O}

where i∈{1,…,H}i \in \{1, \dots, H\} is the index of the head, Qi=hℓ−1WQ,i∈RN×dheadQ_i = h_{\ell-1}W_{Q,i} \in \mathbb{R}^{N \times d_{\text{head}}}, Ki=hℓ−1WK,i∈RN×dheadK_i = h_{\ell-1}W_{K,i} \in \mathbb{R}^{N \times d_{\text{head}}}, and Vi=hℓ−1WV,i∈RN×dheadV_i = h_{\ell-1}W_{V,i} \in \mathbb{R}^{N \times d_{\text{head}}} are the query, key, and value matrices respectively, and WO∈Rdmodel×dmodelW_{O} \in \mathbb{R}^{d_{\text{model}} \times d_{\text{model}}} is the output matrix.

What is important for us here is that softmax(QiKiTdhead)∈RN×N\text{softmax} \left( \frac{Q_{i}K_{i}^T}{\sqrt{d_{\text{head}}}} \right) \in \mathbb{R}^{N \times N} gives us a measure of how each token attends to each other token. It is a row-stochastic (by virtue of the softmax), lower-triangular (due to the causal masking of autoregressive language models) matrix showing exactly how tokens attend back to one another.

(Using these attention maps by themselves is by no means the only way of seeing how tokens attend to one another. See our other blog post on token-based attribution for prompt compression, which looks at which tokens in context are being attended to the least and removes them to cut costs while retaining quality, for more information on attribution methods.)

So we can, by combining these N×NN \times N matrices (perhaps by averaging them across heads and layers), get a sense of what tokens the model is thinking about when making sense of another token!

Prefill / Decode Disaggregation

Prefill is reading text. Every input token is processed in parallel as large matrix multiplications that modern GPUs handle incredibly well. Frontier models such as Kimi-K3 and GLM-5.3 can approach 14,000 tokens per second on an H100 or equivalent. The speed of prefill is why Claude or ChatGPT can process large code files so much faster than they could generate something of the same scale.

Decode is generating text. Each token depends on the preceding one, so generation is sequential. It also relies on vector-matrix multiplications that are poorly optimised for modern GPUs, leaving them largely idle. The result is slow and expensive.

The gap between prefill and decode is so large that the two stages often run on different hardware. Prefill runs on ordinary GPUs. Decode increasingly needs specialised chips like LPUs and Cerebras WSEs, which took years and a tapeout to build, and it still only approaches the speed prefill gets on commodity hardware.

A whole industry has popped up around the fact that we only know how to use a transformer's generated text, created during the inefficient decode stage. But many of the thoughts and insights that make that text valuable are produced while the transformer reads its input, during the cheap and efficient prefill stage. Interpretability opens this part of inference up to practical use by letting us rip out those thoughts.

Applications of Interpretability

This has all been very abstract so far. An LLM may think much faster and more efficiently than it writes, and it has ways of representing information which we can plausibly read, but how can we use this? The applications are near endless, but I have a short list below.

Runtime Monitoring

When a model hallucinates, it runs something like fabrication rather than recall, and concepts like deception and uncertainty are more strongly expressed. If we know how these show up in the residual stream, we can detect hallucinations at runtime!

This is commonly done with probes: small models trained on residual stream states partway through the model. Transformers share surprising commonalities in how they represent information, which makes this generally possible. Anthropic uses similar techniques in production to catch jailbreaks and CBRN-related conversations, and to downgrade conversations from Fable to Opus.

(At Telluvian, we offer this through our monitoring API. See this explainer and the API documentation. We can also detect hallucinations in closed-source models using surrogate interpretability, which is beyond the scope of this piece.)

Extracting Useful Thoughts

Usually we want quite simple things from models. Which parts of this contract do they disagree with? Where would they edit this file? Which tokens are security vulnerabilities?

Getting this out takes a lot of work: careful system prompts, regexes on outputs and, for the labs, intense RL to make models report it accurately. Much of this isn't needed when we can read the model's thoughts. Models have internal representations for disagreement, uncertainty and vulnerabilities. They aren't always easy to find, but once probes are trained they offer a much cheaper and faster way to use models.

Take edit planning in coding agents. We don't need a frontier model to write every edit, just as we don't need a senior engineer to write every line. When reading code, a strong model sees what needs editing, what smells bad and where the problems are. Read its mind as it reads, and those insights can go straight to a cheaper model, with no long plans the cheaper model may not understand.

Thoughts > Words

Models don't always say what they think. They hallucinate instead of admitting uncertainty. They have lied to protect themselves and their values, sometimes in horrendous ways. Anyone who uses coding agents has stories of them cheating their way past tests and misrepresenting their work to please the user. Reading their thoughts directly sidesteps much of this.

Concluding Thoughts

LLMs, now overwhelmingly pre-trained transformers, read far faster than they write. Interpretability lets us extract useful work from that reading, or prefill, stage. We can see where a model thinks you've made a mistake and what it makes of the text it's given. Spotting bugs, security issues, logical errors and factual inaccuracies becomes much faster than otherwise possible.

Interpretability doesn't just show us how models work. It gives us the most efficient way to use them, and we are just getting started.