Fractal fern pattern
All posts
Benedict Mullins

Less is more the same

Everything Telluvian builds relies on reading a model's hidden states, the activations it produces between input and output. They show what a model is doing before it finishes answering. We collect labelled examples, such as truthful and hallucinated text, and train small classifiers, called probes, to tell their hidden states apart. At runtime the probes score each token as it is generated.

Closed models such as GPT and Claude don't expose their internals, so we replay their answers through an open-weight model, such as Gemma, and read its states instead. Replaying a finished answer means the model reads the whole text in one pass. This step, called prefill, is where most of our compute goes. So for us, cheaper prefill means a cheaper product.

Shrinking the model

The obvious lever is quantisation. Gemma's weights ship as 16-bit numbers. Quantised versions store them in 4 bits, with a shared scale factor for each small block so the rounding stays accurate. The model needs a quarter of the memory, and the GPU spends less time moving it about.

NVIDIA's version of this is NVFP4. It rounds weights in groups of 16, so each is measured against its near neighbours and little detail is lost. It is built for NVIDIA's Blackwell cards, which do their arithmetic in it directly rather than converting back to 16-bit first. For prefill, that is what counts.

Smaller weights mean coarser hidden states, and our probes depend on fine detail. If quantisation erased it, the speed would be worthless.

The test

We ran NVIDIA's NVFP4 release of Gemma 4 31B against the original on an RTX PRO 6000, a Blackwell card. For comparison we added two other popular 4-bit versions: AMD's, in MXFP4, a format agreed by the main chipmakers, and Unsloth's, in GGUF, the format used to run models on laptops. For each, we passed our hallucination dataset through the model, extracted layer 40 and trained a fresh probe.

Results

Bar charts of F1, precision, recall, average precision and ROC AUC for probes trained on the Google baseline, NVIDIA NVFP4, AMD MXFP4 and Unsloth GGUF versions of Gemma 4 31B.
Probe scores for each version of Gemma 4 31B. The quantised models match the original on every metric.

The probes performed almost identically. ROC AUC moved by 0.001 at most from the original model, and F1 by about 0.01. Training and test loss were level across all four. NVFP4's probe traded a little precision for recall: it caught more hallucinations (recall of 0.35 against 0.32) but raised slightly more false alarms.

Bar chart of generation throughput relative to the Google baseline: NVIDIA NVFP4 1.64x, AMD MXFP4 0.88x, Unsloth GGUF 0.83x.
Hidden state generation speed on an RTX PRO 6000, relative to the original model.

NVFP4 generated hidden states 1.64 times faster than the original: 7,570 tokens a second against 4,623. Nearly all the gain came from prefill, which fell from 182 to 98 milliseconds per thousand tokens.

Stacked bar chart of generation wall-clock time by stage: Google baseline 13.8 s, NVIDIA NVFP4 8.4 s, AMD MXFP4 15.7 s, Unsloth GGUF 16.6 s, with vLLM prefill the largest stage in each.
Time spent on each stage of generation. Prefill (orange) dominates, and NVFP4 roughly halves it.

NVFP4 takes a minute longer to load (266 seconds against 207), but that cost is paid once per start. MXFP4 and GGUF ran slower than the original, at 0.88x and 0.83x. Both shrink the model, which helps when memory is the limit. Prefill is limited by arithmetic, and in our setup these formats appear to be unpacked to 16-bit before each calculation. The GPU does the original's work plus the unpacking. GGUF fares worst because vLLM, the engine we use, supports it only experimentally. Both our cards are NVIDIA's, which favours its own format; MXFP4 may do better elsewhere.

The probe itself was never the bottleneck. On an H100, scoring took the same time whichever model produced the states.

Bar chart of probe streaming throughput relative to the Google baseline: all four versions between 0.99x and 1.01x.
Probe scoring speed on an H100. The source model makes no difference.

Takeaway

Quantisation costs our probes nothing and, in the right format, makes them far cheaper to run. NVFP4 gives the same detection for about 60% of the GPU time. Unless something else is holding the system back, it is the clear choice.