Back to the post
AppendixAdam Pattenden

Quantisation and surrogate probes

Datasets

CAM. 5,058 coding-agent sessions (78% Claude Opus 4.6), scored 0 to 10 by LLM judges with cited spans. We used overconfidence with laziness, the behaviour with most positives.

GPT-5.4 produced 94% of high-scoring spans, but only judged sessions GPT-5.4-mini had failed, which are longer and more often flagged. Mixing judges would confound behaviour with length and judge, so we kept only GPT-5.4's 2,318 sessions and labels.

  • Positive: a token inside a GPT-5.4 citation scored 7 or above.
  • Transcripts: chat-templated, tool output cut to 500 characters, last 32k tokens kept.
  • Split: 80/20 by session, grouped by repository (1,784 train, 534 eval).
  • Sampling: each positive plus 8 other assistant tokens from flagged sessions; 64 random assistant tokens from clean ones.
A: cited bad spanB: rest of flagged sessionC: cleared sessions
Train tokens36,568301,00066,524
Eval tokens9,35978,89820,687

15% of training sessions are held out for early stopping, and minibatches are oversampled to 50/50. Eval is about 9% positive against a natural 0.36%, so absolute AUPRC is optimistic; paired comparisons are unaffected.

We count both B and C as negative, as in deployment. Dropping B scores higher only because the probe can then tell sessions apart:

SchemeABCLayer 49 AUROC / AUPRCPosition-only AUROC
All negatives (used)positivenegativenegative0.858 / 0.4410.536
Supported onlypositivedroppednegative0.894 / 0.8220.608

Luna. 63,117 GPT-5.6-Luna answers, with GPT-5.6-Terra marking spans Not Supported (positive), Insufficient Information or Supported.

  • Excluded: 1,441 unannotated records, and truncated final spans in 132.
  • Text after the last annotated span (36% of characters) is ignored.
  • Sample: 12,000 records by fixed hash; the full set would need about 650 GB of states.
  • Split: 80/20 by question (9,623 train, 2,377 eval). Eval has 279,058 tokens, 8.0% positive.

← Back to Method

Method

  • Surrogate: Qwen3.8-27B (64 layers, width 5,120). Probes are trained on bf16.
  • Extraction: prefill, every 7th layer from 0 to 63, on one RTX PRO 6000. vLLM 0.25 for standard formats, llama.cpp for GGUF. bf16 differs by at most 0.0003 between the two.
  • Probes: per-token logistic regression; learning rate 1e-5, batch 256, early stopping, 5 seeds.
  • Intervals: 95% bootstrap over sessions (Luna: records), paired when models share tokens.
  • Equivalence: equivalent if the interval lies inside ±0.01, different if entirely outside, otherwise inconclusive.
  • Drift: cosine similarity (1 means same direction) and relative L2, ∥hquant−hbf16∥ / ∥hbf16∥\lVert h_{\text{quant}} - h_{\text{bf16}} \rVert \,/\, \lVert h_{\text{bf16}} \rVert (0 means identical). The bf16 reference is llama.cpp for CAM and vLLM for Luna; they agree (FP8 at layer 49: 0.055 and 0.052).
Per-layer difference in probe AUROC and AUPRC between unquantised bf16 run in llama.cpp and bf16 run in vLLM. Every point sits on zero, well inside the ±0.01 margin.
Engine check: bf16 in llama.cpp minus bf16 in vLLM. The engine contributes no difference (|Δ| ≤ 0.0003).

← Back to Method

Models

NameCheckpointBits per weightEngine
bf16Qwen/Qwen3.8-27B @1d4bf0f216vLLM
FP8Qwen3.8-27B-FP8 @017b9c7a8vLLM
NVFP4NVIDIA NVFP4, mixed 4/8 @482ca0f3~4.5 nominal (~6.5 by file size)vLLM, CUTLASS backend
Q4_K_XLunsloth/Qwen3.8-27B-GGUF UD-Q4_K_XL~4.5 nominal (~5.2 by file size)llama.cpp
Q3_K_XLsame repo, UD-Q3_K_XL~3.9llama.cpp
IQ1_Msame repo, UD-IQ1_M~2.0llama.cpp
Ternary Bonsaiprism-ml/Ternary-Bonsai-2-27B-gguf @b072e1d3 PTQ1_01.7llama.cpp
Bonsai-27B Q1_0prism-ml/Bonsai-27B-gguf @f10afb35 (base Qwen3.6-27B @6a9e13bd)~1.1llama.cpp

Bits per weight is file size × 8 ÷ 27B unless marked nominal.

← Back to Method

Baselines

  • TF-IDF (explainer): word and n-gram counts over the previous 32 tokens.
  • Qwen3-Embedding (8B, 0.6B): an embedding of the previous 32 or 512 tokens, or the turn so far.
  • BM25 kNN (BM25): mean label of the 50 nearest training windows.
  • Position/length: token position and transcript length only.

All but BM25 feed a logistic regression.

CAM

MethodAUROCAUPRC
Probe, layer 490.8580.441
Probe, layer 140.8230.352
TF-IDF, 32 tokens0.7460.250
Probe, layer 00.6740.171
Qwen3-Embedding-8B, 32 / 512 / turn so far0.674 / 0.642 / 0.6700.14 to 0.17
Qwen3-Embedding-0.6B, 32 tokens0.6540.157
BM25 kNN, 32 / 512 / turn so far0.627 / 0.522 / 0.5530.09 to 0.14
Position/length only0.5360.092

Luna

MethodAUROCAUPRC
Probe, layer 490.9090.451
Probe, layer 140.8830.354
Probe, layer 00.8480.273
TF-IDF, 32 tokens0.747 [0.734, 0.760]0.193 [0.180, 0.208]
Position/length only0.572 [0.557, 0.586]0.097 [0.090, 0.104]
Horizontal bar charts of AUROC, AUPRC and precision at 50% recall for every method on the CAM eval set. Probes at layers 49 and 14 lead, followed by TF-IDF, the layer-0 probe, Qwen3 embeddings, BM25 kNN and position/length only.
CAM: every method on the same 108,944 eval tokens.
Paired differences between the layer-49 probe and each baseline, for AUROC and AUPRC. Every difference is positive with a confidence interval that excludes zero.
CAM: layer-49 probe minus each baseline, paired.

← Back to Baselines

Full results below 4 bits

Layer 49 AUPRC change against bf16 (CAM 0.441, Luna 0.451), with verdicts over 10 layers (eq equivalent, inc inconclusive, diff different).

ModelDatasetLayer 49 drift (cosine / rel. L2)ReusedReused verdictsRetrainedRetrained verdicts
Q3_K_XLCAM0.986 / 0.150−0.007AUROC 10 eq; AUPRC 7 eq, 3 inc−0.00320/20 eq
Q3_K_XLLuna0.987 / 0.15−0.00819/20 eq (layer 49 AUPRC inc)−0.00420/20 eq
IQ1_MCAM0.883 / 0.474−0.054AUROC 2 inc, 8 diff; AUPRC 1 inc, 9 diff−0.022AUROC 2 eq, 8 inc; AUPRC 1 eq, 5 inc, 4 diff
Ternary BonsaiCAM0.819 / 0.574−0.034AUPRC −0.02 to −0.06 and AUROC −0.01 to −0.03 across layers−0.007AUROC 8/10 eq; AUPRC −0.001 to −0.013, mostly inc
Ternary BonsaiLuna0.845 / 0.53−0.041AUROC 5 eq, 3 inc, 2 diff; AUPRC 10 diff−0.015AUROC 10 eq; AUPRC 3 eq, 7 inc
Grid of per-layer AUROC and AUPRC changes against bf16 on Luna, one row per model: FP8, NVFP4, Q4_K_XL, Q3_K_XL and ternary Bonsai. The first three sit on zero; Q3_K_XL dips slightly; ternary Bonsai reused AUPRC falls below the band at every layer, with retrained probes closer to zero.
Luna: change against bf16 at every layer, one row per model (IQ1_M was not run on Luna).
Absolute AUROC and AUPRC at layer 49 on Luna for each quantised model, reused and retrained, against a vertical line at bf16. The 4.5-bit-and-above models sit on the line; ternary Bonsai reused probe is well left, its retrained probe closer.
Luna, layer 49. Filled: bf16 probe reused. Hollow: probe retrained on quantised states.

← Back to Below 4 bits

Bonsai-27B Q1_0

There is no 1-bit Bonsai for Qwen3.8, so this is compared with its own base, Qwen3.6-27B, on CAM only. Both Qwen3.6 models were given the Qwen3.8 chat template to keep tokens identical. Qwen3.6 bf16 scores slightly lower (layer 49 AUPRC 0.403 against 0.441).

ProbeLayer 49 AUROC (change)Layer 49 AUPRC (change)AUROC eq/inc/diffAUPRC eq/inc/diffMedian ρ
Reused0.782 (−0.065)0.288 (−0.115)0/0/100/0/10~0.72
Retrained0.834 (−0.012)0.385 (−0.018)2/6/21/6/3~0.77

ρ is the Spearman correlation of token scores with bf16. Layer 49 drift: cosine 0.48, relative L2 0.97.

← Back to Below 4 bits

Linear map diagnostics (ternary Bonsai, Luna)

Maps fit on 1,500 training records, scored with the unchanged bf16 probes.

Layerbf16ReusedPer-dimension rescaleLinear map (eval R²)Linear map recoversCosine, raw → standardised
00.2730.2500.2500.268 (0.94)75%0.996 → 0.813
70.3050.2890.2900.305 (0.81)99%0.985 → 0.849
210.3580.3370.3380.357 (0.80)97%0.953 → 0.832
350.4090.3710.3730.400 (0.77)75%0.901 → 0.780
490.4510.4100.4140.435 (0.78)61%0.845 → 0.768
630.4170.3720.3740.396 (0.76)51%0.768 → 0.741
Three panels by layer: residual-stream RMS norm on a log scale for every model, the ratio of each quantised model's norm to bf16, and the share of the norm that is a shared offset. All models except ternary Bonsai stay at a ratio of about 1; ternary Bonsai is well below 1 through layer 56 and above 1 at layer 63, on both datasets.
Residual-stream norm by layer. Formats at 3 bits or more stay within 3% of bf16; ternary Bonsai does not.

← Back to Why do low-bit models degrade?

Back to the post