Quantisation and surrogate probes
Datasets
CAM. 5,058 coding-agent sessions (78% Claude Opus 4.6), scored 0 to 10 by LLM judges with cited spans. We used overconfidence with laziness, the behaviour with most positives.
GPT-5.4 produced 94% of high-scoring spans, but only judged sessions GPT-5.4-mini had failed, which are longer and more often flagged. Mixing judges would confound behaviour with length and judge, so we kept only GPT-5.4's 2,318 sessions and labels.
- Positive: a token inside a GPT-5.4 citation scored 7 or above.
- Transcripts: chat-templated, tool output cut to 500 characters, last 32k tokens kept.
- Split: 80/20 by session, grouped by repository (1,784 train, 534 eval).
- Sampling: each positive plus 8 other assistant tokens from flagged sessions; 64 random assistant tokens from clean ones.
| A: cited bad span | B: rest of flagged session | C: cleared sessions | |
|---|---|---|---|
| Train tokens | 36,568 | 301,000 | 66,524 |
| Eval tokens | 9,359 | 78,898 | 20,687 |
15% of training sessions are held out for early stopping, and minibatches are oversampled to 50/50. Eval is about 9% positive against a natural 0.36%, so absolute AUPRC is optimistic; paired comparisons are unaffected.
We count both B and C as negative, as in deployment. Dropping B scores higher only because the probe can then tell sessions apart:
| Scheme | A | B | C | Layer 49 AUROC / AUPRC | Position-only AUROC |
|---|---|---|---|---|---|
| All negatives (used) | positive | negative | negative | 0.858 / 0.441 | 0.536 |
| Supported only | positive | dropped | negative | 0.894 / 0.822 | 0.608 |
Luna. 63,117 GPT-5.6-Luna answers, with GPT-5.6-Terra marking spans Not Supported (positive), Insufficient Information or Supported.
- Excluded: 1,441 unannotated records, and truncated final spans in 132.
- Text after the last annotated span (36% of characters) is ignored.
- Sample: 12,000 records by fixed hash; the full set would need about 650 GB of states.
- Split: 80/20 by question (9,623 train, 2,377 eval). Eval has 279,058 tokens, 8.0% positive.
Method
- Surrogate: Qwen3.8-27B (64 layers, width 5,120). Probes are trained on bf16.
- Extraction: prefill, every 7th layer from 0 to 63, on one RTX PRO 6000. vLLM 0.25 for standard formats, llama.cpp for GGUF. bf16 differs by at most 0.0003 between the two.
- Probes: per-token logistic regression; learning rate 1e-5, batch 256, early stopping, 5 seeds.
- Intervals: 95% bootstrap over sessions (Luna: records), paired when models share tokens.
- Equivalence: equivalent if the interval lies inside ±0.01, different if entirely outside, otherwise inconclusive.
- Drift: cosine similarity (1 means same direction) and relative L2, (0 means identical). The bf16 reference is llama.cpp for CAM and vLLM for Luna; they agree (FP8 at layer 49: 0.055 and 0.052).

Models
| Name | Checkpoint | Bits per weight | Engine |
|---|---|---|---|
| bf16 | Qwen/Qwen3.8-27B @1d4bf0f2 | 16 | vLLM |
| FP8 | Qwen3.8-27B-FP8 @017b9c7a | 8 | vLLM |
| NVFP4 | NVIDIA NVFP4, mixed 4/8 @482ca0f3 | ~4.5 nominal (~6.5 by file size) | vLLM, CUTLASS backend |
| Q4_K_XL | unsloth/Qwen3.8-27B-GGUF UD-Q4_K_XL | ~4.5 nominal (~5.2 by file size) | llama.cpp |
| Q3_K_XL | same repo, UD-Q3_K_XL | ~3.9 | llama.cpp |
| IQ1_M | same repo, UD-IQ1_M | ~2.0 | llama.cpp |
| Ternary Bonsai | prism-ml/Ternary-Bonsai-2-27B-gguf @b072e1d3 PTQ1_0 | 1.7 | llama.cpp |
| Bonsai-27B Q1_0 | prism-ml/Bonsai-27B-gguf @f10afb35 (base Qwen3.6-27B @6a9e13bd) | ~1.1 | llama.cpp |
Bits per weight is file size × 8 ÷ 27B unless marked nominal.
Baselines
- TF-IDF (explainer): word and n-gram counts over the previous 32 tokens.
- Qwen3-Embedding (8B, 0.6B): an embedding of the previous 32 or 512 tokens, or the turn so far.
- BM25 kNN (BM25): mean label of the 50 nearest training windows.
- Position/length: token position and transcript length only.
All but BM25 feed a logistic regression.
CAM
| Method | AUROC | AUPRC |
|---|---|---|
| Probe, layer 49 | 0.858 | 0.441 |
| Probe, layer 14 | 0.823 | 0.352 |
| TF-IDF, 32 tokens | 0.746 | 0.250 |
| Probe, layer 0 | 0.674 | 0.171 |
| Qwen3-Embedding-8B, 32 / 512 / turn so far | 0.674 / 0.642 / 0.670 | 0.14 to 0.17 |
| Qwen3-Embedding-0.6B, 32 tokens | 0.654 | 0.157 |
| BM25 kNN, 32 / 512 / turn so far | 0.627 / 0.522 / 0.553 | 0.09 to 0.14 |
| Position/length only | 0.536 | 0.092 |
Luna
| Method | AUROC | AUPRC |
|---|---|---|
| Probe, layer 49 | 0.909 | 0.451 |
| Probe, layer 14 | 0.883 | 0.354 |
| Probe, layer 0 | 0.848 | 0.273 |
| TF-IDF, 32 tokens | 0.747 [0.734, 0.760] | 0.193 [0.180, 0.208] |
| Position/length only | 0.572 [0.557, 0.586] | 0.097 [0.090, 0.104] |


Full results below 4 bits
Layer 49 AUPRC change against bf16 (CAM 0.441, Luna 0.451), with verdicts over 10 layers (eq equivalent, inc inconclusive, diff different).
| Model | Dataset | Layer 49 drift (cosine / rel. L2) | Reused | Reused verdicts | Retrained | Retrained verdicts |
|---|---|---|---|---|---|---|
| Q3_K_XL | CAM | 0.986 / 0.150 | −0.007 | AUROC 10 eq; AUPRC 7 eq, 3 inc | −0.003 | 20/20 eq |
| Q3_K_XL | Luna | 0.987 / 0.15 | −0.008 | 19/20 eq (layer 49 AUPRC inc) | −0.004 | 20/20 eq |
| IQ1_M | CAM | 0.883 / 0.474 | −0.054 | AUROC 2 inc, 8 diff; AUPRC 1 inc, 9 diff | −0.022 | AUROC 2 eq, 8 inc; AUPRC 1 eq, 5 inc, 4 diff |
| Ternary Bonsai | CAM | 0.819 / 0.574 | −0.034 | AUPRC −0.02 to −0.06 and AUROC −0.01 to −0.03 across layers | −0.007 | AUROC 8/10 eq; AUPRC −0.001 to −0.013, mostly inc |
| Ternary Bonsai | Luna | 0.845 / 0.53 | −0.041 | AUROC 5 eq, 3 inc, 2 diff; AUPRC 10 diff | −0.015 | AUROC 10 eq; AUPRC 3 eq, 7 inc |


Bonsai-27B Q1_0
There is no 1-bit Bonsai for Qwen3.8, so this is compared with its own base, Qwen3.6-27B, on CAM only. Both Qwen3.6 models were given the Qwen3.8 chat template to keep tokens identical. Qwen3.6 bf16 scores slightly lower (layer 49 AUPRC 0.403 against 0.441).
| Probe | Layer 49 AUROC (change) | Layer 49 AUPRC (change) | AUROC eq/inc/diff | AUPRC eq/inc/diff | Median ρ |
|---|---|---|---|---|---|
| Reused | 0.782 (−0.065) | 0.288 (−0.115) | 0/0/10 | 0/0/10 | ~0.72 |
| Retrained | 0.834 (−0.012) | 0.385 (−0.018) | 2/6/2 | 1/6/3 | ~0.77 |
ρ is the Spearman correlation of token scores with bf16. Layer 49 drift: cosine 0.48, relative L2 0.97.
Linear map diagnostics (ternary Bonsai, Luna)
Maps fit on 1,500 training records, scored with the unchanged bf16 probes.
| Layer | bf16 | Reused | Per-dimension rescale | Linear map (eval R²) | Linear map recovers | Cosine, raw → standardised |
|---|---|---|---|---|---|---|
| 0 | 0.273 | 0.250 | 0.250 | 0.268 (0.94) | 75% | 0.996 → 0.813 |
| 7 | 0.305 | 0.289 | 0.290 | 0.305 (0.81) | 99% | 0.985 → 0.849 |
| 21 | 0.358 | 0.337 | 0.338 | 0.357 (0.80) | 97% | 0.953 → 0.832 |
| 35 | 0.409 | 0.371 | 0.373 | 0.400 (0.77) | 75% | 0.901 → 0.780 |
| 49 | 0.451 | 0.410 | 0.414 | 0.435 (0.78) | 61% | 0.845 → 0.768 |
| 63 | 0.417 | 0.372 | 0.374 | 0.396 (0.76) | 51% | 0.768 → 0.741 |
