从注意力到证据:一次查询感知提示压缩的研究实践From Attention to Evidence: Research on Query-Aware Prompt Compression
作者:Yutong Wang
研究时间:2026 年 8–9 月
长上下文是大语言模型最有价值、同时也最昂贵的能力之一。输入文本越长,推理时间和计算成本通常越高;与此同时,我们也更难判断模型究竟依赖了哪些信息完成回答。
本文研究一个更具体的问题:
能否根据用户的问题,从长文本中提前筛选出最可能与答案相关的内容,在显著缩短上下文的同时尽量保持最终回答质量?
这里尝试的并不是把原文重新“总结”一遍,而是让模型先看到问题,再从原始上下文中选择最值得保留的证据片段。
实验结果表明,查询注意力(query attention)能够提供有用的证据排序信号。在多组压缩率下,基于注意力的选择方法明显优于随机删除。
本文关注的也不仅是一个单一的压缩分数,而是三个更实际的问题:
- 压缩后是否仍然保留了回答所需的信息;
- 最终问答质量是否保持稳定;
- 每一次删除和保留是否能够被检查和复现。
什么值得留下?
从问题出发,为上下文排序
在模型输入中,上下文位于问题之前。
当模型对问题进行前向计算时,问题中的 token 会通过 attention 机制与前面的上下文 token 建立联系。我们将这些注意力权重作为一种与问题相关性的排序信号。
具体流程包括:
- 提取问题 token 指向上下文 token 的注意力;
- 聚合多个注意力头、多个问题 token 和选定模型层的信号;
- 将 subword 级别的得分重新聚合到完整词;
- 根据得分选择高排名词;
- 按照它们在原文中的顺序恢复文本。
最终得到的不是一段重新生成的摘要,而是一份从原文中直接抽取出来的、可以逐项核查的压缩上下文。
为什么 Attention 值得一试?
Transformer 会将文本拆分为 token,并为 token 之间计算注意力关系。
对于每个 attention head(注意力头),模型首先计算不同位置之间的相关分数,再通过 Softmax 将其转换为总和为 1 的注意力权重。从直觉上看:
如果某个问题 token 对某段上下文分配了较高的注意力权重,那么这段内容可能与回答当前问题有关。
同一层通常包含多个 attention heads,它们相当于并行工作的多组注意力模式。不同 head 可能分别对实体、关系、位置、语义线索或长距离依赖更敏感。因此,我们并不依赖某一个 attention head,而是综合多个 head 和多个问题 token 的信号。
本方法把问题 token 对上下文 token 的注意力权重作为候选证据的排序依据。如果多个问题 token 都对某段上下文赋予较高权重,那么这段内容就更可能与当前问题相关。随后,我们只保留得分靠前的完整词,并恢复原文顺序。因此,这种压缩方式并不是“重新描述原文”,而更接近:
围绕当前问题,从原文中抽取一份可追溯的证据子集。
需要强调的是,attention 并不是模型内部唯一的信息通道,单个高 attention 权重也不能直接等同于“答案证据”。它在这里承担的是一个更有限、也更容易验证的角色:
提供候选证据的排序信号。
它是否真正有价值,最终仍需要通过随机删除、未压缩输入和端到端问答表现进行验证。
四种聚合方式:MaxMax 与 MeanMax 的差别
在实际使用中,我们需要同时处理两个维度:
- H(Head):多个 attention head 的结果如何合并;
- Q(Query):多个问题 token 的结果如何合并。
其中:
- Mean 更接近“共同支持”:如果很多 head 或很多问题 token 都认为某段内容重要,它的分数会更高;
- Max 更接近“任何一个强信号都可以触发”:即使只有一个 head 或一个关键词强烈指向某处,也可以把它保留下来。
| 方法 | 注意力头 | 问题 token | 直观含义 |
|---|---|---|---|
| HMean–QMean | 均值 | 均值 | 保留被问题普遍关注的词 |
| HMean–QMax | 均值 | 最大值 | 任一问题词强烈选择即可保留 |
| HMax–QMean | 最大值 | 均值 | 允许专门头发现证据 |
| HMax–QMax / MaxMax | 最大值 | 最大值 | 任一头连接任一问题词即可保留 |
MeanMax 即 HMean–QMax:注意力头取均值,问题 token 取最大值。
其中,MaxMax 即 HMax–QMax。它的直觉是:
只要某个 attention head 中的某个问题 token 对一个上下文词给出了足够强的信号,就允许这个词进入高优先级候选。
这种规则比较适合捕捉数字、实体名、关系词等局部但关键的信息。但它的代价也很明显:由于 Max 操作只需要一个强信号,因此也更容易受到偶发高权重的影响。这四种组合本身不是一个额外训练出来的模型,而是一种 attention 聚合规则,因此实现简单、结果可解释,也方便逐条检查。
第一次验证:它是否真的比随机删文本好?
早期实验使用 500 个 HotpotQA distractor 样本;关键对照是在保留相同比例文本的情况下,基于查询注意力的选择,是否比随机删除更容易保留正确答案所需的信息?
| 文本保留率 | HMean–QMean | HMean–QMax | HMax–QMean | MaxMax | 随机删除 |
|---|---|---|---|---|---|
| 90% | 69.8% | 69.6% | 69.8% | 70.6% | 63.2% |
| 70% | 68.6% | 70.4% | 67.8% | 70.0% | 57.8% |
| 50% | 66.6% | 66.0% | 65.4% | 68.6% | 48.4% |
| 30% | 62.8% | 60.8% | 61.2% | 62.8% | 36.0% |
| 20% | 53.0% | 55.0% | 52.4% | 53.0% | 27.6% |
| 10% | 31.8% | 31.8% | 33.6% | 32.6% | 15.6% |
所有基于查询注意力的方法都明显优于随机删除。
例如,在仅保留 50% 文本时,MaxMax(68.6%)比随机删除(48.4%)提升约 20 个百分点。
不同聚合规则在不同压缩区间各有优势:
- MaxMax 在 50% 保留率下表现突出;
- MeanMax 在 70% 保留率下更高;
- 在极端压缩下,各种 attention 方法都会明显下降。
因此,这组实验最重要的结论并不是“某一种聚合规则永远最好”,而是:
查询注意力确实包含可以用于证据筛选的稳定信号。
进一步:Multi-hop多跳路径能否找回间接证据?
为什么需要多跳?
一跳 attention 描述的是问题 token 与上下文 token 之间的直接关系,但复杂问题的答案往往并不是通过单一步骤得到的,而是需要经过一个或多个中间实体连接起来。
例如,问题是“《A》的作者出生在哪个国家?”,上下文可能分别提供三条信息:“《A》由 B 创作”“B 出生于巴黎”“巴黎位于法国”。真正的推理链因此是:
问题 → B → 巴黎 → 法国
这就是典型的多跳推理(multi-hop reasoning)。如果只观察问题与“法国”之间的直接 attention,这个关系可能并不突出;但如果允许相关性沿着 attention 关系经过中间 token 传播,就有机会通过 Query → B → 巴黎 → 法国 这样的路径找到间接证据。
因此,多跳方法并不是为了替代直接 attention,而是为了补充那些与问题缺少强直接联系、但可以通过中间事实连接起来的证据。
不过,传播路径越长,相关性也越容易扩散到无关内容。经过三跳、四跳后,越来越多仅有弱关联的 token 可能被带入候选集合,从而引入噪声。为此,可以对更长的路径施加衰减:一跳贡献最高,两跳次之,三跳和四跳进一步降低权重,在保留间接证据信号的同时限制远距离噪声传播。
多跳实验结果
在实验中,我比较了直接 attention(一跳)与通过中间 token 传播的多跳路径。结果显示,两跳在部分压缩区间能够带来收益,但随着路径进一步加深,三跳和四跳在强压缩条件下更容易退化。对长路径加入衰减后,整体表现更加稳定,这说明多跳传播的价值更多来自有限范围内补充间接证据,而不是让相关性无限向外扩散。
| 文本保留率 | MaxMax preservation | 两跳 | 指数衰减 | 随机 |
|---|---|---|---|---|
| 50% | 85.1% | 86.2% | 86.5% | 54.7% |
| 30% | 74.5% | 73.4% | 76.8% | 41.0% |
| 20% | 62.8% | 67.0% | 67.0% | 29.5% |
| 10% | 38.7% | 36.4% | 39.3% | 14.6% |
指数衰减以较小权重加入较长路径,在 20–30% 保留率附近表现最稳定;它为一跳 MaxMax 增加了处理间接证据的能力。
从 Token 保留到最终答案:LongBench v2
前面的实验主要关注“证据是否被保留下来”,但在真实应用中,更重要的问题是:压缩以后,模型还能不能正确回答问题?
因此,下一步需要把压缩器放入完整的问答链路中:
长上下文 + 问题 → 压缩器 → 压缩上下文 → 下游回答模型 → 最终答案评分
在这个流程里,未压缩文本、四种 attention 聚合方法、随机删除以及外部压缩服务,都应尽可能在相同样本、相同回答模型和相同评测条件下进行比较。
132 条 LongBench v2 样本告诉了我们什么?
| 方法 | 实际文本保留率 | LongBench 准确率 |
|---|---|---|
| 未压缩基线 | 100.00% | 62.12% |
| MaxMax | 58.49% | 59.09% |
| HMean–QMax | 58.84% | 58.78% |
| HMax–QMean | 59.32% | 57.58% |
| HMean–QMean | 59.68% | 62.88% |
| 随机删除 | 58.35% | 58.33% |
在这组端到端测试中,大多数压缩方法将输入长度减少了约 40%,而不同 attention 方法的回答准确率仍集中在 57.58%–62.88% 之间。其中 HMean–QMean 达到 62.88%,与未压缩基线的 62.12% 接近。
这个结果并不能证明压缩能够普遍提升回答质量,但至少说明:在显著缩短上下文的情况下,仍有可能保留大部分端到端问答性能。
这也意味着,后续评测不能只看“保留了多少文本”,还需要同时考虑压缩比例、最终回答质量、逐条样本变化,以及不同问题类型下的稳定性。
工程现实:压缩算法快不快,同样重要
一个压缩方法即使在准确率上表现不错,如果压缩本身的成本过高,甚至比直接处理原始上下文还慢,就很难产生实际价值。因此,我进一步测试了不同窗口长度和 batch size 下的计算性能。
这里的 batching 指的是将同一个请求切分出的多个文本窗口同时送入 GPU 计算,而不是生产服务中将多个独立用户请求动态合并的 continuous batching。表中的 p50 表示多次运行延迟的中位数。
| 窗口数(512 token) | 顺序 p50 | 最佳批处理 p50 | 总加速 |
|---|---|---|---|
| 2 | 0.108 s | 0.076 s | 1.46× |
| 4 | 0.217 s | 0.130 s | 1.69× |
| 8 | 0.432 s | 0.277 s | 1.58× |
| 16 | 0.857 s | 0.521 s | 1.67× |
对于 512-token 窗口,批处理可以带来约 1.5–1.7× 的加速。但当窗口变长后,这个结论并不成立。
在 2048-token 窗口下、八个窗口的情况下,batch size 1 的延迟约为 1.041 秒,而 batch size 8 约为 1.136 秒;与此同时,显存占用从约 4.11 GiB 增加到了 10.68 GiB。
也就是说,更大的 batch 并不一定意味着更快。 原因之一是 self-attention 的计算量会随着序列长度近似二次增长。对于 2048-token 的长窗口,单个样本本身已经非常昂贵,此时 GPU 计算和显存带宽更容易成为瓶颈。
性能优化不能以改变结果为代价
批处理实验还暴露了另一个容易被忽略的问题:数值等价性。
在初始的 512-token 测试中,不同 batching 方式得到的词集合重合度约为 0.9245。进一步检查后发现,padding 和 GPU 数值差异可能会改变位于 top-k 边界附近 token 的排名。
通过按真实未 padding 长度进行分桶,并正确应用 attention mask,在后续 2048-token 分桶测试中,最大 attention score 差异降低到了约 3 × 10⁻⁸。
因此,任何性能优化都应该同时验证两个问题:速度是否真的变快,以及输出是否仍然保持等价。 否则,一个“更快”的实现可能实际上已经改变了压缩算法本身。
与外部压缩器比较
除了内部的 attention-based 方法,我还把相同的评测流程连接到了两个外部压缩服务。它们采用了不同的技术路线,因此这里的目的并不是简单判断谁“最好”,而是观察不同压缩范式在相似任务中的实际表现。
服务 A:问题无关压缩与查询引导压缩
在本次可用接口中,服务 A 不接收用户问题,因此属于 query-independent compression;而 MaxMax 同时使用 context + query 进行压缩。
两者因此代表了不同的思路:服务 A 从文本自身结构中判断哪些内容更值得保留,而 MaxMax 则围绕当前问题选择证据。
| 方法 | 实际文本保留率 | 记录数 | 准确率 | 95% CI |
|---|---|---|---|---|
| MaxMax | 57.69% | 50 | 58.0% | [44.2%, 70.6%] |
| 外部服务 A | 62.91% | 50 | 62.0% | [48.2%, 74.1%] |
在这次 50 条样本的比较中,服务 A 的准确率为 62.0%,MaxMax 为 58.0%。但服务 A 同时也保留了更多文本:62.91%,而 MaxMax 为 57.69%。
因此,这个结果提醒我们:压缩方法不能只比较最终准确率,也必须同时考虑它实际保留了多少文本。 一个保留 65% 文本的方法与一个只保留 40% 文本的方法,本身就在不同的质量—成本权衡点上。
服务 B:删除一半文本以后,回答会怎样?
对于服务 B,我完成了完整的评测连接,包括保留率换算、缓存与断点续跑、限流重试、进度记录、实际压缩率和延迟统计,并保存了压缩结果供后续人工检查。
| 方法 | 实际文本保留率 | 可评分记录 | 准确率 |
|---|---|---|---|
| 同轮未压缩基线 | 100.00% | 132 | 59.09% |
| 外部服务 B(coarse 模式) | 50.45% | 131 | 58.78% |
服务 B 将上下文压缩到了原长度的约一半。在共享的 131 条可比较样本中,未压缩输入答对 78 题,服务 B 答对 77 题。
换句话说,在删除约一半文本之后,最终回答表现基本保持一致。这说明在这组样本中,服务 B 找到了一个相当有效的压缩—质量平衡点。
不过,这并不意味着“50% 压缩率一定安全”。真正重要的问题仍然是:它删掉了什么、为什么删,以及当输入结构发生变化时是否仍然稳定。
把黑箱变成可观察对象
对于外部服务,我们看不到内部模型结构,但仍然可以通过系统性的输入扰动来观察它的行为。我设计了 18 次 probing 请求,分别改变用户问题、段落顺序、压缩粒度、重复请求、关键词干扰和输入长度。
| 探测 | 观察 | 可以得到的结论 |
|---|---|---|
| 查询敏感性 | 平均保留段落重合度 0.347 | 输出会明显随问题变化 |
| 确定性 | 5 次相同请求只有 1 个输出 | 本次测试中高度确定 |
| 删除 vs 改写 | 17/18 为原 token 顺序的精确删除;新 token 约 0.21% | coarse mode 主要采用抽取/删除 |
| 段落顺序 | 倒序重合度 0.333;打乱重合度 0.400 | 选择结果对位置变化敏感 |
| 关键词干扰 | 真证据和关键词陷阱均被保留 | 需要进一步对抗性验证 |
这些实验无法告诉我们服务内部真正使用了什么算法,但它们可以回答另一个同样有价值的问题:从外部看,这个系统在面对输入变化时表现出了怎样的规律?
这种黑箱 probing 可以进一步指导后续评测设计,例如证据位置扰动、问题改写、关键词陷阱、重复证据、段落重排以及多用户并发场景。相比单次准确率,这些测试更能帮助我们理解一个压缩系统在真实环境中的稳定性。
从研究原型走向可部署系统
除了离线实验,我还将这套方法封装成了一个可交互的网页 Demo。用户可以粘贴一段长文本、输入一个问题、选择目标文本保留率,并获得压缩后的上下文,同时查看 token 数量变化和压缩比例。
这个 Demo 的目的并不仅仅是展示“文本变短了多少”。更重要的是,它让查询引导压缩变成了一件可以直接观察和检查的事情:用户既可以看到压缩带来的长度变化,也可以检查哪些原始内容被保留、哪些被删除。这对于调试、评测以及理解模型行为都很重要,也使得压缩方法从离线实验进一步向可交互、可验证的实际系统迈进。
下一阶段,我们将重点验证这些方法在更接近真实应用的环境中的表现。除了在统一基线下继续比较不同压缩方案,我们还会加入并发请求、输入扰动和证据位置变化等测试,以评估系统在复杂场景下的稳定性与鲁棒性。最终目标,是从实验室指标进一步走向可复现、可部署的实际性能。
Author: Yutong Wang
Research period: August–September 2026
Long context is one of the most useful capabilities of modern language models, but it is also one of the most expensive. Longer inputs generally mean more compute, higher latency and higher inference costs. They also make it harder to tell which parts of the context the model actually relied on when producing an answer.
This work looks at a narrower question:
Can we use the user's query to identify the parts of a long context that are most likely to matter, shorten the input substantially, and still preserve answer quality?
The aim is not to summarise or rewrite the original text. Instead, the model first sees the question and then selects the parts of the source text that are most likely to contain useful evidence.
Across a range of compression levels, query attention proved to be a useful signal for ranking that evidence, consistently outperforming random deletion.
Rather than reducing the problem to a single compression score, I focused on three practical questions:
- Does the compressed context still contain the information needed to answer the question?
- Does downstream answer quality remain stable?
- Can we inspect and reproduce what was kept and what was removed?
What Should We Keep?
Ranking the Context from the Query
The context is placed before the question in the model input.
During the forward pass, tokens in the question attend back to earlier tokens in the context. I use those attention weights as a signal for query relevance.
The process is straightforward:
- Extract attention from query tokens to context tokens.
- Aggregate across multiple attention heads, query tokens and selected model layers.
- Combine subword-level scores back into whole words.
- Keep the highest-scoring words.
- Restore them in their original order.
The result is not a newly generated summary. It is a smaller subset taken directly from the original text, which means the retained evidence can still be inspected against the source.
Why Attention Is Worth Trying
Transformers split text into tokens and compute attention relationships between them.
Within each attention head, the model first calculates a score between positions and then applies Softmax to turn those scores into attention weights that sum to one. In simple terms:
If a query token assigns a relatively high attention weight to part of the context, that part of the context may be relevant to answering the question.
A transformer layer usually contains several attention heads operating in parallel. Different heads may be more sensitive to different kinds of structure, such as entities, relationships, position, semantic cues or longer-range dependencies.
For that reason, the method does not rely on a single head. Instead, it combines information across several heads and several query tokens.
The attention from query tokens to context tokens is then used to rank candidate evidence. If several query tokens place relatively high attention on the same part of the context, that content becomes a stronger candidate for retention.
Only the highest-scoring whole words are kept, and their original order is preserved. So rather than rewriting the source text, the method is closer to:
Extracting a traceable subset of evidence from the original context, guided by the current query.
Attention is not the only way information moves through a transformer, and a large attention weight does not automatically mean that a token contains the answer.
Here, attention is used in a much narrower sense:
as a ranking signal for candidate evidence.
Whether that signal is useful has to be tested against simpler baselines, including random deletion, uncompressed input and downstream question-answering performance.
Four Aggregation Strategies: MaxMax and MeanMax
There are two dimensions to aggregate:
- H (Head): how attention across different heads is combined.
- Q (Query): how attention across different query tokens is combined.
The distinction is intuitive:
- Mean rewards broad agreement. A token scores highly when many heads or query tokens consistently support it.
- Max allows one strong signal to dominate. A single specialised head or one important query token can be enough to move a token towards the top of the ranking.
| Method | Attention heads | Query tokens | Intuition |
|---|---|---|---|
| HMean–QMean | Mean | Mean | Favours words that are consistently attended to |
| HMean–QMax | Mean | Max | A strong signal from any query token can raise a word's rank |
| HMax–QMean | Max | Mean | Allows a specialised head to surface evidence |
| HMax–QMax / MaxMax | Max | Max | Any strong head–query connection can raise a word's rank |
MeanMax is HMean–QMax: the mean across attention heads and the max across query tokens.
MaxMax is HMax–QMax. Its logic is simple:
If any query token in any attention head gives a strong enough signal to a context word, that word can be treated as a high-priority candidate.
This is useful for local but important information such as numbers, named entities and relation words. The trade-off is that MaxMax is also more exposed to isolated attention spikes, since it only takes one strong signal to raise a token's score.
These four variants are not separately trained models. They are simply different ways of aggregating attention, which makes them lightweight, interpretable and easy to inspect on individual examples.
First Test: Is It Better Than Random Deletion?
The first experiment used 500 HotpotQA distractor examples. The comparison was deliberately simple: if we keep the same amount of text, does query-guided attention preserve useful evidence better than random deletion?
| Text retention | HMean–QMean | HMean–QMax | HMax–QMean | MaxMax | Random deletion |
|---|---|---|---|---|---|
| 90% | 69.8% | 69.6% | 69.8% | 70.6% | 63.2% |
| 70% | 68.6% | 70.4% | 67.8% | 70.0% | 57.8% |
| 50% | 66.6% | 66.0% | 65.4% | 68.6% | 48.4% |
| 30% | 62.8% | 60.8% | 61.2% | 62.8% | 36.0% |
| 20% | 53.0% | 55.0% | 52.4% | 53.0% | 27.6% |
| 10% | 31.8% | 31.8% | 33.6% | 32.6% | 15.6% |
Every query-attention method outperformed random deletion.
At 50% text retention, for example, MaxMax reached 68.6%, compared with 48.4% for random deletion — roughly a 20 percentage-point difference.
No single aggregation rule dominates at every compression level. MaxMax performs particularly well at 50% retention, while MeanMax is strongest at 70%. Under very aggressive compression, all of the attention-based approaches deteriorate.
The main takeaway is therefore not that one aggregation rule is universally best, but that:
query attention contains a useful and reasonably stable signal for selecting evidence.
Going Further: Can Multi-Hop Paths Recover Indirect Evidence?
Why Multi-Hop Reasoning Matters
One-hop attention captures the direct relationship between query tokens and context tokens. More complicated questions, however, often require several pieces of evidence to be linked through intermediate entities.
Suppose the question is:
"In which country was the author of A born?"
The context might tell us that "A was written by B", "B was born in Paris", and "Paris is in France". The reasoning chain is therefore:
Question → B → Paris → France
This is a typical example of multi-hop reasoning.
If we only look at direct attention between the question and "France", the signal may be weak. But if relevance is allowed to propagate through intermediate tokens, the model may recover a path such as Query → B → Paris → France.
The point of multi-hop propagation is not to replace direct attention. It is to recover evidence that may not be strongly connected to the question itself, but can be reached through intermediate facts.
There is an obvious downside. The longer the path, the easier it is for relevance to spread into unrelated parts of the context. By the third or fourth hop, weakly related tokens can begin to enter the candidate set and add noise.
One way to control this is to reduce the contribution of longer paths. One-hop evidence receives the highest weight, two-hop evidence less, and three- or four-hop paths less again. This gives indirect evidence a chance to contribute without allowing long-range propagation to overwhelm the direct signal.
Multi-Hop Results
I compared direct one-hop attention with multi-hop propagation through intermediate tokens. Two-hop paths sometimes improved performance, but three- and four-hop paths were more likely to degrade under aggressive compression.
Applying decay to longer paths produced more stable results, which suggests that multi-hop propagation is most useful when it helps recover nearby indirect evidence, rather than when relevance is allowed to spread indefinitely.
| Text retention | MaxMax preservation | Two-hop | Exponential decay | Random |
|---|---|---|---|---|
| 50% | 85.1% | 86.2% | 86.5% | 54.7% |
| 30% | 74.5% | 73.4% | 76.8% | 41.0% |
| 20% | 62.8% | 67.0% | 67.0% | 29.5% |
| 10% | 38.7% | 36.4% | 39.3% | 14.6% |
Exponential decay was particularly stable around the 20–30% retention range. In practice, it gives one-hop MaxMax some ability to recover indirect evidence without giving longer paths too much influence.
From Token Retention to Final Answers: LongBench v2
The earlier experiments focused mainly on whether useful evidence survived compression. In practice, though, the question that matters most is simpler:
Can the model still answer correctly afterwards?
To test this, I placed the compressor inside a complete question-answering pipeline:
Long context + question → compressor → compressed context → downstream answering model → final answer score
Where possible, uncompressed input, the four attention aggregation methods, random deletion and external compression services were evaluated on the same examples, using the same answering model and the same scoring conditions.
What Do 132 LongBench v2 Examples Tell Us?
| Method | Actual text retention | LongBench accuracy |
|---|---|---|
| Uncompressed baseline | 100.00% | 62.12% |
| MaxMax | 58.49% | 59.09% |
| HMean–QMax | 58.84% | 58.78% |
| HMax–QMean | 59.32% | 57.58% |
| HMean–QMean | 59.68% | 62.88% |
| Random deletion | 58.35% | 58.33% |
Most of the compression methods reduced the input by around 40%, while accuracy across the attention-based variants remained between 57.58% and 62.88%.
HMean–QMean reached 62.88%, close to the 62.12% uncompressed baseline.
This does not show that compression generally improves answer quality. It does suggest, however, that it is possible to remove a substantial amount of context while retaining most of the downstream question-answering performance.
That is why future evaluation should not focus on retention rate alone. It also needs to consider downstream answer quality, per-example changes and how stable the method is across different types of question.
Engineering Reality: Compression Also Has to Be Fast
A compression method can be accurate and still be impractical. If the compression step itself is expensive enough to cancel out the savings from using a shorter context, there is little point in deploying it.
I therefore benchmarked the implementation across different window sizes and batch sizes.
Here, batching means processing several windows from the same request together on the GPU. This is different from production-style continuous batching, where independent requests from different users are grouped together. The p50 figure is the median latency across repeated runs.
| Number of windows (512 tokens) | Sequential p50 | Best batched p50 | Overall speed-up |
|---|---|---|---|
| 2 | 0.108 s | 0.076 s | 1.46× |
| 4 | 0.217 s | 0.130 s | 1.69× |
| 8 | 0.432 s | 0.277 s | 1.58× |
| 16 | 0.857 s | 0.521 s | 1.67× |
For 512-token windows, batching gives roughly a 1.5–1.7× speed-up.
That benefit disappears with longer windows. With eight 2048-token windows, batch size 1 takes about 1.041 seconds, while batch size 8 takes about 1.136 seconds. GPU memory use also rises from roughly 4.11 GiB to 10.68 GiB.
So a larger batch is not automatically faster.
One reason is that self-attention becomes much more expensive as sequence length grows, with compute scaling approximately quadratically with the number of tokens. At 2048 tokens, each individual window is already costly enough that GPU compute and memory bandwidth become much more important bottlenecks.
Faster Is Only Better If the Output Stays the Same
The batching experiments also exposed a less obvious issue: numerical equivalence.
In the initial 512-token tests, the overlap between the sets of words selected by different batching configurations was around 0.9245. Further investigation showed that padding and small numerical differences on the GPU could alter the ordering of tokens close to the top-k selection boundary.
After bucketing samples by their actual unpadded length and applying the attention mask correctly, the maximum difference in attention scores in later 2048-token bucketed tests fell to around 3 × 10⁻⁸.
So any performance optimisation needs to answer two questions at once:
Is it faster, and does it still produce the same result?
Otherwise, an implementation may appear faster simply because it is no longer running exactly the same compression procedure.
Comparing with External Compressors
I also connected the same evaluation pipeline to two external compression services.
They use different approaches, so the aim was not simply to name a winner. The more useful question was how different compression strategies behave under similar evaluation conditions.
Service A: Query-Independent vs Query-Guided Compression
In the interface available for this evaluation, Service A does not accept the user's question, making it a query-independent compressor. MaxMax, by contrast, uses both context + query.
That makes the comparison conceptually interesting. Service A decides what to keep from the text itself, while MaxMax chooses evidence with a particular question in mind.
| Method | Actual text retention | Records | Accuracy | 95% CI |
|---|---|---|---|---|
| MaxMax | 57.69% | 50 | 58.0% | [44.2%, 70.6%] |
| External Service A | 62.91% | 50 | 62.0% | [48.2%, 74.1%] |
Across these 50 examples, Service A achieved 62.0% accuracy compared with 58.0% for MaxMax. It also retained more of the original text: 62.91% rather than 57.69%.
That matters because accuracy alone does not tell the whole story. A compressor retaining 65% of the input and another retaining 40% are operating at different points on the quality–compression trade-off.
Service B: What Happens If We Remove Half the Text?
For Service B, I built the full evaluation integration, including retention-rate conversion, caching and resumable execution, rate-limit retries, progress logging, compression-ratio statistics, latency measurement and storage of compressed outputs for manual inspection.
| Method | Actual text retention | Scorable records | Accuracy |
|---|---|---|---|
| Uncompressed baseline from the same run | 100.00% | 132 | 59.09% |
| External Service B (coarse mode) | 50.45% | 131 | 58.78% |
Service B reduced the context to roughly half its original length.
Across the 131 examples that could be compared directly, the uncompressed input answered 78 correctly, while Service B answered 77 correctly.
So in this particular evaluation, removing roughly half the text had almost no effect on downstream answer accuracy. That represents a strong compression–quality trade-off on this sample set.
It does not mean that 50% retention is safe in general. The more interesting questions are still what was removed, why it was removed, and whether the behaviour remains stable when the input changes.
Making a Black Box Observable
We cannot see the internal architecture of an external compression service, but we can still learn about its behaviour by changing the input in controlled ways.
I ran 18 probing requests that varied the question, paragraph order, compression granularity, repeated requests, distracting keywords and input length.
| Probe | Observation | What we can infer |
|---|---|---|
| Query sensitivity | Mean retained-paragraph overlap: 0.347 | The output changes substantially with the question |
| Determinism | Five identical requests produced one unique output | Highly deterministic in this test |
| Deletion vs rewriting | 17/18 outputs were exact deletions preserving token order; new tokens ≈ 0.21% | Coarse mode is largely extractive |
| Paragraph order | Reverse-order overlap: 0.333; shuffled-order overlap: 0.400 | Selection is sensitive to position |
| Keyword distraction | Both genuine evidence and keyword traps were retained | More adversarial testing is needed |
These experiments cannot tell us what algorithm the service actually uses internally. They can tell us something else that is just as useful: how the system behaves when the input is changed.
That gives us a clearer basis for future robustness tests, including evidence-position shifts, question paraphrasing, keyword traps, duplicated evidence, paragraph reordering and multi-user concurrency.
A single accuracy figure cannot capture all of those behaviours.
From Research Prototype to Deployable System
Alongside the offline experiments, I packaged the method into an interactive web demo.
Users can paste in a long piece of text, enter a question, choose a target retention rate and receive the compressed context, together with token counts and compression statistics.
The demo is not just there to show that the text has become shorter. It makes the compression process inspectable: users can see both how much text has been removed and exactly which parts of the original context were kept.
That is useful for debugging, evaluation and understanding model behaviour, and moves the work beyond a purely offline research experiment towards something more interactive and verifiable.
The next stage is to test these methods under conditions that are closer to real use. That means continuing the comparison under a shared evaluation baseline, while adding concurrent requests, input perturbations and changes in evidence position to measure robustness and stability in more demanding settings.
The longer-term aim is to move beyond isolated benchmark results towards a compression system whose performance is reproducible, observable and practical to deploy.