Official implementation and result release for SurrogateKV: Representation-Preserving KV Cache Compression for Long-Context LLMs (Findings of EMNLP 2026).
Token-level KV compressors ordinarily retain selected entries and discard the rest. SurrogateKV adds a third representation under the same cache-slot budget: a contiguous historical region can be replaced by one independently addressable surrogate KV pair. Salient entries remain exact, admitted regions become surrogates, and the remaining regions are dropped. The packed cache contains ordinary KV pairs and uses standard attention during decoding.
SurrogateKV is evaluated on three base compressors. The resulting variants are SurrogateKV-Snap, SurrogateKV-Dynamic, and SurrogateKV-Ada.
The editable source for the overview is available as a draw.io file. The diagram shows the mean-plus-norm constructor; the released head-wise profile is noted under Variants.
surrogatekv/: allocation, surrogate construction, packing, and runtime APIrun/longbench/: LongBench wrappers and evaluatordata/: machine-readable values for reported tables and selected figurestests/: runtime and registry checksscripts/validate_release.py: consistency checks for released result filesscripts/build_readme_figures.py: README figures generated from released CSVs
Model weights, benchmark datasets, generated responses, and machine-specific logs are not included.
git clone https://github.com/kimjunewon/SurrogateKV.git
cd SurrogateKV
python3 -m pip install -e .Python 3.10 or newer is required. Install the LongBench dependencies with:
python3 -m pip install -e ".[longbench]"The core framework versions reported for the paper experiments are recorded in
docs/PAPER_ENVIRONMENT.md. They are historical
environment metadata; use the standard installation above for ordinary use.
| Variant | Runtime mode | Base allocation |
|---|---|---|
SurrogateKV-Snap |
surrogate_kv |
SnapKV shared-token selection |
SurrogateKV-Dynamic |
surrogate_kv_dynamic_layer |
DynamicKV layer-wise allocation |
SurrogateKV-Ada |
surrogate_kv_ada |
Ada-KV head-wise allocation |
SurrogateKV is an alias of SurrogateKV-Snap. Method names and aliases are
available through SURROGATEKV_METHOD_TO_MODE.
The shared-token and layer-wise modes use mean-pooled, norm-calibrated
surrogate KV pairs. The released head-wise profile uses the highest-salience
representative KV pair in each admitted region. Set
SURKV_HEADWISE_SURROGATE_PROTO=mean to use mean-plus-norm construction in
head-wise mode.
The cache adapter calls SurKVCluster after prefill attention scores are
available. Tensors use the standard [batch, heads, sequence, head_dim]
layout. The current prefill path supports batch size 1.
from surrogatekv import SurKVCluster
cluster = SurKVCluster(
mode="surrogate_kv",
window_size=32,
max_capacity_prompt=512,
kernel_size=7,
)
compressed_k, compressed_v = cluster.update_kv(
key_states,
query_states,
value_states,
attention_mask=None,
num_key_value_groups=num_key_value_groups,
)The paper setting uses allocation atoms of size c = 4, controlled by
SURKV_ATOM_SIZE and used by default. The runtime's chunk_size is a separate
coarse-region parameter and defaults to 32.
The Ada-KV variant preserves a head-specific cache layout and therefore uses
update_kv_headwise() instead of update_kv(). Calling the shared-token entry
point in surrogate_kv_ada mode raises an error.
The result CSVs and saved-prediction evaluator are self-contained. End-to-end
LongBench prediction generation dispatches to the KVCache-Factory adapter used
in the paper, which lives in the companion experiment workspace and is not
vendored into this repository. Set SURKV_WORKSPACE_ROOT to a workspace that
contains tools/run_surkv_longbench.py and repos/KVCache-Factory before using
the prediction wrappers.
export SURKV_WORKSPACE_ROOT=/path/to/SurKV
export LONGBENCH_DATA_DIR=/path/to/LongBench
export MODEL_PATH=/path/to/Meta-Llama-3-8B-Instruct
export METHOD=SurrogateKV
export KV_BUDGETS=128,512
bash run/longbench/scripts/run_llama/run_llama3_8b_instruct_surkv.shSaved predictions can be evaluated independently:
python3 run/longbench/eval.py \
--results_dir runs/longbench/meta-llama-3-8b-instruct_budget_128 \
--datasets qasper,multifieldqa_en,hotpotqa \
--methods SurrogateKV,SurrogateKV-Dynamic,SurrogateKV-AdaEvaluate one model-and-budget directory at a time; prediction outputs are
stored as <model>_budget_<B>/<dataset>/<method>.json.
See docs/REPRODUCIBILITY.md for the evaluation
scope, reported environment, and released artifacts.
LLaMA-3-8B-Instruct at B_KV = 512 (FullKV: 41.92):
| Base allocation | Base | Base score | SurrogateKV variant | Score | Delta |
|---|---|---|---|---|---|
| Shared token | SnapKV | 40.26 | SurrogateKV-Snap | 40.88 | +0.62 |
| Layer-wise | DynamicKV | 40.60 | SurrogateKV-Dynamic | 40.81 | +0.20 |
| Head-wise | Ada-KV | 40.77 | SurrogateKV-Ada | 41.26 | +0.49 |
The six-budget curves and per-dataset scores are in
data/longbench/llama3_8b_instruct/.
The x-axis uses the reported KV budgets on a linear scale.
Mistral-7B-Instruct-v0.2 at B_KV = 128:
| Base | Base score | SurrogateKV variant | Score |
|---|---|---|---|
| SnapKV | 87.51 | SurrogateKV-Snap | 98.84 |
| DynamicKV | 98.46 | SurrogateKV-Dynamic | 98.74 |
| Ada-KV | 90.04 | SurrogateKV-Ada | 98.18 |
The heatmap grids for B_KV = 64 and 128 are released under
data/niah/. The required head-wise evaluation path is recorded
in CORRECTIONS.md.
data/README.md indexes the CSV files for LongBench, NIAH,
motivation and attention diagnostics, ablations, merging comparisons, serving
efficiency, and model scaling. CSV is the canonical tabular format; duplicate
JSON exports are not needed to regenerate the released summaries. Run the data
checks and rebuild the README figures with:
python3 scripts/validate_release.py
python3 -m pip install -e ".[plots]"
python3 scripts/build_readme_figures.pypython3 -m compileall -q surrogatekv run scripts
python3 -m unittest discover -s tests -v
python3 -m ruff check surrogatekv run tests scripts@inproceedings{kim2026surrogatekv,
title = {SurrogateKV: Representation-Preserving KV Cache Compression for Long-Context LLMs},
author = {Kim, Junwon and Ryu, Junghyun and Talibli, Farid and So, Jungmin and Kim, Youngjae},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}Machine-readable citation metadata is available in CITATION.cff.
Copyright (c) 2026 Junwon Kim. All rights reserved. No license is granted for SurrogateKV's original code, documentation, data, figures, or other original material.
The identified third-party-derived portions remain subject to their original
copyright notices and license terms. Those terms do not license SurrogateKV's
original material. See THIRD_PARTY_NOTICES.md for
the applicable files, upstream sources, and full license texts.