Skip to content
AI Atlas
PaperActive

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

arxiv.org/abs/2609.09768

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294GMPXJP1980EYFH1B7788

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2609.09768
T1 · 51 min ago
Category
cs.LG
T1 · 51 min ago

Abstract

In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a longer time to first token (TTFT). While previous works that reuse precomputed key-value (KV) caches effectively reduce TTFT for long-context inputs, it remains unclear whether response quality is preserved when the input context becomes very long. In this paper, we propose a combined approach that (i) fine-tunes the model while taking KV cache concatenation into account and (ii) selectively recomputes a subset of the KV caches. By applying both techniques, we demonstrate improved accuracy for long-context inputs. Experiments on the RULER benchmark show that, for a 124k-token input, our method improves the RULER score by 9.7 point over the baseline that recomputes KV caches only. Moreover, TTFT is reduced by 80% compared with full attention.

Authors 3

Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2609.09768

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.LG, cs.AI, cs.CL

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

51 min ago

Conflicts

None