Skip to content
AI Atlas
PaperActive

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

arxiv.org/abs/2609.11209

Updated 21 min ago · first seen 11 Sept 2026

paper_01M294FP3K3VKPY34S7B4Q01FG

Published
11 Sept 2026
T1 · 21 min ago
arXiv
2609.11209
T1 · 21 min ago
Category
cs.LG
T1 · 21 min ago

Abstract

Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the benefit of shorter prompts. We revisit RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views. We first show that modern compressors have unstable gains over simple truncation and can add substantial inference-time latency. We then propose Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agnostic score store. REVA maps token-level attention to readable word units, aggregates importance across repeated document accesses, and renders budget-specific plain-text views that preserve document order and the standard RAG interface. Across four representative benchmarks and modern LLMs, REVA improves generation quality by 1.0--5.8 points over existing advances, while reducing compression overhead by a factor of 5.3 to 15.6, adding less than 40 ms of latency.

Authors 6

Tuan Nguyen, Qiran Hu, Banruo Liu, Khoa D. Doan, Kok-Seng Wong, Fan Lai

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

arXiv id
2609.11209

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Categories
cs.LG, cs.CL, cs.IR

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

21 min ago

Conflicts

None