REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving
Updated 50 min ago · first seen 11 Sept 2026
paper_01M294FP3K3VKPY34S7B4Q01FG
- Published
- 11 Sept 2026
- T1 · 50 min ago
- arXiv
- 2609.11209
- T1 · 50 min ago
- Category
- cs.LG
- T1 · 50 min ago
Abstract
Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the benefit of shorter prompts. We revisit RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views. We first show that modern compressors have unstable gains over simple truncation and can add substantial inference-time latency. We then propose Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agnostic score store. REVA maps token-level attention to readable word units, aggregates importance across repeated document accesses, and renders budget-specific plain-text views that preserve document order and the standard RAG interface. Across four representative benchmarks and modern LLMs, REVA improves generation quality by 1.0--5.8 points over existing advances, while reducing compression overhead by a factor of 5.3 to 15.6, adding less than 40 ms of latency.
Authors 6
Tuan Nguyen, Qiran Hu, Banruo Liu, Khoa D. Doan, Kok-Seng Wong, Fan Lai
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
- arXiv id
- 2609.11209
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
- Categories
- cs.LG, cs.CL, cs.IR
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 50 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
50 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Tuan Nguyen, Qiran Hu, Banruo Liu
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxivNew paper: REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 50 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 50 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.