Generative Late-Interaction Embeddings For Visual Document Retrieval
Updated 43 min ago · first seen 11 Sept 2026
paper_01M294WYE9ET6SHX55RVAMPEKA
- Published
- 10 Sept 2026
- T2 · 45 min ago
- arXiv
- 2609.11808
- T2 · 45 min ago
Abstract
Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.
Authors 8
Mohamed Eltahir, Talal Aloushan, Rose Khairoalsendi, Jana Shata, Mohammed Alhassan, Leen Alrehaili, Tanveer Hussain, Naeemullah Khan
Specification
- Official page
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 45 min agomedium
- arXiv id
- 2609.11808
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 45 min agomedium
- Github repo
- mohammad2012191/GLIE
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 45 min agomedium
- Hf paper url
- https://huggingface.co/papers/2609.11808
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 45 min agomedium
- Github stars
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 43 min agomedium
- Hf comments
- 1
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 43 min agomedium
- Upvotes
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 43 min agomedium
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 45 min agomedium
- Published
- 10 Sept 2026
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 45 min agomedium
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
11
Source tiers
T211
Freshest observation
43 min ago
Conflicts
None
No models linked to this paper yet.
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Official pageofficial_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2609.11808 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface. | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.11808 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Github repogithub_repo1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| mohammad2012191/GLIE | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Hf paper urlhf_paper_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://huggingface.co/papers/2609.11808 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Github starsmetric.github_stars1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Hf commentsmetric.hf_comments1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 1 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Upvotesmetric.upvotes1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2609.11808 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 10 Sept 2026 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: Generative Late-Interaction Embeddings For Visual Document Retrieval
huggingface
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 43 min ago | 2 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.