Skip to content
AI Atlas
PaperActive

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

arxiv.org/abs/2609.11498

quality89

Updated 56 min ago · first seen 12 Sept 2026

paper_01M29X34NQJ95956756WS9XPG1

Published
12 Sept 2026
T1 · 56 min ago
arXiv
2609.11498
T1 · 56 min ago
Category
cs.AI
T1 · 56 min ago

Abstract

Practical uncertainty quantification (UQ) for large language models must decide, from a single generation, whether a specific answer should be trusted. Existing methods either sample multiple generations, read only output-token probabilities, or reduce the model's internal computation to a single hidden state. We introduce ActMap, a white-box representation that compresses the generation-time hidden- state trajectory (every layer, every generated token) into a fixed $12 \times 32 \times 128$ tensor of temporal-statistic channels that preserves structure across transformer depth and pooled hidden coordinates. The map is captured during the generation pass with no measurable overhead, has a fixed shape across model depths and hidden sizes, and occupies 96 KiB: a compact artifact that can be retained for audit-relevant generations and probed directly, with occlusion analysis localizing the classifier's signal to mid-depth regions of the map. A lightweight classifier, instantiated as a compact Vision Transformer, reads an estimated correctness probability from each map in a fraction of a millisecond; capacity-matched MLPs perform comparably, indicating the representation itself carries the result. Trained and evaluated in-domain on short-answer QA, direct- answer math, and summarization factuality with three instruction-tuned 7-8B models, ActMap consistently outperforms sampling, token-probability, attention, and embedding baselines, and matches ACT-ViT, a detector trained on dense activation tensors $67 \times$ larger, at essentially the same mean AUROC with lower calibration error on ten of twelve pairs. The resulting score supports abstention, routing, and selective verification from a single generation, making it a practical primitive for scalable oversight of deployed models.

Authors 2

Jacopo Dardini (University of Bologna), Roberta Calegari (University of Bologna)

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

arXiv id
2609.11498

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

56 min ago

Conflicts

None