Skip to content
AI Atlas
PaperActive

Legible Failures: Detecting and Repairing In-Context Binding Errors

arxiv.org/abs/2609.11216

Updated 29 min ago · first seen 11 Sept 2026

paper_01M294FP407CJX0TG0GKT81FPY

Published
11 Sept 2026
T1 · 29 min ago
arXiv
2609.11216
T1 · 29 min ago
Category
cs.LG
T1 · 29 min ago

Abstract

A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occurs across 16 public checkpoints, each evaluated with three seeds. We fit a probe on a training fold, select its layer on a validation fold, and report results on a disjoint test fold. On the trials each model gets wrong, probe accuracy exceeds the strict present-obligation baseline, 1/K = 0.125, by +0.196 (95% CI [+0.101, +0.296], bootstrapped over models). A query-entity counterfactual rules out token presence and recency. A score built from the sign of probe-output disagreement improves failure detection over the model's own confidence by +0.079 AUROC (95% CI [+0.036, +0.126]). Raw probe confidence gives no measurable improvement over model confidence. Steering the residual stream toward the probe-decoded binding, with no gold label, raises accuracy on all eight models tested by a mean of +0.168 (95% CI [+0.066, +0.280]). Where recent studies report that probe-detected errors are resistant to interventions, we find that in-context binding is a setting in which probes are actionable.

Authors 3

Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

arXiv id
2609.11216

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

Categories
cs.LG

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 29 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

29 min ago

Conflicts

None