Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
Updated 8 h ago · first seen 11 Sept 2026
paper_01M294GK41Q0R1AMAKX5QG3N90
- Published
- 11 Sept 2026
- T1 · 8 h ago
- arXiv
- 2609.09448
- T1 · 8 h ago
- Category
- cs.AI
- T1 · 8 h ago
Abstract
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decisions. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), our methods consistently outperform surface level generation and sequence-based calibration baselines providing a zero-overhead reliability monitor that requires neither prompt alterations nor multi-sample rollouts.
Authors 3
Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- arXiv id
- 2609.09448
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Categories
- cs.AI
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Primary category
- cs.AI
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
8 h ago
Conflicts
None
No models linked to this paper yet.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperDo Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
New paper: Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 1 h ago | 2 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.