Skip to content
AI Atlas
PaperActive

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

arxiv.org/abs/2609.09448

quality89

Updated 7 h ago · first seen 11 Sept 2026

paper_01M294GK41Q0R1AMAKX5QG3N90

Published
11 Sept 2026
T1 · 7 h ago
arXiv
2609.09448
T1 · 7 h ago
Category
cs.AI
T1 · 7 h ago

Abstract

As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decisions. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), our methods consistently outperform surface level generation and sequence-based calibration baselines providing a zero-overhead reliability monitor that requires neither prompt alterations nor multi-sample rollouts.

Authors 3

Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

arXiv id
2609.09448

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

7 h ago

Conflicts

None