Skip to content
AI Atlas
PaperActive

Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

arxiv.org/abs/2609.10996

quality89

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294G4FCPXCEFDY0WBZCDPF2

Published
11 Sept 2026
T1 · 6 h ago
arXiv
2609.10996
T1 · 6 h ago
Category
cs.CL
T1 · 6 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history · Categories

1 claims · 1 propertiesShow all properties

Categoriescategories1

Claim history for Categories
ValueValid from → toStatusSourceConfidenceExtractor
cs.CLcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →