Skip to content
AI Atlas
PaperActive

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

arxiv.org/abs/2601.08654

Updated 52 min ago · first seen 11 Sept 2026

paper_01M294GPQAAW9RKSZ8X9P74ZTV

Published
11 Sept 2026
T1 · 52 min ago
arXiv
2601.08654
T1 · 52 min ago
Category
cs.CL
T1 · 52 min ago

Abstract

-cross Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the same criteria inconsistently, produce score attributions that are difficult to audit, and map judgments poorly onto human scoring scales. We define this challenge as criteria transfer: translating human rubric intent into a stable, auditable inference-time scoring protocol. We introduce Rulers, which locks a task-level rubric specification, executes it through structured, evidence-grounded judgments, and calibrates the resulting signals to human score boundaries. Across four rubric-governed benchmarks and multiple frozen backbone models, Rulers achieves stronger agreement with human scores in most evaluated settings, while better matching empirical score distributions and remaining more stable under semantically equivalent rubric perturbations. Calibration controls and component ablations show that these gains cannot be attributed to post-hoc alignment alone, but depend on the combination of fixed criteria, traceable evidence, and calibrated score interpretation. These findings suggest that reliable LLM judging requires faithfully operationalizing human evaluation standards rather than relying on prompt-level scoring alone. Our code is available at https://github.com/LabRAI/Rulers.git.

Authors 6

Yihan Hong, Huaiyuan Yao, Bolin Shen, Wanpeng Xu, Hua Wei, Yushun Dong

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

arXiv id
2601.08654

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Categories
cs.CL, cs.AI, cs.LG

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

52 min ago

Conflicts

None