Skip to content
AI Atlas
PaperActive

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Applearxiv.org/pdf/2608.23812

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294AHMK7KAW0QPAZVK69J4V

Published
27 Aug 2026
T1 · 52 min ago
arXiv
2608.23812
T1 · 51 min ago

Abstract

Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the…

Authors 6

Aman Saini, Priyanshu Kumar, Eric Peng, Kai Yuan, Harsh Girase, Wanming Chen

Specification

arXiv id
2608.23812

Source:Apple Machine Learning ResearchT1observed 51 min agohigh

PDF

Source:Apple Machine Learning ResearchT1observed 51 min agohigh

Published
27 Aug 2026

Source:Apple Machine Learning ResearchT1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

6

Source tiers

T16

Freshest observation

51 min ago

Conflicts

None