Skip to content
AI Atlas
PaperActive

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Applearxiv.org/pdf/2608.23812

quality84

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294AHMK7KAW0QPAZVK69J4V

Published
27 Aug 2026
T1 · 6 h ago
arXiv
2608.23812
T1 · 6 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

6 claims · 6 properties

Paperpaper_url1

Claim history for Paper
ValueValid from → toStatusSourceConfidenceExtractor
https://machinelearning.apple.com/research/rubric-based-alignmentcurrentcurrentApple Machine Learning ResearchT1highdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the…currentcurrentApple Machine Learning ResearchT1highdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2608.23812currentcurrentApple Machine Learning ResearchT1highdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Aman Saini, Priyanshu Kumar, Eric Peng, Kai Yuan, Harsh Girase, Wanming ChencurrentcurrentApple Machine Learning ResearchT1highdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2608.23812currentcurrentApple Machine Learning ResearchT1highdeterministic

Publishedpublished_at1

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
27 Aug 2026currentcurrentApple Machine Learning ResearchT1highdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →