From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers
Updated 1 h ago · first seen 11 Sept 2026
paper_01M294AHMK7KAW0QPAZVK69J4V
- Published
- 27 Aug 2026
- T1 · 2 h ago
- arXiv
- 2608.23812
- T1 · 1 h ago
Abstract
Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the…
Authors 6
Aman Saini, Priyanshu Kumar, Eric Peng, Kai Yuan, Harsh Girase, Wanming Chen
Specification
- Paper
Source:Apple Machine Learning ResearchT1observed 2 h agohigh
- arXiv id
- 2608.23812
Source:Apple Machine Learning ResearchT1observed 1 h agohigh
Source:Apple Machine Learning ResearchT1observed 1 h agohigh
- Published
- 27 Aug 2026
Source:Apple Machine Learning ResearchT1observed 2 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
6
Source tiers
T16
Freshest observation
1 h ago
Conflicts
None
No models linked to this paper yet.
- Published by
- Apple
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Published
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 27 Aug 2026 | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperFrom Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge AnswersApple
New paper: From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers (Apple)
apple_ml
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Apple Machine Learning Research | machinelearning.apple.com/rss.xml | feed | T1· Official | 1 h ago | 1 |
| Apple Machine Learning Research | machinelearning.apple.com/research/rubric-based-alignment | paper_page | T1· Official | 1 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.