Skip to content
AI Atlas
PaperActive

PROOF-Gen: From Optimized Data to Better Distillation

Applearxiv.org/pdf/2608.23911

quality84

Updated 5 h ago · first seen 11 Sept 2026

paper_01M294AHMSE8H14PM7NQQ719F9

Published
26 Aug 2026
T1 · 5 h ago
arXiv
2608.23911
T1 · 5 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

6 claims · 6 properties

Paperpaper_url1

Claim history for Paper
ValueValid from → toStatusSourceConfidenceExtractor
https://machinelearning.apple.com/research/proof-gen-optimized-distillationcurrentcurrentApple Machine Learning ResearchT1highdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…currentcurrentApple Machine Learning ResearchT1highdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2608.23911currentcurrentApple Machine Learning ResearchT1highdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Anh Ta, Junjie Zhu, Shahin ShayandehcurrentcurrentApple Machine Learning ResearchT1highdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2608.23911currentcurrentApple Machine Learning ResearchT1highdeterministic

Publishedpublished_at1

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
26 Aug 2026currentcurrentApple Machine Learning ResearchT1highdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →