Skip to content
AI Atlas
PaperActive

PROOF-Gen: From Optimized Data to Better Distillation

Applearxiv.org/pdf/2608.23911

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294AHMSE8H14PM7NQQ719F9

Published
26 Aug 2026
T1 · 52 min ago
arXiv
2608.23911
T1 · 51 min ago

Abstract

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…

Authors 3

Anh Ta, Junjie Zhu, Shahin Shayandeh

Specification

arXiv id
2608.23911

Source:Apple Machine Learning ResearchT1observed 51 min agohigh

PDF

Source:Apple Machine Learning ResearchT1observed 51 min agohigh

Published
26 Aug 2026

Source:Apple Machine Learning ResearchT1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

6

Source tiers

T16

Freshest observation

51 min ago

Conflicts

None