Updated 51 min ago · first seen 11 Sept 2026
paper_01M294AHMSE8H14PM7NQQ719F9
- Published
- 26 Aug 2026
- T1 · 52 min ago
- arXiv
- 2608.23911
- T1 · 51 min ago
Abstract
Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…
Authors 3
Anh Ta, Junjie Zhu, Shahin Shayandeh
Specification
- Paper
Source:Apple Machine Learning ResearchT1observed 52 min agohigh
- arXiv id
- 2608.23911
Source:Apple Machine Learning ResearchT1observed 51 min agohigh
Source:Apple Machine Learning ResearchT1observed 51 min agohigh
- Published
- 26 Aug 2026
Source:Apple Machine Learning ResearchT1observed 52 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
6
Source tiers
T16
Freshest observation
51 min ago
Conflicts
None
No models linked to this paper yet.
- Published by
- Apple
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Paperpaper_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://machinelearning.apple.com/research/proof-gen-optimized-distillation | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone… | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2608.23911 | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2608.23911 | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 26 Aug 2026 | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: PROOF-Gen: From Optimized Data to Better Distillation (Apple)
apple_ml
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Apple Machine Learning Research | machinelearning.apple.com/rss.xml | feed | T1· Official | 49 min ago | 1 |
| Apple Machine Learning Research | machinelearning.apple.com/research/proof-gen-optimized-distillation | paper_page | T1· Official | 51 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.