When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
Updated 5 h ago · first seen 11 Sept 2026
paper_01M294FQ1X35ZMAABV1N1TV9W4
- Published
- 11 Sept 2026
- T1 · 5 h ago
- arXiv
- 2609.10750
- T1 · 5 h ago
- Category
- cs.IR
- T1 · 5 h ago
Abstract
LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the performance on OOD skills retrieval but also improve the retrieval on synthetic in-distribution skills by 13.98\% for 0.6B Qwen retriever and reranker. Our results provide a practical benchmark and a robust fine-tuning recipe for scarce, multi-positive supervision.
Authors 5
Syed Shariyar Murtaza, Yifan Nie, Utkarsh Soni, Eugene Wen, Arvid Frydenlund
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- arXiv id
- 2609.10750
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Categories
- cs.IR, cs.AI, cs.LG
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Primary category
- cs.IR
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
5 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Syed Shariyar Murtaza, Yifan Nie, Utkarsh Soni
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · PDF
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2609.10750 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperWhen Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
New paper: When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 3 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.