Skip to content
AI Atlas
PaperActive

When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents

arxiv.org/abs/2609.10750

quality89

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294FQ1X35ZMAABV1N1TV9W4

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2609.10750
T1 · 4 h ago
Category
cs.IR
T1 · 4 h ago

Abstract

LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the performance on OOD skills retrieval but also improve the retrieval on synthetic in-distribution skills by 13.98\% for 0.6B Qwen retriever and reranker. Our results provide a practical benchmark and a robust fine-tuning recipe for scarce, multi-positive supervision.

Authors 5

Syed Shariyar Murtaza, Yifan Nie, Utkarsh Soni, Eugene Wen, Arvid Frydenlund

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

arXiv id
2609.10750

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Categories
cs.IR, cs.AI, cs.LG

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Primary category
cs.IR

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

4 h ago

Conflicts

None