Skip to content
AI Atlas
PaperActive

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

arxiv.org/abs/2609.11677

quality89

Updated 2 h ago · first seen 12 Sept 2026

paper_01M29X34ZF4TVNNZTQMK3PWYX3

Published
12 Sept 2026
T1 · 2 h ago
arXiv
2609.11677
T1 · 2 h ago
Category
cs.SE
T1 · 2 h ago

Abstract

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded generalization to unseen tasks. We identify the lack of principled failure diagnosis as a key bottleneck in harness evolution: an observed failure can reflect either model-specific deficiencies or systematic harness deficiencies, and directly optimizing against individual failures can lead to unnecessary model-specific accommodation. We therefore propose Ecdysis, an efficient and effective framework that distinguishes model-specific accommodation from harness-level repair and biases adaptation toward systematic harness deficiencies by identifying recurring cross-task failure patterns. Ecdysis adopts a batch-level cross-instance failure aggregation paradigm to jointly analyze failure evidence from multiple task instances and further introduces Failure-Driven Collaborative Refinement to diagnose failure causes and iteratively refine harness modification specifications. By combining cross-instance failure analysis with multi-role diagnosis, Ecdysis enables more effective harness evolution with lower training time. Experiments show that Ecdysis achieves up to a 1.84x speedup in harness training compared with existing harness evolution methods, while improving the reasoning accuracy of the resulting harnesses by 18.56%.

Authors 14

Baohan Huang, Cong Zuo, Haibin Zhang, Ruiqing Yue, Sicheng Pan, Ting Li, Tingyu Li, Wenzhuo Zhu, Xianhong Xue, Yi Chen, Yifei Liu, Yu Cui, Zhe Cui, Zhuoyu Sun

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2609.11677

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.AI, cs.SE

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.SE

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None