Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
Published 18 Sept 2026arXiv:2609.19799
Updated 4 h ago · first seen 18 Sept 2026
paper_01M2SEG31FCD21WEM04QMJKAZB
Abstract
LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterations (depth) changes with the strategy, the task, and the total budget. Furthermore, we observe that the ranking of strategies also changes with the budget. On one task the strategy that looks worst at one seed is best at forty seeds. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score. We provide a measurement protocol that reports the seeds-by-iterations frontier and practical guidance for using it.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 3
Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxivEvolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxivNew paper: Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
arxiv
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.