The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech
Published 15 Sept 2026arXiv:2609.13150
Updated 29 h ago · first seen 15 Sept 2026
paper_01M2JK0C7J7Y74JTZKHCEKDGM5
Abstract
Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingly adopted as reward signals for preference optimization. Both roles presuppose that the predicted score tracks human preference. In this work, we test this assumption across six human-rated corpora spanning the quality range from artifact-rich to defect-free TTS, evaluating each predictor on a pairwise task that asks whether the clip it scores higher is the clip listeners prefer, and we subject interpretable prosodic and signal-processing features to the same protocol. When one clip carries audible defects the predictors tend to agree with listeners. Once both clips are clean, no single predictor reliably identifies the preferred sample, and several fall below the accuracy of simply picking the longest-duration clip. A calibrated composite of complementary signals is the strongest evaluator we test, though on the cleanest audio it recovers only part of the gap to the human ceiling. Additionally, using even an equal-weighted ensemble of metrics helps as a post-training reward, where no calibration data is available. Optimizing a single score with policy optimization induces reward hacking, driving the metric toward its optimum while independent held-out judges and a human listening test deteriorate. The composite reward resists this behavior and tends to improve the model. Our contribution is the evaluation protocol, the predictor scores across these corpora, and the diagnosis of when and why single scores fail.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 1
- New paperPaperThe Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech
New paper: The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech
arxiv
Sources
Sources 1
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.