Skip to content
AI Atlas
PaperActive

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

arxiv.org/abs/2609.11545

quality89

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294G4RBWPHCY7JR711H3X8J

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2609.11545
T1 · 4 h ago
Category
cs.CL
T1 · 4 h ago

Abstract

Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, and content consistency using regular test sentences, while providing limited insight into how multilingual TTS systems fail when handling challenging inputs such as numbers, dates, named entities, long sentences, code-switched expressions, and punctuation-related structures. This paper proposes a complex-text robustness diagnosis framework for low-resource multilingual TTS. We evaluate robustness from three dimensions: content consistency, language consistency, and generation stability. A multilingual robustness testing scheme is designed for Thai, Vietnamese, Swahili, and Indonesian, covering ordinary sentences and multiple types of complex text inputs. We further introduce automatic diagnostic metrics, including character error rate, language identification accuracy, and duration abnormal rate. To support input-level risk analysis before speech generation, we propose a lightweight Text Risk Score (TRS), which estimates synthesis risk from interpretable text features without manual annotation or model training. Experiments on three representative multilingual TTS systems, including OmniVoice, VoxCPM2, and MMS-TTS, show that complex text inputs expose systematic failure patterns that are not fully reflected by ordinary short-sentence evaluation. Different systems exhibit distinct vulnerabilities in number normalization, named entity handling, long-text generation, and code-switched input processing. Furthermore, TRS shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.

Authors 5

Tianlun Zuo, Ziyu Zhang, Tingzhi Mao, Zhonghua Fu, Lei Xie

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

arXiv id
2609.11545

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Categories
cs.CL, cs.SD

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

4 h ago

Conflicts

None