Skip to content
AI Atlas
PaperActive

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

arxiv.org/abs/2609.09263

Updated 52 min ago · first seen 11 Sept 2026

paper_01M294GKZQQTSH356PF9Q7DNPM

Published
11 Sept 2026
T1 · 52 min ago
arXiv
2609.09263
T1 · 52 min ago
Category
cs.SD
T1 · 52 min ago

Abstract

Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the words shift the perceived gender of the output voice (voice rendering)? And when the model states the speaker's gender, does it follow the voice or the content (gender attribution)? We answer both with one controlled experiment crossing male and female voices with masculine-, neutral-, and feminine-stereotyped passages, on five open- and closed-source models in English, Spanish, and Mandarin. The rendered voice shows no stereotype drift. But every model decides the speaker's gender from the content, not the voice. Making the content one step more feminine (masculine -> neutral -> feminine) multiplies the odds of a "female" judgment by 1.7-24. When the content clashes with the voice, the worst model misgenders the speaker in 90% of cases. When they agree, it misgenders in only 2%. The bias thus hides in gender attribution, where fixed-voice evaluation cannot see, and where audits must look as S2S systems increasingly speak for real people.

Authors 4

Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

arXiv id
2609.09263

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Categories
cs.SD, cs.AI

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Primary category
cs.SD

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

52 min ago

Conflicts

None