Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Published 16 Sept 2026arXiv:2609.14896
Updated 12 h ago · first seen 15 Sept 2026
paper_01M2JK0CRZ7FBJE92D3VC320J5
Abstract
A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring open-ended exploration and pluralistic perspectives, such as scientific ideation and creative writing. We present MoDA (Mode-conditioned Diversity Alignment), an online post-training RL algorithm that jointly optimizes generation quality and diversity, inspired by the coordination perspective in multi-agent reinforcement learning (MARL). MoDA trains a single shared LLM policy conditioned on abstract numbered roles, where each role acts as an agent competing to produce outputs distinct from the others. This formulation encourages mode-conditioned agents to explore complementary regions of the high-quality output space without requiring hand-crafted personas or architectural modifications. MoDA employs a prompt-adaptive quality gating mechanism that calibrates a reference quality threshold and grants diversity rewards only to responses that meet the threshold, preventing reward-hacking behaviors that compromise response quality. To study quality-diversity tradeoffs, we evaluate MoDA on a comprehensive suite of benchmarks spanning seven general capability tasks and four domain-specific diversity tasks in scientific ideation and creative writing. MoDA improves SBERT diversity by 265% on the Infinite-Chat held-out prompts, while increasing average general capability pass@1 by 10.3% over the Qwen3-8B baseline. Compared with the strongest DivPO baseline, MoDA improves SBERT diversity from 0.274 to 0.482 (+75.9%) and E-Vendi from 2.86 to 4.4 (+53.8%), while improving average general capability pass@1 by 7.0%. Overall, MoDA provides a drop-in alternative to standard post-training methods that preserves and expands the model's expressive output space while improving quality.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
- Property changedPaperForty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - Property changedPaperForty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperForty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperForty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
New paper: Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
arxiv
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.