DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Published 16 Sept 2026arXiv:2605.16342
Updated 9 h ago · first seen 17 Sept 2026
paper_01M2RME12ZPZPRW4EYPY1D1X2F
Abstract
Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO…
Organizations
Organizations 2
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 1
- New paperPaperDACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language ModelsApple
New paper: DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models (Apple)
apple_ml
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.