Skip to content
AI Atlas
Paper

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Published 16 Sept 2026arXiv:2605.16342

data quality84

Updated 9 h ago · first seen 17 Sept 2026

paper_01M2RME12ZPZPRW4EYPY1D1X2F

Abstract

Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO…

Authors

Authors 7

Amin Karimi MonsefiDominic CulverIrina BelousovaLokesh BoominathanManuel R. CiosiciNikhil BhendawadeYizhe Zhang

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 2

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 1

Full timeline →

Sources

Sources 2

Source documents
SourceDocumentTypeTierLast observedSnapshots
Apple Machine Learning Researchmachinelearning.apple.com/research/denoising-aware-credit-assignment paper_pageT1· Official9 h ago1
Apple Machine Learning Researchmachinelearning.apple.com/rss.xml feedT1· Official9 h ago2

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.