Skip to content
AI Atlas
PaperActive

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

arxiv.org/abs/2609.09783

Updated 52 min ago · first seen 11 Sept 2026

paper_01M294GMQQQZT52D563DMDS98R

Published
11 Sept 2026
T1 · 52 min ago
arXiv
2609.09783
T1 · 52 min ago
Category
cs.LG
T1 · 52 min ago

Abstract

Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE improves mean@1 on BrowseComp-Plus by $2.4\%$ over the strongest baseline, runs $2.46\times$ faster per step than synchronous training, and remains stable $50$ updates off-policy.

Authors 6

Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong, Zeyu Chen

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

arXiv id
2609.09783

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Categories
cs.LG, cs.AI

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

52 min ago

Conflicts

None