BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
Updated 4 h ago · first seen 11 Sept 2026
paper_01M294GMQQQZT52D563DMDS98R
- Published
- 11 Sept 2026
- T1 · 4 h ago
- arXiv
- 2609.09783
- T1 · 4 h ago
- Category
- cs.LG
- T1 · 4 h ago
Abstract
Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE improves mean@1 on BrowseComp-Plus by $2.4\%$ over the strongest baseline, runs $2.46\times$ faster per step than synchronous training, and remains stable $50$ updates off-policy.
Authors 6
Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong, Zeyu Chen
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- arXiv id
- 2609.09783
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Categories
- cs.LG, cs.AI
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
4 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Guanqun Zhao, Zijun Xie, Binbin Zheng
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · arXiv id
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.09783 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 2 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.