Skip to content
AI Atlas
Papercs.LG

The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards

Published 15 Sept 2026arXiv:2609.13866

data quality89

Updated 29 h ago · first seen 15 Sept 2026

paper_01M2JK0BSFKCW72ACR3TZKAF0W

Abstract

Group-relative policy optimization (GRPO and descendants) can discard no-contrast rollout groups through dynamic sampling, while practical implementations expose a configurable filter metric. We identify and quantify a metric-predicate mismatch under composite shaped rewards. When filtering follows the shaped training score rather than the task outcome, all-fail groups retain nonzero within-group spread and pass the predicate; standard-deviation normalization then promotes shaping differences among failures to full-size phantom advantages. In a controlled GSM8K comparison (Qwen2.5-1.5B, LoRA), no filtering and shaped-score filtering end at EM 0.080 +/- 0.112 and 0.040 +/- 0.008, whereas binary-outcome filtering holds 0.754 +/- 0.005 across four runs per arm (three default-seed reruns and one seed-123 run; mean +/- sample SD). On verl's native recipe/dapo trainer, holding model, data, reward and trainer fixed and changing only the metric, the score arm requires no batch refill in any of 40 observed steps and ends at EM 0.160; the accuracy arm refills in 29/40 steps and ends at 0.763. Both use the same custom shaped-reward hook and unmodified trainer/filter code. Prior work established shaping-induced amplification and all-fail filtering; our contribution isolates the metric-predicate semantic mismatch and directly instruments native deletion/refill telemetry. Across tested positive coefficients lambda in {0.1, 0.3, 0.5}, unsafe arms collapse; exploratory one-run cells reproduce the failure at 1.5B/7B on MATH and under GSPO, while disabling standard-deviation normalization avoids the observed collapse. Filtering under a composite reward should use a task-outcome signal whose semantics are independent of shaping.

Authors

Authors 1

Juntao Yu

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 0

No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 1

Full timeline →

Sources

Sources 1

Source documents
SourceDocumentTypeTierLast observedSnapshots
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.LG feedT1· Official21 h ago3

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.