Skip to content
AI Atlas
PaperActive

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

arxiv.org/abs/2605.20555

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294GQ2CS3VSRRJ4SDXM56GJ

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2605.20555
T1 · 51 min ago
Category
cs.LG
T1 · 51 min ago

Abstract

-cross Abstract: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimization (GRPO). In contrast to Reinforcement Learning with Verifiable Rewards (RLVR) methods, our proposal does not involve a Kullback Leibler (KL) regularization or critic; the trainable policy and the reference anchor are coupled through the logit averaging structure to leverage the reasoning expertise of the trainable policy while maintaining the formatting advantage of SFT. Our method is evaluated on MATH, cn-k12, and MMLU, and the results show a higher accuracy or at least comparable accuracy relative to the canonical KL-regularized GRPO.

Authors 2

Xingwei Gan, Ying Zhu

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2605.20555

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.LG, cs.AI

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

51 min ago

Conflicts

None