StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
Published 16 Sept 2026arXiv:2609.16340
Updated 11 h ago · first seen 16 Sept 2026
paper_01M2MD8BFDA9741MVEJ4E9E1VK
Abstract
Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may already surpass. Moreover, collecting fresh post-edits for every new model is prohibitively expensive. We call this the Stale Preference problem. Standard DPO can fail in this setting: it may increase the likelihood of inferior post-edits, erode the model's existing quality, and fail to provide the per-token control needed to correct localized errors. We introduce StalePO, an objective derived from three requirements this regime imposes. Likelihood movement must be downward on both responses, the policy must be anchored to its own base response, and the KL constraint must apply at the token level. These requirements are jointly necessary. In ablations, each mechanism in isolation leaves the model's performance indistinguishable from the base model, and only their combination converts stale feedback into gains. On English-to-Hindi and English-to-Turkish localization data, StalePO improves the fraction of segments passing all LLM-as-judge MQM quality checks by 14.9 and 4.6 percentage points, respectively, with gains concentrated on style and fluency. A human evaluation under the same framework confirms these gains on English-to-Hindi, raising the fraction of segments passing all seven human checks by 13.8 percentage points.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperStalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperStalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
New paper: StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.