TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On
Published 15 Sept 2026arXiv:2609.13259
Updated 29 h ago · first seen 15 Sept 2026
paper_01M2JK197EFYW0Q6WNGGTRKHWV
Abstract
Virtual Try-On (VTON) aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turning this preference-oriented goal into an actionable objective relies on a scoring function aligned with human taste. However, classic fidelity metrics exhibit weak correlation with human judgments, and generic VLMs fail to provide the discriminative granularity demanded by try-on quality evaluation, which hinges on faithfully preserving garment and person details. This shortcoming is further exacerbated in the reinforcement fine-tuning (RFT) optimization and leads to severe reward hacking. To this end, we present TryOnReward, a fine-grained reward model tailored for VTON. Built on a vision-language backbone, it adopts a foveation calibration objective that grounds each quality dimension in the relevant region to avoid global shortcut learning. Meanwhile, TryOnReward jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision, leveraging both relative and absolute quality signals. For model training and evaluation, we build TryOnReward-100K, a human-annotated per-dimension rating dataset, alongside TryOn-Bench and TryOnRewardBench, two benchmarks covering diverse real scenarios. Extensive experiments confirm that TryOnReward significantly outperforms generic judges in human preference alignment, and when serving as the RFT reward function, it consistently yields human-preferred try-on results across multiple baselines.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperTryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On
TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperTryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On
New paper: TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.