MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
Published 15 Sept 2026arXiv:2601.20433
Updated 24 h ago · first seen 14 Sept 2026
paper_01M2F50EGGHDW1QKFM71ZEQ578
Abstract
Deepfake detection is a widely researched topic that is crucial for combating the spread of malicious content, with existing methods mainly modeling the problem as classification or spatial localization. The rapid advancements in generative models impose new demands on Deepfake detection. In this paper, we propose multimodal alignment and reinforcement for explainable Deepfake detection via vision-language models, termed MARE, which aims to enhance the accuracy and reliability of Vision-Language Models (VLMs) in Deepfake detection and reasoning. Specifically, MARE designs comprehensive reward functions, incorporating reinforcement learning from human feedback (RLHF), to incentivize the generation of text-spatially aligned reasoning content that adheres to human preferences. Besides, MARE introduces a forgery disentanglement module to capture intrinsic forgery traces from high-level facial semantics, thereby improving its authenticity detection capability. We conduct thorough evaluations on the reasoning content generated by MARE. Both quantitative and qualitative experimental results demonstrate that MARE achieves state-of-the-art performance in terms of accuracy and reliability.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperMARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models: published at changed from 2026-09-14T04:00:00+00:00 to 2026-09-15T04:00:00+00:00
Published14 Sept 2026→15 Sept 2026arxiv - New paperPaperMARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
New paper: MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
arxiv
Sources
Sources 1
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.