Skip to content
AI Atlas
PaperActive

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Applearxiv.org/pdf/2609.09973

Updated 52 min ago · first seen 11 Sept 2026

paper_01M294AHKESSJTHSANFDEAFPMD

Published
11 Sept 2026
T1 · 52 min ago
arXiv
2609.09973
T1 · 52 min ago

Abstract

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…

Authors 8

Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Simon Wang

Specification

arXiv id
2609.09973

Source:Apple Machine Learning ResearchT1observed 52 min agohigh

PDF

Source:Apple Machine Learning ResearchT1observed 52 min agohigh

Published
11 Sept 2026

Source:Apple Machine Learning ResearchT1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

6

Source tiers

T16

Freshest observation

52 min ago

Conflicts

None