Skip to content
AI Atlas
PaperActive

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Applearxiv.org/pdf/2609.09973

quality84

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294AHKESSJTHSANFDEAFPMD

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2609.09973
T1 · 4 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

6 claims · 6 properties

Paperpaper_url1

Claim history for Paper
ValueValid from → toStatusSourceConfidenceExtractor
https://machinelearning.apple.com/research/video-caption-qualitycurrentcurrentApple Machine Learning ResearchT1highdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…currentcurrentApple Machine Learning ResearchT1highdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2609.09973currentcurrentApple Machine Learning ResearchT1highdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Simon WangcurrentcurrentApple Machine Learning ResearchT1highdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2609.09973currentcurrentApple Machine Learning ResearchT1highdeterministic

Publishedpublished_at1

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
11 Sept 2026currentcurrentApple Machine Learning ResearchT1highdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →