Skip to content
AI Atlas
PaperActive

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Applearxiv.org/pdf/2609.09973

quality84

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294AHKESSJTHSANFDEAFPMD

Published
11 Sept 2026
T1 · 2 h ago
arXiv
2609.09973
T1 · 2 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history · Abstract

1 claims · 1 propertiesShow all properties

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…currentcurrentApple Machine Learning ResearchT1highdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →