Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
Updated 2 h ago · first seen 11 Sept 2026
paper_01M294AHKESSJTHSANFDEAFPMD
- Published
- 11 Sept 2026
- T1 · 2 h ago
- arXiv
- 2609.09973
- T1 · 2 h ago
Abstract
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…
Authors 8
Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Simon Wang
Specification
- Paper
Source:Apple Machine Learning ResearchT1observed 2 h agohigh
- arXiv id
- 2609.09973
Source:Apple Machine Learning ResearchT1observed 2 h agohigh
Source:Apple Machine Learning ResearchT1observed 2 h agohigh
- Published
- 11 Sept 2026
Source:Apple Machine Learning ResearchT1observed 2 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
6
Source tiers
T16
Freshest observation
2 h ago
Conflicts
None
No models linked to this paper yet.
- Published by
- Apple
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · arXiv id
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.09973 | → current | current | Apple Machine Learning ResearchT1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperPutting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question AnsweringApple
New paper: Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering (Apple)
apple_ml
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Apple Machine Learning Research | machinelearning.apple.com/rss.xml | feed | T1· Official | 1 h ago | 1 |
| Apple Machine Learning Research | machinelearning.apple.com/research/video-caption-quality | paper_page | T1· Official | 2 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.