Skip to content
AI Atlas
PaperActive

TempCloze: Can Video-LLMs Identify the Missing Middle?

arxiv.org/abs/2609.01515

quality59

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294WYDYBCN8974K609HMSN8

Published
1 Sept 2026
T2 · 5 h ago
arXiv
2609.01515
T2 · 5 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

11 claims · 11 properties

Official pageofficial_url1

Claim history for Official page
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/abs/2609.01515currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2609.01515currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang DucurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github repogithub_repo1

Claim history for Github repo
ValueValid from → toStatusSourceConfidenceExtractor
CedricPei/Temporal-ClozecurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf paper urlhf_paper_url1

Claim history for Hf paper url
ValueValid from → toStatusSourceConfidenceExtractor
https://huggingface.co/papers/2609.01515currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github starsmetric.github_stars1

Claim history for Github stars
ValueValid from → toStatusSourceConfidenceExtractor
5currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf commentsmetric.hf_comments1

Claim history for Hf comments
ValueValid from → toStatusSourceConfidenceExtractor
1currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Upvotesmetric.upvotes1

Claim history for Upvotes
ValueValid from → toStatusSourceConfidenceExtractor
6currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2609.01515currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Publishedpublished_at1

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
1 Sept 2026currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →