Updated 31 min ago · first seen 11 Sept 2026
paper_01M294WYDYBCN8974K609HMSN8
- Published
- 1 Sept 2026
- T2 · 2 h ago
- arXiv
- 2609.01515
- T2 · 2 h ago
Abstract
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.
Authors 7
Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du
Specification
- Official page
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- arXiv id
- 2609.01515
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Github repo
- CedricPei/Temporal-Cloze
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Hf paper url
- https://huggingface.co/papers/2609.01515
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Github stars
- 5
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 31 min agomedium
- Hf comments
- 1
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 31 min agomedium
- Upvotes
- 6
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 31 min agomedium
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Published
- 1 Sept 2026
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
11
Source tiers
T211
Freshest observation
31 min ago
Conflicts
None
No models linked to this paper yet.
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: TempCloze: Can Video-LLMs Identify the Missing Middle?
huggingface
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 31 min ago | 3 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.