Skip to content
AI Atlas
PaperActive

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

arxiv.org/abs/2609.10016

quality77

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294GN2ZGGQ72ZAB9NHD2RHS

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2609.10016
T1 · 4 h ago
Category
cs.LG
T1 · 4 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history · Published

4 claims · 1 properties3 conflictingShow all properties

Publishedpublished_at4conflicting claims

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
9 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
9 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
9 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
11 Sept 2026currentcurrentarXiv (Atom API + RSS)T1conflicteddeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →