Skip to content
AI Atlas
PaperActive

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

arxiv.org/abs/2609.10016

quality77

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294GN2ZGGQ72ZAB9NHD2RHS

Published
11 Sept 2026
T1 · 5 h ago
arXiv
2609.10016
T1 · 5 h ago
Category
cs.LG
T1 · 5 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

18 claims · 14 properties4 conflicting

Official pageofficial_url1

Claim history for Official page
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/abs/2609.10016currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Arxiv announce typearxiv_announce_type1

Claim history for Arxiv announce type
ValueValid from → toStatusSourceConfidenceExtractor
crosscurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2609.10016currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Authorsauthors3conflicting claims

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Remco HendrikscurrentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
Remco HendrikscurrentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
Remco Hendriks (Continker)currentcurrentarXiv (Atom API + RSS)T1conflicteddeterministic

Categoriescategories1

Claim history for Categories
ValueValid from → toStatusSourceConfidenceExtractor
cs.LG, cs.AI, cs.CLcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

Github repogithub_repo1

Claim history for Github repo
ValueValid from → toStatusSourceConfidenceExtractor
continker/metrollm-benchcurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf paper urlhf_paper_url1

Claim history for Hf paper url
ValueValid from → toStatusSourceConfidenceExtractor
https://huggingface.co/papers/2609.10016currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github starsmetric.github_stars1

Claim history for Github stars
ValueValid from → toStatusSourceConfidenceExtractor
0currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf commentsmetric.hf_comments1

Claim history for Hf comments
ValueValid from → toStatusSourceConfidenceExtractor
2currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Upvotesmetric.upvotes1

Claim history for Upvotes
ValueValid from → toStatusSourceConfidenceExtractor
3currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2609.10016currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Primary categoryprimary_category1

Claim history for Primary category
ValueValid from → toStatusSourceConfidenceExtractor
cs.LGcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

Publishedpublished_at3conflicting claims

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
9 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
9 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
11 Sept 2026currentcurrentarXiv (Atom API + RSS)T1conflicteddeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →