MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
Updated 2 h ago · first seen 11 Sept 2026
paper_01M294GN2ZGGQ72ZAB9NHD2RHS
- Published
- 11 Sept 2026
- T1 · 2 h ago
- arXiv
- 2609.10016
- T1 · 2 h ago
- Category
- cs.LG
- T1 · 2 h ago
Abstract
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.
Authors 1
Remco Hendriks (Continker)
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- arXiv id
- 2609.10016
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Categories
- cs.LG, cs.AI, cs.CL
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Github repo
- continker/metrollm-bench
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Hf paper url
- https://huggingface.co/papers/2609.10016
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Github stars
- 0
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Hf comments
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
- Upvotes
- 3
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
14
Source tiers
T1T29 / 5
Freshest observation
2 h ago
Conflicts
4 flagged
No models linked to this paper yet.
- Authors
- Remco Hendriks (Continker)
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 2 min ago | 1 |
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 2 h ago | 2 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.