Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Published 14 Sept 2026arXiv:2609.13005
Updated 2 h ago · first seen 14 Sept 2026
paper_01M2F4ZG0WB6CKG3HAV1Q6XAQ7
Abstract
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines. In this paper, we introduce Tasks over Application Manuals (TAM), a benchmark for evaluating long-horizon procedural reasoning. We construct TAM by curating real-world tasks from two domains: ICD-10-CM clinical coding (mapping medical conditions to diagnostic codes) and U.S. federal sentencing (computing crime sentencing guideline outcomes, specifically offense levels), with human-validated labels. Each task requires following an authoritative manual with tens of thousands of rules and executing a sequence of interdependent steps across different sections to produce an exact answer. We evaluate general-purpose prompting approaches, including retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5, and find that the best exact-match performance remains extremely low: 1% on ICD-10-CM coding and 15.5% on sentencing tasks. These results show that current benchmarks may overestimate LLM reasoning ability and miss a key challenge: reliably following long, rule-based procedures. The complete TAM data and code are publicly available.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperTasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - New paperPaperTasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
New paper: Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.