LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
Updated 9 h ago · first seen 11 Sept 2026
paper_01M294GK9AXN4MZXZD0EXGWB3X
- Published
- 11 Sept 2026
- T1 · 9 h ago
- arXiv
- 2609.09754
- T1 · 9 h ago
- Category
- cs.AI
- T1 · 9 h ago
Abstract
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414 instances across 17 legal categories and 6 task types. Each instance is annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses, covering both substantive errors and agent-procedural failures. We further design fine-grained metrics that quantify to what extent and localize how each failure occurs along an agent's execution path. Our evaluation across 18 proprietary and open-source agents uncovers a Right-Answer-Wrong-Reason effect and reveals that hallucination subclasses cluster rather than scatter, forming distinct agentic framework, legal task, and category profiles. These findings, invisible to outcome-level evaluation, validate the diagnostic power of LexAgentHallu for evaluating agentic hallucination in law.
Authors 7
Yujin Zhou, Mingxuan Zheng, Chuxue Cao, Huang Yidan, Jiale Chen, Yike Guo, Sirui Han
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- arXiv id
- 2609.09754
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Categories
- cs.AI
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Primary category
- cs.AI
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
9 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Yujin Zhou, Mingxuan Zheng, Chuxue Cao
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Categories
Categoriescategories1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.AI | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 7 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.