Updated 7 h ago · first seen 11 Sept 2026
paper_01M294G5P6WXM39MEKMAW71CJC
- Published
- 11 Sept 2026
- T1 · 7 h ago
- arXiv
- 2602.22787
- T1 · 7 h ago
- Category
- cs.CL
- T1 · 7 h ago
Abstract
Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relying on knowledge conflicts. Probes trained on AttriWiki achieve up to 0.96 Macro-$F_1$ on Llama-3.1-8B, Mistral-7B, and Qwen-7B, transfer to SQuAD and WebQuestions with 0.94-0.99 Macro-$F_1$, and generalise zero-shot to Tighidet et al. (2024)'s benchmark, outperforming their probe on conflicting settings without retraining. Furthermore, attribution mismatches raise error rates by up to 70%, though correct attribution does not guarantee correct answers, pointing to the need for broader detection frameworks.
Authors 3
Ivo Brink, Alexander Boer, Dennis Ulmer
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- arXiv id
- 2602.22787
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Categories
- cs.CL, cs.AI
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
7 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Ivo Brink, Alexander Boer, Dennis Ulmer
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Primary category
Primary categoryprimary_category1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.CL | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: Probing for Knowledge Attribution in Large Language Models
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 5 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.