"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
Updated 1 h ago · first seen 11 Sept 2026
paper_01M294G5FG894SGZDBTWEWXV6Q
- Published
- 11 Sept 2026
- T1 · 1 h ago
- arXiv
- 2508.05830
- T1 · 1 h ago
- Category
- cs.CL
- T1 · 1 h ago
Abstract
Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition). LLMs were prompted to predict depression scores in each condition. As expected, Mirror evaluations were near-perfect. However, Non-Mirror evaluations also displayed prediction sizes considered outstanding in psychology. Further, both Mirror and Non-Mirror predictions correlated with Patient Health Questionnaire-9 scores at similar sizes, suggesting the Mirror condition's advantage collapses when predicting an independent depression measurement. Topic modeling revealed differing depression-related themes across interview types. Mirror evaluations are better considered as reliability evaluations than as validity evaluations. Incorporating Non-Mirror approaches in LLM depression assessment may support more valid and clinically-relevant applications. Keywords: large language models, psychological assessment, psychopathology, depression, reliability, validity, criterion contamination
Authors 4
Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
- arXiv id
- 2508.05830
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
- Categories
- cs.CL, cs.CY
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
1 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Tong Li, Rasiq Hussain, Mehak Gupta
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · PDF
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2508.05830 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 1 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.