Skip to content
AI Atlas
PaperActive

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

arxiv.org/abs/2508.05830

Updated 40 min ago · first seen 11 Sept 2026

paper_01M294G5FG894SGZDBTWEWXV6Q

Published
11 Sept 2026
T1 · 40 min ago
arXiv
2508.05830
T1 · 40 min ago
Category
cs.CL
T1 · 40 min ago

Abstract

Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition). LLMs were prompted to predict depression scores in each condition. As expected, Mirror evaluations were near-perfect. However, Non-Mirror evaluations also displayed prediction sizes considered outstanding in psychology. Further, both Mirror and Non-Mirror predictions correlated with Patient Health Questionnaire-9 scores at similar sizes, suggesting the Mirror condition's advantage collapses when predicting an independent depression measurement. Topic modeling revealed differing depression-related themes across interview types. Mirror evaluations are better considered as reliability evaluations than as validity evaluations. Incorporating Non-Mirror approaches in LLM depression assessment may support more valid and clinically-relevant applications. Keywords: large language models, psychological assessment, psychopathology, depression, reliability, validity, criterion contamination

Authors 4

Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

arXiv id
2508.05830

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

Categories
cs.CL, cs.CY

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 40 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

40 min ago

Conflicts

None