Skip to content
AI Atlas
PaperActive

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

arxiv.org/abs/2609.09428

quality89

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294GK3N4TYC27NK3ZP1HPF0

Published
11 Sept 2026
T1 · 6 h ago
arXiv
2609.09428
T1 · 6 h ago
Category
cs.AI
T1 · 6 h ago

Abstract

Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.

Authors 4

Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

arXiv id
2609.09428

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

6 h ago

Conflicts

None