Skip to content
AI Atlas
PaperActive

SoK: Privacy Attacks on Machine Learning via Explainable AI

arxiv.org/abs/2609.10627

quality89

Updated 1 h ago · first seen 11 Sept 2026

paper_01M294FPS46W4AMF2Y9EXQXW59

Published
11 Sept 2026
T1 · 1 h ago
arXiv
2609.10627
T1 · 1 h ago
Category
cs.CR
T1 · 1 h ago

Abstract

Machine learning explanations reveal model behavior beyond predictions, creating attack surfaces for model confidentiality and data privacy. We systematize 25 studies that exploit explanations for model extraction, membership inference, and model inversion, treating attribute inference as partial inversion. Existing work is often labeled only black- or white-box, obscuring substantial differences in what explanation signal reaches an adversary. We therefore separate model knowledge from explanation acquisition and identify five paths: target-released, attacker-derived, secondary disclosure, privileged access, and released global artifacts. Across these paths, explanations reduce extraction cost, expose membership signals through explanation statistics, recourse distance, and explanation-guided robustness, and support spatial or algebraic reconstruction of private inputs. We compare system and threat models, explanation signals, auxiliary knowledge, target models, modalities, query budgets, evaluation metrics, reported performance, and defenses. Our analysis shows that no explanation family is uniformly unsafe and no defense is uniformly effective. Risk depends on which signal is exposed, how it is acquired, which asset is targeted, and what the attacker already knows. We argue that explanation privacy should therefore be evaluated as an end-to-end disclosure problem, with defenses matched to the acquisition path and protected asset.

Authors 3

Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

arXiv id
2609.10627

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Categories
cs.CR, cs.LG

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Primary category
cs.CR

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

1 h ago

Conflicts

None