MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
Updated 55 min ago · first seen 11 Sept 2026
paper_01M294FP54WN3G26SRFJRHP0BH
- Published
- 11 Sept 2026
- T1 · 55 min ago
- arXiv
- 2609.11253
- T1 · 55 min ago
- Category
- cs.LG
- T1 · 55 min ago
Abstract
Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a deterministic, dual counterfactual mapping. Bridging the gap between theoretical optimality and practical representation learning, we design an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern language models. Our framework enables seamless navigation between concept erasure and counterfactual generation. We empirically demonstrate its efficacy in improving downstream algorithmic fairness and generating counterfactual texts.
Authors 1
Antoine Saillenfest
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- arXiv id
- 2609.11253
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Categories
- cs.LG, cs.CL
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
55 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Antoine Saillenfest
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Official page
Official pageofficial_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2609.11253 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxivNew paper: MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 55 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 55 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.