Skip to content
AI Atlas
PaperActive

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

arxiv.org/abs/2608.22230

quality89

Updated 5 h ago · first seen 11 Sept 2026

paper_01M294G65TXEVPMJQN0BTMDJ7S

Published
11 Sept 2026
T1 · 5 h ago
arXiv
2608.22230
T1 · 5 h ago
Category
cs.CL
T1 · 5 h ago

Abstract

Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.

Authors 11

Junyu Lu, Kaiyuan Liu, Kaichun Wang, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu, Qi Zhu, Bo Xu, Liang Yang, Hongfei Lin

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

arXiv id
2608.22230

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

5 h ago

Conflicts

None