Skip to content
AI Atlas
PaperActive

Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

arxiv.org/abs/2609.10410

Updated 50 min ago · first seen 11 Sept 2026

paper_01M294GNJH9862SHQXF6G9M69K

Published
11 Sept 2026
T1 · 50 min ago
arXiv
2609.10410
T1 · 50 min ago
Category
cs.CL
T1 · 50 min ago

Abstract

The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.

Authors 9

Ayan Majumdar, Shounak Paul, Pushpdeep Singh, Ines Abdelaziz, Sayeh Jarollahi, Seungeon Lee, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

arXiv id
2609.10410

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Categories
cs.CL, cs.AI, cs.CY

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

50 min ago

Conflicts

None