Skip to content
AI Atlas
PaperActive

An Empirical Measurement of Jailbreaking Evaluators

arxiv.org/abs/2609.10594

Updated 31 min ago · first seen 11 Sept 2026

paper_01M294FPN1MH5V7TGKD8Y42WMS

Published
11 Sept 2026
T1 · 31 min ago
arXiv
2609.10594
T1 · 31 min ago
Category
cs.CR
T1 · 31 min ago

Abstract

Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their chosen evaluator independently, repeatedly spending resources on similar evaluation efforts while making results across papers difficult to compare. Different evaluators also encode different definitions of jailbreak success, meaning that reported attack strength and apparent progress can depend substantially on which evaluator is used. We systematically compare six evaluators that recur in recent jailbreak attack and defense research: HarmBench, JailbreakBench, JailbreakRadar, StrongReject, JADES, and JailMeter. To our knowledge, no prior study has evaluated all six on the same human-labeled data under a controlled setup. We evaluate them on JailbreakQR and JailMeter-Eva, using human judgments as the reference, and measure agreement with humans, error types, and consistency across attack families. For evaluators that require a general-purpose LLM judge, we use a shared backbone to control for model-specific variation. We found that JADES exhibits the best overall performance, while HarmBench and StrongReject also demonstrate good performance.

Authors 1

Yujie Mu

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

arXiv id
2609.10594

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Categories
cs.CR, cs.LG

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Primary category
cs.CR

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

31 min ago

Conflicts

None