Skip to content
AI Atlas
PaperActive

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

arxiv.org/abs/2609.09647

quality89

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294GK6QY6JJ59JE3TBFKADA

Published
11 Sept 2026
T1 · 6 h ago
arXiv
2609.09647
T1 · 6 h ago
Category
cs.AI
T1 · 6 h ago

Abstract

Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities. We present a systematic black-box framework for risk-aware agent evaluation requiring only basic system descriptions. Our approach introduces: (1) a seven-domain taxonomy mapping observable behaviors to risk categories, (2) fully automated SAGE-RT red teaming producing 120 adversarial scenarios per domain, and (3) human-validated evaluation using LLM judges. Empirical validation across two agent architectures (CrewAI and AutoGen) with four base models reveals alarming patterns: 56.25\% average governance risk, 65\% privacy risk in multi-agent configurations, and agent behavior vulnerabilities reaching 85\%. Our black-box approach effectively identifies critical architectural vulnerabilities without privileged access, providing a scalable path toward safer agent deployments.

Authors 5

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

arXiv id
2609.09647

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

6 h ago

Conflicts

None