Skip to content
AI Atlas
PaperActive

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

arxiv.org/abs/2609.05903

quality59

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294WYDRESKYVPQA3VZ64QGN

Published
5 Sept 2026
T2 · 7 h ago
arXiv
2609.05903
T2 · 7 h ago

Abstract

Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domains differ in the effects, state, and action sequences that must be governed. A harness that is strict enough for one model may over-block another, and a policy that transfers across domains may miss application-specific safety relations. We present EvoSafeHarness, a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain. It jointly searches a natural-language policy and executable code logic, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules. Across four agent benchmark families, EvoSafeHarness achieves a stronger safety-utility frontier than fixed expert-designed defenses. On DecodingTrust-Agent, it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost and achieves the best score in 14 of 15 cells. On AgentDojo, it reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at the same operating point, and transfers unchanged to unseen AgentDyn suites. It also achieves the best score on Agent-SafetyBench for every victim and keeps mean ASR below 20% under adaptive PAIR attacks with a refinement budget of 16. Analysis shows that domain semantics determine which safety relations and trajectory state are needed, while model and runtime behavior determine how and where those relations should be enforced.

Authors 7

Nanxi Li, Yingzi Ma, Yulong Cao, Edward Suh, Bo Li, Dawn Song, Chaowei Xiao

Specification

arXiv id
2609.05903

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium

Github repo
SaFo-Lab/EvoSafeHarness

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium

Hf paper url
https://huggingface.co/papers/2609.05903

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium

Github stars
1

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium

Hf comments
2

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium

Upvotes
36

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 2 h agomedium

Published
5 Sept 2026

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

11

Source tiers

T211

Freshest observation

2 h ago

Conflicts

None