EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Updated 3 h ago · first seen 11 Sept 2026
paper_01M294WYDRESKYVPQA3VZ64QGN
- Published
- 5 Sept 2026
- T2 · 9 h ago
- arXiv
- 2609.05903
- T2 · 9 h ago
Abstract
Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domains differ in the effects, state, and action sequences that must be governed. A harness that is strict enough for one model may over-block another, and a policy that transfers across domains may miss application-specific safety relations. We present EvoSafeHarness, a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain. It jointly searches a natural-language policy and executable code logic, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules. Across four agent benchmark families, EvoSafeHarness achieves a stronger safety-utility frontier than fixed expert-designed defenses. On DecodingTrust-Agent, it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost and achieves the best score in 14 of 15 cells. On AgentDojo, it reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at the same operating point, and transfers unchanged to unseen AgentDyn suites. It also achieves the best score on Agent-SafetyBench for every victim and keeps mean ASR below 20% under adaptive PAIR attacks with a refinement budget of 16. Analysis shows that domain semantics determine which safety relations and trajectory state are needed, while model and runtime behavior determine how and where those relations should be enforced.
Authors 7
Nanxi Li, Yingzi Ma, Yulong Cao, Edward Suh, Bo Li, Dawn Song, Chaowei Xiao
Specification
- Official page
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 9 h agomedium
- arXiv id
- 2609.05903
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 9 h agomedium
- Github repo
- SaFo-Lab/EvoSafeHarness
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 9 h agomedium
- Hf paper url
- https://huggingface.co/papers/2609.05903
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 9 h agomedium
- Github stars
- 1
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium
- Hf comments
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium
- Upvotes
- 36
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 9 h agomedium
- Published
- 5 Sept 2026
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 9 h agomedium
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
11
Source tiers
T211
Freshest observation
3 h ago
Conflicts
None
No models linked to this paper yet.
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Abstract
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domains differ in the effects, state, and action sequences that must be governed. A harness that is strict enough for one model may over-block another, and a policy that transfers across domains may miss application-specific safety relations. We present EvoSafeHarness, a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain. It jointly searches a natural-language policy and executable code logic, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules. Across four agent benchmark families, EvoSafeHarness achieves a stronger safety-utility frontier than fixed expert-designed defenses. On DecodingTrust-Agent, it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost and achieves the best score in 14 of 15 cells. On AgentDojo, it reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at the same operating point, and transfers unchanged to unseen AgentDyn suites. It also achieves the best score on Agent-SafetyBench for every victim and keeps mean ASR below 20% under adaptive PAIR attacks with a refinement budget of 16. Analysis shows that domain semantics determine which safety relations and trajectory state are needed, while model and runtime behavior determine how and where those relations should be enforced. | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
huggingface
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 3 h ago | 6 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.