Skip to content
AI Atlas
PaperActive

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

arxiv.org/abs/2609.11758

quality89

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294G4TBTF3WTTANXFVXJV1T

Published
11 Sept 2026
T1 · 3 h ago
arXiv
2609.11758
T1 · 3 h ago
Category
cs.CL
T1 · 3 h ago

Abstract

Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overall safety of the generated responses, when prompted for harmful or dangerous content. A clearer understanding of the mechanisms leading to this result is needed, as increasing numbers of end users turn to RAG to incorporate corporate documents and knowledge bases into LLM-based systems. We introduce RAG-Safety-Bench, a benchmark to measure the safety impact of RAG on LLM models. By removing the confounding effect of retriever quality, and cleanly separating the problem into four conditions -- non-RAG, RAG with an oracle document containing the answer to the harmful request, RAG with documents related to the harmful request but without the specific answer, and RAG with random, safe documents -- the benchmark isolates the impacts of different factors in the observed safety degradation. We report results across five open-source LLMs, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-enabled systems.

Authors 2

Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

arXiv id
2609.11758

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Categories
cs.CL, cs.IR

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

3 h ago

Conflicts

None