How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
Published 16 Sept 2026arXiv:2512.01037
Updated 12 h ago · first seen 15 Sept 2026
paper_01M2JK0TPK76WRTQKWP4P4T762
Abstract
As safety alignment becomes standard in large language models, refusal behavior has become an important part of model reliability. However, models may still reject benign prompts, especially when the wording resembles risky content. Existing evaluations usually report global scores, such as false rejection rate or compliance rate. These scores are useful, but they treat each prompt independently. As a result, they miss local inconsistency, where a model accepts one phrasing of an intent but rejects a close paraphrase. This makes it difficult to understand whether refusals are only frequent or also semantically unstable. We address this gap by introducing Semantic Confusion, a failure mode that captures contradictory refusal decisions across meaning-preserving paraphrases. We build ParaGuard, a 10k-prompt corpus of controlled paraphrase clusters that keep intent fixed while varying surface form. We also propose three model-agnostic token-level metrics: Confusion Index, Confusion Rate, and Confusion Depth. These metrics compare each rejected prompt with its nearest accepted neighbors using token embeddings, next-token probabilities, and perplexity signals. Experiments across diverse model families show that global false rejection rate can hide important structure in the refusal boundary. Our metrics reveal cases where confusion is spread broadly, cases where it appears only in specific semantic regions, and cases where stricter refusal does not lead to more local inconsistency. These findings show that refusal evaluation should measure not only how often a model refuses, but also how consistently it refuses across nearby paraphrases. This gives developers a practical signal for reducing false refusals while preserving safety.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperHow Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - New paperPaperHow Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
New paper: How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.