Skip to content
AI Atlas

How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries

Published 16 Sept 2026arXiv:2512.01037

data quality89

Updated 12 h ago · first seen 15 Sept 2026

paper_01M2JK0TPK76WRTQKWP4P4T762

Abstract

As safety alignment becomes standard in large language models, refusal behavior has become an important part of model reliability. However, models may still reject benign prompts, especially when the wording resembles risky content. Existing evaluations usually report global scores, such as false rejection rate or compliance rate. These scores are useful, but they treat each prompt independently. As a result, they miss local inconsistency, where a model accepts one phrasing of an intent but rejects a close paraphrase. This makes it difficult to understand whether refusals are only frequent or also semantically unstable. We address this gap by introducing Semantic Confusion, a failure mode that captures contradictory refusal decisions across meaning-preserving paraphrases. We build ParaGuard, a 10k-prompt corpus of controlled paraphrase clusters that keep intent fixed while varying surface form. We also propose three model-agnostic token-level metrics: Confusion Index, Confusion Rate, and Confusion Depth. These metrics compare each rejected prompt with its nearest accepted neighbors using token embeddings, next-token probabilities, and perplexity signals. Experiments across diverse model families show that global false rejection rate can hide important structure in the refusal boundary. Our metrics reveal cases where confusion is spread broadly, cases where it appears only in specific semantic regions, and cases where stricter refusal does not lead to more local inconsistency. These findings show that refusal evaluation should measure not only how often a model refuses, but also how consistently it refuses across nearby paraphrases. This gives developers a practical signal for reducing false refusals while preserving safety.

Authors

Authors 3

Md Labid Al NahiyanMd Tanvir HassanRiad Ahmed Anonto

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 0

No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 2

Full timeline →

Sources

Sources 2

Source documents
SourceDocumentTypeTierLast observedSnapshots
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.AI feedT1· Official7 h ago5
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.CL feedT1· Official7 h ago4

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.