Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
Published 18 Sept 2026arXiv:2609.19472
Updated 4 h ago · first seen 18 Sept 2026
paper_01M2SEG2ZPK5MM8CAY9R4BRKWK
Abstract
Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperSafety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperSafety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
New paper: Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.