Skip to content
AI Atlas
PaperActive

SpecGuard: Inference-Time Backdoor Detection For Free

arxiv.org/abs/2609.11799

quality89

Updated 1 h ago · first seen 11 Sept 2026

paper_01M294G5BHDX6S0SVTA708HXRB

Published
11 Sept 2026
T1 · 1 h ago
arXiv
2609.11799
T1 · 1 h ago
Category
cs.CR
T1 · 1 h ago

Abstract

Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger appears. While backdoors can be audited before deployment, runtime monitoring remains important for models that are frequently updated. The challenge is that LLM serving is latency-sensitive: existing inference-time detectors either rely on assumptions about the trigger form, which can fail on stealthy attacks, or require extra model computation, such as input perturbations or an additional generation pass. We introduce SpecGuard, an inference-time backdoor detector that repurposes speculative decoding at zero added model-computation cost. Speculative decoding speeds up inference by using a small draft model to propose tokens and a target model to verify them. We observe that this verification process already exposes a useful signal: when a backdoor is triggered, the target model shifts toward the attacker's behavior, while a clean draft model does not predict this shift, causing the draft-token acceptance rate to change. We formalize when this signal appears and show that an attacker who suppresses it must also weaken the backdoor. Across diverse backdoor types and model families, SpecGuard reliably detects triggered behavior, including stealthy cases where input-level filters are blind, while avoiding the extra generation cost of existing runtime detectors. Speculative decoding therefore doubles as a free, always-on signal for detecting backdoored LLM behavior.

Authors 5

Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

arXiv id
2609.11799

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Categories
cs.CR, cs.CL

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Primary category
cs.CR

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

1 h ago

Conflicts

None