Skip to content
AI Atlas
PaperActive

Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting

arxiv.org/abs/2609.10613

quality89

Updated 1 h ago · first seen 11 Sept 2026

paper_01M294FPR93HPSR7MWMEMEE6ZM

Published
11 Sept 2026
T1 · 1 h ago
arXiv
2609.10613
T1 · 1 h ago
Category
cs.CR
T1 · 1 h ago

Abstract

In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating over competing behavioral modes, and interprets in-context demonstrations as inference-time evidence that dynamically shifts the model's posterior preference between safe and harmful behaviors. This view formalizes jailbreak as a process of evidence accumulation, yielding predictive scaling laws with respect to demonstration count, harmful ratio, adversarial strength, and semantic diversity. Guided by this framework, we introduce a posterior-aware inference-time defense that adaptively injects benign counter-evidence based on estimated risk, effectively suppressing harmful posterior drift while preserving model utility. Compared to existing in-context defenses, our method achieves a significantly improved robustness-utility trade-off under a fixed intervention budget. Together, our results establish posterior reweighting as a unifying and predictive framework for understanding and mitigating ICL jailbreak in MLLMs.

Authors 4

Xu Zhang, Dev Mistry, Xiang Xu, Ren Wang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

arXiv id
2609.10613

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Categories
cs.CR, cs.LG

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Primary category
cs.CR

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

1 h ago

Conflicts

None