Skip to content
AI Atlas
PaperActive

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

arxiv.org/abs/2609.09553

Updated 50 min ago · first seen 11 Sept 2026

paper_01M294GMA0N5RSX4XXFDXFZ5KX

Published
11 Sept 2026
T1 · 50 min ago
arXiv
2609.09553
T1 · 50 min ago
Category
cs.CR
T1 · 50 min ago

Abstract

Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak and have previously been demonstrated against the fine-tuning APIs of commercial models. In these attacks, target models are trained on a corpus of encrypted harmful questions and responses and subsequently respond to harmful requests through the learned encryption scheme. In this paper, we show that newer frontier models do not require fine-tuning to acquire cipher-based communication skills. Instead, they can learn these skills through prompting and, when necessary, through in-context learning. Furthermore, model alignment is significantly weakened or entirely bypassed when communication occurs through the learned cipher. To the best of our knowledge, this constitutes a novel attack vector against commercial black-box large language models. We demonstrate successful jailbreaks against frontier models developed by Anthropic, Google, and OpenAI. Our attack bypasses commercial harmfulness classifiers because harmful content is encrypted and therefore appears as nonsensical text or gibberish.

Authors 1

Thomas Rivasseau

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

arXiv id
2609.09553

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Categories
cs.CR, cs.AI

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Primary category
cs.CR

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

50 min ago

Conflicts

None