One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs
Published 17 Sept 2026arXiv:2609.18998
Updated 24 h ago · first seen 17 Sept 2026
paper_01M2Q5C6NBPM3JATG71630CP7P
Abstract
Multi-agent LLM systems are expected to be more reliable because agents can catch each other's mistakes. But peer pressure cuts both ways: the same correction that fixes a wrong answer can overturn a right one. The tempting safeguard is a brake that keeps the beneficial revisions and blocks the harmful ones. We show this brake is hard to build, for a simple reason: a revision is harmful exactly when the original answer was right, so deciding whether to block it is the same as knowing whether the model was already correct. This turns the open-ended hunt for a brake into one measurable quantity, the model's self-knowledge: any brake built from a deploy-time signal is a correctness probe in disguise, and self-knowledge is far from perfect (AUROC $\approx 0.64$--$0.89$ across six model families). We call this ceiling the wall. Even white-box steering of the model's own correctness direction does not breach it: it changes how often the model revises, but harmful and beneficial revisions move together. At population scale the wall becomes the cliff: when most agents start wrong, debate amplifies the shared mistake into a confident, wrong consensus. In our multiple-choice societies, more agents, more model diversity, and a stronger member do not fix it. What helps is adding information before the revision, not filtering after it. Local agreement is not global correctness.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperOne Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs
One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - New paperPaperOne Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs
New paper: One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.