Skip to content
AI Atlas
PaperActive

Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

arxiv.org/abs/2609.11319

quality89

Updated 56 min ago · first seen 12 Sept 2026

paper_01M29X34KSZPPS617CS0715DS5

Published
12 Sept 2026
T1 · 56 min ago
arXiv
2609.11319
T1 · 56 min ago
Category
cs.AI
T1 · 56 min ago

Abstract

Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathematical reasoning. Restraining LLMs to informal reasoning misses out on the opportunity to use the discrete verification abilities that machines offer through machine-checkable proofs. In this paper, we bridge the gap between informal and formal reasoning by integrating Lean signals into the informal reasoning process. We introduce Magenta, a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof. A statement judge verifies whether the formalisation preserves the original problem, while an error-attribution judge routes failed attempts either to mathematical re-derivation or local Lean repair. Magenta achieves 100% accuracy across all evaluated olympiad benchmarks, including AIME 2025, AIME 2026, and HMMT February 2026. When paired with the open-weight K2-Horizon-7B reasoner, it solves all six IMO 2026 problems. Our analysis shows that statement adjudication is essential for preventing false certificates and that feedback-guided correction outperforms independent resampling on difficult problems.

Authors 9

Eleonora Giunchiglia, Erix Xing, Haonan Li, Joshua Ong Jun Leang, Shay Cohen, Wenda Li, Xinyi Shang, Zheng Zhao, Zhengzhong Liu

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

arXiv id
2609.11319

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

56 min ago

Conflicts

None