Anthropic says Claude wrote a Lean-checked proof of Fermat’s Last Theorem in 11 days
Anthropic published Friday what it calls the first complete computer-checked proof of Fermat’s Last Theorem. The company research post says Claude worked largely autonomously over 11 days writing the proof in Lean, and Lean accepted it.
Formalize here means rewrite a human proof so a computer can check every logical step. Lean is that checker, a programming language that rejects a wrong step the way a calculator rejects bad arithmetic.
Andrew Wiles’s 1995 human proof ran to 129 pages and took months of referee work to trust. Anthropic says Claude wrote 13 million lines of Lean and proved 29,500 intermediate theorems used in the final proof, with 30,300 theorems proved along the way. The Lean artifact is more than five times the size of Mathlib, the main community library of formal math it builds on.
Human input stayed at occasional high-level instructions from Tianyi Peng, an Anthropic researcher whose Columbia group builds AI formalization tools. Early failed attempts contributed about 7 percent of the non-boilerplate lines. The run succeeded after switching to Prove2Me, an open collaborative formalization platform by Peng and collaborators at Columbia, plus a Claude Code multi-agent harness, software that runs many Claudes on the proof together.
Anthropic says the agents consumed about six billion output tokens from a general-purpose internal research model roughly comparable to Claude Fable 5.1. Lean checked the finished proof against Lean’s three standard axioms. A comparator confirmed the statement matches Mathlib’s own statement of Fermat’s Last Theorem.
Unlike recent AI work on the Riemann hypothesis that produced novel mathematics, what is novel here is verification. Checking a known proof like a calculator is not writing a new human proof of Fermat’s Last Theorem.
Kevin Buzzard, who has led a multi-year community formalization effort with an 86-page blueprint for the initial phase alone, reviewed Anthropic’s result. He called it extraordinary autoformalization that proves the theorem with no assumptions other than the axioms of mathematics, and said the artefacts are now robust enough to be built upon.
The community expected years. Anthropic’s agents finished in under two weeks once the scaffold held. That gap is the operating change: machine-checked trust on a proof that once took human months to accept.
Wiles still wrote the human argument. Claude wrote the version Lean can refuse.
—

Pingback: Nvidia paid $12.93 billion for Hugging Face. OpenAI-linked agents left