AI is starting to do real mathematics, and the hype is running ahead of it
In 2026 an AI model disproved a conjecture that had stood since 1946, and a Fields Medallist said the write-up was fit for a top journal. It is a genuine milestone. It is also narrower than the headlines. Here is the honest line between the two.

Something real happened in mathematics this year. An internal OpenAI model disproved a conjecture the great Paul Erdős posed in 1946, finding an arrangement of points no mathematician had, and the Fields Medallist Timothy Gowers said he would have waved the write-up into a top journal. Just this month the same lab went further, publishing ten more proofs with machine-checkable certificates. It is a genuine milestone. It is also narrower than the headlines: the model recombined known ideas, humans verified and rewrote the proof, the flashiest claims still arrive wrapped in marketing, and most "AI solved maths" stories run well ahead of the substance. Here is the honest version.
Every few weeks now, a headline announces that an artificial intelligence has "solved" an unsolved mathematics problem. Last autumn a senior OpenAI executive claimed GPT-5 had cracked ten open problems posed by the legendary Paul Erdős, then quietly deleted the post once a mathematician pointed out the model had merely dug up solutions that already existed in the literature. This month the same company published ten genuinely new proofs, each with a machine-checkable certificate, and still drew accusations of overselling them. Some of these stories point at something real and important. Others blur a verified result into a looser claim until you cannot tell which is which. Since telling them apart is the whole point of covering this honestly, here is what has actually been proven, what has not, and why serious mathematicians are both excited and uneasy.
Did an AI actually solve an unsolved maths problem?
Yes, one genuinely important one. In May 2026 an internal OpenAI model disproved a conjecture about the Erdős unit-distance problem, a question the Hungarian mathematician Paul Erdős first posed in 1946 and that had resisted serious effort ever since. The problem asks, in effect, how many pairs of points you can place so that exactly one unit of distance separates them, and for decades the assumption was that neat, grid-like arrangements were the best you could do.
The model found a counterexample: an arrangement that packs in more unit-distance pairs than a square grid, for infinitely many sizes, showing the long-standing belief was wrong. Be precise about what fell, though: it was Erdős's specific conjecture about the best possible arrangement, not the whole problem. The true maximum is still unknown, so the model proved the old guess too low without settling where the real answer lies. According to the mathematicians who examined it, the result came with minimal human intervention beyond the initial prompt. OpenAI published that prompt along with a rewritten, human-readable summary of the model's long chain of reasoning, though not the raw, unedited transcript itself.
The reaction that carried the most weight came from Timothy Gowers, a Fields Medallist and one of the most respected mathematicians alive. He wrote that if a human researcher had submitted the paper to the prestigious journal Annals of Mathematics, he would have recommended acceptance "without any hesitation." When a mathematician of that stature says a machine's output is journal-ready, it is no longer a demo.
What did the model actually do?
The clever part is the construction. The model built a high-dimensional grid and projected it down into two dimensions using algebraic integers, which let it pack unit distances more densely than the classic square-grid approach. That is a real mathematical idea, and it is the thing no human had assembled in nearly eighty years of trying.
What happened next matters just as much. Human mathematicians verified the result and then wrote a cleaner, extended proof, a "short, digested, human-verified version" of the machine's counterexample, credited to a lineup of nine serious mathematicians, among them Noga Alon, Thomas Bloom, Gowers himself, Daniel Litt, Jacob Tsimerman and Melanie Matchett Wood. The AI supplied the breakthrough construction; humans confirmed it was correct and turned it into something the field could read and build on. That collaboration, not a machine working entirely alone, is the actual shape of the milestone.
Is this real research or just hype?
Both are true at once, which is exactly why it gets misreported. The careful analyses of the result make an important point: it played to AI's strengths. Two things suited a machine here. First, the solution required stitching together techniques from distant corners of mathematics, and an AI trained on essentially all of published maths carries a broader (if shallower) map than any single specialist. Second, the winning strategy was a grind, one that "consumes much time and frequently doesn't work out," the kind of long, unglamorous search most humans abandon, but that a tireless model can push through, exploring many dead ends cheaply.
Crucially, the model did not invent a genuinely new mathematical technique. It recombined existing ideas in a way no one had. That is a real contribution, but it is a different thing from the science-fiction picture of an AI opening whole new branches of mathematics on its own. Gowers, tellingly, admitted he was relieved the model had produced a disproof rather than a proof of the conjecture, because a full proof would have hinted at something closer to replacing mathematicians. His considered view was measured: AI, he said, will "soon reach a high level at other activities such as building theories, formulating definitions and asking interesting questions." Soon, not yet.
Then OpenAI did it again, at ten times the scale
Just before this was published, the same company raised the stakes. On 1 August 2026 OpenAI revealed Astra, an unreleased next-generation model, by dropping solutions to ten problems that had each been open for years, ranging from high-dimensional sphere packing and coding theory to a counterexample to Connes's rigidity conjecture on von Neumann algebras and one of Erdős's Ramsey problems. What makes this batch hard to wave away is the receipts. For every problem, OpenAI released a proof certificate written in Lean, a proof-checking system a computer can verify line by line, published openly on GitHub alongside a 249-page manuscript. A Lean certificate is not a press release; it is a formal object that either checks out or does not. Even sharp critics conceded the underlying mathematics is real.
And yet Astra is also the cleanest illustration of the hype problem in this whole story. OpenAI's headline number, that generating all ten solutions would cost only about $2,000, is token cost alone, priced at the cheaper rates of an already-released model, and it pointedly leaves out the salaries of the humans who framed the problems and checked the work. The 249 pages say almost nothing about how the model actually operated, how many problems it was handed before it found ten it could crack, or where people stepped in. "Astra is amazing. No denying that," wrote the AI critic Gary Marcus, before adding that "not one page is about how the model works ... whether any of the proposed proofs had errors," with no control group and no accounting of the failures. A real result and an oversold telling, in the same week's news. That gap is the whole subject of this article.
What is FrontierMath, and how much has AI really solved?
If you want a sober scoreboard rather than headlines, the best one is FrontierMath, a benchmark run by the research group Epoch AI. Its Open Problems track is a collection of 50 genuine, research-level questions that, in its own words, "have resisted serious attempts by professional mathematicians," where an AI solution "would meaningfully advance the state of human mathematical knowledge."
So far, AI systems have solved three of the 50. That is the value of a fixed, pre-registered set: unlike a lab announcing its own wins, where you never see how many attempts quietly failed, three-of-fifty is a number you cannot cherry-pick. Three real open problems is a genuinely striking number for a technology that, two years ago, routinely failed at arithmetic. It is also, plainly, three out of fifty. Both halves of that sentence are the story: real, and early.
What do mathematicians actually think?
Mixed feelings, and honestly so. The excitement is real: a tool that can find counterexamples, spot patterns across fields and help turn a hand-wavy argument into a machine-checkable proof is a genuine gift to research. In April 2026, for instance, a team at Carnegie Mellon cracked an open problem in Ramsey theory by combining automated solvers, code written by a language model and formal proof verification, exactly the kind of human-plus-machine workflow that is becoming normal.
The unease is just as real. Some worry about a flood of plausible-looking "results" that are subtly wrong and expensive to check. Others simply dislike how quickly a careful, human-verified achievement gets flattened into "AI solves maths" on the way to a headline. The healthiest position in the field right now is neither the boosters' nor the sceptics': it is that these systems have become serious research collaborators, and that a collaborator is not the same as a replacement.
Notable AI-in-mathematics results of 2026
| Result | Who | What the AI did | Where it stands |
|---|---|---|---|
| Erdős unit-distance conjecture (1946) disproved | OpenAI internal model | Found a denser point construction as a counterexample | Human-verified; write-up called journal-ready by a Fields Medallist |
| FrontierMath Open Problems | Various frontier models | Solved research-level open questions | 3 of 50 solved so far (Epoch AI benchmark) |
| Ramsey-theory problem | Carnegie Mellon team + models | Model-written code inside a solver-and-proof pipeline | Solved via human-plus-machine workflow (April 2026) |
| OpenAI "Astra" — ten proofs (Aug 2026) | OpenAI internal model | Solved 10 long-open problems, published Lean certificates + a 249-page paper | Real and machine-checkable, but criticised as oversold (no method disclosed; $2,000 = token cost only) |
| "GPT-5 solved 10 Erdős problems" (Oct 2025) | Claimed by an OpenAI executive | Claimed to solve 10 open Erdős problems | Debunked: the model only surfaced solutions already in the literature; the post was deleted |
So, can AI replace mathematicians?
Not now, and not in the way the loudest headlines imply. What genuinely changed in 2026 is the kind of thing these models do: they crossed from "help me with a step" to "here is a new result a human can verify." That is a real threshold, and it is why even careful sceptics are paying attention.
But the pattern underneath every honest example is similar. The machine supplies a construction, a candidate, a grind no human wanted to do; then humans, or a formal proof-checker they trust, confirm it holds up, and humans still decide whether it was the right question and what the answer means. For now, AI is the most powerful research assistant mathematics has ever had, not its successor. When you next see a headline announcing that an AI has "solved maths," the useful question is not whether it is impressive. It is which of those two things actually happened.
For more, see the AI section, and our guides to the best AI chatbot and the best AI coding tools, where the same models are already changing everyday work.


