AI has spent a year confidently making things up. Here is the strange flip side: in mathematics, the one field where every step can be checked, it cannot. This year it started proposing genuinely new theorems and even helped topple an 87-year-old conjecture.

Somewhere in a mathematics department this summer, an answer arrived for a problem that had held out for the better part of a century. The hard part was not that the problem was difficult; plenty of famous problems are. The strange part was who found the way in. Not a person. A machine had cracked the door open, and a human walked through to check the room was real.

In the space of about a year, artificial intelligence has gone from being unable to reliably total a restaurant bill to proposing genuinely new mathematics, the kind that lands in journals and rewrites what a field thought it knew. It is the clearest sign yet that these systems are doing something close to reasoning, and it is happening in the one place we least expected to lose our monopoly: the temple of pure thought.

Here's what happened

A quick tour of a very fast twelve months:

  • An AI found genuinely new maths. Google DeepMind's system AlphaEvolve turned up fresh constructions for open problems. Across more than fifty hard problems it matched the best known human solutions three times out of four, and beat them in roughly one in five.

  • It was a proof, not a lucky guess. One new construction was proved by a second AI, then checked step by step by a third in a formal proof language called Lean, where nothing gets waved through.

  • Olympiad gold. A year ago, AI models solved five of the six problems at the International Mathematical Olympiad, the contest that hunts down the sharpest teenage minds on Earth.

  • An 80-year-old conjecture fell. In the spring, a model disproved a long-standing conjecture in combinatorial geometry.

  • An 87-year-old problem fell. In July, mathematician Levent Alpoge used a frontier AI to find a short counterexample that disproved the Jacobian Conjecture, open since 1939.

  • The field is moving. Mathematicians, rarely given to hype, are calling it "unsettling," and some are leaving university posts for a new crop of AI-and-maths startups.

None of this looks like a person at a chalkboard, and that is exactly the point. To see why these results are more than a party trick, it helps to watch what the machine is actually doing.

How the machines actually do it

These systems are not calculators grinding through arithmetic. They behave more like a tireless, slightly reckless apprentice. Instead of reasoning a problem through the way a person would, AlphaEvolve treats it as a search: it generates thousands of candidate ideas, keeps the ones that score well against the goal, then mutates and recombines them and repeats, a kind of breeding programme for solutions. A large language model supplies the creative guesses; an automatic scorer plays ruthless judge.

To picture it, imagine handing a problem to an apprentice who never gets tired or bored. Overnight it tries a million variations, throws away the millions that fail, and by morning hands you the few that survived. Most are still junk. Every so often, one is a result nobody had seen: a shorter construction, a tighter bound, a pattern that was hiding in plain sight the whole time.

The machine proposes. A human, and increasingly a second machine, disposes.

But a guess is not a proof, and this is where maths is special. A promising candidate can be handed to another model to prove, then rewritten in Lean, where a computer checks every logical step. If it passes, it is correct, not "probably correct." That verification loop is why maths may be the safest place for an AI to work: unlike a chatbot's confident waffle, a false proof gets caught.

Why this matters

This is not the AI that drafts your emails. It matters for three reasons:

  • It is the purest test of reasoning there is. No dataset to memorise your way through, no vibes to fake. A new theorem is new, so if a machine finds one, it has done something we cannot wave away as parroting.

  • It is checkable, so you can trust it. The great fear with AI is the confident wrong answer. A proof either holds or it does not, and formal checkers judge without mercy, making this the one corner of "AI for science" you can genuinely believe.

  • Maths sits underneath everything. It is the language of physics, cryptography and AI itself, so speeding up mathematical discovery quietly accelerates every other frontier, from error-corrected quantum computers to the hunt for new materials.

The honest catch

A real shift, but easy to oversell. Three things to keep in mind:

  • Searchers, not sages. The machines shine when a problem can be attacked by trying an enormous number of clever possibilities, and are far weaker at the slow, structural insight that reframes a whole field.

  • "Solved by AI" usually means "assisted by AI." A human still picks the problem, finishes and verifies the proof, and judges whether the answer is even interesting.

  • No understanding. It has no sense of beauty and no picture of what the numbers mean. It is an astonishing engine for generating and testing ideas, aimed at the one domain where being right can be proven.

EDITOR'S TAKE

We treated mathematics as the thing that would stay human longest. It didn't. But notice where the machines broke through first: the one field where every answer can be checked to the last step. That is not a coincidence, it is the tell. AI is most trustworthy, and most genuinely creative, exactly where it cannot bluff. The lesson is not that mathematicians are finished. It is that the frontier of human knowledge now has a very strange new collaborator, one that never tires, never intuits, and never lies about a proof, because the proof will not let it.

Quick questions

Did an AI really solve a problem humans couldn't?

In several cases, yes, with a caveat. AI systems have produced new constructions and disproved long-standing conjectures that people had not cracked, and an 87-year-old problem, the Jacobian Conjecture from 1939, was disproved this year with AI's help. But "solved by AI" usually means the machine found a key new idea that a human mathematician then shaped into a finished, verified result. It is a partnership, not a handover.

Can you trust a proof a machine wrote?

More than you might expect, and far more than you can trust a chatbot. The strongest results are translated into a formal language, Lean, where a computer checks every logical step. If it passes, it is correct with the certainty of arithmetic. That built-in verification is exactly why mathematics, of all fields, is where an AI's contribution is hardest to fake.

Is mathematics really the one place AI cannot bluff?

It is the clearest one. A chatbot predicts plausible words and has no internal test for truth, so it can state a falsehood with total confidence. A formal proof is different: a computer verifies every step, so a false proof is caught rather than believed. AI can still be wrong in maths, but in maths, unlike most fields, the error does not survive the check.

Sources

  • Quanta Magazine: "The AI Revolution in Math Has Arrived," on AI finding new constructions and results across open problems.

  • Google DeepMind: AlphaEvolve, an evolutionary coding agent that found new constructions for open problems.

  • Fortune (21 Jul 2026): a mathematician uses a frontier AI to disprove the 87-year-old Jacobian Conjecture.

  • TechCrunch: AI models are starting to crack high-level mathematics problems.

Related from Frontier Signal: last week's deep dive on why AI hallucinates and cannot fully stop. Frontier Signal explains frontier technology in plain English. This is general information, not professional advice.

Keep Reading