AI Cracked a Famous Math Problem—But Humans Did Most of the Work Before the Final Step
The proof touted as AI's breakthrough had its core components built by humans over months; the model merely completed the final step first.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
AI must outsource mathematical verification to non-AI systems to be trustworthy
Models can now generate proofs and interact with proof checkers, but they're poor at verifying correctness. A key shift last year: coupling models with Lean, a programming language that translates conjectures into code for external software to verify step-by-step against logical axioms. What makes AI math results more trustworthy isn't model intelligence—it's the division between AI generation and independent software verification.
— Justin SolomonMathematicians who prove theorems face more disruption than theory-builders
Mathematics splits into theory-builders (proposing what to research) and theorem-provers (solving established problems). Theorem-provers face greater impact: many spent decades on a single, well-defined conjecture—the exact kind of problem AI excels at, working the final stretch of something nearly solved. Whether a problem deserves research or contains real insight remains deeply dependent on human judgment.
— Justin SolomonMathematical proofs shift from scarce to abundant, upending how we value mathematicians
When proofs were rare and laborious, metrics for a mathematician's impact—where published, output volume, novelty, prizes—were tightly correlated. In Terence Tao's framing, we've entered an age of proof surplus where these metrics are decoupling: AI companies discovered that proving a major theorem itself is good press, and prompt-writing has become a new, economically valued skill. This blurs who truly contributed insight.
— Justin SolomonAcademic validation has become algorithmic spam from AI-using researchers
Justin recently received floods of email from outsiders who used AI tools to prove a ‘new result,’ were told by AI it matters, and sought MIT faculty verification. His own tests show AI hasn't solved any real open problems in his field, but after proving a trivial conclusion, helpfully offers to format it for journal submission. Peer review faces the same crisis: ICLR submissions grew from 1–2k to 60k annually—unsustainable review, with high schoolers infiltrating the reviewing pool.
— Justin SolomonThe Navier-Stokes counterexample credits the wrong hero in its headlines
The widely publicized counterexample that appeared to refute Navier-Stokes had its core construction hand-built by a Spanish mathematics team over months. AI completed the key final step, but two competing AI groups squabbled over credit. This reveals an unresolved attribution problem: when humans build most of the foundation and the model fills the final gap, how should credit be divided?
— Justin SolomonWhoever has the most compute power wins in mathematics research today
When word spread that someone was close to a Navier-Stokes proof, OpenAI reportedly poured compute at the problem to finish first and publish. OpenAI later noted it didn't need Clay's million-dollar prize—their two-minute compute bill exceeded the award. Ordinary researchers lack such budgets. Elite labs also control internal models a generation ahead of public versions. The same dynamic exists in education: students with pricier AI subscriptions likely produce better work.
— Justin SolomonAI proofs stay inside existing knowledge; they haven't found the new frontier yet
By Nestor Guillen's analysis, most AI-generated proofs cleverly recombine known facts and fill computational gaps—they stay within the convex hull of known mathematics, never venturing into fully novel theory. This marks the crucial unsettled question: can AI eventually be recursive and self-improving, proposing its own next research questions? Or can it only answer what humans already asked?
— Justin SolomonMIT reframed coursework as strength-building, not task completion, to verify learning
For the past couple of years, nearly all grades hovered near perfect—AI silently does the work. MIT didn't ban AI; instead, it redefined assignments as ‘exercises’ where the task itself isn't the goal; strength is. It added a follow-up quiz mimicking the homework to verify genuine understanding, plus oral defense requirements to ensure assessment targets the student, not Claude's capability.
— Justin SolomonIn their own words · checked verbatim
And these AI tools are sort of at least good at sort of cherry picking the ones that are somehow close to being proved and taking that last mile in some cases.
Justin Solomon23:31
The way that Terence Tao put it is that it used to be we were in this era of proof scarcity, that mathematicians, it was really hard. And it was this very bespoke object that took a lot of craftwork and training to produce. Now we're in this era of proof abundance.
Justin Solomon25:32
what I've noticed lately is I get a lot of emails from outside folks that sort of, they've been playing with their AI system, they've proven a new result, and the AI told them it was really important and they need an MIT professor to validate, which is absolutely fascinating.
Justin Solomon26:35
the construction of that spoon arguably was done in large part by a team in Spain of human mathematicians on pen and paper.
Justin Solomon34:45
OpenAI was quick to say that they don't need the million dollars, because of course they don't. In two minutes, they're burning through a million dollars worth of compute anyway.
Justin Solomon39:48
they're arguably in what mathematicians might call the convex hole of existing knowledge, meaning there's sort of a bunch of stuff we already know.
Justin Solomon44:55
What we observed in the last year or two is that our students are getting nearly 100% on all their homeworks.
Justin Solomon50:02
Figures
| Millennium Prize | $1,000,000 | 38:48 |
| Premium AI subscription cost | ~$200/person/month | 41:51 |
Glossary
- Lean
- A programming language that encodes mathematical proofs as code for automated, step-by-step verification against logical axioms.
- Navier-Stokes equations
- A system of equations describing fluid motion; whether solutions always exist and remain bounded is one of mathematics' famous unsolved conjectures.
- Clay Math Institute
- Institution that established the Millennium Prize Problems, offering million-dollar prizes for proofs of designated major mathematical conjectures.
- convex hull
- Here: AI proofs remain within existing knowledge boundaries, recombining known facts rather than generating truly novel mathematics.
How to listen
Anyone skeptical whether AI math breakthrough headlines are oversold; university faculty, research administrators, and investors assessing AI company research narratives.
The first 15 minutes discuss Good Will Hunting and Navier-Stokes background; core content begins at 17 minutes.