AI can prove theorems, but it still can't ask questions worth asking
Daniel Litt, a mathematician at the University of Toronto, says models are best at applying known techniques to their limit — but the real value of mathematics is understanding, and understanding still lives in human heads.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
The unit distance problem is what changed his mind most
Daniel says his favorite fully autonomous AI result so far is the solution to the Erdős unit distance problem, announced around mid-May. He values it because the result is somewhat creative: it borrows techniques from classic 1960s ideas, but those techniques are new to people who study planar point configurations. And it proved productive afterward — a group of mathematicians took these ideas and constructed counterexamples to other open problems, such as the sum-product conjecture over the reals. He calls it the only autonomous result he knows of that, in hindsight, genuinely produced new understanding.
— Daniel LittThe model's chain of thought looks very human
The host asks whether this is "non-human" logical reasoning, and Daniel pushes back: he thinks the argument is very human. Some of the chains of thought OpenAI has released look completely recognizable to him — if he imagined his own chain of thought while solving a problem, it would look roughly like that. He says that's true of most results he has studied: there's no brilliant "move 37," it's just a human mathematician doing the mathematics humans do. The only way the model seems non-human is that it never gets tired and it knows a lot.
— Daniel LittThe claim that math is verifiable is overrated
Daniel says many people argue that math is a verifiable domain, so AI progresses fast there. But his judgment is that models are mainly extending informal reasoning, not Lean proofs, so these techniques will very likely generalize to other fields. He also points out that a lot of mathematical training data doesn't reflect how mathematicians think — papers are polished, textbooks almost never show motivation, so to keep up with a research direction you usually have to talk to the researchers directly.
— Daniel LittModels are strong at applying known techniques, weak at intuition
Daniel says model results have a fixed flavor: they lean on long computation, or they stitch together technical ideas from many fields and many papers. But they're weaker at intuition and the big picture. He describes how he works: first a very imprecise philosophy that this thing is analogous to that thing, then an effort to make it precise. And in the best model results so far, he sees little of that kind of reasoning. He stresses that being very good at applying all known techniques is itself powerful, and that some mathematicians do high-quality work this way — but that's only a narrow band of what mathematicians care about.
— Daniel LittWhen humans can't prove it, a better theorem comes out
Daniel tells the story of his only paper where a model was useful. He wanted to prove a lemma, and every frontier model failed to prove it. So he worked through a pile of examples himself, realized there might be a better statement, and once he found it the model proved that better version quickly. He says that if you now hand the original lemma to ChatGPT 5.6 Pro, it outputs the worst proof you've ever seen — ten pages of brutal computation, zero insight. The proof itself isn't wrong, but it won't lead to the prettier conceptual explanation. He says he realized the ugly proof would work but simply couldn't bring himself to do it, so he went looking for another argument.
— Daniel LittCurrent incentives are mass-producing junk papers
Daniel says the goal of mathematics isn't to produce papers, it's to produce understanding. But the current incentive structure doesn't encourage people to really commit. He describes an experiment: have Codex go online, find five recent conjectures in algebraic geometry and prove them, and after a few rounds, within an hour, you get three terrible but correct papers, now sitting on his hard drive. He says arXiv submissions have exploded, much of it obviously low quality, with the same proof, the same theorem, appearing three to five times within days. He thinks this isn't necessarily a problem that goes away as models get better.
— Daniel LittA thousand mathematicians, or one mathematician copied a thousand times
Daniel says mathematical progress comes from letting a thousand flowers bloom, people following their own curiosity, the frontier of knowledge expanding in a relatively even way, and applications then appearing opportunistically. What worries him: if mathematical exploration is subordinated to what a model wants to pursue, what you get may be one mathematician copied a thousand times, rather than a million mathematicians doing a million different things. The host adds that if the proofs labs put out all converge on similar things, because they technically depend on the same body of literature, then it's unclear where that diversity of intuition would come from.
— Daniel LittModel proofs are short because they can't check long ones
Some people praise model proofs for being short and beautiful; Daniel's judgment is the opposite: they don't produce long, complex proofs because they can't — their ability to check correctness isn't good enough. He says that even when you have a model write a short proof, if you then ask it whether the proof is right, it often says no. The problem with long proofs is that the model may not know it's wrong. He gives an example: the ten problems OpenAI recently released are all formalized in Lean, which is strong evidence they're true; he has no doubt there are more that can't be formalized, because the prerequisite lemmas aren't in Mathlib yet. He also mentions someone posting an 800-page AI-generated proof of resolution of singularities, and says there's no way it's correct.
— Daniel LittIn their own words · checked verbatim
I actually think the argument was very human.
Daniel Litt6:11
It's like a human mathematician doing math. It's like a human mathematician doing certain types of math.
Daniel Litt7:12
the goal of mathematics is not to produce mathematics papers. Um, like it's to produce some kind of understanding.
Daniel Litt34:44
a lot of progress in mathematics comes from like letting you know, a thousand different flowers bloom and people pursue their own curiosity
Daniel Litt37:45
if what you're getting is like one mathematician duplicated a thousand times or like actually you know a million different mathematicians doing a million different things.
Daniel Litt38:46
they will still sometimes produce things that are just wrong and they know they're wrong.
Daniel Litt53:13
there's no way it's correct. Like, this will be a major result.
Daniel Litt54:13
Figures
| Announcement date of the autonomous AI solution to the Erdős unit distance problem | mid-May | 3:09 |
| Number of papers Codex produced in one hour | 3 terrible but correct papers | 35:45 |
| Number of problems OpenAI released formalized | 10, all formalized in Lean | 53:13 |
| Page count of the AI-generated resolution of singularities proof | 800 pages | 54:13 |
| Range of numbers Daniel's daughter can compute reliably | reliable up to 30, half-reliable up to 50 | 59:31 |
Glossary
- Lean
- A programming language and verifier for formalizing mathematical proofs, in which proofs can be checked line by line by machine.
- Mathlib
- The largest library of formalized mathematical theorems in the Lean ecosystem; prerequisite lemmas must be in it before a new proof can be formalized.
- Erdős unit distance problem
- A classic problem about how often unit distances occur among planar point configurations, long believed to have an upper bound, later overturned by counterexamples.
- sum-product conjecture
- A conjecture over the reals that addition and multiplication cannot both keep a set small, later disproved by construction.
- resolution of singularities
- A classic theorem in algebraic geometry turning singular varieties into smooth ones, long unresolved in full in positive characteristic.
How to listen
Engineers and researchers who need a read on AI's mathematical ability, and technical decision-makers who want to know where models are actually strong and where they're weak.
The last ten minutes cover his daughter learning math and an icosahedron toy — unrelated to the main argument.