ATRIUMsearch → argument graph
Audio · 2026-06-30 · 6 moments

Grant Sanderson – AI and the future of math

Watch now (94 mins) | Math is where we’ll see superintelligence first. What will it look like? ✦ AI generated

timeline · colored by role

01
Claim

Getting gold at the IMO or even solving a Millennium Prize problem does not mean an AI can do white-collar work, because mathematics sits at a spiky frontier of AI capability and the rate-limiter for deep mathematical insight (drawing lightning-bolt connections between expert domains or building entirely new mountains of theory) is different from what white-collar work requires.

Grant argues that mathematical progress sits on a spiky, fractal frontier of AI capability. Solving IMO problems or even a Millennium Prize problem could happen through domain-bridging 'lightning bolts' or 'mountain building' — capacities distinct from what automation of white-collar work actually needs, so such achievements would not self-evidently imply general economic capability.

transcript

Grant Sanderson: It's an interesting question because it's hard to answer without knowing what the solution looks like ahead of time. If we take the IMO, the spirit of your question three years ago was in looking at how some of the solutions to these problems really seem to require creativity. The designers of these problems try to come up with things that you can't train for as easily. The dirty secret with the IMO is that you really can train for a lot of them. With the whole AI and math project underway, as you point out, one of the reasons it's interesting at all is that there's a spiky frontier to AI, and math is just right there in one of the spikes. But there's a fractal nature to that spikiness, because when you zoom into the specific progress within math, you have some things that are a lot easier than others. If we focus on the Riemann hypothesis, what would it look like to solve that? These things are extremely good at a specific domain of knowledge, knowing it very deeply, and then knowing another domain, and another. It's bizarre to have something with this superhuman breadth that knows all the fields so well, and yet isn't finding those lightning bolts that connect them. I think we're starting to see sparks of it actually finding connections between the things it's an expert at. If the nature of the solution to the Riemann hypothesis was something like that, that feels pretty distinct to me from what's necessary to get good at white-collar work. If the solution to the Riemann hypothesis involved building a new mountain, that's a kind of skill—the ability to come up with the right new ideas—that feels sufficiently different from the character of how they're intelligent right now. But if it's capable of building mountains that are the correct new theory crystallizing how we should be thinking about a subject, that's just such a level of intelligence that it would be surprising if it didn't permeate into other aspects of the economy besides just the mountain-building for math itself.

02
Claim

The next benchmark in AI mathematics is not solving more problems, but generating good conjectures and definitions — yet this cannot be made into a clean benchmark, and its arrival will instead show up as a tone shift in how working mathematicians describe their interactions with AI.

Grant endorses Dwarkesh's framing that conjecture-generation and definition-generation are the 'premium-tier' mathematics. But he argues such ability resists benchmark formulation — there is no clean goalpost like a theorem — and its presence will be felt as a subjective tone shift in mathematicians' accounts rather than a headline result.

transcript

Grant Sanderson: Many mathematicians get really new insights from proving theorems by hand. There's this quote: 'good mathematicians prove theorems, great mathematicians come up with conjectures, and the greatest mathematicians come up with definitions.' That's more or less exactly your framing here. We need the conjecture generator and then the definition generator. That's the premium-tier mathematician. I don't understand how exactly you'd make that a benchmark. Usually, when I think of the word benchmark, I'm thinking of something that is a goalpost. The ball is through the goal or it's not. You can clearly say, 'Yes, this is done.' Partly that's to be able to do things like RLVR, but also partly just to know that you haven't moved the goalpost in answering. OpenAI can have their headline on disproving the unit distance conjecture because it's a clear, distinct thing. It did it. Whereas imagine trying to have a headline on GPT-5.4 coming up with a really good conjecture. 'We promise, everyone thinks it's a good conjecture.' It just doesn't land the same way. But maybe that doesn't negate the fact that it's the right thing to be thinking about. I think the way you'd measure conjecture-generating ability is going to be more subjective, based on that tone shift. It will be mathematicians saying they're not just using it to solve their problems, but that as they step back and decide what their research field should even be, a conversation with such-and-such model was genuinely helpful for that.

explains mechanism · 1

03
Mechanism

The value of a conceptual breakthrough like Galois theory is verified on a timescale of a century, through a chain of human judgment rather than any immediate reward — meaning current RLVR-style reward functions are structurally unable to recognize the 'Galois instinct'.

Grant walks through the history from Lagrange's questions about symmetry, to Abel's proof, to Galois's rejected papers, to Liouville and Jordan, and finally Gell-Mann's group-theoretic prediction of quarks. He argues that recognizing a great new conceptualization is a roughly century-long verification loop through many human judgments — the opposite of an RLVR feedback signal.

transcript

Grant Sanderson: What makes Galois theory such an interesting example is that you literally have this hundred-year segment of an idea that flows through many different people's heads before it settles into something the math community agrees is good. With Einstein and GR, people could feel this was a good theory right away. You have to ask, what is the way of measuring progress that's not based on solving a problem, but that is somehow capturing the instinct inside Galois's mind when he says, 'I think there's something here'? What's the instinct inside Lagrange's mind when he says, 'I think this is the right way to think about it'? What's the instinct inside Liouville's mind when he says, 'These scattered notes from this long-dead youngster might have something to them'? It's so hard to put a finger on that.

04
Prediction

How well humans can digest an AI proof of the Riemann hypothesis depends entirely on which of three forms the proof takes — a parsable lightning bolt between fields, an alien 'new mountain' of theory, or a raw thousand-page chain of reasoning — and the biggest risk is an alien mountain that turns out to be wrong like the abc conjecture.

Grant breaks out the three candidate shapes an AI Riemann-hypothesis proof could take. The field-bridging 'lightning bolt' form is very human-parsable; building a new mountain of theory could be an alien, hard-to-digest mathematics — and the abc-conjecture episode shows the catastrophic case where an AI-style alien theory looks right but isn't.

transcript

Grant Sanderson: If we break down the three possible ways of solving the Riemann hypothesis… The other big one from this year was a certain Erdős problem numbered 1196, about these things called primitive sets. It had that character of bringing an idea from a seemingly different field. You have this very small idea that has the form of expertise in one field and expertise in another, drawing a little lightning bolt between them. Those are going to be very human-parsable, because all you have to do is show the start and end point of what those connections are. If the character of it is mountain building, you have to put in a lot more time to understand that new mountain that was built, because it's a new thread, not just a lightning bolt between them. And if the nature of the progress was just raw hustle—a super long chain of reasoning with no new theories—then you would have that worry of this whole digestion process. The biggest fear would be that an AI does that, and then much like the abc conjecture, people work for years to go up the mountain, and they're like, 'Dang it. This just isn't right.' If it turns out to be wrong, but it really looked right. Even if it was right, there's just a lot of effort to hike up a new mountain.

05
Mechanism

Mathematics and coding progress faster than other fields not primarily because they are verifiable, but because they are 'grindable' — you can containerize environments, spin up deterministic parallel rollouts, and cleanly solve the credit assignment problem, which is impossible in the mutable real world.

Dwarkesh argues that verifiability alone (as in computer use) is insufficient — the decisive factor is grindability: deterministic, containerizable environments like code and math that allow many parallel rollouts and clear credit assignment. Grant agrees and extends this to the vision of an AI endlessly growing Lean/Mathlib's 'tree of logic' with no human check-ins.

transcript

Dwarkesh Patel: Dwarkesh: What computer use lacks is grindability. Because websites have bot detectors—and it takes a tremendous amount of compute to run parallel rollouts—it's very hard to run a thousand parallel rollouts of the same checkout flow on Amazon. The reason you currently need to do so many parallel rollouts to learn a skill with deep learning is that we haven't solved sample efficiency. With code, you can containerize a given level of progress in a repository and then spin out hundreds of parallel containers and say, 'Try to implement this feature,' and it's totally deterministic. Because it's deterministic, you can solve the credit assignment problem because you know that whatever caused this rollout to succeed and this one to fail, the diff is the thing that worked. Math, of course, is the exception, and I feel like this is an important driver of progress in this domain and also in coding. Grant: That's a very unique thing that math has that nothing else has, where you could press go and just pour compute at it, look away for ten years, and then come back and say, 'What do you have?' There's going to be something. Then there's a question: is it useful or not? How do you suss that out? That's just an interesting thing to be able to do. It would be very surprising if that didn't yield some sort of interesting mathematical insight from it.

06
Prediction

Even in a world where AIs solve and explain mathematics better than humans, mathematical work and teaching will not disappear: they will shift toward relational curation and mentoring — the museum-curator and teacher role — because motivation and trust are fundamentally social phenomena.

Across the close of the conversation, Grant argues that the lasting human role in mathematics is curation — deciding what ideas are worth pursuing — and mentoring, which remains a social, coaching function that explanations alone cannot replace. He advises students to understand where their value and funding actually come from, and to lean into the teaching role, which he calls one of the most stable post-AGI careers.

transcript

Grant Sanderson: One interesting take that I've heard about what mathematicians will end up being is that it's actually more analogous to art museum curators than anything else. The AI solved the thing, so the art exists. They even know how to explain it really well. But you still want someone to help you navigate this nearly infinite space of what ideas are worth engaging with. Even if AIs were in some sense better at that, I think we would always still prefer a human that we had a relationship with, because the way we get motivated to be interested in things is a social phenomenon. So my role, and arguably that of other mathematicians, might actually just shift subtly into that curation direction of what ideas are worth pursuing. I actually think teaching is one of the most stable post-AGI jobs that there is, because it's so relational. This is where parents want to spend their money if they have an abundance of wealth: on good teaching and good educating. It goes so far beyond explanations. Even if LLMs are good explainers, the thing that a teacher is doing is such a social, coaching, mentor-type thing that that's probably one of the most stable careers that's going to exist over the next fifty years.

Highlight slides
Gold at the IMO ≠ White-Collar Capability✦ from: Getting gold at the IMO or even solving a Millennium Prize problem does not mean an AI can do white-collar work, because mathematics sits at a spiky frontier of AI capability and the rate-limiter for deep mathematical insight (drawing lightning-bolt connections between expert domains or building entirely new mountains of theory) is different from what white-collar work requires.Bridging vs. Building: Three Distinct Skills✦ from: Getting gold at the IMO or even solving a Millennium Prize problem does not mean an AI can do white-collar work, because mathematics sits at a spiky frontier of AI capability and the rate-limiter for deep mathematical insight (drawing lightning-bolt connections between expert domains or building entirely new mountains of theory) is different from what white-collar work requires.The Century-Long Verification of Galois Theory✦ from: The value of a conceptual breakthrough like Galois theory is verified on a timescale of a century, through a chain of human judgment rather than any immediate reward — meaning current RLVR-style reward functions are structurally unable to recognize the 'Galois instinct'.Contrast: Instant vs. Generational Recognition✦ from: The value of a conceptual breakthrough like Galois theory is verified on a timescale of a century, through a chain of human judgment rather than any immediate reward — meaning current RLVR-style reward functions are structurally unable to recognize the 'Galois instinct'.Why RLVR Cannot Capture the 'Galois Instinct'✦ from: The value of a conceptual breakthrough like Galois theory is verified on a timescale of a century, through a chain of human judgment rather than any immediate reward — meaning current RLVR-style reward functions are structurally unable to recognize the 'Galois instinct'.
Related episodes