The next benchmark in AI mathematics is not solving more problems, but generating good conjectures and definitions — yet this cannot be made into a clean benchmark, and its arrival will instead show up as a tone shift in how working mathematicians describe their interactions with AI.
Grant endorses Dwarkesh's framing that conjecture-generation and definition-generation are the 'premium-tier' mathematics. But he argues such ability resists benchmark formulation — there is no clean goalpost like a theorem — and its presence will be felt as a subjective tone shift in mathematicians' accounts rather than a headline result. ✦ AI generated
Grant Sanderson · Dwarkesh Podcast · 2026-06-30 · original ↗
plays this moment only · 0:00 — 11:32
“There are a couple of candidate ideas here. One could be coming up with interesting problems in the first place, and the other is coming up with new kinds of objects or conceptualizations that create or unify fields.”
Many mathematicians get really new insights from proving theorems by hand. There's this quote: 'good mathematicians prove theorems, great mathematicians come up with conjectures, and the greatest mathematicians come up with definitions.' That's more or less exactly your framing here. We need the conjecture generator and then the definition generator. That's the premium-tier mathematician. I don't understand how exactly you'd make that a benchmark. Usually, when I think of the word benchmark, I'm thinking of something that is a goalpost. The ball is through the goal or it's not. You can clearly say, 'Yes, this is done.' Partly that's to be able to do things like RLVR, but also partly just to know that you haven't moved the goalpost in answering. OpenAI can have their headline on disproving the unit distance conjecture because it's a clear, distinct thing. It did it. Whereas imagine trying to have a headline on GPT-5.4 coming up with a really good conjecture. 'We promise, everyone thinks it's a good conjecture.' It just doesn't land the same way. But maybe that doesn't negate the fact that it's the right thing to be thinking about. I think the way you'd measure conjecture-generating ability is going to be more subjective, based on that tone shift. It will be mathematicians saying they're not just using it to solve their problems, but that as they step back and decide what their research field should even be, a conversation with such-and-such model was genuinely helpful for that.
verbatim transcript · starts at 0:00
0:00– AI is discovering new proofs. Is that AGI?
11:32– The verification loop on conceptual breakthroughs can be a century long
0:00– AI is discovering new proofs. Is that AGI?
11:32– The verification loop on conceptual breakthroughs can be a century long