ATRIUMsearch → argument graph
Audio · 2025-06-15 · 3h 24m · 6 moments

#472 – Terence Tao: Hardest Problems in Mathematics, Physics & the Future of AI

Terence Tao is widely considered to be one of the greatest mathematicians in history. He won the Fields Medal and the Breakthrough Prize in Mathematics, and has contributed to a wide range of fields from fluid dynamics with Navier-Stokes equations to mathematical physics & quantum mechanics, prime numbers & analytics number theory, harmonic analysis, compressed sensing, random matrix theory, combinatorics, and progress on many of the hardest problems in the history of mathematics. Thank you for ✦ AI generated

timeline · colored by role

01
Claim

The hardest problems in mathematics sit on the boundary between what existing techniques can easily do and what is hopeless, where the existing techniques can do 90% of the job and you need the remaining 10%.

Terence Tao explains that the most interesting mathematical problems are not the impossibly hard ones, but those just on the boundary where existing techniques can do most of the work but the final step remains elusive.

transcript

Terence Tao: What's really interesting are the problems just on the boundary between what we can do easily and what are hopeless. But what are problems where existing techniques can do like 90% of the job and then you just need that remaining 10%?

02
Example

The Kakeya problem — how little area you need to rotate a needle in the plane to point in every direction — connects to wave propagation, partial differential equations, number theory, and the Navier-Stokes regularity problem.

Tao describes the Kakeya needle problem, from its simple puzzle origins to its surprisingly deep connections to wave concentration, blow-up phenomena, and the Navier-Stokes millennium problem.

transcript

Terence Tao: I think as a PhD student, the Kakeya problem certainly caught my eye. And it just got solved, actually. It's a problem I've worked on a lot in my early research. Historically, it came from a little puzzle by the Japanese mathematician Soichi Kakeya in like 1918 or so. So the puzzle is that you have a needle on the plane, or think like driving on a road, and you want to execute a U-turn, you want to turn the needle around, but you want to do it in as little space as possible. So you want to use this little area in order to turn it around. So, but the needle is infinitely maneuverable. So, you can imagine just spinning it around its center, and I think that gives you a disc of area, I think pi over 4. Or you can do a three-point U-turn, which is what we teach people in their driving schools to do, and that actually takes area pi over 8. So it's a little bit more efficient than a rotation. And so for a while, people thought that was the most efficient way to turn things around. But Bazakovich showed that, in fact, you could actually turn the needle around using as little area as you wanted. So 0.001, there was some really fancy multi-back and forth U-turn thing that you could do. that you could turn a needle around, and in so doing, it would pass through every intermediate direction. So we understand everything in two dimensions. So the next question is what happens in three dimensions? So suppose like the Hubble Space Telescope is tube in space. And you want to observe every single star in the universe. So you want to rotate the telescope to reach every single direction. And here's the unrealistic part. Suppose that space is at a premium, which totally is not. You want to occupy as little volume as possible in order to rotate your needle around in order to see every single star in the sky. How small a volume do you need to do that? And so you can modify Bezekovich's construction. And so if your telescope has zero thickness, then you can use as little volume as you need. That's a simple modification of the two-dimensional construction. But the question is that if your telescope is not zero thickness, but just very, very thin, some thickness delta, what is the minimum volume needed to be able to see every single direction as a function of delta? So as delta gets smaller, as the needle gets thinner, the volume should go down, but how fast does it go down? And the conjecture was that it goes down very, very slowly, like logarithmically, roughly speaking. And that was proved after a lot of work. So this seems like a puzzle, why is it interesting? So it turns out to be surprisingly connected to a lot of problems in partial differential equations, in number theory, in geometry, combinatorics. For example, in wave propagation, you splash some water around, you create water waves and they travel in various directions. But waves exhibit both particle and wave type behavior. So you can have what's called a wave packet, which is like a very localized wave that is localized in space and moving a certain direction in time. And so if you plot it into space and time, it occupies a region which looks like a tube. And so what can happen is that you can have a wave which initially is very dispersed, but it all focuses at a single point later in time. Like you can imagine dropping a pebble into a pond and the rubble spread out. But then if you time reverse that scenario, and the equations of wave motion are time reversible, you can imagine ripples that are converging to a single point, and then a big splash occurs, maybe even a singularity.

03
Mechanism

The Navier-Stokes regularity problem is difficult because of a 'Maxwell's demon' type possibility — the fluid could conspire to concentrate all its energy into smaller and smaller scales, forming a singularity in finite time.

Tao explains why proving that Navier-Stokes never blows up is so hard: the fluid's energy could, in principle, be transferred into ever smaller scales fast enough to overcome viscosity, like a 'demon' pushing energy into a singularity.

transcript

Terence Tao: Short answer is Maxwell's demon. So Maxwell's demon is a concept in thermodynamics. Like if you have a box with two gases, oxygen and nitrogen, and maybe you start with all the oxygen on one side and nitrogen on the other side, but there's no barrier between them, right? Then they will mix. And they should stay mixed. There's no reason why they should unmix. But in principle, because of all the collisions between them, there could be some sort of weird conspiracy like maybe there's a microscopic demon called Maxwell's demon that will, every time an oxygen and nitrogen atom collide, they will bounce off in such a way that the oxygen sort of drifts onto one side and then nitrogen goes to the other. And you could have an extremely improbable configuration emerge, which we never see. And statistically, it's extremely unlikely. But mathematically, it's possible that this can happen, and we can't rule that out. And this is a situation that shows up a lot in mathematics. A basic example is the digits of pi, 3.14159, and so forth. The digits look like they have no pattern, and we believe they have no pattern. On the long term, you should see as many ones and twos and threes as fours and fives and sixes. There should be no preference in the digits of pi to favor, let's say, 7/8. But maybe there's some demon in the digits of pi that like every time you compute more and more digits, it sort of biases 1 digit to another. And this is a conspiracy that should not happen. There's no reason it should happen. But there's no way to prove it with our current technology. Okay, so getting back to Navier-Stokes, a fluid has a certain amount of energy. And because the fluid is in motion, the energy gets transported around. And water is also viscous. So if the energy is spread out over many different locations, the natural viscosity of the fluid will just damp out the energy and it will go to zero. And this is what happens when we actually experiment with water. You splash around, there's some turbulence and waves and so forth, but eventually it settles down and the lower the amplitude, the smaller the velocity, the more calm it gets. But potentially there is some sort of demon that keeps pushing the energy of the fluid into a smaller and smaller scale. And it will move faster and faster. And at faster speeds, the effect of viscosity is relatively less. And so it could happen that it creates some sort of what's called a self-similar blob scenario where, you know, the energy of the fluid starts off at some large scale and then it all sort of transfers its energy into a smaller region of the fluid, which then at a much faster rate moves into an even smaller region and so forth. And each time it does this, it takes maybe half as long as the previous one. And then you could actually converge to all the energy concentrating in one point in a finite amount of time. And that's an always good finite time blow up.

gives example · 1provides context · 1

04
Prediction

By constructing a 'liquid computer' from the Navier-Stokes equations — a self-replicating fluid machine that transfers energy to smaller and smaller copies of itself — one could prove that finite-time blow-up is possible for the actual equations.

Tao describes a visionary idea inspired by Conway's Game of Life: building a fluid analog of a Turing machine out of water, which would create a self-replicating cascade that proves finite-time blow-up is possible for Navier-Stokes.

transcript

Terence Tao: So what I realized is that if you could pull the same thing off for the actual equations, so if the equations of water supported computation, so if you can imagine kind of a steampunk, but it's really water punk type of thing where, so modern computers are electronic, they're powered by electrons passing through very tiny wires and interacting with other electrons and so forth. But instead of electrons, you can imagine these pulses of water moving at a certain velocity, and maybe it's, they're two different configurations corresponding to a bit being up or down. probably that if you had two of these moving bodies of water collide, they would come out with some new configuration, which would be something like an AND gate or OR gate, that the output would depend in a very predictable way on the inputs. And like you could chain these together and maybe create a Turing machine, and then you have computers which are made completely out of water. And if you have computers, then maybe you can do robotics, you know, hydraulics and so forth. And so you could create some machine, which is basically a fluid analog of what's called a von Neumann machine. So von Neumann proposed, if you want to colonize Mars, the sheer cost of transporting people and machines to Mars is just ridiculous. But if you could transport one machine to Mars, and this machine had the ability to mine the planet, create some more materials, smelt them, and build more copies of the same machine, then you could colonize the whole planet over time. So if you could build a fluid machine, which, so it's a fluid robot. And what it would do, it's purpose in life, it's programmed so that it would create a smaller version of itself in some sort of cold state. It wouldn't start just yet. Once it's ready, the big robot configured water would transfer all its energy into the smaller configuration and then power down. And then I clean itself up. And then what's left is this newest state, which would then turn on and do the same thing, but smaller and faster. And then the equation has a certain scaling symmetry. Once you do that, it can just keep iterating. So this in principle would create a blow up for the actual Navier-Stokes. And this is what I managed to accomplish for this average Navier-Stokes. So it provided this sort of roadmap to solve the problem. Now, this is a pipe dream because there's so many things that are missing for this to actually be a reality. So I can't create these basic logic gates. I don't have these in these special configurations of water. I mean, there's candidates that include vortex rings that might possibly work, but also, analog computing is really nasty compared to digital computing. I mean, because there's always errors. You have to do a lot of error correction along the way. I don't know how to completely power down the big machine so that it doesn't interfere with the running of a smaller machine. But everything in principle can happen, like it doesn't contradict any of the laws of physics. So it's sort of evidence that this thing is possible.

explains mechanism · 1gives example · 1provides context · 1

05
Claim

The structure-randomness dichotomy is a core organizing principle in mathematics — most objects are random, but the few structured ones can be understood through 'inverse theorems' that connect partial patterns to fully structured objects.

Tao explains the dichotomy between structure and randomness in mathematics: most mathematical objects are random, but the rare structured ones can be analyzed via inverse theorems, as illustrated by Szemerédi's theorem on arithmetic progressions.

transcript

Terence Tao: This is a recurring challenge in mathematics that I call the dichotomy between structure and randomness, that most objects that you can generate in mathematics are random. They look like random, like the digits of pi. Well, we believe is a good example. But there's a very small number of things that have patterns. But now you can prove something as a pattern by just constructing, like if something has a simple pattern and you have a proof that it does something like repeat itself every so often, you can do that. And you can prove that, for example, you can prove that most sequences of digits have no pattern. So like if you just pick digits randomly, there's something called the low large numbers. It tells you you're going to get as many ones as twos in the long run. But we have a lot few. If you are tools to, if I give you a specific pattern like the digits of pi, how can I show that this doesn't have some weird pattern to it? Some other work that I spend a lot of time on is to prove what are called structure theorems or inverse theorems that give tests for when something is very structured. So some functions are what's called additive, like if you have a function of natural numbers, the natural numbers. So maybe, you know, two maps to four, three maps to six, and so forth. Some functions are what's called additive, which means that if you add 2 inputs together, the output gets added as well. For example, I'm multiplying by a constant. If you multiply a number by 10, if you multiply A + B by 10, that's the same as multiplying A by 10 and B by 10 and then adding them together. So some functions are additive. Some functions are kind of additive, but not completely additive. So for example, if I take a number N, I multiply by the square root of 2, and I take the integer part of that. So 10 by square root 2 is like 14 point something, so 10 up to 14. 20 up to 28. So in that case, additivity is true then. So 10 plus 10 is 20 and 14 plus 14 is 28. But because of this rounding, sometimes there's round-off errors and sometimes when you add A plus B, this function doesn't quite give you the sum of the two individual outputs, but the sum plus or minus 1. So it's almost additive, but not quite additive. So there's a lot of useful results in mathematics, and I've worked a lot on developing things like this to the effect that if a function exhibits some structure like this, then it's basically, there's a reason for why it's true, and the reason is because there's some other nearby function which is actually completely structured, which is explaining this sort of partial pattern that you have. And so if you have these sort of inverse theorems, it creates this sort of dichotomy that either the objects that you study are either have no structure at all, or they are somehow related to something that is structured. And in either way, in either case, you can make progress. A good example of this is that there's this old theorem in mathematics called Zemeredi's theorem, proven in the 1970s. It concerns trying to find a certain type of pattern in a set of numbers, the patterns have progression. Things like 3, 5, and 7, or 10, 15, and 20. And Zemeredi, Andrei Zemeredi proved that any set of numbers that are sufficiently big, what's called positive density, has arithmetic progressions in it of any length you wish. So for example, the odd numbers have a set of density one-half, and they contain arithmetic progressions of any length. So in that case, it's obvious because the odd numbers are really, really structured. I can just take 11, 13, 15, 17. I can easily find arithmetic progressions in that set. But Zermanism also applies to random sets. If I take the set of all numbers and I flip a coin for each number and I only keep the numbers for which I got a heads, because I just flip coins, I just randomly take out half the numbers, I keep one half. So that's a set that has no patterns at all. But just from random fluctuations, you will still get a lot of arithmetic progressions in that set.

06
Context

The most incomprehensible thing about the universe is that it is comprehensible — the 'unreasonable effectiveness of mathematics' reflects the phenomenon of universality, where complex micro-scale interactions produce simple macro-scale laws.

Tao reflects on the 'unreasonable effectiveness of mathematics' through the lens of universality: complex systems at the micro scale produce simple, compressible laws at the macro scale, as exemplified by the central limit theorem and the limitations revealed by the 2008 financial crisis.

transcript

Terence Tao: In fact, one of the great surprises of our universe and of everything in it is that it's compressible at all. It's the unreasonable effectiveness of mathematics. Yeah, Einstein had a quote like that, the most incomprehensible thing about the universe is that it is comprehensible. Right, and not just comprehensible. You can do an equation like E equals MC squared. There is actually some mathematical possible explanation for that. So there's this phenomenon in mathematics called universality. So many complex systems at the macro scale are coming out of lots of tiny new interactions at the macro scale. And normally because of the common form of explosion, you would think that the macro scale equations must be infinitely, exponentially more complicated than the macro scale ones. And they are, if you want to solve them completely exactly. Like if you want to model all the atoms in a box of air. That's like Avogadro's number is humongous. There's a huge number of particles. If you actually have to track each one, it'll be ridiculous. But certain laws emerge at the microscopic scale that almost don't depend on what's going on at the microscale, or only depend on a very small number of parameters. So if you want to model a gas of, you know, frontilian particles in a box. You just need to know its temperature and pressure and volume and a few parameters, like 506, and it models almost everything you need to know about these 10 to 23 or whatever particles. So we have We don't understand universality anywhere new as we would like mathematically, but there are much simpler toy models where we do have a good understanding of why universality occurs. Most basic one is the central limit theorem. That explains why the bell curve shows up everywhere in nature. But so many things are distributed by what's called a Gaussian distribution, famous bell curve. There's now even a meme with this curve. And even the meme applies broadly. The universality to the meme. Yes, you can go meta if you like, but there are many, many processes, for example, you can take lots and lots of independent random variables and average them together in various ways. You can take a simple average or more complicated average, and we can prove in various cases that these bell curves, these Gaussians emerge. And it is a satisfying explanation. Sometimes they don't. So if you have many different inputs and they're all correlated in some systemic way, then you can get something very far from the bell curve show up. And this is also important to know whether this element fails. So universality is not a 100% reliable thing to rely on. The global financial crisis was a famous example of this. People thought that mortgage defaults had this sort of Gaussian type behavior that if you ask if a population of, you know, 100,000 Americans with mortgages, ask what proportion of them would default on their mortgages. If everything was decorrelated, it could be an asset bill curve and you can manage risk with options and derivatives and so forth. And it is a very beautiful theory. But if there are systemic shocks in the economy that can push everybody to default at the same time, that's very non-Gaussian behavior. And this wasn't fully accounted for in 2008.

explains mechanism · 1

Highlight slides
Related episodes