Forecasting is uniquely valuable as an AI evaluation method because it is the only capability where you can generate limitless, arbitrarily hard questions and eventually get exact, unambiguous ground truth simply by waiting for the future to arrive.
Dan Schwarz argues that unlike coding or other domains where humans struggle to generate problems harder than the models themselves, forecasting offers an endless supply of hard questions with guaranteed, verifiable ground truth once time passes. ✦ AI generated
Dan Schwarz · The Cognitive Revolution · 2026-07-07 · original ↗
starts at this moment · 92:41
“what what what should they be doing from your perspective?”
Forecasting uh has this beautiful property that you basically get ground truth by waiting... you basically have a completely limitless set of extremely hard basically impossible questions where you get exact ground truth. And there is no other email like this... Forecasting, I think, is the only completely and utterly renewable source of this.
verbatim transcript · starts at 92:41
92:41some question about the future and basically an impossibly hard question, a question that even an AGI, an oracle, a god could never really say because of chaos theory. Imagine just trying to predict, you know, like cubic meter weather 3 weeks in the future. Like you'd never be able to do it. Um, but if you just wait, then you will see what that weather was in that cubic meter
93:03three weeks in the future. And so you basically have a completely limitless set of extremely hard basically impossible questions where you get exact ground truth. And there is no other email like this. There is if you want to say improve a coding harness, you just need to have more and more hard coding problems that are not in the training data for which you can say this is
93:23definitely the correct answer so that you can do some training on it. And that's hard. I think uh human experts, doctors, lawyers, engineers, financers, whoever who are trying to make evals to try to produce data for the frontier labs are finding that they are not smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that
93:43correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting as the, you know, the the kind of Elon Musk tweet the quips that forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it
94:02is. Again, I think coding intelligence, AI, R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier Labs like Canon should be paying attention to. Do you have a sense for how much irreducible chaos there is? Like in the limit of super intelligence, how much visibility
94:31do we get into the future? >> Yeah, that is a great question. Uh David Mannheim actually just uh posted a tweet that has a graph that has a claim to estimate this that I was meaning to dig into because I don't know where those lines are coming from. Um I would say without evidence behind this my sense is very Ukowskian like I think there is a lot of detail in reality that is far
- ·Forecasting yields exact ground truth just by waiting
- ·Endless supply of arbitrarily hard questions
- ·No human effort needed to generate harder problems
- ·Unlike coding, ground truth arrives automatically
- ·Humans struggle to write questions harder than models
- ·Coding lacks a renewable, limitless problem source
- ·Forecasting alone is 'completely and utterly renewable'