Forecasting is the only truly renewable source of ground-truth-verified hard evaluation data for AI, since you can pose arbitrarily hard questions about the future and just wait for reality to grade them.
Dan Schwarz argues forecasting is uniquely valuable as an AI eval because, unlike expert-authored benchmarks (which AI is starting to outsmart), the future eventually resolves every forecasting question with objective ground truth, making it an inexhaustible source of hard test cases. ✦ AI generated
Dan Schwarz · The Cognitive Revolution · 2026-07-09 · original ↗
starts at this moment · 57:08
Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting as the the kind of Elon Musk tweet the quips that forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is.
verbatim transcript · starts at 57:08
56:49smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting as the the kind of Elon Musk
57:08tweet the quips that forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is. Again, I think coding intelligence, AI, R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier Labs like Canon should be paying attention to.
57:30One of the concerns a lot of people have about AI super forecasting is that it's too in distribution. I actually heard this from one of the very best forecasters I've ever had the pleasure of working with in my career. He basically said he believes that a system like future search would beat him head-to-head in a forecasting tournament about kind of near-term outcomes of things that are within distribution. But
57:49if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the AIS for exactly the reason that you gave. they are trained to try to predict things that have actually happened and when things get wonky you need some kind of uh creative lateral thinking. I think uh the rate of AI improvement is uh so
58:07astounding that I think that even the kind of lateral thinking like trying to imagine a completely different scenario will fall to the AIs. One unfortunate thing about it is it's hard to test this. Um, so I think the more that AI continues doing strange things to the world and we wake up and see strange things in the news and those strange things are metaculous questions and on
58:25forecast bench and teams like mine are trying to predict them better, we will actually get more evidence. But if there's a kind of a step change in the nature of the world, if we enter some sort of AGI transformative AI type of world, you know, we've got these geniuses and data centers as people say or anything like AI 2027 happens, then I think it's uh it's kind of going to be
58:45wild west. I will say I don't think humans are doing particularly great at imagining transformative AI. So the bar is a bit lower. Really when you play with these AI forecasters, you will find them to be quite human and how they structure their reasoning. And again, this is not an accident. Like they're trained on how humans have structured their reasoning before. So a human forecaster would love to say, okay, was
59:06the last 10 times something like this happened? What were the outcomes of those 10 times? And now I can make a distribution and say it's probably going to be something like this. the fact that an AI will do that. Is it because it independently is arriving at the same conclusion? Is it because it's trained on humans doing that? Is it because it just thinks like a human? I don't think
- ·Expert-authored benchmarks: AI is starting to outsmart them
- ·Forecasting questions: reality itself grades every answer
- ·Future always resolves, providing objective ground truth
- ·Arbitrarily hard questions can be posed, then verified later
- ·Only 'completely and utterly renewable' source of hard eval data
- ·New hard questions can always be posed about the future
- ·Forecasting framed as the 'ultimate measure of intelligence'