Forecasting is the only truly renewable source of evaluation questions with guaranteed ground truth, since human experts can no longer out-forecast the models they're supposed to be grading, making forecasting close to the 'ultimate eval' for AI even if not the ultimate form of intelligence itself.
Dan Schwarz argues forecasting is uniquely suited as an AI eval because, unlike coding or expert-authored benchmarks, it supplies an endless stream of hard questions whose correctness is guaranteed once time passes—and human experts can no longer game or outsmart it. ✦ AI generated
Dan Schwarz · The Cognitive Revolution · 2026-07-09 · original ↗
starts at this moment · 56:49
Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting as the the kind of Elon Musk tweet the quips that forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is. Again, I think coding intelligence, AI, R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval.
verbatim transcript · starts at 56:49
56:30need to have more and more hard coding problems that are not in the training data for which you can say this is definitely the correct answer so that you can do some training on it. And that's hard. I think uh human experts, doctors, lawyers, engineers, financers, whoever who are trying to make evals to try to produce data for the frontier labs are finding that they are not
56:49smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting as the the kind of Elon Musk
57:08tweet the quips that forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is. Again, I think coding intelligence, AI, R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier Labs like Canon should be paying attention to.
57:30One of the concerns a lot of people have about AI super forecasting is that it's too in distribution. I actually heard this from one of the very best forecasters I've ever had the pleasure of working with in my career. He basically said he believes that a system like future search would beat him head-to-head in a forecasting tournament about kind of near-term outcomes of things that are within distribution. But
57:49if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the AIS for exactly the reason that you gave. they are trained to try to predict things that have actually happened and when things get wonky you need some kind of uh creative lateral thinking. I think uh the rate of AI improvement is uh so
58:07astounding that I think that even the kind of lateral thinking like trying to imagine a completely different scenario will fall to the AIs. One unfortunate thing about it is it's hard to test this. Um, so I think the more that AI continues doing strange things to the world and we wake up and see strange things in the news and those strange things are metaculous questions and on
58:25forecast bench and teams like mine are trying to predict them better, we will actually get more evidence. But if there's a kind of a step change in the nature of the world, if we enter some sort of AGI transformative AI type of world, you know, we've got these geniuses and data centers as people say or anything like AI 2027 happens, then I think it's uh it's kind of going to be
58:45wild west. I will say I don't think humans are doing particularly great at imagining transformative AI. So the bar is a bit lower. Really when you play with these AI forecasters, you will find them to be quite human and how they structure their reasoning. And again, this is not an accident. Like they're trained on how humans have structured their reasoning before. So a human forecaster would love to say, okay, was
- ·Forecasting yields endless, renewable evaluation questions
- ·Ground truth is guaranteed once time passes
- ·Human experts can no longer out-forecast models
- ·Coding and expert benchmarks lack this renewability
- ·Forecasting isn't the sole measure of intelligence
- ·Coding, R&D, interpersonal skills remain vital
- ·Still, forecasting is close to the ultimate eval