AI forecasting systems, evaluated via 'past casting' (testing on pre-training-cutoff internet snapshots to avoid hindsight bias), are now competitive with human superforecasters and teams of humans.
Dan Schwarz explains that Future Search's 'past casting' method — evaluating models on snapshots from before their training cutoff to get instant, hindsight-free ground truth — shows that over the past year AI forecasters have become competitive with, or better than, human superforecasters and teams. ✦ AI generated
Dan Schwarz · The Cognitive Revolution · 2026-07-09 · original ↗
starts at this moment · 54:04
“But how do you evaluate a forecaster without waiting months for the future to arrive?”
if you read Stat's article, you will see that over the last 12 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. AI is competitive with humans and even teams of humans working together.
verbatim transcript · starts at 54:04
53:45is taking a snapshot of the internet from some months ago and using the training window cut off of models to basically trick them into forecasting without the hindsight bias. This is very useful for us because we can evaluate things immediately. So when Fable came out the first time that Claude Fable came out, uh we were able to evaluate it within 24 hours and it was the best
54:04single agent forecaster on our leaderboard, everyone else had to wait weeks or months to find out how good Claude Fable actually was. So we internally using the benchmark we call bench to the future saw this progression kind of in real time. The rest of the world is seeing it kind of some months behind. And so if you if you read Stat's article, you will see that over the last
54:2512 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. AI is competitive with humans and even teams of humans working together. Whether it's better is you got to synthesize a whole bunch of different disparate sources of evidence. I would say if you're curious about this or if
54:44you have forecasting needs in your life, you really should try it. So just go to future search. You get $20 free. So you can try some Frontier forecasts immediately and I think you should judge for yourself whether you think they're good. [music] >> We asked what the Frontier Lab should do with a forecaster this good. [music] >> Yeah. So there's kind of two questions to this. One is what should they be
55:03doing with forecasting as a capability and what should they be doing with forecasting as an eval? So forecasting as a capability is kind of a business decision. What does say OpenAI care whether Chat GBT is a good forecaster? I think that question is based on whether they their consumers care about it as a good forecaster. If you're anthropic, I think you probably care more about the enterprise case. Like when people are
55:24using claude to do white collar work, do they care how good it is as a forecaster? Are people trying to use claude to make say financial forecast in an Excel spreadsheet? Is that something they care about? So that's a business decision and I can't really weigh in on that. I think again people will be discovering over time just how important forecasting is in everything, but it's going to be a slow process for humans to
55:44notice that. I think um from an eval side it's very different. Forecasting uh has this beautiful property that you basically get ground truth by waiting. So if I ask some question about the future and basically an impossibly hard question, a question that even an AGI an oracle a god could never really say because of chaos theory. Imagine just trying to predict you know like cubic meter weather 3 weeks in the future.
- ·Models evaluated on pre-training-cutoff internet snapshots
- ·Method yields instant, hindsight-free ground truth
- ·Used by Future Search to score AI forecasters
- ·Past 12 months: strong evidence AI has closed the gap
- ·Past 6 months: live tournaments, prediction markets tested
- ·AI competitive with individual humans and human teams