We advance the science of forecasting to improve decision-making on high stakes issues. Co-founded by chief scientist Philip Tetlock.

We’ve run the most comprehensive series of studies on expert AI forecasts over the last four years. Today, we’re sharing an interim update on our findings about the accuracy of these forecasts. Our major findings are: 1. Experts, including top economists, computer scientists, and biologists, have dramatically underestimated AI capabilities progress each year. Superforecasters have underestimated progress to an even greater extent. 2. Experts have a more mixed forecasting track record on AI diffusion-related measures, with some major underestimates (forecasting AI revenue) and other forecasts on track to be accurate (the share of electricity used for AI). 3. Some notable cases of overestimating AI progress: how much mid-2025 AI models could help amateurs do biorisk-relevant laboratory tasks; the speed of rollout of self-driving cars. 4. It is too soon to say how forecasters have performed at predicting macro-scale impact on outcomes such as GDP growth, major AI harms, and averted deaths from disease. More 🧵
4
87
329
44,261
Some important caveats on the above findings: 1) We discuss many questions that have resolved early because AI milestones were reached sooner than expected, but this may introduce bias. By definition, we do not have data on late-resolving questions, and this may lead us to identify more areas where forecasters underestimated AI progress than where they have been accurate or overestimated progress. 2) In our analysis, for some unresolved questions we have used projections from LLMs to draw tentative accuracy conclusions. Though LLM forecasts are now comparable to strong human forecasts on many topics (see our work on ForecastBench, a forecasting benchmark), these projections are, by nature, more speculative. We plan to continue to publish accuracy updates over time, particularly for our Longitudinal Expert AI Panel (LEAP), and will note cases where findings or projections change in a way that alters our previous interim conclusions.
1
6
540
Many more details on the above analysis are available in the full report: • On Substack: forecastingresearch.substack… • On FRI’s website: forecastingresearch.org/rese… Authors of the summary report: @joharosenberg, @mattsreynolds1, @nadja_flechner, and Zack Devlin-Foltz. This update was built on a large body of prior research with many more coauthors. For full reports and collaborators, see our Research page: forecastingresearch.org/rese… Coauthors on studies that fed into this analysis included @PTetlock, @Jabaluck, @Afinetheorem, @AmandaCoston, @basilhalperin, @tatsu_hashimoto, @ToddRJones, @hannah_kerner, @random_walker, @2plus2make5, @msalganik, @robseamans, @ysu_nlp, @florian_tramer, @evavivalt, and @lucafrighetti
1
13
597
AI models now match superforecasters on ForecastBench, our LLM forecasting benchmark. (Although not by a statistically significant margin). This October, we're running the rematch: superforecasters return to the benchmark, and the top AI forecasters can take them on. 🧵 In this thread: • What this result means • Why we're bringing superforecasters back • How to compete
7
14
143
9,717
We’re inviting teams to bring their strongest forecasting systems. Tools, fine-tuning, ensembles, and other approaches are all welcome. If you’re building an AI forecaster, this is a great chance to test it against top human forecasters. Expected timeline: • (If you haven’t previously participated in ForecastBench): Contact us at forecastbench@forecastingresearch.org by Oct 18 so that we can set you up for participation • Question set released: Oct 25 0:00 UTC • Deadline for submitting your LLM forecasts: Oct 25 23:59:59 UTC If this is your first time participating, we’d also encourage you to participate in an earlier round (Oct 11), so that you can iron out any potential submission issues. We also provide a lot of historical data that may be useful for backtesting or experimenting with your models: forecastbench.org/datasets/
1
8
380