Some important caveats on the above findings:
1) We discuss many questions that have resolved early because AI milestones were reached sooner than expected, but this may introduce bias. By definition, we do not have data on late-resolving questions, and this may lead us to identify more areas where forecasters underestimated AI progress than where they have been accurate or overestimated progress.
2) In our analysis, for some unresolved questions we have used projections from LLMs to draw tentative accuracy conclusions. Though LLM forecasts are now comparable to strong human forecasts on many topics (see our work on ForecastBench, a forecasting benchmark), these projections are, by nature, more speculative.
We plan to continue to publish accuracy updates over time, particularly for our Longitudinal Expert AI Panel (LEAP), and will note cases where findings or projections change in a way that alters our previous interim conclusions.