Published: Sep 28, 2026
Editorial
  • Editorial

Improving the Accuracy and Relevance of AI Forecasts We Publish

Improving the Accuracy and Relevance of AI Forecasts We Publish
Josh Rosenberg, Matt Reynolds, Nadja Flechner, and Zack Devlin-Foltz
Published: Sep 28, 2026
Josh Rosenberg, Matt Reynolds, Nadja Flechner, and Zack Devlin-Foltz

Introduction

In the second post in our series on the accuracy of AI progress forecasts, we look at a few ways we are aiming to improve the accuracy and relevance of LEAP forecasts in the future.

In the previous post we looked at AI progress forecasts from across FRI’s work. Broadly, we found that the median expert, superforecaster, and member of the public have consistently underestimated AI progress on near-term capability questions such as benchmarks. They also underestimated revenue growth at AI companies but have a more mixed track record on other adoption indicators. We don’t have enough evidence on forecasts about macroeconomics or other broad social impact forecasts to draw conclusions about their track record in those areas.

Our findings suggest that median forecasts from these groups are not reliable estimates for predicting the trajectory of AI on many questions—particularly those concerning benchmarks and AI capabilities.

This raises an obvious question: If we want to get a better grasp of the trajectory of AI and its impacts on society, whose forecasts should we listen to?

We’re working on three major changes to the Longitudinal Expert AI Panel (LEAP) to address this question:

  1. We will share forecasts from a subsample of panelists who expect fast AI progress by 2040.1
  2. We are adding regularly updated LLM forecasts to LEAP findings.
  3. We will highlight forecasts from the most accurate LEAP participants as soon as we have sufficient data to make this judgment rigorously.

We’re also planning on putting more emphasis on the rationales of forecasters who make particularly insightful arguments as part of their forecasts, and will make some changes to our elicitation process, like allowing forecasters to update after seeing each other’s forecasts and reasoning.

In the rest of the post, we go into more detail on these changes and give a glimpse of what LEAP forecasts might look like in the future with these new additions.

Sharing forecasts from a “fast AI progress” worldview

As we saw in the last post, the majority of the LEAP panel so far has underrated AI progress on benchmarks and other capability indicators, as well as some near-term proxies of AI adoption such as revenue at top AI companies. This was highlighted recently with the news that OpenAI has likely settled the Navier–Stokes problem, surprising LEAP experts and superforecasters who gave median forecasts of 10% and 5.4% that a Millennium Prize Problem would be solved with AI assistance before the end of 2027.

However, a subset of the LEAP panel predicts much faster AI progress and, because of this, has been more accurate on the near-term questions we highlighted in the first post in this series. This subsample puts more weight on the possibility of transformative AI arriving soon than most LEAP forecasters, and we’re interested in what this group thinks about other AI progress questions. By breaking out this sample, we’ll address a core critique of typical LEAP experts’ forecasts: that they greatly underrate how fast AI progress is moving and therefore are giving an inaccurate picture of how the world will look going forward.

Perhaps the “fast AI progress” subgroup will be more accurate in general, or maybe they will be particularly good at forecasting capabilities while doing worse on societal impacts, for example. Our goal is to highlight this group’s forecasts for people who want to put more weight on them, and to explicitly track where this worldview is more and less accurate.

We explore several alternative ways to select this subsample in more detail in the brief appendix below. In our forthcoming work, we plan to share more about what this group forecasts about the future of AI and why.

LLM forecasts of LEAP questions

Having LLMs forecast LEAP questions could be a powerful way to make our forecasts more accurate and timely. We know from our forecasting benchmark, ForecastBench, that top LLMs are at parity with superforecasters on specific kinds of forecasting questions, and we expect LLMs’ forecasting abilities to increase as models become more capable.

LLMs also provide an advantage in forecasting areas with rapidly changing information, such as AI progress. Since we run LEAP surveys every month—and ask most questions only once—awkward timing can sometimes mean that our panel’s forecasts are based on information that quickly becomes out of date. 

For example, in Wave 11 of LEAP we asked participants to forecast the annualized revenue of Anthropic and OpenAI. At the time of forecasting (July 13–August 11, 2026), Anthropic’s reported annualized revenue was $47 billion, but just a few days after the survey closed, this figure was updated to $65 billion. While it’d be tricky to re-ask questions of human forecasters, it is relatively trivial to ask LLMs the same questions and see how their forecasts change over time.

We plan to run regular LLM forecasts on key LEAP questions, allowing us to provide high-quality, up-to-date forecasts across key questions. In a later post, we’ll describe how we plan to prompt and ensemble our LLMs to provide the most useful forecasts.

We see these LLM forecasts as a supplement to—not a replacement for—human forecasts. LLM forecasts still have many limitations, and we plan to regularly assess their strengths and weaknesses, including by building new benchmarks that assess LLM skill at AI-focused questions, low-probability questions, conditional questions, and more.

Highlighting more accurate LEAP forecasters 

At the moment, we’re unable to confidently identify the most accurate LEAP forecasters because of the small number of questions that have definitively resolved. We’ve made this easier by adding more questions that resolve in the near term, and by the end of 2026 we will have roughly 16 additional resolved questions that should enable us to start identifying more accurate forecasters.

In the new year we plan to publish a thorough analysis of LEAP accuracy so far, and as part of that analysis we will identify a group of more accurate LEAP forecasters. We’ll break out this more accurate sample in our LEAP reports to draw particular attention to their forecasts, and may also explore ways to increase the prominence of these more accurate forecasters within the LEAP sample.

Example of future LEAP forecasts

The charts below are examples of the kinds of forecasts we plan to share in future LEAP waves. The “fast AI progress” sample represented here is from a rough subsample that fulfills the inclusion criteria outlined in the appendix. The exact composition of this group and the criteria we use to select them may change, although the forecasts below are a good indication of how the views of this group differ from those of experts and superforecasters. The LLM forecasts are the median of 160 pooled runs from eight models, run in May 2026. As with the “fast AI progress” sample, the exact methodology we use to derive and display our LLM forecasts may change.

We’ve also recruited a new group of LEAP experts from people who work on mitigating catastrophic risk from AI. The forecasts from this group are not included in the graphs below as they have only forecast a small number of LEAP questions, but this subsample will be referred to in future LEAP reports. We include more information on how this group was selected in the appendix below.

Figure 1: An example of a LEAP forecast from Wave 8 incorporating LLM forecasts and a subsample of respondents who expect fast AI progress. The Technological Richter Scale (TRS) ranks technologies on a logarithmic scale by societal impact. For example, Level 9 (technology of the millennium) includes agriculture and the wheel, while Level 10 (technology of the epoch) includes events that alter the fate of the planet, such as the rise of humans. Note that the “fast AI progress” sample in this example was selected for having at least 50% probability on TRS Level 9 or 10.
Figure 2: An example of a LEAP forecast from Wave 9 incorporating LLM forecasts and a subsample of respondents who expect fast AI progress. 

Appendix on the “fast AI progress” subsample

There are several ways to select a subsample of LEAP panelists who place more weight on transformative AI arriving soon than other panelists. 

One approach is to do this relatively. We could select a set of questions that measure expectations about AI progress, then select the top 20% or 10% of panelists based on how quickly they expect AI to progress. For example, we could select the 20% of panelists who gave the highest median forecasts on a set of benchmark or other capability questions across LEAP waves.

This approach is easy to describe and would capture a large subsample of LEAP participants. It would also make it easy to define comparison groups, as we could compare this group with the 20% of forecasters who expect the slowest progress, and so on, through the different quintiles. However, this approach does not guarantee the inclusion of participants who believe that AI will be transformative on relatively short timescales and would not define the subsample in terms that believers in the relevant hypothesis would necessarily agree with. (E.g., if our sample happens to not include many very fast AI progress forecasters, choosing the top 20% could still lead to conservative estimates of AI progress.)

A second approach would be to define this group by relevant professional expertise. We could highlight the median forecasts of a group of experts who work in the field of responding to risks from transformative AI. Understanding this group’s worldview is an important task, particularly as these views are becoming more prominent in AI policy discussions. Because of this, we’ve recruited a subsample that represents these views, and we will present these as a distinct and separate subsample in future LEAP waves. We call this subsample our “AI risk expert” subsample.

However, this subsample is not necessarily a good proxy for people who believe that AI progress will be very rapid. For example, median AI progress forecasts of the AI risk expert subsample were not dramatically higher than those from other experts. They fell short of the implied forecasts from senior AI company staff members, for example, suggesting that this group does not fully capture a truly “fast AI progress” worldview. We plan to report median forecasts from the AI risk expert subsample, but will make it clear that this is a separate sample from our “fast AI progress” group. We’ll share more details on how we recruited this group and what they believe about the future of AI in a forthcoming post.

The final approach is to select forecasters in more absolute terms. In the case of this sample, we could include candidates whose forecasts meet all three of a set of criteria, for example:

  • A median probability of at least 90% that by 2100 more than 50% of LEAP panelists will agree that AGI exists.
  • Conditional on AGI existing, a median forecasts that it arrives in or before 2040.
  • A combined probability of at least 50% of AI reaching either Level 9 (technology of the millennium) or 10 (technology of the epoch) on the Technological Richter Scale by 2040.

The absolute approach involves selecting somewhat arbitrary cutoffs for who belongs in vs. out of the subsample. But it also sets a higher bar for beliefs about AI progress and therefore provides a cleaner test of the “fast AI progress” hypothesis. We want to choose a group that people who hold a fast AI progress worldview would feel properly represents them. 

We will share more details about our final decision on how to select this sample, how many forecasters are in the “fast AI progress” subsample, and what they believe about the future of AI in forthcoming work.

Notes

  1. We’re also adding an “AI risk expert” subsample to future LEAP waves, which is separate from this group. For more on this group, see the appendix below. ↩︎

    Related Research
    Editorial
    How Accurate Have AI Progress Forecasts Been So Far?
    Sep 22, 2026
    Project
    ForecastBench
    Ongoing
    Technical report
    Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks
    Aug 12, 2026
    Working paper
    Forecasting AI Cyber Risks and Capabilities: Results of a 2025 Pilot Study
    Jul 23, 2026