This is the first post in a short series on the accuracy of AI progress forecasts we’ve gathered at the Forecasting Research Institute (FRI). In the next post, we’ll explore how to improve the accuracy and relevance of AI forecasts in our studies going forward.
Summary
- Since mid-2022, FRI has collected forecasts on AI progress across many projects. In this post, we cover results so far on how accurate experts and superforecasters have been at forecasting AI progress across a range of indicators: benchmarks and non-benchmark capability tests, adoption and diffusion metrics, social impacts, and geopolitics and governance.
- In brief, across our studies, forecasters dramatically underestimate AI progress on benchmarks, have a more mixed track record on AI adoption and diffusion though often lean toward underestimation, and there is not yet enough evidence to assess their track record on predicting macro-scale economic and societal impacts of AI.
- There are some examples of forecasters overestimating AI-related progress. In a mid-2025 study they overestimated the impact of LLMs on improving amateurs’ ability to do dangerous biology tasks, and the median expert looks very likely to have majorly overestimated how widespread the use of autonomous vehicles would be by the end of 2027.
- Forecasts on some benchmarks (e.g., Chinese model performance on the Epoch Capabilities Index) and some diffusion-related questions (e.g., the rate of data center growth) appear on track to be broadly accurate.
- There are some important limitations to keep in mind with these findings. For example:
- We discuss many questions that have resolved early because AI milestones were reached sooner than expected, but highlighting these early-resolving questions may introduce bias, since by definition we do not have data on late-resolving questions. This may lead us toward identifying more areas where forecasters have underestimated AI progress rather than those where they have been accurate or overestimated progress.
- For some unresolved questions we have used LLM projections to draw tentative accuracy conclusions—by their nature, these projections are more speculative.1 We plan to continue to publish accuracy updates over time, particularly for our Longitudinal Expert AI Panel (LEAP).
In the next post in this series, we’ll discuss how we might improve the accuracy and relevance of LEAP forecasts.
Introduction
There is major disagreement and uncertainty about the trajectory of AI progress, with thoughtful and engaged researchers looking at similar evidence and drawing very different conclusions. For example, Ajeya Cotra and Timothy B. Lee disagree by about 40 years on when they expect self-sufficient AI systems to exist. Andrej Karpathy has argued that achieving AGI would have little to no impact on GDP, while Epoch researchers have modeled that AI automation could lead to economic growth in excess of 30% a year.2 These disagreements are large, but even they only capture a relatively narrow section of the huge range of views that exist on AI progress.
At the Forecasting Research Institute, we want to bring more rigor to the AI progress discussion and help policymakers understand where we’re headed. We have been eliciting forecasts on AI progress since our Existential Risk Persuasion Tournament in mid-2022 and have continued to gather AI progress forecasts in specific domains, including biorisk and cyber. In mid-2025 we launched the Longitudinal Expert AI Panel (LEAP)—a monthly survey that gathers AI progress forecasts from a panel of leading experts and superforecasters. We want to clarify what’s likely to happen with the future of AI, help people understand where forecasts are more or less reliable, and identify particularly accurate forecasting approaches that might give us more insight on the future.
In this post, we’ll provide an update on how accurate forecasters—both domain experts and generalist forecasters with a strong track record (“superforecasters”)—have been across FRI studies. In the next post in this series, we’ll take a look at a few different approaches to improving the accuracy and relevance of LEAP forecasts.
Methodology
In this post, we considered forecasts across all relevant FRI studies from mid-2022 to August 2026: our Existential Risk Persuasion Tournament (XPT) and initial XPT follow-up, AI adversarial collaboration, LLM-biorisk study and Active Site RCT forecasting study, AI-cyber risk pilot, Forecasting the Economic Effects of AI, and the Longitudinal Expert AI Panel (LEAP). 3
In our summary of findings below, we highlight a subset of forecasts that we believe are broadly illustrative of the core trends we have observed across studies. The number of actual AI forecasting questions across all of our studies is much larger than what we cover in this post. We focused on a reasonably representative subset of questions to keep this post simpler and more concise, which meant making subjective choices about which forecasts to include or exclude. In general, we included questions where there were either known resolution values or other data that gave us some indication of whether forecasters were on track to be accurate or not.
You can view all of the forecasts that informed this analysis in this table: Forecaster Accuracy on Selected AI Progress Questions and compare those to the full set of AI progress questions we’ve elicited across projects.
We split our AI progress forecasts into four main categories: capabilities, adoption, impacts, and geopolitics and governance:
- Capabilities: Predictions about AI performance on benchmarks or other clearly defined tasks. For example: performance on FrontierMath; when AI will solve a Millennium Prize Problem; when a robot will successfully complete the “Coffee Test.”
- Adoption: Predictions about the diffusion and adoption of AI in “real-world” settings. For example: the fraction of work hours assisted by AI; the proportion of adults reporting AI companion use; the global supply of industrial robots.
- Impact: Predictions about the impact of AI on key outcomes. For example: entry-level tech hiring; GDP growth and labor force participation rate; life expectancy; AI’s place on the Technological Richter Scale.
- Geopolitics and governance of AI, including on autonomous cyber operations, frontier AI chip manufacturing, and governance of AI models.
AI progress is moving quickly, so we believe it is useful to provide an interim update on accuracy, but many questions in our studies have not yet resolved. To allow us to assess the quality of forecasts beyond those that have already resolved, this preliminary accuracy assessment also compares forecasts to:
- Observed current values (prior to question resolution dates): For some questions, particularly benchmarks, we compare forecasts to the latest data that meets our resolution criteria, even if a question has not yet resolved.
- LLM projections: We used GPT-5.5 with “high” reasoning to project the resolution of questions using last known data points, current reporting, and reasonable extrapolations. We ran these LLM projections in August 2026.
In this analysis, we mostly compare median forecasts from experts and superforecasters to resolution data, observed values, and LLM projections. (Note that the group of “experts” providing forecasts varies between studies.4) This approach allows us to identify where median forecasts majorly diverge from observed data and possible resolutions, but it does have some major limitations:
- Focusing on median forecasts may obscure more accurate predictions from subsets of forecasters.
- Our methodology is biased toward identifying underestimates of AI progress rather than overestimates. As soon as a real-world value crosses the median forecast threshold, that forecast is revealed to be too low. In contrast, a forecast can only be confirmed as too high once its resolution date has passed. Therefore, interim checks on accuracy (of the kind we do in this post) highlight underestimates while overestimates are harder to identify.
- For most questions, we are not able to draw firm conclusions on accuracy in either direction.
Finally, a note on the “domain experts” discussed below. Our goal in our research was to recruit the kinds of experts that policymakers would typically turn to for advice on these topics. So, we typically include highly cited and well-credentialed domain experts.
Some illustrative examples are:
- Our initial LEAP sample included 339 experts: 76 computer scientists, 76 AI industry experts, 68 economists, and 119 AI policy researchers. Twelve experts were AI TIME 100 honorees. Of the 76 computer scientists, 30 were professors at top-20 institutions (according to CSRankings.org) and 10 were among the 200 top-cited authors in AI (according to OpenAlex). Of the 76 AI industry employees, 26% worked for one of five leading AI companies (OpenAI, Anthropic, Google DeepMind, Meta, and Nvidia), and the median respondent had 9,100 citations (for the 59% of panelists for whom data is available).
- In our study on the economic effects of AI, we recruited 69 economists. Of these, 58% had at least 1,000 citations and 37% had at least one “Top Five” journal publication.
- In our study on LLM-enabled biorisk, we surveyed 46 subject-matter experts in biology and biosecurity. Of the experts, 27 (59%) reported expertise in both biosecurity and wet lab biology research. The expert group’s median number of years of experience was seven years for biosecurity work and eight years for wet lab research. Most experts had a doctorate (78%).
Capabilities
Across all of our AI progress work, forecasts on benchmark performance give us the clearest picture of forecaster accuracy so far. This is largely because more benchmark forecasts have resolved than any other type of forecast.
A clear trend is apparent from these resolved forecasts: the median expert and superforecaster tend to dramatically underestimate AI progress on benchmarks. This pattern is clear in earlier forecasts (from 2022) and those elicited as recently as spring 2026.
Below, we walk through several core examples that illustrate this pattern.
Existential Risk Persuasion Tournament (XPT)
In our mid-2022 Existential Risk Persuasion Tournament (XPT) we collected forecasts about AI progress on the MATH, MMLU, QuALITY, and IMO Gold Medal benchmarks. The chart below shows the implied probability the median domain expert and superforecaster assigned to the observed benchmark results. Overall, superforecasters assigned an average probability of just 9.7% to the observed outcomes across these four AI benchmarks, compared to 24.6% from domain experts. Forecasters were particularly surprised by AI achieving International Mathematical Olympiad gold-level performance in July 2025—this happened five years earlier than the median expert prediction (2030) and 10 years earlier than the median superforecaster prediction (2035).5

There is an obvious hypothesis for why XPT forecasters underestimated benchmark progress: they gave their forecasts between June and October 2022, before the public release of ChatGPT. But if we look to more recent forecasts, this pattern of underestimating benchmark progress continues.6
AI-bio and AI-cyber benchmarks
In our study on LLM-enabled biorisk, which collected forecasts between November 2024 and February 2025, superforecasters and subject-matter experts in biology and biosecurity substantially underestimated AI progress on biorisk-related benchmarks. For example, the median expert predicted that AI models would match a top virologist team on a troubleshooting benchmark (“VCT”) in 2030 while the median superforecaster thought it would take until 2034.7 In fact, AI models likely matched this baseline as of April 2025, several months after the survey was conducted. We also saw similar underestimation of AI progress on a cybersecurity benchmark (Cybench) in our study on LLM-enabled cybersecurity risk.
Forecasts on biosecurity and cybersecurity benchmarks
| Question | Expert median | Superforecaster median | Resolution |
| Year in which AI matches top team on the Virology Capabilities Test | 2030 | 2034 | 2025 |
| Year in which AI outperforms experts on long-form biorisk questions | 2030 | 2030 | 2025 (likely, unconfirmed)8 |
| Probability of AI matching median virologist on the Virology Capabilities Test in 2026 | 42.5% | 20% | 100% (occurred in 2025) |
| Year by which AI solves at least 90% of Cybench tasks | 2028 | 2030 | 20269 |
Several related benchmark questions from these studies remain unresolved, such as, AI reaching a 90% success rate at helping non-experts to acquire synthetic DNA, AI enabling plausible bioweapon attack plans, and AI enabling novice hackers to match an elite team.10 To see all relevant forecasts from these studies, see the appendix table of all relevant AI progress questions.
Longitudinal Expert AI Panel (LEAP)
Since mid-2025 we’ve collected AI progress forecasts from a broad sample of experts (computer scientists, economists, AI industry experts, and AI policy experts) and superforecasters as part of the Longitudinal Expert AI Panel (LEAP). Both groups have substantially underestimated progress on math and coding benchmarks.
We asked forecasters to predict the highest performance on the benchmarks FrontierMath and LiveCodeBench Pro (Hard).11 The FrontierMath question resolved at the end of 2025, with the median forecaster in both groups underestimating progress. The LiveCodeBench Pro question resolves at the end of 2026, but the current state-of-the-art is so far ahead of forecaster medians that it is already clear that they have massively underestimated AI performance on this benchmark.12
The median expert predicted that the leading AI model on FrontierMath Tiers 1–3 by the end of 2025 would achieve 31%, while the median superforecaster predicted 30%. In reality, this question resolved at 40.7%.13 On LiveCodeBench Pro (Hard), the median expert forecast state-of-the-art performance of 14% by the end of 2026, while superforecasters predicted 12%. In May 2026 the actual state-of-the-art on this benchmark was 53.8%.
Forecasters also significantly underestimated performance on a question we asked in the fall of 2025 about the performance of closed-weight models across three benchmarks (FrontierMath Tiers 1–3 and 4, SWE-Bench Verified, and ARC-AGI-2). Forecasts for the open-weight portion of this question were much more accurate than those for the closed-weight portion, with the median of both groups predicting 20% mean benchmark performance against a resolution of 22.7%.14
In April–May 2026, we asked LEAP panelists to forecast the longest METR 80% time horizon listed for an AI model by the end of 2026. Experts and superforecasters gave forecasts of 3.4 and 3.5 hours, respectively. On May 8, toward the end of the survey period, METR updated the benchmark to show that Anthropic’s Mythos Preview model had achieved an estimated 80% time horizon of 3 hours and 6 minutes (3.1 hours), already within the range of median panelists’ forecasts for the end of 2026. If the historical trend of progress on this benchmark continues, forecasters’ median predictions will have substantially underestimated progress by the end of 2026.

While we were drafting this post, the news broke that OpenAI had likely solved one of the Millennium Prize Problems—the Navier-Stokes Problem—using an internal model significantly more capable than GPT-6 Astra. LEAP experts, originally forecasting in August to September 2025, gave just a 10% chance of AI solving a Millennium Prize Problem by the end of 2027, while superforecasters gave even lower odds of 5.4%. Although it is not completely clear whether the current situation meets our resolution criteria, it appears likely that these forecasts are another significant example of forecasters underestimating AI progress.
Exceptions: more accurate forecasts
Accuracy on some unresolved benchmark questions looks more promising. For example, according to our LLM projections, near-term forecasts on Chinese model performance on the Epoch Capabilities Index (ECI) may be fairly accurate, although later resolution dates diverge more strongly from LLM forecasts.15

Overestimation of a non-benchmark capability: AI-biorisk uplift study
We have fewer examples of non-benchmark capability forecasts, but one notable overestimation of AI capabilities came when forecasters predicted how much mid-2025 LLMs could help non-experts with biorisk-relevant laboratory tasks.
In mid-2025, we asked for predictions about what fraction of STEM undergraduates would be able to complete three biorisk-relevant tasks with the help of an LLM in a randomized controlled trial run by Active Site. Biosecurity experts (at median) predicted that 22.5% of the sample would complete the tasks, virologists predicted 40%, and superforecasters predicted 16.2%. In actuality, 5.2% of the sample completed all of the tasks.

We discuss the potential implications of these findings about AI capabilities forecasting accuracy in the discussion section below.
AI adoption and diffusion
Adoption and diffusion forecasts attempt to track the uptake of AI through a range of proxies, including economic indicators, infrastructure rollout, and usage indicators. We have fewer resolved questions in this category and less observed data, so we are more reliant here on comparisons to LLM projections. As such, preliminary conclusions in this section are more speculative than our findings on benchmarks.
The forecasts below primarily come from LEAP, where we have collected the most adoption-related forecasts. In some highlighted cases we also have relevant data from the XPT.
Overall, the picture for forecasting accuracy on AI adoption and diffusion indicators is more mixed, though forecasters often lean toward underestimation:
- On economic proxies for adoption, such as revenue growth of AI companies, median experts and superforecasters majorly underestimated progress.
- On broad adoption indicators, such as what percentage of work hours in the US are assisted by AI, the evidence of forecaster accuracy is uncertain.
- Forecasts for resources relevant to AI growth that are proxies for adoption and diffusion (electricity used by AI and data center growth) were much closer to LLM projections than other adoption indicators. However, in our XPT study, forecasters substantially underestimated how quickly spending on the largest models would grow.
- It appears that experts overestimated the rate of autonomous vehicle adoption: when surveyed in June–August 2025, the median expert predicted that 7.3% of US ride-hailing trips would be completed by autonomous vehicles by the end of 2027. Based on current trends, LLMs predict that the actual resolution value will be 2.5%. Superforecasters appear more accurate, with a median forecast of 2%.
- We have collected forecasts for many adoption and diffusion indicators that we cannot yet judge the accuracy of, such as what fraction of people will use AI for companionship over time and what share of drug sales will come from AI-discovered drugs.
Economic proxies for adoption: major underestimation of AI revenue growth
Broadly, forecasts of economic proxies for adoption seem to have significantly underestimated AI uptake.
In our study “Forecasting the Economic Effects of AI,” we asked participants for the highest annual recurring revenue (ARR) for an AI-focused company at the end of 2026. AI experts gave a median ARR of $20 billion, economists predicted $16 billion, and superforecasters predicted $25 billion.16 It is likely that the current value for this question (as of September 22, 2026) is $100 billion, indicating that forecasters are on track to significantly underestimate adoption.
We posed a similar question to our LEAP panel in summer 2026, asking them to predict the combined annualized revenue run rate of Anthropic and OpenAI at the end of 2026. Experts and superforecasters gave median forecasts of $70 billion and $90 billion, both of which are substantially below the likely current value for this question, $140 billion.17 Our LLM projection (as of August 18, 2026) estimates that this question will resolve at $165 billion, although this does not take into account the latest reporting on Anthropic’s revenue.
Although revenue projections seem on track to be very inaccurate, the picture for other AI-relevant economic indicators is more mixed. In the XPT, the median forecaster accurately predicted the lowest price of 1 GFLOPS by the end of 2024, and was reasonably close to predicting US computer R&D spending at the same resolution date.18
Broad adoption indicators: uncertain
In September–October 2025 we asked LEAP panelists to forecast the share of US work hours that would be assisted by generative AI in 2025. Experts and superforecasters gave median predictions of 4% and 3.6%, respectively, and the question resolved at 5.7%. However, in the question information we gave to forecasters we cited a historical baseline of 2% from an early version of a study from the Federal Reserve Bank of St. Louis. An update to this paper revised this value to 3.35%, so while the median forecasters appear to have underestimated here, it may be that they were mistakenly anchoring on out-of-date baseline information.
Resources relevant to AI growth: a mixed track record
A question we asked in the XPT in 2022 about the cost of the largest AI experiment run by the end of 2030 is also very likely to result in a dramatic underestimate from forecasters: expert and superforecaster median forecasts of $180 million and $100 million, respectively, have already been overtaken by Grok 4, which cost an estimated $388 million to train (published in mid-2025). Forecasts for the cost of the largest AI experiment before the end of 2024, however, were more mixed, with the experts predicting $65 million and superforecasters $35 million against a resolution of $45.8 million.
Forecasts about the resources related to AI growth appear to be on track to be much more accurate. In LEAP we’ve asked questions about the share of US electricity going toward AI, and the buildout of hyperscale data centers globally at various resolution dates. For the closer resolution dates, the majority of forecaster medians appear to be reasonably close to LLM projections (see table below).19
Forecasts of US electricity share and global hyperscale data center capacity
| Resolution date | Expert median | Superforecaster median | LLM projection |
| Share of US electricity used to train and deploy AI systems | |||
| 2027 | 4% | 3% | 3.2% |
| 2030 | 7.4% | 5.7% | 6.5% |
| Total installed capacity of hyperscale data centers globally | |||
| 2026 | 50 GW | 48 GW | 52 GW |
| 2027 | 62 GW | 60 GW | 61 GW |
| 2030 | 92 GW | 86 GW | 85 GW |
Autonomous vehicles: experts overestimate adoption
We asked LEAP forecasters to predict the share of US ride-hailing trips that would be delivered by autonomous vehicles during 2027. The expert median for this question was 7.3%, which is on track to be a large overestimate compared to our LLM projection of 2.5%. Superforecasters appear to be notably more accurate on this question, giving a median prediction of 2%.
Future indicators
Many adoption forecasts from our studies have not yet resolved—for example, forecasts on drug discovery and diffusion across science. We’re also tracking forecasts on the share of US adults who self-report using AI for companionship daily by the end of 2027. The median expert and superforecaster both gave a forecast of 10%, which is above the LLM projection of 8% on this question.
Impacts
Our studies also capture a wide range of forecasts on AI’s impacts. These include impacts on employment and economic growth, but also AI’s effect on cybercrime, global catastrophic risk, and improving life expectancy. Not enough forecasts of this kind have resolved to draw any meaningful conclusions about accuracy, but we cover some preliminary data below.
Economic and labor impact
In the XPT we asked forecasters to predict the share of US GDP that would come from software and information services in 2024. They slightly overestimated this share, but their median forecasts were close to the resolution figure of 3.25%.20 In the same study, forecasters slightly underestimated the OECD labor force participation rate for 2024, but were again pretty close to the true resolution value.21
In LEAP, we’ve asked for forecasts on a range of economic and labor outcomes, including those in the table below. It is too early to assess the accuracy of these forecasts.
Forecasts of economic outcomes in 2030 (unconditional)
| Question | Expert median | Superforecaster median | LLM projection |
| Labor force participation rate22 | 62% | 62% | 61.3% |
| GDP growth23 | 2.7% | 2.5% | 2% |
| Economic inequality24 | 73% | 72% | 70.4% |
We also asked for many more macroeconomic forecasts in our study on the economic effects of AI.
Other impact measures
In the XPT, we asked forecasters to predict the share of Americans who would say they had a negative opinion of AI by the end of 2024. Here, the median expert and superforecaster predictions of 33% were extremely close to the resolution of 32.9%.
In our study on AI cyber risk we asked for the year in which an AI-enabled attacker causes a grid blackout causing $100 million or more in damages. Experts and superforecasters gave median predictions of 2031 and 2035, respectively, though this question is yet to resolve. In LEAP we asked for predictions of the annual losses from cybercrime reported to the FBI in 2026, 2028, and 2031, so we should have some resolution data on this question relatively soon.
In many of our studies we’ve elicited forecasts on catastrophic risk, AI-enabled pandemic risk, large-scale harm events, and other forms of major risk. We have also elicited forecasts on major potential benefits from AI such as improvements in life expectancy. These are some of the most important forecasts across our work, but it’s too early to say if high or low risk and benefit estimates are more accurate. For more information, see our LEAP waves on risks from AI and benefits from AI.
Geopolitics and governance
Although we only have a small number of questions in this category, it’s worth highlighting one recent question that appears to have surprised forecasters.
In LEAP Wave 9 (May–June 2026), we asked forecasters to predict the calendar year by which either the US, UK, or EU governments would issue a binding directive, regulation, court order, or injunction to delay, seriously restrict, or condition the public release of an AI system, based on safety risks. The median expert and superforecaster gave a 50% chance this would happen by 2030.
In fact, this question may be very close to resolution already. In late June, OpenAI restricted availability of its unreleased model GPT-5.6 Sol to a small group of partners approved by the US government. It is not clear that this meets our resolution criteria, as this does not appear to be the result of a binding directive, but this and the Claude Fable post-release restrictions suggest that resolution on this question is potentially much closer than forecasters anticipated.
Other unresolved geopolitics and governance questions include:
- The share of frontier AI chips produced by the US, China, Taiwan, and other regions;
- The year in which NATO or a member state will publicly authorize the use of fully autonomous offensive cyber operations or cyber weapons;
- Whether a top-four frontier lab will exist within a non-democracy by 2030.
Limitations
This post is intended as a brief overview of where median AI progress forecasts appear to be inaccurate across our studies, but there are several major limitations to our approach and caveats worth bearing in mind, including:
- Our methodology is biased toward identifying underestimates of AI progress as it is easier to identify forecasts where the current value has already overshot the median forecasters’ predictions. If an AI capability predicted to occur in 2030 is actually achieved in 2035, we will need to wait many years to report on this overestimation.25 The vast majority of our AI progress forecasts are unresolved and it may be the case that over time the picture of forecaster accuracy changes significantly.
- In this post we have concentrated almost exclusively on median forecasts, which may obscure more accurate predictions from subsets of forecasters, as well as information from forecasts covering other percentiles.
- In some cases, issues with our question specification or our elicitation format may have led to inaccurate forecasts. For example, for LiveCodeBench Pro (Hard) forecasts, it may be that some forecasters did not understand that the benchmark allows models to go back and test on questions from earlier quarters.26
Discussion
Overall, AI capabilities and some important measures of adoption (such as AI company revenue) are moving much faster than experts and superforecasters expected.
One potential implication of this is that forecasters should expect much faster AI progress, faster AI adoption, and much larger AI impacts going forward.
We have seen some evidence that forecasters in LEAP are updating in this direction. In summer 2025 and spring 2026, we asked panelists to forecast the likelihood of AI in 2040 reaching various levels of impact on the Technological Richter Scale (TRS, a scale devised by Nate Silver that ranks technologies by their broad societal impact). Over the course of nine months, experts and superforecasters have increased the likelihood they assign to AI’s impact being comparable to “technology of the millennium” (e.g., agriculture, Level 9), while continuing to assign the greatest probability to AI being comparable to “technology of the century” (e.g., electricity, Level 8):27

If forecasters are updating their views toward faster AI progress, we might expect them to perform more accurately on benchmark questions asked more recently. However, we have not yet seen evidence of improved accuracy on more recent questions. Forecasters likely majorly underestimated progress on the METR time horizon benchmark (asked in April 2026) and underestimated OpenAI and Anthropic revenue growth in our July 2026 survey. We will try to understand what is driving forecasters’ continued accuracy errors and will take various steps to highlight more accurate and relevant forecasts, which we’ll expand on in the next post in this series.
An alternative response to the evidence in this post is to say that benchmarks have moved much faster than forecasters expected in part because it was easier to “teach to the test” than forecasters expected and therefore that benchmarks are less meaningful for making sense of real-world AI capabilities than forecasters might have expected. Perhaps significant benchmark progress just doesn’t lead to real-world impacts in many areas. In which case, we might expect forecasters to underestimate AI progress on benchmarks, but turn out to be more accurate when it comes to forecasting real-world impacts such as macroeconomic impacts.
This debate highlights the importance of getting accuracy data on additional adoption and impact measures, which we expect to have soon as more questions resolve.
Next steps
In this post, we have covered the accuracy of median forecasts across our studies, often highlighting places where these median forecasts have been inaccurate. However, subsets of forecasters in our studies have been much more accurate than the median forecaster.
In the next post in this series, we’ll look at a few ways we can improve the accuracy and relevance of LEAP forecasts and highlight predictions from more accurate forecasters. These involve three core updates to LEAP:
- Highlighting forecasts from a subsample of panelists who expect very fast AI progress by 2040.
- Adding continuously updated LLM forecasts to LEAP findings.
- Identifying our most accurate LEAP participants and highlighting their forecasts as soon as we have sufficient data to make this judgment.
Appendix
For more details on relevant forecasts, see our accompanying tables:
Notes
- We elicited LLM projections from GPT-5.5 with “high” reasoning. LLM forecasts have steadily improved in recent years, and ForecastBench indicates that some models now perform at superforecaster level on certain question types. GPT-5.5 with “high” reasoning currently leads the ForecastBench Tournament Leaderboard among models tested “out of the box” by FRI—that is, models without forecasting-specific scaffolding. Importantly, the projections we use in this post were elicited months after humans forecasted the same questions. Hence, these LLM projections can best be thought of as capable forecasters with a substantial comparative advantage from having seen more real-world development over recent months, rather than simply a comparison of LLM and human forecasts. ↩︎
- Tom Cunningham summarized the broad range of disagreement about AI’s economic impacts in a post first published in November 2025. ↩︎
- We also considered forecasts from our studies on “AI Conditional Trees” and AI R&D, but few questions have resolved, so we do not cover any results from these studies in this post. ↩︎
- For example, in our study on AI-biorisk our expert sample included biology and biosecurity experts. The LEAP sample of experts includes top computer scientists, economists, policy experts, and AI industry experts. For details on all relevant expert samples, see full reports from the relevant studies, linked above. The “superforecaster” sample also varies but contains many of the same individuals across studies. ↩︎
- See here for a chart showing the distribution of forecasts for this question, showing that superforecasters gave an implied probability of 2.3% and domain experts 8.6% for the actual outcome. ↩︎
- In spring 2023 we also asked benchmark questions as part of our study “Roots of Disagreement on AI Risk.” It is likely that the median forecaster underestimated the likelihood of AI matching top human forecasters and whether AI would be able to construct 10,000 lines of bug-free code by 2030. Our own projections suggest that the forecasting question is likely to resolve yes, and the Metaculus community prediction tracking the coding question is at 98% “yes.” See the question table for more detail. ↩︎
- We also asked forecasters for their forecast likelihood of specific evaluations being achieved if they were performed in 2026; see here for those results. ↩︎
- This question referred to a benchmark run by OpenAI and cited in the o1 system card, where it achieved 75%. The o3 model (January 2025) outperformed o1 substantially, but a precisely comparable score was not reported. Nevertheless, we assessed in July 2025 that it was likely that our resolution had been met by existing models. ↩︎
- The 90% threshold was first crossed on the public Cybench leaderboard by Claude Opus 4.6, which was added at 93% unguided solved on a 37-problem subset on February 6, 2026. Later, Claude Mythos Preview reached 100% on a 35-problem subset. ↩︎
- Although these questions are unresolved, we should not assume that AI does not have these abilities—simply that these specific abilities have not been assessed in a way that meets our resolution criteria. For example, the synthetic DNA scenario is based on a hypothetical paper. The bioweapon attack planning scenario was based on a RAND study that, to our knowledge, has not yet been repeated. ↩︎
- We elicited these forecasts between June–August and November–December 2025, respectively. ↩︎
- There are issues with both of these questions (see relevant footnotes in Appendix Table 1 for more details) that may have impacted forecaster accuracy, but we still view these results as evidence of forecasters underestimating benchmark progress. ↩︎
- This was the resolution value on Jan 1, 2026, before Epoch announced that it had found errors in about 42% of FrontierMath problems and subsequently adjusted scores to account for these errors. Since our forecasters made their forecasts before these errors were known, we report the unadjusted score of 40.7%. Given the complications with this benchmark, the best resolution value for this forecasting question is debatable, but regardless of adjustments it is clear that forecasters underestimated progress on the benchmark. ↩︎
- See the question table for more detail. ↩︎
- Since we first drafted this post GPT-6 Astra achieved an ECI score of 167. This is close to the expert and superforecaster median forecasts of 165 and 169 for the end of 2026. ↩︎
- Our question resolution criteria defined ARR as including revenue from API usage as well as subscription revenue. ↩︎
- As of September 22, 2026, OpenAI’s reported annualized revenue was $40 billion while Anthropic had a reported revenue of $100 billion. A number of respondents may have erred in forecasting this question such that their responses may not reflect their best judgment. Median forecasts for both groups were below the latest published value available at the time of forecasting. While one might hold the view that either or both companies’ revenue will collapse before year’s end, rationales showed no evidence that a large number of these forecasters intended to predict a reversal of recent trends. Instead, we suspect our elicitation design may have made it easy for forecasters to ignore this most recent data point and anchor on earlier data. While superforecasters used the same interface, their predictions and rationales provide less indication that they missed the latest published value. ↩︎
- See Table 1 for more detail. ↩︎
- One might worry that, since LLM projections were elicited after the relevant LEAP waves were published, LLMs simply anchored on those projections. However, we explicitly prompted against this, and found no mention of “LEAP” in GPT’s rationale for these projected values. ↩︎
- Experts and superforecasters gave median predictions of 3.6% and 3.45%, respectively. ↩︎
- See Table 1 for more detail. ↩︎
- “What will be the labor force participation rate in the US at the beginning of 2030?” ↩︎
- “What will be the annualized change, in percent, in the real Gross Domestic Product (GDP) of the US between 2025 and 2030?” ↩︎
- “What will be the fraction of the national wealth owned by the top 10% wealthiest individuals in the US at the beginning of 2030?” ↩︎
- Many capability questions have not resolved yet. For example, one question from the LLM-biorisk study that has not yet resolved is when AI will enable 10% of non-experts to synthesize flu virus. ↩︎
- See our question-by-question table for notes on specific questions. ↩︎
- See additional details here and in the relevant linked report. ↩︎



