Published: Aug 12, 2026
Technical report
  • Technical report

Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks

Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks
A study to better understand the effectiveness of Anthropic's AI Safety Level 3 measures at reducing biosecurity risks.
Bridget Williams1,2, Rebecca Ceppas de Castro1, Dan Mayland1, Ria Viswanathan1, Zack Devlin-Foltz1, Jordan Canedy1, Victoria Schmidt1, Philip E. Tetlock1,3, Ezra Karger1,4, Josh Rosenberg1
,
1 Forecasting Research Institute
2 University of Oxford
3 University of Pennsylvania
4 Federal Reserve Bank of Chicago
Published: Aug 12, 2026
Bridget Williams1,2, Rebecca Ceppas de Castro1, Dan Mayland1, Ria Viswanathan1, Zack Devlin-Foltz1, Jordan Canedy1, Victoria Schmidt1, Philip E. Tetlock1,3, Ezra Karger1,4, Josh Rosenberg1
Acknowledgments

This research would not be possible without the thoughtful participation of our survey respondents. We are grateful to Mrinank Sharma and Jerry Wei for helpful feedback on survey design. We are also grateful to Anna Bartlett, Coralie Consigny, Johan S Daniel, and Josh Connor for research assistance, and to Otto Kuusela and Simas Kučinskas for helpful feedback on research design.

Disclaimers

This report was originally produced as part of an internal project funded by and completed for Anthropic PBC. Anthropic provided input and advice in the development of the report, but all findings and claims represent the views of the authors. Due to the general interest of the findings, we are releasing a version of the report with confidential information removed.

The views expressed in this paper do not necessarily reflect the views of the Federal Reserve Bank of Chicago or the Federal Reserve System.

Summary

In May 2025, Anthropic announced that it had activated AI Safety Level 3 (ASL-3) Deployment and Security Standards as a precautionary measure, in response to Claude Opus 4’s improved CBRN-related capabilities [1]. The deployment standards include the implementation of real-time classifier guards, offline monitoring, access controls, a bug bounty program, threat intelligence, and rapid response.

The Forecasting Research Institute conducted a study to better understand the effectiveness of these measures—particularly the ASL-3 deployment standards—at mitigating biosecurity risks. We surveyed 22 people with expertise in biosecurity, national security, and/or terrorism studies (“experts”), and 20 top generalist forecasters (“superforecasters”), all of whom completed the survey between March 4 and March 29, 2026.

We asked respondents to forecast the total expected financial damages from large-scale human-caused outbreaks (outbreaks leading to at least $100 million in total worldwide damages) occurring between April 2026 and the end of 2028, and how this value would change across four scenarios: a world without general-purpose AI (GPAI); a world where no GPAI models had any safeguards; a world where all GPAI models were protected by ASL-3; and a world where frontier GPAI models were protected by ASL-3.

Because we did not include a status-quo scenario—one in which many frontier models already carry their developers’ own misuse protections—our estimates do not capture the marginal benefit of extending ASL-3 beyond what other developers’ existing safeguards already provide. We also asked respondents to assume that the capabilities of GPAI models do not improve before December 31, 2028. We made these stipulations because we wanted respondents’ views on ASL-3’s ability to mitigate harm from current models, rather than their views on other companies’ safeguards or likely AI progress.

Key Takeaways

  • Experts and superforecasters generally believe that implementation of ASL-3 safeguards across all general-purpose AI (GPAI) models can substantially reduce the risk of large-scale, human-caused outbreaks, relative to a world with no safeguards. Under the ‘no safeguards’ scenario, the median forecasted probability of human-caused outbreaks, occurring between April 2026 and December 2028, causing at least $100 million in damages was 16% (95% CI: 9.2%, 32.8% ). This dropped to 9.2% (95% CI: 4% , 16.9%) under the ‘ASL-3 for all models’ scenario.
  • Forecasters estimate that ASL-3 safeguards, if implemented across all GPAI models, would capture most of the risk reduction that is possible to achieve via safeguarding. For the median respondent, implementing ASL-3 across all GPAI models achieved 71.7% (95% CI: 65.7%, 97.2%) of the risk reduction associated with a hypothetical scenario where no GPAI exists.
  • Most respondents believe protecting frontier models (those at least as capable as Sonnet 4.5) with ASL-3 would substantially reduce risk relative to a ‘no safeguards’ baseline, although not by as much as extending it to all GPAI, regardless of model size or capability. For the median respondent, ‘ASL-3 for frontier models’ achieved 36.3% (95% CI: 22.9%, 68.4%) of the risk reduction associated with the ‘no GPAI’ scenario.
  • The median respondent thought that variations to ASL-3—making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models—would reduce risk by an amount similar to applying ASL-3 to all GPAI models. The largest median risk reductions were associated with patching jailbreaks 5x faster and with a further 30% drop in jailbroken models’ performance relative to non-jailbroken models. For the median respondent, each of these two variants achieved 73.5% of the risk reduction associated with a hypothetical scenario where no GPAI exists (95% CI: 52.3%, 100% for faster patching; 47.8%, 97.3% for reduced jailbroken performance). Many respondents noted that such variations may matter more in a world where ASL-3 covers all GPAI models, and many saw trusted user exemptions as a potential weak spot in ASL-3.

Key Results

ASL-3 safeguards are seen as having high efficacy: if applied across all GPAI models, respondents estimate that ASL-3 would capture most of the risk reduction achievable through safeguarding.

  • When respondents assumed all models had ASL-3 protections, the median forecasted probability of human-caused outbreaks causing at least $100 million in damages dropped by roughly 40% relative to the ‘no safeguards’ scenario (from 16% (95% CI: 9.2%, 32.8%) to 9.2% (95% CI: 4%, 16.9%)). For comparison, when asked to consider a hypothetical world where GPAI was never developed, the median forecast was 9.3% (95% CI: 3.8%, 13.2%). (See Figure 1.)
  • Regarding expected financial damages, the result was similar, although the ‘no GPAI’ scenario was generally associated with lower damage estimates than the ASL-3 scenarios. Compared to a world with no safeguards, a world where all models have ASL-3 safeguards reduced median predicted financial damages of human-caused outbreaks between now and 2028 by 40% (95% CI: 17%, 63%). Under the ‘no GPAI’ scenario, the median participant reduced their forecast of expected damages by 49% (95% CI: 19%, 71%).
  • Overall, the forecasts suggest that respondents see ‘ASL-3 for all models’ as reducing risks of a large-scale human-caused outbreak to a risk level slightly above a hypothetical world with no GPAI. The median respondent’s forecasts suggest that ‘ASL-3 for all models’ is associated with a reduction in risk that is 71.7% (95% CI: 65.7%, 97.2%) of the risk reduction associated with the ‘no GPAI’ scenario.
  • Rationales showed that respondents generally believed that ASL-3 would effectively block non-experts. Many respondents were reassured by the pace of jailbreak patching and the difficulty in finding jailbreaks, suggesting that this likely requires substantial cybersecurity expertise. Several also noted the deterrence effects of safeguards.
Figure 1: Probability of human-caused outbreaks that start between April 2026 and December 2028, causing at least $100 million in economic damages. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals.

Real-world impact likely depends on breadth of adoption: protecting only frontier models still reduces risk substantially, but less than comprehensive coverage.

  • We also asked how respondents’ forecasts would change if ASL-3 were only applied to ‘frontier’ GPAI models, which we defined as those at least as capable as Sonnet 4.5. Under this scenario, the median respondent forecast for the probability of human-caused outbreaks starting between now and 2028, causing at least $100 million in damages, was 10.6% (95% CI: 5.5%, 18.1%). The median estimate of expected damages was 20% (95% CI: 9%, 52%) lower under this scenario than under the ‘no safeguards’ scenario.
  • For the median participant, the expected damages in the ‘ASL-3 for all models’ scenario was 4% (95% CI: 1%, 16%) lower than the expected damages in the ‘ASL-3 for frontier models’ scenario.
  • Most respondents placed the frontier-only ASL-3 scenario’s risk between ‘ASL-3 for all models’ and ‘no safeguards’, but disagreed on where, with some treating it as nearly equivalent to ‘ASL-3 for all models’ and others as only marginally better than ‘no safeguards’. Much of this disagreement hinged on the capability and accessibility of near-frontier models, particularly Chinese open-weight models.

Making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models were generally expected to improve ASL-3 performance, with faster patching expected to be the most impactful.

  • We asked respondents to consider six variants of ASL-3 safeguards that differed from the current safeguards’ performance in one of three ways: time required to find a jailbreak, time taken to patch a jailbreak, and performance of jailbroken models relative to the base model. For the median respondent, each of the variations associated with improvement—making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models—was expected to reduce risk to a greater extent than the ‘ASL-3 for frontier models’ scenario.
  • The largest median deviation from the ‘ASL-3 for frontier models’ expected damages baseline was an 8% decrease (95% CI: 5%, 18%), associated with a scenario of a 5x decrease in the time taken to patch a jailbreak. The median respondent’s forecasts suggest that ASL-3 for frontier models with 5x faster patching is associated with a reduction in risk that is 73.5% (95% CI: 52.3%, 100%) of the risk reduction associated with the ‘no GPAI’ scenario.

ASL-3 shifts the threat landscape but does not address all major risk pathways.

  • Respondents expected ASL-3 to disproportionately screen out non-expert and lone-wolf actors, shifting the relative probability of damages toward state actors and accidental releases from laboratories.
  • Superforecasters saw trusted user exemptions as a notable vulnerability, with some noting that they combine AI access with proximity to the physical infrastructure needed for an attack and exploit human factors that technical safeguards cannot fully address. Many respondents emphasized that lab accidents represent a substantial share of baseline human-caused outbreak risk that AI safeguards do not address.

1. Introduction

In May 2025, Anthropic announced that it had activated AI Safety Level 3 (ASL-3) Deployment and Security Standards as described in Version 2 of Anthropic’s Responsible Scaling Policy (RSP) [1]. This activation came in conjunction with the launch of Claude Opus 4, and was described as a precautionary measure. Anthropic determined that due to continued improvements in CBRN-related knowledge and capabilities, it was not possible to rule out that Claude Opus 4 had crossed its CBRN-3 capability threshold. This threshold is defined in the RSP as: “The ability to significantly help individuals or groups with basic technical backgrounds (e.g., undergraduate STEM degrees) create/obtain and deploy CBRN weapons” [2].

Anthropic’s announcement noted that ASL-3 defenses would be focused on biological weapons risks. The possibility that AI could enable more actors to develop biological weapons has been raised by heads of AI companies and experts in biosecurity [3, 4, 5]. There is debate about the significance of the risks, with some believing that AI does little to overcome the key barriers to biosecurity risks [6, 7]. Prior work has found that most experts and top generalist forecasters expect biological risks to increase with AI advances in AI capabilities [8].

The ASL-3 Deployment Standard consists of requirements across the following areas: threat modeling, defense in depth, red-teaming, rapid remediation, monitoring, access controls for trusted users, and third-party environments. To meet these requirements, Anthropic implemented the following mitigations: real-time classifier guards, offline monitoring, access controls, a bug bounty program, threat intelligence, and rapid response. Further detail on Anthropic’s implementation of ASL-3 Deployment and Security Standards is available in Appendix A.

Although Anthropic can evaluate aspects of its mitigations (for example, the time taken to patch an identified jailbreak), it is unclear how well these measures achieve the overarching goal of reducing biosecurity risks posed by general-purpose AI (GPAI) models. To investigate this, we conducted a forecasting survey of experts in biosecurity, national security, or terrorism studies, as well as top generalist forecasters. Respondents were given information on Anthropic’s ASL-3 standards and the measures taken to meet them. This included confidential data from red-teaming efforts, the bug bounty program, and other safeguards testing conducted by Anthropic. Although respondents’ forecasts included consideration of both the Deployment and Security Standards that make up ASL-3, the focus of this survey was the Deployment Standards.

This survey aimed to answer four research questions:

  • How effective are ASL-3 safeguards?
  • What would be the impact of more models being protected by ASL-3?
  • How would variations in ASL-3 influence this impact?
  • How do ASL-3 safeguards influence threat models related to biosecurity risks?

In this report, we briefly describe the survey methods before presenting the results of the survey that speak to each of these research questions.

2. Methods and participants

2.1 Survey contents

Description of the main outcome

The main outcome we asked participants to forecast was:

The total worldwide damages attributable to human-caused outbreaks that start between April 1, 2026, and December 31, 2028, and that each individually cause at least $100 million in total worldwide damages (in 2026 USD).

We asked respondents to make a few important assumptions when answering this question. First, we asked them to assume that until December 31, 2028, the capabilities of general-purpose AI models do not improve, as if a ‘pause on general-purpose AI development’ were in place. That is, for that period, no models outperform Claude 4.6, Gemini 3, or GPT-5, and this slowdown is not indicative of a more general slowdown in technological progress. We made these stipulations because we were interested in their views on ASL-3’s ability to mitigate harm from current models, rather than their views on likely AI progress.

We asked respondents to forecast the global economic costs attributable to the outbreaks, beginning with the isolation of the pathogen and extending for the full life of an outbreak, and including direct economic losses and monetized mortality. We asked participants to use a value of statistical life of $2.2 million, regardless of geographic location. For full details of this question, including the resolution criteria, please see Appendix B.

Description of scenarios

  • No GPAI: In this scenario, respondents were asked to assume that GPAI models never existed and will not exist. They were asked to assume that all other features of the world remain the same, e.g., they should not assume that the absence of GPAI models indicates that technological progress is slower across all domains.
  • No safeguards: In this scenario, respondents were asked to assume that no safeguards are applied to all GPAI models. In other words, to assume that any system safeguards, refusal behaviors, and other safety-focused constraints have been removed or disabled, and that such AI models are optimized for helpfulness (efficient and accurate task performance) without safety protections.
  • ASL-3 safeguards for all GPAI models: In this scenario, respondents were asked to assume that all GPAI models—regardless of size or capability—are protected by Anthropic’s ASL-3 safeguards.
  • ASL-3 safeguards for frontier GPAI models: In this scenario, respondents were asked to assume that “frontier” GPAI models are protected by Anthropic’s ASL-3 safeguards. “Frontier” GPAI models were defined as all those performing as well or better than Claude Sonnet 4.5 on the Epoch Capabilities Index (ECI) [9]. At the time of the survey, Claude Sonnet 4.5 scored 147 on the ECI, and 13 models scored at or above that level: Gemini 3 Pro, GPT-5.2, Claude Opus 4.6, Gemini 3 Flash, GPT-5 Pro, Claude Opus 4.5, GPT-5, GPT-5.1, o3-pro, Kimi K2.5, Grok 4, o3, and Claude Sonnet 4.5.

Additional questions

In addition to forecasts of the main outcome, we asked questions to understand respondents’ views on pathways to human-caused outbreaks, and how these are influenced by ASL-3. Specifically, we asked respondents to assume that a human-caused outbreak that caused more than $100 million in damages had occurred, and to say what probability they placed on different types of actors being the primary cause of the outbreak. We asked participants to answer these questions for two scenarios: ‘no safeguards’ and ‘ASL-3 for all GPAI models’. For the ‘ASL-3 for all GPAI models’ scenario, we also asked about the probability that a GPAI model had been used in an unauthorized way to assist the process of causing the outbreak, and the most likely pathways to unauthorized use. We also asked participants for some basic details on their demographics and expertise.

Calibration modules

Before completing any forecasting questions, participants completed two short calibration modules. One module asked participants to estimate low probabilities. An example question is, “What is the probability of being struck by lightning in any year?” (Answer: 8.1e-7). The other module asked for values relevant to the survey subject matter. An example question is, “In the 1984 Rajneeshee bioterror attack in Oregon (the largest bioterrorism attack in US history), how many people were infected with Salmonella?” (Answer: 751). The primary purpose of these modules was to give participants an opportunity to test their own calibration.

2.2 Data Analysis

Data analysis was conducted using R after aggregating and cleaning the data submitted by participants. For details of data cleaning, see Appendix B.

For the main outcome, we asked participants to forecast the probability that human-caused outbreaks would lead to different levels of financial damage by 2028. We then calculated expected financial damages for each scenario. To calculate these values, we assumed that damages within each bin were all equally likely (uniformly distributed). We then multiplied the probability assigned to each range by the midpoint dollar amount of that range and summed these probability-weighted contributions across all bins to obtain the overall value of expected damage. For the highest bin (more than $100 trillion), we used $1,000 trillion as the upper bound for calculations.

This approach could be a limitation, especially for wider damage bins or upper-tail outcomes where damages may not be evenly distributed within a range. To reduce the impact of this issue, we used a relatively large number of bins and had participants review the calculated expected damages. Participants were asked to adjust their probability forecasts if the expected damages calculation did not align with their expectations, so these values reflect calculations that participants reviewed and had the opportunity to revise.

To describe forecasts on the main outcome, we report several metrics: the probability that damages from human-caused outbreaks, occurring from now until the end of 2028, will be at least $100 million; the ratio of calculated expected damages values between different scenarios; and a metric we define as “share of GPAI-attributable risk mitigated”.

We define “GPAI-attributable risk” in relative terms as the proportional increase in expected damages between the ‘no safeguards’ scenario and the ‘no GPAI’ scenario. Formally, we define this as

GPAI-attributable risk=Eno sgEno AI1GPAI\text{-attributable risk} = \frac{E_{\text{no sg}}}{E_{\text{no AI}}} – 1

where (Eno AI) denotes expected damages with no GPAI and (Eno sg) denotes expected damages with no safeguards.

This represents our estimate of the additional risk posed by the existence of unsafeguarded GPAI models. We then measure the share of this GPAI-attributable risk that is mitigated by ASL-3 safeguards by comparing expected damages under the ‘no safeguards’ scenario to those under the relevant ASL-3 scenario (Es). Specifically, we compute

Proportion of risk removed=Eno sgEs1Eno sgEno AI1\text{Proportion of risk removed} = \frac{\frac{E_{\text{no sg}}}{E_s} – 1}{\frac{E_{\text{no sg}}}{E_{\text{no AI}}} – 1}

This quantity captures the fraction of the proportional increase in expected damages attributable to unsafeguarded GPAI that is eliminated under ASL-3 safeguards. Although we refer to this quantity as a “share” of risk mitigated, it is more precisely a normalized relative-risk reduction measure based on multiplicative differences in expected damages, rather than an additive share of the total risk.

A caveat applies to this definition: the ‘no GPAI’ scenario removes not only the misuse risks of GPAI but also any defensive benefits (e.g., contributions to pandemic preparedness and response). A ‘no GPAI’ world therefore has higher baseline outbreak risk than a hypothetical world with perfectly safeguarded AI would, making it a more lenient benchmark. The share mitigated by ASL-3, measured against this benchmark, is best interpreted as an upper bound.

A further caveat applies to a small subset of four participants. These participants were excluded from aggregate calculations and figures of this metric because they assessed expected damages to be lower in the presence of AI than in the ‘no GPAI’ scenario. In these cases, the calculations of the share of GPAI-attributable risk reduced do not yield interpretable measures since their baseline implies that AI itself reduces risk.

We focus on the relative change in risk because, given the difficulty of forecasting expected damages, we believe these relative changes are likely to be more reliable than the absolute values of damages. However, we present the details of absolute values in Appendix C.

2.3 Participants

The results represent responses from a total of 42 participants: 22 participants were recruited for their expertise in biosecurity, national security, or terrorism studies (hereafter, “experts”), and 20 participants were recruited as top generalist forecasters (hereafter, “superforecasters”). Participants completed the survey between March 4 and March 29, 2026. On average, participants spent 6 hours completing the survey.

The median expert participant reported 11 years of experience relevant to the survey. The expert sample was recruited across three domains — biosecurity, national security, and terrorism studies — but most experts had experience spanning multiple domains. Although it was not a requirement for their participation in the study, seven of the superforecaster participants reported experience in at least one of the biosecurity, national security, or terrorism studies. See Appendix B for details of recruitment.

More than half (55%) of the expert participants reported having been granted access to restricted national security information (now or in the past). Only 20% of superforecasters reported the same. Most participants reported using LLMs regularly (at least once a week). More details on the participants are available in Appendix C.

3. How effective are ASL-3 safeguards?

To understand participants’ views on the effectiveness of ASL-3 at reducing biosecurity misuse risks of GPAI models, we can first review the difference in the probability of human-caused outbreaks occurring between now and the end of 2028 and causing at least $100 million in damages under the different scenarios. Figure 1 in the Summary shows the probability of this outcome under the four scenarios. Figure 2 shows the relative reduction in the probability of this outcome under the ‘no GPAI’ and ASL-3 scenarios, relative to ‘no safeguards’. For the median respondent, the ‘no GPAI’ scenario is associated with the largest reduction in risk [39% (95% CI: 16%, 50%)], but this is only slightly greater than the risk reduction associated with the ‘ASL-3 for all models’ scenario [32% (95% CI: 15%, 41%)].

Figure 3 shows how the calculated expected damages under the two ASL-3 scenarios and the ‘no GPAI’ scenario compare to the ‘no safeguards’ scenario. Although there was substantial variation, most believed that ASL-3 protections for all GPAI models would significantly reduce the expected damages from large-scale human-caused outbreaks, relative to a world where GPAI had no safeguards. This was particularly true for the expert respondents. The median respondent’s forecasts suggested that ASL-3 protecting all models would reduce expected damages by 40% (95% CI: 17%, 63%).

Figure 2: Ratio of forecasts of probability of human-caused outbreaks that start between April 2026 and December 2028, causing at least $100 million in economic damages, relative to ‘no safeguards’ scenario. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals. Values below 1 indicate lower P(≥$100M) than the no-safeguards scenario.
Figure 3: Ratio of expected damages for each protection scenario compared to the ‘no safeguards’ baseline. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals. Values below 1 indicate lower expected damages than the no-safeguards scenario.

GPAI-attributable risk

The scenario that asked respondents to assume GPAI models did not exist can provide an upper bound on the risk reduction that could be achieved by safeguards against GPAI model misuse, as it demonstrates the risk of human-caused outbreaks that would persist regardless of GPAI models. We can also express these results in terms of the share of GPAI-attributable risk mitigated by ASL-3. Aggregates presented here exclude four participants whose responses imply that GPAI reduces risk, making this metric uninterpretable for them. See Section 2.2 for a definition of the metric and more detail on the filtering. The median participant’s forecasts imply that ASL-3 for all models would mitigate 71.7% (95% CI: 65.7%, 97.2%) of GPAI-attributable risk. Figure 4 shows this comparison.

Figure 4: Share of GPAI-attributable risk mitigated by each ASL-3 scenario. GPAI-attributable risk is defined as the multiplicative gap in expected damages between the ‘no safeguards’ and ‘no GPAI’ scenarios. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals.

An important limitation of this comparison deserves emphasis. The ‘no GPAI’ scenario serves as a proxy for perfect safeguards, but it is an imperfect one. Perfect safeguards would eliminate misuse risk while preserving the defensive benefits of GPAI: its contributions to pandemic surveillance, pathogen characterization, drug and vaccine development, and outbreak response. The ‘no GPAI’ scenario removes all of these. This means the ‘no GPAI’ world is likely more vulnerable to outbreaks (including human-caused ones) than a world with perfectly safeguarded GPAI, making it a lenient benchmark.

The practical consequence is that the 71.7% efficacy figure—the share of maximum risk reduction achieved by ASL-3 for all models—is likely biased upward. If GPAI’s defensive contributions are small relative to the misuse risk, the bias is minor. If they are large, the bias could be substantial. We did not directly elicit respondents’ views on the magnitude of GPAI’s defensive benefits, though several rationales mentioned them. For instance, some respondents noted that GPAI could accelerate the development of medical countermeasures or improve biosurveillance. Future work could address this limitation by eliciting forecasts under a ‘perfect safeguards’ scenario directly—one where GPAI exists and provides defensive benefits but cannot be misused—or by asking respondents to separately estimate the defensive and offensive contributions of GPAI to outbreak risk.

Insights from rationales1

Aligning with quantitative results, text rationales showed that most respondents believed comprehensive ASL-3 protection would meaningfully reduce risk relative to the helpful-only scenario. There was broad agreement that ASL-3’s greatest value is in blocking non-expert and lone-wolf actors. Many respondents were reassured by the difficulty of finding jailbreaks and the pace of jailbreak patching. Several noted the deterrence effect of safeguards. On the other hand, multiple respondents noted that ASL-3 does nothing to reduce accidental releases, and one argued that ASL-3 would increase risk by weakening collective defensive capabilities.

“With safeguards in place, the risk of a human-caused outbreak should be around the same as with no GPAI models around at all. In particular, the safeguards should significantly reduce the likelihood that an untrained non-expert would use an AI system to create a dangerous pathogen.”

“Given the available information, it doesn’t seem like the ASL-3 security standard is a panacea. The patching of jailbreaks is slower than I thought. … It seems rather unbalanced that only $26k are awarded to an individual hacker for discovering a bug that, not only could cost the company millions (plus the potential negative publicity), but also requires so much effort to fix. There’s probably a black market in the dark web for these universal jailbreaks, and I wouldn’t be surprised if the pay is competitive.”

Most respondents viewed the helpful-only (‘no safeguards’) scenario as increasing biorisk, but physical bottlenecks—not information access—were seen as the primary constraint on bioweapons development. Many cited the Active Site study [10] — which found that helpful-only LLM access didn’t aid participants much with practical virology tasks relative to internet-only access — and stressed that access to labs and equipment, along with crucial hands-on skills, would remain the real constraints. Others, however, weren’t so sanguine, arguing that current frontier models can produce detailed technical guidance and that stripping away refusal behaviors makes that knowledge accessible to a much wider set of actors. Several challenged the relevance of the Active Site study, noting it used older models, and identified less direct pathways to increased risk: more people being inspired to try, lab accidents driven by AI-enabled overconfidence, and the volatile mix of helpful-only models with ongoing geopolitical conflicts.

“Active Site’s recent work on this topic showed that overall LLM-aided participants performed fairly similarly to internet-only in a pseudo-virological workflow. This also fits my strong prior, having worked in laboratory science/bio for 7 years, that converting theoretical knowledge and/or text-based protocols into actual physical actions in the real world is most efficiently and accurately performed with real-time human expert guidance…Until we either have AI-powered VR lab training or completely closed lab-in-the-loop type setups for the full stack of virology/bacteriology workflows, a 2026-level LLM isn’t going to aid the lone-wolf DIY bio actor significantly over and above the Internet.”

“Recent studies have shown that the bottleneck is not the AI. It is the lab equipment and more importantly the implicit knowledge of how to do protocols…Conceptually this is similar to the nuclear field. Although access to the core material on how to create a nuclear bomb is widely available…the practical knowledge and tooling has proven to be a barrier for all but the most determined nation states.”

Lab accidents were often identified as the most likely source of human-caused outbreaks and largely independent of AI conditions. Some respondents pointed to the proliferation of BSL-3/4 facilities globally, ongoing gain-of-function research, and variable biosafety standards as key drivers of lab accidents.

“Laboratory accident risk reflects global BSL-3/4 infrastructure expansion—approximately 70 BSL-4 and 1,800 BSL-3 facilities [outside of the U.S.] operational, with substantial growth in China, India, and developing nations with variable biosafety and biosecurity standards (mostly substandard).”

Much of the variation in baseline forecasts of expected damages may have been driven by which historical incidents respondents counted when assessing a base rate for human-caused outbreaks. The 2001 anthrax attacks served as a common anchor, but beyond that, the incidents considered diverged. The breadth of each respondent’s historical audit was a strong predictor of their final base rate, with a wider breadth being associated with a higher base rate.

“When I look at the incidents post anthrax (600M estimated damages), I see several incidents that would also have likely crossed the 100M threshold given the way damages are calculated…Claude estimates (in 2026 USD) that the SARS lab escapes was 80-360M in damages, UK foot and mouth (escaped from the Pirbright laboratory complex) at 600-900M, Lanzhou Brucellosis at 60-250M, and then there’s COVID-19, which call it 25% chance of lab leak. So past 25 years, 4 confirmed incidents that probably hit the 100M-1B bin given the method of counting damages, and a 25% chance that there was an incident that hit the 10T-100T bin. All told, a higher base rate than I thought I’d find.”’

4. What would be the impact of more models being protected by ASL-3?

It is implausible that all GPAI models (regardless of size or capability) would be protected by ASL-3 safeguards. It seems more realistic for all models at the frontier of capabilities to be protected by ASL-3 or equivalent safeguards, even if that is not the case today. To understand the potential impact of all frontier models being protected by ASL-3 safeguards, we asked respondents to forecast the main outcome under this scenario.

We defined “frontier” GPAI models as all those performing as well or better than Claude Sonnet 4.5 on the Epoch Capabilities Index (ECI). The ECI combines scores from more than 40 different AI benchmarks into a single general capability score [9]. At the time of the survey, Claude Sonnet 4.5 scored 147 on the ECI, and 13 models scored at or above that level: Gemini 3 Pro, GPT-5.2, Claude Opus 4.6, Gemini 3 Flash, GPT-5 Pro, Claude Opus 4.5, GPT-5, GPT-5.1, o3-pro, Kimi K2.5, Grok 4, o3, and Claude Sonnet 4.5.

Under the ‘ASL-3 for frontier models’ scenario, the median probability of human-caused outbreaks occurring between now and the end of 2028 and leading to at least $100 million in damages was 10.6% (95% CI: 5.5%, 18.1%). Compared to the ‘no safeguards’ scenario, the ‘ASL-3 for frontier models’ scenario was associated with an 18% (95% CI: 9%, 30%) reduction in this risk. The median respondent thought that ‘ASL-3 for frontier models’ only would be associated with a 20% (95% CI: 9%, 52%) reduction in expected damages due to large-scale human-caused outbreaks relative to the ‘no safeguards’ scenario (see Figure 3). ASL-3 for frontier models mitigated 36.3% of GPAI-attributable risk (see Figure 4). Most respondents placed the ‘ASL-3 for frontier models’ scenario’s risk between ‘ASL-3 for all models’ and ‘no safeguards’, but disagreed on where, with some treating it as nearly equivalent to ‘ASL-3 for all models’ and others as only marginally better than ‘no safeguards’.

Insights from rationales

Forecasters disagreed on the degree to which powerful near-frontier models—particularly Chinese ones—would undermine frontier-only ASL-3 coverage. In their text rationales, some emphasized that frontier models are where the real capability uplift lies, with some citing the Active Site study as evidence that unrestricted models would likely provide limited practical uplift for virology tasks. Some also noted that near-frontier models would likely retain some proprietary safeguards even without ASL-3. Others worried that a community of practice could develop around non-frontier models and that the mere knowledge of capable unsafeguarded models could lower the psychological barrier to attempting an attack.

“Even if they don’t have ASL-3 safeguards, the American models still have quite good safeguards. Unfortunately, this is not the case for most of the Chinese models though. So I think in this world, the easiest path to harm for someone with malicious intent would be to use a Chinese LLM to help them.”

“My guess here, which is strongly influenced by the Active Site study conducted in the summer of 2025, is that no uplift will be provided by models on the list that currently aren’t classified as frontier.”

5. How would variations in ASL-3 influence this impact?

We asked respondents to consider the impact of variations on ASL-3 in a scenario where frontier models are protected by ASL-3. Respondents were asked to forecast the main outcome conditional on frontier models being protected by the following hypothetical variations on the ASL-3 scenario:

Time required to identify jailbreaks

  • A1: The average time required to find a universal jailbreak is 5x longer
  • A2: The average time required to find a universal jailbreak is 5x shorter

Time to jailbreak patching

  • B1: It takes a fifth (0.2x) of the time to patch all universal jailbreaks
  • B2: It takes up to 2x as long to patch all universal jailbreaks

Capabilities loss associated with jailbreaks

  • C1: All universal jailbreaks significantly reduce the model’s GPQA score. Compared to jailbroken models’ current performance relative to the “no jailbreak” baseline, jailbroken models’ relative performance decreases by roughly 30%.
  • C2: All universal jailbreaks only minimally reduce the model’s GPQA score. Compared to jailbroken models’ current performance relative to the “no jailbreak” baseline, jailbroken models’ relative performance increases by roughly 30%.

Figure 5 shows the median change in expected damages relative to ‘ASL-3 for frontier models’ under the variant scenarios. For each of the scenarios associated with improvement from current ASL-3 safeguards, the median expected change was similar to that associated with ‘ASL-3 for all models’. The largest median change—an 8% (95% CI: 5%, 18%) decrease from the ‘ASL-3 for frontier models’ scenario—was associated with a faster time to jailbreak patching (B1).

Figure 5: Relative difference in expected financial damages under each variant of ASL-3 and the ‘ASL-3 for all models’ scenario. Numbers show medians from each group. Values below 1 represent a decrease in expected damages compared to a scenario with ASL-3 for frontier models.

Figure 6 shows the share of GPAI-attributable risk mitigated by the variant scenarios for frontier models, ‘ASL-3 for frontier models’, and ‘ASL-3 for all models’ scenarios. For the median respondent, the ‘5x faster patching’ (B1) and ‘30% decrease in relative performance’ (C1) scenarios were seen as comparable to the ‘ASL-3 for all models’ scenario on this metric. Relative to superforecasters, experts expected variants to ASL-3 to have a greater impact.

More details are available in Appendix C.

Figure 6: Share of GPAI-attributable risk mitigated by each ASL-3 scenario. GPAI-attributable risk is defined as the multiplicative gap in expected damages between the ‘no safeguards’ and ‘no GPAI’ scenarios. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals.

Insights from rationales

Several respondents expressed concern that if non-frontier models remain unprotected and are nearly as capable, then making frontier jailbreaks harder, faster to patch, or more capability-degrading has limited marginal value.

Most respondents agreed that making jailbreaks harder to find would help, but several argued that it would mainly screen out under-resourced actors who were unlikely to succeed anyway. Many also noted that more organized groups would be able to adapt. Making jailbreaks easier to find was generally assessed to be more dangerous because it would expand access to a much larger pool of actors, however unsophisticated. A few respondents noted the interaction of this variant with patching: discovery time is largely irrelevant if jailbreaks persist for months once found. One flagged that AI agents capable of autonomous software research could erode the barrier further.

Forecasters noted that faster patching would more effectively disrupt the months-long efforts required to develop a bioweapon. Credible bioweapon efforts typically involve iterative troubleshooting over weeks or months and that fast patching could disrupt this workflow by closing the window before a project could be completed—and could also reduce the economic incentive to discover and sell jailbreaks. Slow patching was predicted to do the opposite in that it could make jailbreaks more useful and give bad actors ample time to extract everything they need.

Forecasters differed in opinion on the impact of model performance degradation. Some respondents argued that a significant degradation would eliminate the incentive to pursue jailbreaks, since the degraded model would offer little advantage over unprotected non-frontier alternatives. Others argued that even a significantly degraded frontier model would still be remarkably capable by historical standards.

We asked respondents to answer several questions related to pathways to a large-scale human-caused outbreak. This included the type of actor that is most likely to cause such an event, the probability that unauthorized use of a GPAI model was involved in the event, and the most likely pathway to unauthorized model use.

6.1 Actor type

We asked what type of actor is the most likely primary cause of a human-caused outbreak that causes more than $100 million in damages, and how this would change depending on whether GPAI models had no safeguards or were all protected by ASL-3 safeguards. The mean responses for each group of participants are shown in Figure 7. Both groups of respondents thought that state actors were the most likely cause of such an event, and that they would be an even more likely cause under the ASL-3 safeguards scenario. Among the expert respondents, the next most likely group was individual expert actors, although this group’s probability of being the primary cause of a large-scale human-caused outbreak fell under the ASL-3 safeguards scenario. Compared to experts, superforecasters generally put a higher probability on a different type of actor (an “Other” category) being the primary cause. Respondents who put a large probability in this category suggested that groups of experts or non-experts that did not count as terrorist groups would fit into this category.

Figure 7: Probability of the given actor being the primary cause of a large-scale human-caused outbreak. Numbers show group averages.

Insights from rationales

Across both the ‘no-safeguards’ and ‘ASL-3 for all models’ conditions, a dominant theme in respondent rationales was a capability-intent asymmetry. State actors have the resources to develop bioweapons but face strong disincentives (e.g., attribution risk, self-harm), while non-state actors may have motivation but lack capability. Many respondents assign significant probability to accidental release pathways rather than deliberate attacks, shifting the expected actor profile toward individual experts and state labs. Without safeguards, AI is expected to partially close the capability gap for non-state actors, but practical barriers—such as wet lab access, weaponization, and delivery—remain and are largely unaffected by AI. Some argue that pursuing bioweapons is irrational for most actors given the availability of simpler alternatives, and that the “Other” category captures important scenarios (e.g., corporate negligence, criminal groups, insider threats, etc.) outside the listed actor types.

ASL-3 safeguards shift relative probability toward better-resourced actors. State actors and large terrorist organizations gain conditional share because they can overcome or bypass safeguards through in-house expertise, cyber capabilities, and sustained jailbreak efforts, while individual non-experts—the primary beneficiaries of unrestricted AI—lose the most. Jailbreak difficulty is the key mechanism: successful jailbreaking requires specialized cybersecurity skills (compounding the expertise requirement), jailbreaks are transient due to patching, and even successful jailbreaks may produce degraded outputs that less capable actors cannot independently verify.

With deliberate misuse pathways partially blocked, the accidental release pathway gains further relative prominence. Several respondents also flag trusted user exemptions (insiders with legitimate access who misuse it or whose credentials are compromised) as a new attack surface, and that such actors with reduced safeguards may develop over-dependence on AI, increasing the risk of accidental misuse or unintended consequences.

6.2 Unauthorized model use

When asked about the probability of unauthorized frontier model use contributing to a human-caused outbreak that causes more than $100 million in damages, the median expert forecasted a 30% (95% CI: 15 , 50) probability, and the median superforecaster a 23.5% (95% CI: 10.2, 52) probability. The distribution of these forecasts is shown in Figure 8.

Figure 8: Probability of unauthorized use of frontier models, conditional on a large-scale human-caused outbreak occurring. Numbers show group medians and black lines show their bootstrapped 95% confidence intervals.

We then asked respondents to assume that unauthorized model use had contributed to a human-caused outbreak causing more than $100 million in damages, and say how likely they thought different pathways were to have contributed to that unauthorized model use. The responses are shown in Figure 9. Each of the pathways we asked about had a median probability of 15% or higher, suggesting that most respondents found all pathways plausible. Relative to experts, superforecasters generally thought a trusted user exemption was more likely to be used and the actor generating their own jailbreak was less likely. It is worth noting that these median values hide substantial variation in responses.

Figure 9: Distribution of probabilities assigned to each pathway to gain unauthorized access. Numbers show group medians. The pathways were non-exclusive, and each individual was allowed to assign more than 100% across them.

Insights from rationales

Whether respondents weighted accidental versus deliberate pathways more heavily strongly predicted their estimates of unauthorized model use. Those who saw accidents as dominant gave low estimates, since lab leaks were thought to rarely involve AI in the causal chain. Those focused on deliberate attacks gave high estimates, since any motivated attacker would exploit every available tool.

Trusted user exemptions were identified as a critical vulnerability because they exploit human frailties that technical safeguards cannot fully address. Bribery, blackmail, and social engineering can compromise exemption holders regardless of how robust the underlying model defenses are. Moreover, exemption holders are precisely those with proximity to pathogens and lab infrastructure, meaning this type of compromise combines AI access with all the physical infrastructure needed for an attack. However, a countervailing view holds that the exemption pool is small, misuse would be highly attributable, and that screening should filter most threats.

Black market and nonpublic jailbreaks are seen as the path of least resistance for actors lacking technical skills. These markets effectively commoditize safeguard bypass via dark web markets that are harder for labs to monitor and take down than public jailbreaks. Public jailbreaks were seen as the most accessible, but also the most short-lived, limiting their utility for the sustained interaction biological development requires.

Creating novel jailbreaks is the highest-barrier pathway, constrained by a skill mismatch between jailbreaking and biology expertise. As one respondent put it, “most actors who can create a jailbreak, wouldn’t have the biology background to create a serious virus.” This restricts the route mainly to large organizations with resources to recruit across both domains.

Real-world attacks would likely chain multiple methods rather than relying on one, and that the “Other” category is a catch-all for important vectors like open-weight model manipulation, API vulnerabilities, and “unknown unknowns.”

7. Conclusion

7.1 Summary of findings

This study asked 22 domain experts and 20 superforecasters to estimate how ASL-3 safeguards would change the expected financial damages from large-scale human-caused outbreaks between now and 2028. Four findings stand out.

Respondents believe ASL-3 safeguards demonstrate high efficacy. If applied to all GPAI models, the median respondent’s forecasts imply that ASL-3 would mitigate roughly 70% of GPAI-attributable risk. This figure likely overstates true efficacy, because the ‘no GPAI’ benchmark used to define GPAI-attributable risk also removes GPAI’s defensive contributions. Nonetheless, even accounting for this bias, the results suggest that ASL-3 captures a substantial share of the achievable risk reduction when applied.

Protecting only frontier models with ASL-3 reduces risk, but not as much as protecting all GPAI models. Protecting only frontier models captured less of the risk reduction than protecting all GPAI models, though respondents disagreed substantially on the size of the gap. This disagreement hinged on how capable and accessible near-frontier models—particularly Chinese open-weight models—are judged to be. The difference underscores that the value of ASL-3 as a mitigation depends not just on its technical performance but on how broadly equivalent safeguards are adopted across the ecosystem.

Making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models were generally expected to improve ASL-3 performance. For the median respondent, the effects of these variants of ASL-3 for frontier models were roughly similar to the effect of ASL-3 for all models. Of the variants we asked about, the largest effect was associated with faster patching, consistent with the logic that bioweapon development requires sustained iterative access.

ASL-3 shifts the threat landscape toward better-resourced actors and accidental pathways. Respondents expected ASL-3 to disproportionately screen out non-expert and lone-wolf actors, shifting relative probability toward state actors, well-resourced organizations, and accidental releases from laboratories. Trusted user exemptions were identified as a notable vulnerability, because they combine AI access with proximity to the physical infrastructure needed for an attack and exploit human factors that technical safeguards cannot fully address. Many respondents emphasized that lab accidents represent a substantial share of baseline risk that AI safeguards do not address.

7.2 Limitations

Beyond the caveats noted above, several broader limitations should inform interpretation.

Forecasting low-probability, high-consequence events is inherently difficult. We focus on relative changes across scenarios rather than absolute damage estimates for this reason, but even relative judgments may be unreliable when base rates are poorly constrained. The wide variation in respondents’ baseline forecasts—driven largely by which historical incidents they counted—illustrates this challenge.

The ‘no safeguards’ baseline is more extreme than the status quo. This scenario asks respondents to imagine all GPAI models are optimized purely for helpfulness with no safety constraints. In practice, most commercial models retain some safety training even without ASL-3. The measured gap between ‘no safeguards’ and the ASL-3 scenarios therefore captures the value of ASL-3 above a minimal baseline, not its marginal value above current industry-standard practices.

The capability freeze assumption bounds the shelf life of the findings. The study asked respondents to assume no improvement in GPAI capabilities through 2028. This isolates views on current-generation safeguards, but it means respondents were evaluating ASL-3 against a threat landscape that will almost certainly change. More capable models may provide greater uplift to malicious actors, may be harder to safeguard, or may shift the balance between information access and physical bottlenecks. These findings are best understood as a snapshot of expert views on ASL-3’s efficacy against today’s models.

The sample is modest and skewed toward Western biosecurity institutions. While the expert group had broad experience—14 of 22 reported expertise in more than one of biosecurity, national security and terrorism studies—perspectives from researchers in regions with rapidly expanding BSL infrastructure and those facing the highest burden of terrorist attacks are underrepresented.

Little information was provided on ASL-3 Security Standards. Forecasters were asked to consider the effects of ASL-3 as a whole—including both the Deployment Standards and Security Standards. We provided detailed information on the effectiveness of the Deployment Standards but did not have comparable information on the effectiveness of the Security Standards.

7.3 Implications

These results provide some evidence that deployment safeguards of the kind specified by ASL-3 can meaningfully reduce biosecurity risk. Respondents judged that comprehensive ASL-3 coverage would reduce most of the risk added by GPAI models, and they attributed this largely to ASL-3’s ability to block non-expert actors. To the extent this judgment is accurate, it suggests that investment in deployment-stage safeguards is a worthwhile component of biosecurity risk management. An important caveat is that respondents were evaluating the effectiveness of safeguards for early-2026 models. As such, this should not be taken as evidence that ASL-3 would be similarly effective for substantially more capable systems.

The value of these safeguards depends on how broadly they, or equivalent measures, are adopted. Respondents saw a substantial gap between protecting all GPAI models and protecting only frontier models, and the size of that gap hinged on how capable and accessible near-frontier and open-weight models were judged to be. This suggests that the risk reduction achievable by any single developer is bounded by the safeguards adopted by other developers. This strengthens the case for industry-wide standards and for governance attention to models that fall outside the reach of any individual developer’s deployment controls.

The findings also provide some insight into views on pathways to improve safeguards. For most respondents, the time required to find a jailbreak, speed of patching jailbreaks, and the performance associated with jailbreaks were all expected to reduce risk. These offer pathways to improving the effectiveness of ASL-3. The speed of patching jailbreaks was generally considered to have the greatest impact on risk. Respondents reasoned that bioweapon development requires sustained, iterative access over weeks or months, so closing jailbreaks quickly can disrupt an effort even after a jailbreak is found. Some respondents suggested that trusted user exemptions may also present a potential weak point in ASL-3, which may be a useful point of intervention.

Finally, the findings indicate that safeguards reshape the threat landscape rather than uniformly suppressing it. As ASL-3 screens out less sophisticated actors, risk concentrates among well-resourced state actors and accidental laboratory releases. Because a substantial share of baseline risk lies outside the influence of AI safeguards, AI-focused measures are best understood as complementary to established biosecurity interventions such as laboratory biosafety and biosecurity, and nucleic-acid synthesis screening.

7.4 Future directions

Two methodological improvements would most strengthen future iterations of this work. First, introducing a ‘perfect safeguards’ scenario—where GPAI exists and provides defensive benefits but cannot be misused—would allow a cleaner estimate of ASL-3’s efficacy by disentangling GPAI’s offensive and defensive contributions to outbreak risk. Second, developing an estimate of baseline risk under the status quo—where most, but not all, frontier models have some protections—would give insight into the impact of ASL-3 under real world conditions and the potential benefits of improving frontier model protections from what exists currently. However, developing this baseline would require more detailed information on the protections applied to other frontier models.

It would also be valuable to see how these findings change as model capabilities advance: repeating this elicitation periodically would track whether the efficacy of ASL-3, or future safeguards, holds or erodes as the models it protects become more capable. More broadly, this study demonstrates the feasibility of structured expert elicitation as a tool for evaluating AI safety measures against diffuse, hard-to-observe threats. Similar approaches could be applied to other risk domains covered by ASL-3 or future safety levels.

Notes

  1. Participant rationales quoted in this report are included as originally shared and have not been edited or fact-checked. We do not necessarily endorse the views or information expressed. ↩︎

1 Forecasting Research Institute
2 University of Oxford
3 University of Pennsylvania
4 Federal Reserve Bank of Chicago
    Related Research
    Working paper
    Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards
    Jul 1, 2025