{"id":2265,"date":"2026-08-12T15:55:00","date_gmt":"2026-08-12T15:55:00","guid":{"rendered":"https:\/\/forecastingresearch.org\/?post_type=research&#038;p=2265"},"modified":"2026-08-12T15:16:05","modified_gmt":"2026-08-12T15:16:05","slug":"impacts-of-asl-3-safeguards-on-biorisks","status":"publish","type":"research","link":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks","title":{"rendered":"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks"},"content":{"rendered":"\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-constrained wp-block-group-is-layout-constrained\">\n<details class=\"wp-block-details is-layout-flow wp-block-details-is-layout-flow\"><summary>Acknowledgments<\/summary>\n<p class=\"wp-block-paragraph\">This research would not be possible without the thoughtful participation of our survey respondents. We are grateful to Mrinank Sharma and Jerry Wei for helpful feedback on survey design. We are also grateful to Anna Bartlett, Coralie Consigny, Johan S Daniel, and Josh Connor for research assistance, and to Otto Kuusela and Simas Ku\u010dinskas for helpful feedback on research design.<\/p>\n<\/details>\n\n\n\n<details class=\"wp-block-details is-layout-flow wp-block-details-is-layout-flow\"><summary>Disclaimers<\/summary>\n<p class=\"wp-block-paragraph\">This report was originally produced as part of an internal project funded by and completed for Anthropic PBC. Anthropic provided input and advice in the development of the report, but all findings and claims represent the views of the authors. Due to the general interest of the findings, we are releasing a version of the report with confidential information removed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The views expressed in this paper do not necessarily reflect the views of the Federal Reserve Bank of Chicago or the Federal Reserve System.<\/p>\n<\/details>\n<\/div><\/div>\n\n\n\n<h2 id=\"summary\" class=\"wp-block-heading\">Summary<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In May 2025, Anthropic announced that it had activated AI Safety Level 3 (ASL-3) Deployment and Security Standards as a precautionary measure, in response to Claude Opus 4&#8217;s improved CBRN-related capabilities [1]. The deployment standards include the implementation of real-time classifier guards, offline monitoring, access controls, a bug bounty program, threat intelligence, and rapid response.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Forecasting Research Institute conducted a study to better understand the effectiveness of these measures\u2014particularly the ASL-3 deployment standards\u2014at mitigating biosecurity risks. We surveyed 22 people with expertise in biosecurity, national security, and\/or terrorism studies (\u201cexperts\u201d), and 20 top generalist forecasters (\u201csuperforecasters\u201d), all of whom completed the survey between March 4 and March 29, 2026.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We asked respondents to forecast the total expected financial damages from large-scale human-caused outbreaks (outbreaks leading to at least $100 million in total worldwide damages) occurring between April 2026 and the end of 2028, and how this value would change across four scenarios: a world without general-purpose AI (GPAI); a world where no GPAI models had any safeguards; a world where all GPAI models were protected by ASL-3; and a world where frontier GPAI models were protected by ASL-3.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because we did not include a status-quo scenario\u2014one in which many frontier models already carry their developers&#8217; own misuse protections\u2014our estimates do not capture the marginal benefit of extending ASL-3 beyond what other developers&#8217; existing safeguards already provide. We also asked respondents to assume that the capabilities of GPAI models do not improve before December 31, 2028. We made these stipulations because we wanted respondents&#8217; views on ASL-3&#8217;s ability to mitigate harm from current models, rather than their views on other companies&#8217; safeguards or likely AI progress.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-left is-layout-flex wp-container-core-buttons-is-layout-61db0649 wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button is-style-fill\"><a class=\"btn orange\" href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">View the full PDF report <svg width=\"7\" height=\"9\" viewBox=\"0 0 7 9\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n  <path d=\"M0.000156283 8.60806L4.22416 4.33606V4.24006L0.000156283 6.10352e-05H1.80816L6.06416 4.28806L1.80816 8.60806H0.000156283Z\" fill=\"#102B23\"\/>\n<\/svg>\n<svg width=\"8\" height=\"10\" viewBox=\"0 0 8 10\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n  <path d=\"M0.601719 8.85794L4.82572 4.58594V4.48994L0.601719 0.249939H2.40972L6.66572 4.53794L2.40972 8.85794H0.601719Z\" fill=\"#102B23\"\/>\n<\/svg><\/a><\/div>\n<\/div>\n\n\n\n<h3 id=\"key-takeaways\" class=\"wp-block-heading\">Key Takeaways<\/h3>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Experts and superforecasters generally believe that\nimplementation of ASL-3 safeguards across all general-purpose AI (GPAI)\nmodels can substantially reduce the risk of large-scale, human-caused\noutbreaks, relative to a world with no safeguards. Under the \u2018no\nsafeguards\u2019 scenario, the median forecasted probability of human-caused\noutbreaks, occurring between April 2026 and December 2028, causing at\nleast $100 million in damages was 16% (95% CI: 9.2%, 32.8% ). This\ndropped to 9.2% (95% CI: 4% , 16.9%) under the \u2018ASL-3 for all models\u2019\nscenario.<\/li>\n\n\n\n<li>Forecasters estimate that ASL-3 safeguards, if implemented across all GPAI models, would capture most of the risk reduction that is possible to achieve via safeguarding. For the median respondent, implementing ASL-3 across all GPAI models achieved 71.7% (95% CI: 65.7%, 97.2%) of the risk reduction associated with a hypothetical scenario where no GPAI exists.<\/li>\n\n\n\n<li>Most respondents believe protecting frontier models (those at\nleast as capable as Sonnet 4.5) with ASL-3 would substantially reduce\nrisk relative to a \u2018no safeguards\u2019 baseline, although not by as much as\nextending it to all GPAI, regardless of model size or capability. For\nthe median respondent, \u2018ASL-3 for frontier models\u2019 achieved 36.3% (95%\nCI: 22.9%, 68.4%) of the risk reduction associated with the \u2018no GPAI\u2019\nscenario.<\/li>\n\n\n\n<li>The median respondent thought that variations to ASL-3\u2014making\njailbreaks harder to find, patching them faster, or further degrading\nthe performance of jailbroken models\u2014would reduce risk by an amount\nsimilar to applying ASL-3 to all GPAI models. The largest median risk\nreductions were associated with patching jailbreaks 5x faster and with a\nfurther 30% drop in jailbroken models&#8217; performance relative to\nnon-jailbroken models. For the median respondent, each of these two\nvariants achieved 73.5% of the risk reduction associated with a\nhypothetical scenario where no GPAI exists (95% CI: 52.3%, 100% for\nfaster patching; 47.8%, 97.3% for reduced jailbroken performance). Many\nrespondents noted that such variations may matter more in a world where\nASL-3 covers all GPAI models, and many saw trusted user exemptions as a\npotential weak spot in ASL-3.<\/li>\n<\/ul>\n\n\n\n<h3 id=\"key-results\" class=\"wp-block-heading\">Key Results<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>ASL-3 safeguards are seen as having high efficacy: if applied across all GPAI models, respondents estimate that ASL-3 would capture most of the risk reduction achievable through safeguarding.<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>When respondents assumed all models had ASL-3 protections, the median forecasted probability of human-caused outbreaks causing at least $100 million in damages dropped by roughly 40% relative to the \u2018no safeguards\u2019 scenario (from 16% (95% CI: 9.2%, 32.8%) to 9.2% (95% CI: 4%, 16.9%)). For comparison, when asked to consider a hypothetical world where GPAI was never developed, the median forecast was 9.3% (95% CI: 3.8%, 13.2%). (See <a href=\"#fig-01\" id=\"#fig-01\">Figure 1<\/a>.)<\/li>\n\n\n\n<li>Regarding expected financial damages, the result was similar,\nalthough the \u2018no GPAI\u2019 scenario was generally associated with lower\ndamage estimates than the ASL-3 scenarios. Compared to a world with no\nsafeguards, a world where all models have ASL-3 safeguards reduced\nmedian predicted financial damages of human-caused outbreaks between now\nand 2028 by 40% (95% CI: 17%, 63%). Under the \u2018no GPAI\u2019 scenario, the\nmedian participant reduced their forecast of expected damages by 49%\n(95% CI: 19%, 71%).<\/li>\n\n\n\n<li>Overall, the forecasts suggest that respondents see \u2018ASL-3 for\nall models\u2019 as reducing risks of a large-scale human-caused outbreak to\na risk level slightly above a hypothetical world with no GPAI. The\nmedian respondent\u2019s forecasts suggest that \u2018ASL-3 for all models\u2019 is\nassociated with a reduction in risk that is 71.7% (95% CI: 65.7%, 97.2%)\nof the risk reduction associated with the \u2018no GPAI\u2019 scenario.<\/li>\n\n\n\n<li>Rationales showed that respondents generally believed that ASL-3\nwould effectively block non-experts. Many respondents were reassured by\nthe pace of jailbreak patching and the difficulty in finding jailbreaks,\nsuggesting that this likely requires substantial cybersecurity\nexpertise. Several also noted the deterrence effects of\nsafeguards.<\/li>\n<\/ul>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter is-resized\" id=\"fig-01\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-01.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 1:<\/strong> Probability of human-caused outbreaks that start between April 2026 and December 2028, causing at least $100 million in economic damages. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals.<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\"><strong>Real-world impact likely depends on breadth of adoption: protecting only frontier models still reduces risk substantially, but less than comprehensive coverage.<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>We also asked how respondents\u2019 forecasts would change if ASL-3\nwere only applied to \u2018frontier\u2019 GPAI models, which we defined as those\nat least as capable as Sonnet 4.5. Under this scenario, the median\nrespondent forecast for the probability of human-caused outbreaks\nstarting between now and 2028, causing at least $100 million in damages,\nwas 10.6% (95% CI: 5.5%, 18.1%). The median estimate of expected damages\nwas 20% (95% CI: 9%, 52%) lower under this scenario than under the \u2018no\nsafeguards\u2019 scenario.<\/li>\n\n\n\n<li>For the median participant, the expected damages in the \u2018ASL-3\nfor all models\u2019 scenario was 4% (95% CI: 1%, 16%) lower than the\nexpected damages in the \u2018ASL-3 for frontier models\u2019 scenario.<\/li>\n\n\n\n<li>Most respondents placed the frontier-only ASL-3 scenario\u2019s risk\nbetween \u2018ASL-3 for all models\u2019 and \u2018no safeguards\u2019, but disagreed on\nwhere, with some treating it as nearly equivalent to \u2018ASL-3 for all\nmodels\u2019 and others as only marginally better than \u2018no safeguards\u2019. Much\nof this disagreement hinged on the capability and accessibility of\nnear-frontier models, particularly Chinese open-weight models.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models were generally expected to improve ASL-3 performance, with faster patching expected to be the most impactful.<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>We asked respondents to consider six variants of ASL-3 safeguards\nthat differed from the current safeguards\u2019 performance in one of three\nways: time required to find a jailbreak, time taken to patch a\njailbreak, and performance of jailbroken models relative to the base\nmodel. For the median respondent, each of the variations associated with\nimprovement\u2014making jailbreaks harder to find, patching them faster, or\nfurther degrading the performance of jailbroken models\u2014was expected to\nreduce risk to a greater extent than the \u2018ASL-3 for frontier models\u2019\nscenario.<\/li>\n\n\n\n<li>The largest median deviation from the \u2018ASL-3 for frontier models\u2019\nexpected damages baseline was an 8% decrease (95% CI: 5%, 18%),\nassociated with a scenario of a 5x decrease in the time taken to patch a\njailbreak. The median respondent\u2019s forecasts suggest that ASL-3 for\nfrontier models with 5x faster patching is associated with a reduction\nin risk that is 73.5% (95% CI: 52.3%, 100%) of the risk reduction\nassociated with the \u2018no GPAI\u2019 scenario.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>ASL-3 shifts the threat landscape but does not address all major risk pathways.<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Respondents expected ASL-3 to disproportionately screen out\nnon-expert and lone-wolf actors, shifting the relative probability of\ndamages toward state actors and accidental releases from\nlaboratories.<\/li>\n\n\n\n<li>Superforecasters saw trusted user exemptions as a notable\nvulnerability, with some noting that they combine AI access with\nproximity to the physical infrastructure needed for an attack and\nexploit human factors that technical safeguards cannot fully address.\nMany respondents emphasized that lab accidents represent a substantial\nshare of baseline human-caused outbreak risk that AI safeguards do not\naddress.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"introduction\" class=\"wp-block-heading\">1. Introduction<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In May 2025, Anthropic announced that it had activated AI Safety Level 3 (ASL-3) Deployment and Security Standards as described in Version 2 of Anthropic\u2019s Responsible Scaling Policy (RSP) [1]. This activation came in conjunction with the launch of Claude Opus 4, and was described as a precautionary measure. Anthropic determined that due to continued improvements in CBRN-related knowledge and capabilities, it was not possible to rule out that Claude Opus 4 had crossed its CBRN-3 capability threshold. This threshold is defined in the RSP as: \u201cThe ability to significantly help individuals or groups with basic technical backgrounds (e.g., undergraduate STEM degrees) create\/obtain and deploy CBRN weapons\u201d [2].<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Anthropic\u2019s announcement noted that ASL-3 defenses would be focused on biological weapons risks. The possibility that AI could enable more actors to develop biological weapons has been raised by heads of AI companies and experts in biosecurity [3, 4, 5]. There is debate about the significance of the risks, with some believing that AI does little to overcome the key barriers to biosecurity risks [6, 7]. Prior work has found that most experts and top generalist forecasters expect biological risks to increase with AI advances in AI capabilities [8].<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The ASL-3 Deployment Standard consists of requirements across the following areas: threat modeling, defense in depth, red-teaming, rapid remediation, monitoring, access controls for trusted users, and third-party environments. To meet these requirements, Anthropic implemented the following mitigations: real-time classifier guards, offline monitoring, access controls, a bug bounty program, threat intelligence, and rapid response. Further detail on Anthropic\u2019s implementation of ASL-3 Deployment and Security Standards is available in <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=30\" target=\"_blank\" rel=\"noreferrer noopener\">Appendix A<\/a>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Although Anthropic can evaluate aspects of its mitigations (for example, the time taken to patch an identified jailbreak), it is unclear how well these measures achieve the overarching goal of reducing biosecurity risks posed by general-purpose AI (GPAI) models. To investigate this, we conducted a forecasting survey of experts in biosecurity, national security, or terrorism studies, as well as top generalist forecasters. Respondents were given information on Anthropic\u2019s ASL-3 standards and the measures taken to meet them. This included confidential data from red-teaming efforts, the bug bounty program, and other safeguards testing conducted by Anthropic. Although respondents&#8217; forecasts included consideration of both the Deployment and Security Standards that make up ASL-3, the focus of this survey was the Deployment Standards.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This survey aimed to answer four research questions:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>How effective are ASL-3 safeguards?<\/li>\n\n\n\n<li>What would be the impact of more models being protected by\nASL-3?<\/li>\n\n\n\n<li>How would variations in ASL-3 influence this impact?<\/li>\n\n\n\n<li>How do ASL-3 safeguards influence threat models related to\nbiosecurity risks?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">In this report, we briefly describe the survey methods before presenting the results of the survey that speak to each of these research questions.<\/p>\n\n\n\n<h2 id=\"methods-and-participants\" class=\"wp-block-heading\">2. Methods and participants<\/h2>\n\n\n\n<h3 id=\"survey-contents\" class=\"wp-block-heading\">2.1 Survey contents<\/h3>\n\n\n\n<h4 id=\"description-of-the-main-outcome\" class=\"wp-block-heading\">Description of the main\noutcome<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The main outcome we asked participants to forecast was:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">The total worldwide damages attributable to human-caused outbreaks that start between April 1, 2026, and December 31, 2028, and that each individually cause at least $100 million in total worldwide damages (in 2026 USD).<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">We asked respondents to make a few important assumptions when answering this question. First, we asked them to assume that until December 31, 2028, the capabilities of general-purpose AI models do not improve, as if a \u2018pause on general-purpose AI development\u2019 were in place. That is, for that period, no models outperform Claude 4.6, Gemini 3, or GPT-5, and this slowdown is not indicative of a more general slowdown in technological progress. We made these stipulations because we were interested in their views on ASL-3\u2019s ability to mitigate harm from current models, rather than their views on likely AI progress.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We asked respondents to forecast the global economic costs attributable to the outbreaks, beginning with the isolation of the pathogen and extending for the full life of an outbreak, and including direct economic losses and monetized mortality. We asked participants to use a value of statistical life of $2.2 million, regardless of geographic location. For full details of this question, including the resolution criteria, please see <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=33\" target=\"_blank\" rel=\"noreferrer noopener\">Appendix B<\/a>.<\/p>\n\n\n\n<h4 id=\"description-of-scenarios\" class=\"wp-block-heading\">Description of scenarios<\/h4>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>No GPAI<\/strong>: In this scenario, respondents were\nasked to assume that GPAI models never existed and will not exist. They\nwere asked to assume that all other features of the world remain the\nsame, e.g., they should not assume that the absence of GPAI models\nindicates that technological progress is slower across all\ndomains.<\/li>\n\n\n\n<li><strong>No safeguards<\/strong>: In this scenario, respondents\nwere asked to assume that no safeguards are applied to all GPAI models.\nIn other words, to assume that any system safeguards, refusal behaviors,\nand other safety-focused constraints have been removed or disabled, and\nthat such AI models are optimized for helpfulness (efficient and\naccurate task performance) without safety protections.<\/li>\n\n\n\n<li><strong>ASL-3 safeguards for all GPAI models<\/strong>: In this\nscenario, respondents were asked to assume that all GPAI\nmodels\u2014regardless of size or capability\u2014are protected by Anthropic\u2019s\nASL-3 safeguards.<\/li>\n\n\n\n<li><strong>ASL-3 safeguards for frontier GPAI models:<\/strong> In\nthis scenario, respondents were asked to assume that \u201cfrontier\u201d GPAI\nmodels are protected by Anthropic\u2019s ASL-3 safeguards. \u201cFrontier\u201d GPAI\nmodels were defined as all those performing as well or better than\nClaude Sonnet 4.5 on the <a href=\"https:\/\/epoch.ai\/benchmarks\/eci?view=graph&amp;tab=release-date\"><u>Epoch\nCapabilities Index<\/u><\/a> (ECI) [9]. At the time of the survey, Claude\nSonnet 4.5 scored 147 on the ECI, and 13 models scored at or above that\nlevel: Gemini 3 Pro, GPT-5.2, Claude Opus 4.6, Gemini 3 Flash, GPT-5\nPro, Claude Opus 4.5, GPT-5, GPT-5.1, o3-pro, Kimi K2.5, Grok 4, o3, and\nClaude Sonnet 4.5.<\/li>\n<\/ul>\n\n\n\n<h3 id=\"additional-questions\" class=\"wp-block-heading\">Additional questions<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In addition to forecasts of the main outcome, we asked questions to understand respondents\u2019 views on pathways to human-caused outbreaks, and how these are influenced by ASL-3. Specifically, we asked respondents to assume that a human-caused outbreak that caused more than $100 million in damages had occurred, and to say what probability they placed on different types of actors being the primary cause of the outbreak. We asked participants to answer these questions for two scenarios: \u2018no safeguards\u2019 and \u2018ASL-3 for all GPAI models\u2019. For the \u2018ASL-3 for all GPAI models\u2019 scenario, we also asked about the probability that a GPAI model had been used in an unauthorized way to assist the process of causing the outbreak, and the most likely pathways to unauthorized use. We also asked participants for some basic details on their demographics and expertise.<\/p>\n\n\n\n<h3 id=\"calibration-modules\" class=\"wp-block-heading\">Calibration modules<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Before completing any forecasting questions, participants completed two short calibration modules. One module asked participants to estimate low probabilities. An example question is, \u201cWhat is the probability of being struck by lightning in any year?\u201d (Answer: 8.1e-7). The other module asked for values relevant to the survey subject matter. An example question is, \u201cIn the 1984 Rajneeshee bioterror attack in Oregon (the largest bioterrorism attack in US history), how many people were infected with Salmonella?\u201d (Answer: 751). The primary purpose of these modules was to give participants an opportunity to test their own calibration.<\/p>\n\n\n\n<h3 id=\"data-analysis\" class=\"wp-block-heading\">2.2 Data Analysis<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Data analysis was conducted using R after aggregating and cleaning the data submitted by participants. For details of data cleaning, see <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=33\" target=\"_blank\" rel=\"noreferrer noopener\">Appendix B<\/a>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For the main outcome, we asked participants to forecast the probability that human-caused outbreaks would lead to different levels of financial damage by 2028. We then calculated expected financial damages for each scenario. To calculate these values, we assumed that damages within each bin were all equally likely (uniformly distributed). We then multiplied the probability assigned to each range by the midpoint dollar amount of that range and summed these probability-weighted contributions across all bins to obtain the overall value of expected damage. For the highest bin (more than $100 trillion), we used $1,000 trillion as the upper bound for calculations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This approach could be a limitation, especially for wider damage bins or upper-tail outcomes where damages may not be evenly distributed within a range. To reduce the impact of this issue, we used a relatively large number of bins and had participants review the calculated expected damages. Participants were asked to adjust their probability forecasts if the expected damages calculation did not align with their expectations, so these values reflect calculations that participants reviewed and had the opportunity to revise.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To describe forecasts on the main outcome, we report several metrics: the probability that damages from human-caused outbreaks, occurring from now until the end of 2028, will be at least $100 million; the ratio of calculated expected damages values between different scenarios; and a metric we define as \u201cshare of GPAI-attributable risk mitigated\u201d.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We define \u201cGPAI-attributable risk\u201d in relative terms as the proportional increase in expected damages between the \u2018no safeguards\u2019 scenario and the \u2018no GPAI\u2019 scenario. Formally, we define this as<\/p>\n\n\n\n<div class=\"wp-block-math\"><math display=\"block\"><semantics><mrow><mi>G<\/mi><mi>P<\/mi><mi>A<\/mi><mi>I<\/mi><mtext>-attributable&nbsp;risk<\/mtext><mo>=<\/mo><mfrac><msub><mi>E<\/mi><mtext>no&nbsp;sg<\/mtext><\/msub><msub><mi>E<\/mi><mtext>no&nbsp;AI<\/mtext><\/msub><\/mfrac><mo>\u2212<\/mo><mn>1<\/mn><\/mrow><annotation encoding=\"application\/x-tex\">GPAI\\text{-attributable risk} = \\frac{E_{\\text{no sg}}}{E_{\\text{no AI}}} &#8211; 1<\/annotation><\/semantics><\/math><\/div>\n\n\n\n<p class=\"wp-block-paragraph\">where <em>(E<sub>no AI<\/sub>)<\/em> denotes expected damages with no GPAI and <em>(E<sub>no sg<\/sub>)<\/em> denotes expected damages with no safeguards.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This represents our estimate of the additional risk posed by the existence of unsafeguarded GPAI models. We then measure the share of this GPAI-attributable risk that is mitigated by ASL-3 safeguards by comparing expected damages under the \u2018no safeguards\u2019 scenario to those under the relevant ASL-3 scenario (<span class=\"math inline\"><em>E<\/em><sub><em>s<\/em><\/sub><\/span>). Specifically, we compute<\/p>\n\n\n\n<div class=\"wp-block-math\"><math display=\"block\"><semantics><mrow><mtext>Proportion&nbsp;of&nbsp;risk&nbsp;removed<\/mtext><mo>=<\/mo><mfrac><mrow><mfrac><msub><mi>E<\/mi><mtext>no&nbsp;sg<\/mtext><\/msub><msub><mi>E<\/mi><mi>s<\/mi><\/msub><\/mfrac><mo>\u2212<\/mo><mn>1<\/mn><\/mrow><mrow><mfrac><msub><mi>E<\/mi><mtext>no&nbsp;sg<\/mtext><\/msub><msub><mi>E<\/mi><mtext>no&nbsp;AI<\/mtext><\/msub><\/mfrac><mo>\u2212<\/mo><mn>1<\/mn><\/mrow><\/mfrac><\/mrow><annotation encoding=\"application\/x-tex\">\\text{Proportion of risk removed} = \\frac{\\frac{E_{\\text{no sg}}}{E_s} &#8211; 1}{\\frac{E_{\\text{no sg}}}{E_{\\text{no AI}}} &#8211; 1}<\/annotation><\/semantics><\/math><\/div>\n\n\n\n<p class=\"wp-block-paragraph\">This quantity captures the fraction of the proportional increase in expected damages attributable to unsafeguarded GPAI that is eliminated under ASL-3 safeguards. Although we refer to this quantity as a \u201cshare\u201d of risk mitigated, it is more precisely a normalized relative-risk reduction measure based on multiplicative differences in expected damages, rather than an additive share of the total risk.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A caveat applies to this definition: the \u2018no GPAI\u2019 scenario removes not only the misuse risks of GPAI but also any defensive benefits (e.g., contributions to pandemic preparedness and response). A \u2018no GPAI\u2019 world therefore has higher baseline outbreak risk than a hypothetical world with perfectly safeguarded AI would, making it a more lenient benchmark. The share mitigated by ASL-3, measured against this benchmark, is best interpreted as an upper bound.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A further caveat applies to a small subset of four participants. These participants were excluded from aggregate calculations and figures of this metric because they assessed expected damages to be lower in the presence of AI than in the \u2018no GPAI\u2019 scenario. In these cases, the calculations of the share of GPAI-attributable risk reduced do not yield interpretable measures since their baseline implies that AI itself reduces risk.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We focus on the relative change in risk because, given the difficulty of forecasting expected damages, we believe these relative changes are likely to be more reliable than the absolute values of damages. However, we present the details of absolute values in <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=40\" target=\"_blank\" rel=\"noreferrer noopener\">Appendix C<\/a>.<\/p>\n\n\n\n<h3 id=\"participants\" class=\"wp-block-heading\">2.3 Participants<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The results represent responses from a total of 42 participants: 22 participants were recruited for their expertise in biosecurity, national security, or terrorism studies (hereafter, \u201cexperts\u201d), and 20 participants were recruited as top generalist forecasters (hereafter, \u201csuperforecasters\u201d). Participants completed the survey between March 4 and March 29, 2026. On average, participants spent 6 hours completing the survey.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The median expert participant reported 11 years of experience relevant to the survey. The expert sample was recruited across three domains \u2014 biosecurity, national security, and terrorism studies \u2014 but most experts had experience spanning multiple domains. Although it was not a requirement for their participation in the study, seven of the superforecaster participants reported experience in at least one of the biosecurity, national security, or terrorism studies. See <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=33\" target=\"_blank\" rel=\"noreferrer noopener\">Appendix B<\/a> for details of recruitment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">More than half (55%) of the expert participants reported having been granted access to restricted national security information (now or in the past). Only 20% of superforecasters reported the same. Most participants reported using LLMs regularly (at least once a week). More details on the participants are available in <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=40\" target=\"_blank\" rel=\"noreferrer noopener\">Appendix C<\/a>.<\/p>\n\n\n\n<h2 id=\"how-effective-are-asl-3-safeguards\" class=\"wp-block-heading\">3. How effective are ASL-3\nsafeguards?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To understand participants\u2019 views on the effectiveness of ASL-3 at reducing biosecurity misuse risks of GPAI models, we can first review the difference in the probability of human-caused outbreaks occurring between now and the end of 2028 and causing at least $100 million in damages under the different scenarios. <a href=\"#fig-01\" id=\"#fig-01\">Figure 1<\/a> in the Summary shows the probability of this outcome under the four scenarios. <a href=\"#fig-02\" id=\"#fig-02\">Figure 2<\/a> shows the relative reduction in the probability of this outcome under the \u2018no GPAI\u2019 and ASL-3 scenarios, relative to \u2018no safeguards\u2019. For the median respondent, the \u2018no GPAI\u2019 scenario is associated with the largest reduction in risk [39% (95% CI: 16%, 50%)], but this is only slightly greater than the risk reduction associated with the \u2018ASL-3 for all models\u2019 scenario [32% (95% CI: 15%, 41%)].<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"#fig-03\" id=\"#fig-03\">Figure 3<\/a> shows how the calculated expected damages under the two ASL-3 scenarios and the \u2018no GPAI\u2019 scenario compare to the \u2018no safeguards\u2019 scenario. Although there was substantial variation, most believed that ASL-3 protections for all GPAI models would significantly reduce the expected damages from large-scale human-caused outbreaks, relative to a world where GPAI had no safeguards. This was particularly true for the expert respondents. The median respondent\u2019s forecasts suggested that ASL-3 protecting all models would reduce expected damages by 40% (95% CI: 17%, 63%).<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-02\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-02.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 2:<\/strong> Ratio of forecasts of probability of human-caused outbreaks that start between April 2026 and December 2028, causing at least $100 million in economic damages, relative to \u2018no safeguards\u2019 scenario. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals. Values below 1 indicate lower P(\u2265$100M) than the no-safeguards scenario.<\/figcaption><\/figure>\n<\/div>\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-03\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-03.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 3:<\/strong> Ratio of expected damages for each protection scenario compared to the \u2018no safeguards\u2019 baseline. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals. Values below 1 indicate lower expected damages than the no-safeguards scenario.<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 id=\"gpai-attributable-risk\" class=\"wp-block-heading\">GPAI-attributable risk<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The scenario that asked respondents to assume GPAI models did not exist can provide an upper bound on the risk reduction that could be achieved by safeguards against GPAI model misuse, as it demonstrates the risk of human-caused outbreaks that would persist regardless of GPAI models. We can also express these results in terms of the share of GPAI-attributable risk mitigated by ASL-3. Aggregates presented here exclude four participants whose responses imply that GPAI reduces risk, making this metric uninterpretable for them. See Section 2.2 for a definition of the metric and more detail on the filtering. The median participant\u2019s forecasts imply that ASL-3 for all models would mitigate 71.7% (95% CI: 65.7%, 97.2%) of GPAI-attributable risk. <a href=\"#fig-04\" id=\"#fig-04\">Figure 4<\/a> shows this comparison.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-04\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-04.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 4:<\/strong> Share of GPAI-attributable risk mitigated by each ASL-3 scenario. GPAI-attributable risk is defined as the multiplicative gap in expected damages between the \u2018no safeguards\u2019 and \u2018no GPAI\u2019 scenarios. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals.<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">An important limitation of this comparison deserves emphasis. The \u2018no GPAI\u2019 scenario serves as a proxy for perfect safeguards, but it is an imperfect one. Perfect safeguards would eliminate misuse risk while preserving the defensive benefits of GPAI: its contributions to pandemic surveillance, pathogen characterization, drug and vaccine development, and outbreak response. The \u2018no GPAI\u2019 scenario removes all of these. This means the \u2018no GPAI\u2019 world is likely more vulnerable to outbreaks (including human-caused ones) than a world with perfectly safeguarded GPAI, making it a lenient benchmark.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The practical consequence is that the 71.7% efficacy figure\u2014the share of maximum risk reduction achieved by ASL-3 for all models\u2014is likely biased upward. If GPAI\u2019s defensive contributions are small relative to the misuse risk, the bias is minor. If they are large, the bias could be substantial. We did not directly elicit respondents\u2019 views on the magnitude of GPAI\u2019s defensive benefits, though several rationales mentioned them. For instance, some respondents noted that GPAI could accelerate the development of medical countermeasures or improve biosurveillance. Future work could address this limitation by eliciting forecasts under a \u2018perfect safeguards\u2019 scenario directly\u2014one where GPAI exists and provides defensive benefits but cannot be misused\u2014or by asking respondents to separately estimate the defensive and offensive contributions of GPAI to outbreak risk.<\/p>\n\n\n\n<h4 id=\"insights-from-rationales\" class=\"wp-block-heading\">Insights from rationales<sup class=\"fn\" data-fn=\"ad5737ad-a5b7-4e0e-b1e2-5892ca002e62\"><a href=\"#ad5737ad-a5b7-4e0e-b1e2-5892ca002e62\" id=\"ad5737ad-a5b7-4e0e-b1e2-5892ca002e62-link\">1<\/a><\/sup><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Aligning with quantitative results, text rationales showed that most respondents believed comprehensive ASL-3 protection would meaningfully reduce risk relative to the helpful-only scenario. There was broad agreement<\/strong> that ASL-3\u2019s greatest value is in blocking non-expert and lone-wolf actors. Many respondents were reassured by the difficulty of finding jailbreaks and the pace of jailbreak patching. Several noted the deterrence effect of safeguards. On the other hand, multiple respondents noted that ASL-3 does nothing to reduce accidental releases, and one argued that ASL-3 would increase risk by weakening collective defensive capabilities.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cWith safeguards in place, the risk of a human-caused outbreak should be around the same as with no GPAI models around at all. In particular, the safeguards should significantly reduce the likelihood that an untrained non-expert would use an AI system to create a dangerous pathogen.\u201d<\/p>\n<\/blockquote>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cGiven the available information, it doesn\u2019t seem like the ASL-3 security standard is a panacea. The patching of jailbreaks is slower than I thought. &#8230; It seems rather unbalanced that only $26k are awarded to an individual hacker for discovering a bug that, not only could cost the company millions (plus the potential negative publicity), but also requires so much effort to fix. There\u2019s probably a black market in the dark web for these universal jailbreaks, and I wouldn\u2019t be surprised if the pay is competitive.\u201d<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Most respondents viewed the helpful-only (\u2018no safeguards\u2019) scenario as increasing biorisk, but physical bottlenecks\u2014not information access\u2014were seen as the primary constraint on bioweapons development. Many cited the Active Site s<\/strong>tudy [10] \u2014 which found that helpful-only LLM access didn\u2019t aid participants much with practical virology tasks relative to internet-only access \u2014 and stressed that access to labs and equipment, along with crucial hands-on skills, would remain the real constraints. Others, however, weren\u2019t so sanguine, arguing that current frontier models can produce detailed technical guidance and that stripping away refusal behaviors makes that knowledge accessible to a much wider set of actors. Several challenged the relevance of the Active Site study, noting it used older models, and identified less direct pathways to increased risk: more people being inspired to try, lab accidents driven by AI-enabled overconfidence, and the volatile mix of helpful-only models with ongoing geopolitical conflicts.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cActive Site\u2019s recent work on this topic showed that overall LLM-aided participants performed fairly similarly to internet-only in a pseudo-virological workflow. This also fits my strong prior, having worked in laboratory science\/bio for 7 years, that converting theoretical knowledge and\/or text-based protocols into actual physical actions in the real world is most efficiently and accurately performed with real-time human expert guidance\u2026Until we either have AI-powered VR lab training or completely closed lab-in-the-loop type setups for the full stack of virology\/bacteriology workflows, a 2026-level LLM isn\u2019t going to aid the lone-wolf DIY bio actor significantly over and above the Internet.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cRecent studies have shown that the bottleneck is not the AI. It is the lab equipment and more importantly the implicit knowledge of how to do protocols\u2026Conceptually this is similar to the nuclear field. Although access to the core material on how to create a nuclear bomb is widely available\u2026the practical knowledge and tooling has proven to be a barrier for all but the most determined nation states.\u201d<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Lab accidents were often identified as the most likely source of human-caused outbreaks and largely independent of AI conditions. Some respondents poi<\/strong>nted to the proliferation of BSL-3\/4 facilities globally, ongoing gain-of-function research, and variable biosafety standards as key drivers of lab accidents.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cLaboratory accident risk reflects global BSL-3\/4 infrastructure expansion\u2014approximately 70 BSL-4 and 1,800 BSL-3 facilities [outside of the U.S.] operational, with substantial growth in China, India, and developing nations with variable biosafety and biosecurity standards (mostly substandard).\u201d<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Much of the variation in baseline forecasts of expected damages may have been driven by which historical incidents respondents counted when assessing a base rate for human-caused outbreaks. The 2001 anthrax attacks served<\/strong> as a common anchor, but beyond that, the incidents considered diverged. The breadth of each respondent\u2019s historical audit was a strong predictor of their final base rate, with a wider breadth being associated with a higher base rate.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cWhen I look at the incidents post anthrax (600M estimated damages), I see several incidents that would also have likely crossed the 100M threshold given the way damages are calculated\u2026Claude estimates (in 2026 USD) that the SARS lab escapes was 80-360M in damages, UK foot and mouth (escaped from the Pirbright laboratory complex) at 600-900M, Lanzhou Brucellosis at 60-250M, and then there\u2019s COVID-19, which call it 25% chance of lab leak. So past 25 years, 4 confirmed incidents that probably hit the 100M-1B bin given the method of counting damages, and a 25% chance that there was an incident that hit the 10T-100T bin. All told, a higher base rate than I thought I\u2019d find.\u201d\u2019<\/p>\n<\/blockquote>\n\n\n\n<h2 id=\"what-would-be-the-impact-of-more-models-being-protected-by-asl-3\" class=\"wp-block-heading\">4.\nWhat would be the impact of more models being protected by ASL-3?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">It is implausible that all GPAI models (regardless of size or capability) would be protected by ASL-3 safeguards. It seems more realistic for all models at the frontier of capabilities to be protected by ASL-3 or equivalent safeguards, even if that is not the case today. To understand the potential impact of all frontier models being protected by ASL-3 safeguards, we asked respondents to forecast the main outcome under this scenario.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We defined \u201cfrontier\u201d GPAI models as all those performing as well or better than Claude Sonnet 4.5 on the <a href=\"https:\/\/epoch.ai\/benchmarks\/eci?view=graph&amp;tab=release-date\"><u>Epoch Capabilities Index<\/u><\/a> (ECI). The ECI combines scores from more than 40 different AI benchmarks into a single general capability score [9]. At the time of the survey, Claude Sonnet 4.5 scored 147 on the ECI, and 13 models scored at or above that level: Gemini 3 Pro, GPT-5.2, Claude Opus 4.6, Gemini 3 Flash, GPT-5 Pro, Claude Opus 4.5, GPT-5, GPT-5.1, o3-pro, Kimi K2.5, Grok 4, o3, and Claude Sonnet 4.5.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Under the \u2018ASL-3 for frontier models\u2019 scenario, the median probability of human-caused outbreaks occurring between now and the end of 2028 and leading to at least $100 million in damages was 10.6% (95% CI: 5.5%, 18.1%). Compared to the \u2018no safeguards\u2019 scenario, the \u2018ASL-3 for frontier models\u2019 scenario was associated with an 18% (95% CI: 9%, 30%) reduction in this risk. The median respondent thought that \u2018ASL-3 for frontier models\u2019 only would be associated with a 20% (95% CI: 9%, 52%) reduction in expected damages due to large-scale human-caused outbreaks relative to the \u2018no safeguards\u2019 scenario (see <a href=\"#fig-03\" id=\"#fig-03\">Figure <u>3<\/u><\/a>). ASL-3 for frontier models mitigated 36.3% of GPAI-attributable risk (see <a href=\"#fig-04\" id=\"#fig-04\">Figure <u>4<\/u><\/a>). Most respondents placed the \u2018ASL-3 for frontier models\u2019 scenario\u2019s risk between \u2018ASL-3 for all models\u2019 and \u2018no safeguards\u2019, but disagreed on where, with some treating it as nearly equivalent to \u2018ASL-3 for all models\u2019 and others as only marginally better than \u2018no safeguards\u2019.<\/p>\n\n\n\n<h4 id=\"insights-from-rationales-1\" class=\"wp-block-heading\">Insights from rationales<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Forecasters disagreed on the degree to which powerful near-frontier models\u2014particularly Chinese ones\u2014would undermine frontier-only ASL-3 coverage. In their text rati<\/strong>onales, some emphasized that frontier models are where the real capability uplift lies, with some citing the Active Site study as evidence that unrestricted models would likely provide limited practical uplift for virology tasks. Some also noted that near-frontier models would likely retain some proprietary safeguards even without ASL-3. Others worried that a community of practice could develop around non-frontier models and that the mere knowledge of capable unsafeguarded models could lower the psychological barrier to attempting an attack.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cEven if they don\u2019t have ASL-3 safeguards, the American models still have quite good safeguards. Unfortunately, this is not the case for most of the Chinese models though. So I think in this world, the easiest path to harm for someone with malicious intent would be to use a Chinese LLM to help them.\u201d<\/p>\n<\/blockquote>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">\u201cMy guess here, which is strongly influenced by the Active Site study conducted in the summer of 2025, is that no uplift will be provided by models on the list that currently aren\u2019t classified as frontier.\u201d<\/p>\n<\/blockquote>\n\n\n\n<h2 id=\"how-would-variations-in-asl-3-influence-this-impact\" class=\"wp-block-heading\">5. How\nwould variations in ASL-3 influence this impact?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We asked respondents to consider the impact of variations on ASL-3 in a scenario where frontier models are protected by ASL-3. Respondents were asked to forecast the main outcome conditional on frontier models being protected by the following hypothetical variations on the ASL-3 scenario:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Time required to identify jailbreaks<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>A1: The average time required to find a universal jailbreak is 5x\nlonger<\/li>\n\n\n\n<li>A2: The average time required to find a universal jailbreak is 5x\nshorter<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Time to jailbreak patching<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>B1: It takes a fifth (0.2x) of the time to patch all universal\njailbreaks<\/li>\n\n\n\n<li>B2: It takes up to 2x as long to patch all universal\njailbreaks<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Capabilities loss associated with jailbreaks<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>C1: All universal jailbreaks significantly reduce the model\u2019s\nGPQA score. Compared to jailbroken models\u2019 current performance relative\nto the &#8220;no jailbreak&#8221; baseline, jailbroken models\u2019 relative performance\ndecreases by roughly 30%.<\/li>\n\n\n\n<li>C2: All universal jailbreaks only minimally reduce the model\u2019s\nGPQA score. Compared to jailbroken models\u2019 current performance relative\nto the &#8220;no jailbreak&#8221; baseline, jailbroken models\u2019 relative performance\nincreases by roughly 30%.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"#fig-05\" id=\"#fig-05\">Figure <u>5<\/u><\/a> shows the median change in expected damages relative to \u2018ASL-3 for frontier models\u2019 under the variant scenarios. For each of the scenarios associated with improvement from current ASL-3 safeguards, the median expected change was similar to that associated with \u2018ASL-3 for all models\u2019. The largest median change\u2014an 8% (95% CI: 5%, 18%) decrease from the \u2018ASL-3 for frontier models\u2019 scenario\u2014was associated with a faster time to jailbreak patching (B1).<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-05\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-05.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 5:<\/strong> Relative difference in expected financial damages under each variant of ASL-3 and the \u2018ASL-3 for all models\u2019 scenario. Numbers show medians from each group. Values below 1 represent a decrease in expected damages compared to a scenario with ASL-3 for frontier models.<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\"><a href=\"#fig-06\" id=\"#fig-06\">Figure 6<\/a> shows the share of GPAI-attributable risk mitigated by the variant scenarios for frontier models, \u2018ASL-3 for frontier models\u2019, and \u2018ASL-3 for all models\u2019 scenarios. For the median respondent, the \u20185x faster patching\u2019 (B1) and \u201830% decrease in relative performance\u2019 (C1) scenarios were seen as comparable to the \u2018ASL-3 for all models\u2019 scenario on this metric. Relative to superforecasters, experts expected variants to ASL-3 to have a greater impact.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">More details are available in <a href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=40\">Appendix C<\/a>.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-06\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-06.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 6:<\/strong> Share of GPAI-attributable risk mitigated by each ASL-3 scenario. GPAI-attributable risk is defined as the multiplicative gap in expected damages between the \u2018no safeguards\u2019 and \u2018no GPAI\u2019 scenarios. Numbers show medians from each group and black lines show their bootstrapped 95% confidence intervals.<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 id=\"insights-from-rationales-2\" class=\"wp-block-heading\">Insights from rationales<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Several respondents expressed concern that if non-frontier models remain unprotected and are nearly as capable, then making frontier jailbreaks harder, faster to patch, or more capability-degrading has limited marginal value.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Most respondents agreed that making jailbreaks harder to find would help, but several argued that it would mainly screen out under-resourced actors who were unlikely to succeed anyway. Many also noted that more organi<\/strong>zed groups would be able to adapt. Making jailbreaks easier to find was generally assessed to be more dangerous because it would expand access to a much larger pool of actors, however unsophisticated. A few respondents noted the interaction of this variant with patching: discovery time is largely irrelevant if jailbreaks persist for months once found. One flagged that AI agents capable of autonomous software research could erode the barrier further.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Forecasters noted that faster patching would more effectively disrupt the months-long efforts required to develop a bioweapon. Credible bioweapo<\/strong>n efforts typically involve iterative troubleshooting over weeks or months and that fast patching could disrupt this workflow by closing the window before a project could be completed\u2014and could also reduce the economic incentive to discover and sell jailbreaks. Slow patching was predicted to do the opposite in that it could make jailbreaks more useful and give bad actors ample time to extract everything they need.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Forecasters differed in opinion on the impact of model performance degradation. Some respon<\/strong>dents argued that a significant degradation would eliminate the incentive to pursue jailbreaks, since the degraded model would offer little advantage over unprotected non-frontier alternatives. Others argued that even a significantly degraded frontier model would still be remarkably capable by historical standards.<\/p>\n\n\n\n<h2 id=\"how-do-asl-3-safeguards-influence-threat-models-related-to-biosecurity-risks\" class=\"wp-block-heading\">6.\nHow do ASL-3 safeguards influence threat models related to biosecurity\nrisks?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We asked respondents to answer several questions related to pathways to a large-scale human-caused outbreak. This included the type of actor that is most likely to cause such an event, the probability that unauthorized use of a GPAI model was involved in the event, and the most likely pathway to unauthorized model use.<\/p>\n\n\n\n<h3 id=\"actor-type\" class=\"wp-block-heading\">6.1 Actor type<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We asked what type of actor is the most likely primary cause of a human-caused outbreak that causes more than $100 million in damages, and how this would change depending on whether GPAI models had no safeguards or were all protected by ASL-3 safeguards. The mean responses for each group of participants are shown in <a href=\"#fig-07\">Figure 7<\/a>. Both groups of respondents thought that state actors were the most likely cause of such an event, and that they would be an even more likely cause under the ASL-3 safeguards scenario. Among the expert respondents, the next most likely group was individual expert actors, although this group\u2019s probability of being the primary cause of a large-scale human-caused outbreak fell under the ASL-3 safeguards scenario. Compared to experts, superforecasters generally put a higher probability on a different type of actor (an \u201cOther\u201d category) being the primary cause. Respondents who put a large probability in this category suggested that groups of experts or non-experts that did not count as terrorist groups would fit into this category.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter is-resized\" id=\"fig-07\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-07.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 7:<\/strong> Probability of the given actor being the primary cause of a large-scale human-caused outbreak. Numbers show group averages.<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 id=\"insights-from-rationales-3\" class=\"wp-block-heading\">Insights from rationales<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Across both the \u2018no-safeguards\u2019 and \u2018ASL-3 for all models\u2019 conditions, a dominant theme in respondent rationales was a capability-intent asymmetry. State actors have the r<\/strong>esources to develop bioweapons but face strong disincentives (e.g., attribution risk, self-harm), while non-state actors may have motivation but lack capability. Many respondents assign significant probability to accidental release pathways rather than deliberate attacks, shifting the expected actor profile toward individual experts and state labs. Without safeguards, AI is expected to partially close the capability gap for non-state actors, but practical barriers\u2014such as wet lab access, weaponization, and delivery\u2014remain and are largely unaffected by AI. Some argue that pursuing bioweapons is irrational for most actors given the availability of simpler alternatives, and that the \u201cOther\u201d category captures important scenarios (e.g., corporate negligence, criminal groups, insider threats, etc.) outside the listed actor types.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>ASL-3 safeguards shift relative probability toward better-resourced actors. State ac<\/strong>tors and large terrorist organizations gain conditional share because they can overcome or bypass safeguards through in-house expertise, cyber capabilities, and sustained jailbreak efforts, while individual non-experts\u2014the primary beneficiaries of unrestricted AI\u2014lose the most. Jailbreak difficulty is the key mechanism: successful jailbreaking requires specialized cybersecurity skills (compounding the expertise requirement), jailbreaks are transient due to patching, and even successful jailbreaks may produce degraded outputs that less capable actors cannot independently verify.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>With deliberate misuse pathways partially blocked, the accidental release pathway gains further relative prominence. Several respon<\/strong>dents also flag trusted user exemptions (insiders with legitimate access who misuse it or whose credentials are compromised) as a new attack surface, and that such actors with reduced safeguards may develop over-dependence on AI, increasing the risk of accidental misuse or unintended consequences.<\/p>\n\n\n\n<h3 id=\"unauthorized-model-use\" class=\"wp-block-heading\">6.2 Unauthorized model use<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">When asked about the probability of unauthorized frontier model use contributing to a human-caused outbreak that causes more than $100 million in damages, the median expert forecasted a 30% (95% CI: 15 , 50) probability, and the median superforecaster a 23.5% (95% CI: 10.2, 52) probability. The distribution of these forecasts is shown in <a href=\"#fig-08\">Figure 8<\/a>.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-08\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-08.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 8:<\/strong> Probability of unauthorized use of frontier models, conditional on a large-scale human-caused outbreak occurring. Numbers show group medians and black lines show their bootstrapped 95% confidence intervals.<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">We then asked respondents to assume that unauthorized model use had contributed to a human-caused outbreak causing more than $100 million in damages, and say how likely they thought different pathways were to have contributed to that unauthorized model use. The responses are shown in <a href=\"#fig-09\" id=\"#fig-09\">Figure 9<\/a>. Each of the pathways we asked about had a median probability of 15% or higher, suggesting that most respondents found all pathways plausible. Relative to experts, superforecasters generally thought a trusted user exemption was more likely to be used and the actor generating their own jailbreak was less likely. It is worth noting that these median values hide substantial variation in responses.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\" id=\"fig-09\"><img decoding=\"async\" src=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/technical-report_2026-06-15_impacts-of-asl-3-safeguards-on-biorisks_fig-09.png\" alt=\"\" style=\"width:700px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 9:<\/strong> Distribution of probabilities assigned to each pathway to gain unauthorized access. Numbers show group medians. The pathways were non-exclusive, and each individual was allowed to assign more than 100% across them.<\/figcaption><\/figure>\n<\/div>\n\n\n<h4 id=\"insights-from-rationales-4\" class=\"wp-block-heading\">Insights from rationales<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Whether respondents weighted accidental versus deliberate pathways more heavily strongly predicted their estimates of unauthorized model use. Those who saw accid<\/strong>ents as dominant gave low estimates, since lab leaks were thought to rarely involve AI in the causal chain. Those focused on deliberate attacks gave high estimates, since any motivated attacker would exploit every available tool.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Trusted user exemptions were identified as a critical vulnerability because they exploit human frailties that technical safeguards cannot fully address. Bribery, blackmail, a<\/strong>nd social engineering can compromise exemption holders regardless of how robust the underlying model defenses are. Moreover, exemption holders are precisely those with proximity to pathogens and lab infrastructure, meaning this type of compromise combines AI access with all the physical infrastructure needed for an attack. However, a countervailing view holds that the exemption pool is small, misuse would be highly attributable, and that screening should filter most threats.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Black market and nonpublic jailbreaks are seen as the path of least resistance for actors lacking technical skills. These markets effec<\/strong>tively commoditize safeguard bypass via dark web markets that are harder for labs to monitor and take down than public jailbreaks. Public jailbreaks were seen as the most accessible, but also the most short-lived, limiting their utility for the sustained interaction biological development requires.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Creating novel jailbreaks is the highest-barrier pathway, constrained by a skill mismatch between jailbreaking and biology expertise. As one respondent p<\/strong>ut it, &#8220;most actors who can create a jailbreak, wouldn\u2019t have the biology background to create a serious virus.&#8221; This restricts the route mainly to large organizations with resources to recruit across both domains.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Real-world attacks would likely chain multiple methods rather than relying on one, and that the \u201c<\/strong>Other\u201d category is a catch-all for important vectors like open-weight model manipulation, API vulnerabilities, and \u201cunknown unknowns.\u201d<\/p>\n\n\n\n<h2 id=\"conclusion\" class=\"wp-block-heading\">7. Conclusion<\/h2>\n\n\n\n<h3 id=\"summary-of-findings\" class=\"wp-block-heading\">7.1 Summary of findings<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This study asked 22 domain experts and 20 superforecasters to estimate how ASL-3 safeguards would change the expected financial damages from large-scale human-caused outbreaks between now and 2028. Four findings stand out.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Respondents believe ASL-3 safeguards demonstrate high efficacy. If appl<\/strong>ied to all GPAI models, the median respondent\u2019s forecasts imply that ASL-3 would mitigate roughly 70% of GPAI-attributable risk. This figure likely overstates true efficacy, because the \u2018no GPAI\u2019 benchmark used to define GPAI-attributable risk also removes GPAI\u2019s defensive contributions. Nonetheless, even accounting for this bias, the results suggest that ASL-3 captures a substantial share of the achievable risk reduction when applied.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Protecting only frontier models with ASL-3 reduces risk, but not as much as protecting all GPAI models. Protecting only fr<\/strong>ontier models captured less of the risk reduction than protecting all GPAI models, though respondents disagreed substantially on the size of the gap. This disagreement hinged on how capable and accessible near-frontier models\u2014particularly Chinese open-weight models\u2014are judged to be. The difference underscores that the value of ASL-3 as a mitigation depends not just on its technical performance but on how broadly equivalent safeguards are adopted across the ecosystem.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models were generally expected to improve ASL-3 performance.<\/strong> For the median respondent, the effects of these variants of ASL-3 for frontier models were roughly similar to the effect of ASL-3 for all models. Of the variants we asked about, the largest effect was associated with faster patching, consistent with the logic that bioweapon development requires sustained iterative access.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>ASL-3 shifts the threat landscape toward better-resourced actors and accidental pathways. Respondent<\/strong>s expected ASL-3 to disproportionately screen out non-expert and lone-wolf actors, shifting relative probability toward state actors, well-resourced organizations, and accidental releases from laboratories. Trusted user exemptions were identified as a notable vulnerability, because they combine AI access with proximity to the physical infrastructure needed for an attack and exploit human factors that technical safeguards cannot fully address. Many respondents emphasized that lab accidents represent a substantial share of baseline risk that AI safeguards do not address.<\/p>\n\n\n\n<h3 id=\"limitations\" class=\"wp-block-heading\">7.2 Limitations<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Beyond the caveats noted above, several broader limitations should inform interpretation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Forecasting low-probability, high-consequence events is inherently difficult. We focu<\/strong>s on relative changes across scenarios rather than absolute damage estimates for this reason, but even relative judgments may be unreliable when base rates are poorly constrained. The wide variation in respondents\u2019 baseline forecasts\u2014driven largely by which historical incidents they counted\u2014illustrates this challenge.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The \u2018no safeguards\u2019 baseline is more extreme than the status quo. This scenar<\/strong>io asks respondents to imagine all GPAI models are optimized purely for helpfulness with no safety constraints. In practice, most commercial models retain some safety training even without ASL-3. The measured gap between \u2018no safeguards\u2019 and the ASL-3 scenarios therefore captures the value of ASL-3 above a minimal baseline, not its marginal value above current industry-standard practices.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The capability freeze assumption bounds the shelf life of the findings. The study as<\/strong>ked respondents to assume no improvement in GPAI capabilities through 2028. This isolates views on current-generation safeguards, but it means respondents were evaluating ASL-3 against a threat landscape that will almost certainly change. More capable models may provide greater uplift to malicious actors, may be harder to safeguard, or may shift the balance between information access and physical bottlenecks. These findings are best understood as a snapshot of expert views on ASL-3\u2019s efficacy against today\u2019s models.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The sample is modest and skewed toward Western biosecurity institutions. While the e<\/strong>xpert group had broad experience\u201414 of 22 reported expertise in more than one of biosecurity, national security and terrorism studies\u2014perspectives from researchers in regions with rapidly expanding BSL infrastructure and those facing the highest burden of terrorist attacks are underrepresented.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Little information was provided on ASL-3 Security Standards.<\/strong> Forecasters were asked to consider the effects of ASL-3 as a whole\u2014including both the Deployment Standards and Security Standards. We provided detailed information on the effectiveness of the Deployment Standards but did not have comparable information on the effectiveness of the Security Standards.<\/p>\n\n\n\n<h3 id=\"implications\" class=\"wp-block-heading\">7.3 Implications<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">These results provide some evidence that deployment safeguards of the kind specified by ASL-3 can meaningfully reduce biosecurity risk. Respondents judged that comprehensive ASL-3 coverage would reduce most of the risk added by GPAI models, and they attributed this largely to ASL-3&#8217;s ability to block non-expert actors. To the extent this judgment is accurate, it suggests that investment in deployment-stage safeguards is a worthwhile component of biosecurity risk management. An important caveat is that respondents were evaluating the effectiveness of safeguards for early-2026 models. As such, this should not be taken as evidence that ASL-3 would be similarly effective for substantially more capable systems.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The value of these safeguards depends on how broadly they, or equivalent measures, are adopted. Respondents saw a substantial gap between protecting all GPAI models and protecting only frontier models, and the size of that gap hinged on how capable and accessible near-frontier and open-weight models were judged to be. This suggests that the risk reduction achievable by any single developer is bounded by the safeguards adopted by other developers. This strengthens the case for industry-wide standards and for governance attention to models that fall outside the reach of any individual developer&#8217;s deployment controls.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The findings also provide some insight into views on pathways to improve safeguards. For most respondents, the time required to find a jailbreak, speed of patching jailbreaks, and the performance associated with jailbreaks were all expected to reduce risk. These offer pathways to improving the effectiveness of ASL-3. The speed of patching jailbreaks was generally considered to have the greatest impact on risk. Respondents reasoned that bioweapon development requires sustained, iterative access over weeks or months, so closing jailbreaks quickly can disrupt an effort even after a jailbreak is found. Some respondents suggested that trusted user exemptions may also present a potential weak point in ASL-3, which may be a useful point of intervention.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Finally, the findings indicate that safeguards reshape the threat landscape rather than uniformly suppressing it. As ASL-3 screens out less sophisticated actors, risk concentrates among well-resourced state actors and accidental laboratory releases. Because a substantial share of baseline risk lies outside the influence of AI safeguards, AI-focused measures are best understood as complementary to established biosecurity interventions such as laboratory biosafety and biosecurity, and nucleic-acid synthesis screening.<\/p>\n\n\n\n<h3 id=\"future-directions\" class=\"wp-block-heading\">7.4 Future directions<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Two methodological improvements would most strengthen future iterations of this work. First, introducing a \u2018perfect safeguards\u2019 scenario\u2014where GPAI exists and provides defensive benefits but cannot be misused\u2014would allow a cleaner estimate of ASL-3\u2019s efficacy by disentangling GPAI\u2019s offensive and defensive contributions to outbreak risk. Second, developing an estimate of baseline risk under the status quo\u2014where most, but not all, frontier models have some protections\u2014would give insight into the impact of ASL-3 under real world conditions and the potential benefits of improving frontier model protections from what exists currently. However, developing this baseline would require more detailed information on the protections applied to other frontier models.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It would also be valuable to see how these findings change as model capabilities advance: repeating this elicitation periodically would track whether the efficacy of ASL-3, or future safeguards, holds or erodes as the models it protects become more capable. More broadly, this study demonstrates the feasibility of structured expert elicitation as a tool for evaluating AI safety measures against diffuse, hard-to-observe threats. Similar approaches could be applied to other risk domains covered by ASL-3 or future safety levels.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Notes<\/h2>\n\n\n<ol class=\"wp-block-footnotes\"><li id=\"ad5737ad-a5b7-4e0e-b1e2-5892ca002e62\">Participant rationales quoted in this report are included as originally shared and have not been edited or fact-checked. We do not necessarily endorse the views or information expressed. <a href=\"#ad5737ad-a5b7-4e0e-b1e2-5892ca002e62-link\" aria-label=\"Jump to footnote reference 1\">\u21a9\ufe0e<\/a><\/li><\/ol>\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"btn orange\" href=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/pdf\/impacts-of-asl-3-safeguards-on-biorisks.pdf#page=30\" target=\"_blank\" rel=\"noreferrer noopener\">The Appendix is provided in the full PDF report <svg width=\"7\" height=\"9\" viewBox=\"0 0 7 9\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n  <path d=\"M0.000156283 8.60806L4.22416 4.33606V4.24006L0.000156283 6.10352e-05H1.80816L6.06416 4.28806L1.80816 8.60806H0.000156283Z\" fill=\"#102B23\"\/>\n<\/svg>\n<svg width=\"8\" height=\"10\" viewBox=\"0 0 8 10\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n  <path d=\"M0.601719 8.85794L4.82572 4.58594V4.48994L0.601719 0.249939H2.40972L6.66572 4.53794L2.40972 8.85794H0.601719Z\" fill=\"#102B23\"\/>\n<\/svg><\/a><\/div>\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"A study to better understand the effectiveness of Anthropic&#8217;s AI Safety Level 3 measures at reducing biosecurity risks.","protected":false},"featured_media":2291,"template":"","meta":{"footnotes":"[{\"content\":\"Participant rationales quoted in this report are included as originally shared and have not been edited or fact-checked. We do not necessarily endorse the views or information expressed.\",\"id\":\"ad5737ad-a5b7-4e0e-b1e2-5892ca002e62\"}]"},"research_type":[22,6],"class_list":["post-2265","research","type-research","status-publish","has-post-thumbnail","hentry","research_type-technical-report","research_type-report"],"acf":[],"yoast_head":"<title>Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks &#8211; Forecasting Research Institute<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks &#8211; Forecasting Research Institute\" \/>\n<meta property=\"og:description\" content=\"A study to better understand the effectiveness of Anthropic&#039;s AI Safety Level 3 measures at reducing biosecurity risks.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks\" \/>\n<meta property=\"og:site_name\" content=\"Forecasting Research Institute\" \/>\n<meta property=\"og:image\" content=\"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/illustration_Midjourney_anthropic-asl-3.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"2560\" \/>\n\t<meta property=\"og:image:height\" content=\"1607\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks\",\"url\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks\",\"name\":\"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks &#8211; Forecasting Research Institute\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/forecastingresearch.org\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/illustration_Midjourney_anthropic-asl-3.jpg\",\"datePublished\":\"2026-08-12T15:55:00+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks#primaryimage\",\"url\":\"https:\\\/\\\/forecastingresearch.org\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/illustration_Midjourney_anthropic-asl-3.jpg\",\"contentUrl\":\"https:\\\/\\\/forecastingresearch.org\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/illustration_Midjourney_anthropic-asl-3.jpg\",\"width\":2560,\"height\":1607},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/research\\\/impacts-of-asl-3-safeguards-on-biorisks#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/forecastingresearch.org\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/forecastingresearch.org\\\/#website\",\"url\":\"https:\\\/\\\/forecastingresearch.org\\\/\",\"name\":\"Forecasting Research Institute\",\"description\":\"\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/forecastingresearch.org\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>","yoast_head_json":{"title":"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks &#8211; Forecasting Research Institute","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks","og_locale":"en_US","og_type":"article","og_title":"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks &#8211; Forecasting Research Institute","og_description":"A study to better understand the effectiveness of Anthropic's AI Safety Level 3 measures at reducing biosecurity risks.","og_url":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks","og_site_name":"Forecasting Research Institute","og_image":[{"width":2560,"height":1607,"url":"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/illustration_Midjourney_anthropic-asl-3.jpg","type":"image\/jpeg"}],"twitter_card":"summary_large_image","schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks","url":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks","name":"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks &#8211; Forecasting Research Institute","isPartOf":{"@id":"https:\/\/forecastingresearch.org\/#website"},"primaryImageOfPage":{"@id":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks#primaryimage"},"image":{"@id":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks#primaryimage"},"thumbnailUrl":"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/illustration_Midjourney_anthropic-asl-3.jpg","datePublished":"2026-08-12T15:55:00+00:00","breadcrumb":{"@id":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks#primaryimage","url":"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/illustration_Midjourney_anthropic-asl-3.jpg","contentUrl":"https:\/\/forecastingresearch.org\/wp-content\/uploads\/2026\/06\/illustration_Midjourney_anthropic-asl-3.jpg","width":2560,"height":1607},{"@type":"BreadcrumbList","@id":"https:\/\/forecastingresearch.org\/research\/impacts-of-asl-3-safeguards-on-biorisks#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/forecastingresearch.org\/"},{"@type":"ListItem","position":2,"name":"Forecasting the Impacts of ASL-3 Safeguards on Biosecurity Risks"}]},{"@type":"WebSite","@id":"https:\/\/forecastingresearch.org\/#website","url":"https:\/\/forecastingresearch.org\/","name":"Forecasting Research Institute","description":"","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/forecastingresearch.org\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/research\/2265","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/research"}],"about":[{"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/types\/research"}],"version-history":[{"count":49,"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/research\/2265\/revisions"}],"predecessor-version":[{"id":2529,"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/research\/2265\/revisions\/2529"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/media\/2291"}],"wp:attachment":[{"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/media?parent=2265"}],"wp:term":[{"taxonomy":"research_type","embeddable":true,"href":"https:\/\/forecastingresearch.org\/api\/wp\/v2\/research_type?post=2265"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}