Appendices
8.1. Appendix: All open questions
2: Why pace?
- How much, and in what ways, would more time allow us to better manage various AI risks? What are the bottlenecks to mitigation or adaptation of different AI risks? What factors other than time influence AI risk management? What risk management efforts can be taken now, and which can only be taken once certain AI capability or adoption thresholds are crossed? Related work:
- MacAskill & Moorhouse (2025), Preparing for the Intelligence Explosion Explicitly sorts “grand challenges” by whether they need calendar time, human deliberation, or just more AI.
- Hobbhahn (2025), What’s the short timeline plan? A concrete inventory of which safety measures are ready to deploy now, and which need years of preparation.
- How might pacing become more or less difficult over time? What investments can be made now to preserve optionality over pacing in the future? In particular, under what circumstances does pacing now make future pacing more or less feasible? Related work:
- Rahman (2026), Does Distributed Training Undermine Compute Governance?
- Sastry et al. (2024), Computing Power and the Governance of Artificial Intelligence. Maps compute governance options and their readiness.
- How do actors in this space make decisions about pacing? What evidence do they currently consider and what assumptions do they currently make? What pathways exist for external research to inform such decisions, e.g. in government or lab leadership, and what makes that information transfer more effective?
- METR (2025), Common Elements of Frontier AI Safety Policies
- Which AI risks are most likely to motivate a pacing intervention, now and in the future? Where do different dangerous capabilities sit on the offense/defense balance, and how will that change over time? Related work:
- Garfinkel & Dafoe (2019), How Does the Offense-Defense Balance Scale?
- Shevlane & Dafoe (2020), The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?
- Esvelt (2022), Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics
- How does transparency about capabilities affect coordination? When does common knowledge of research progress intensify or weaken race dynamics, for example by revealing that a rival is close behind or that progress is possible? Related work:
- Bostrom (2017), Strategic Implications of Openness in AI Development
- Armstrong, Bostrom & Shulman (2016), Racing to the Precipice: a Model of Artificial Intelligence Development Better information about rivals’ capabilities can increase danger when teams are close, because it removes uncertainty that induces caution.
- What are the risks and benefits of titration (“deploy it and learn the risks empirically”)? Have past release strategies from frontier labs succeeded in managing known AI risks? What might change in the future risk landscape? Given our uncertainty about the true risks, what criteria should gate the deployment of new frontier models?
- Shevlane et al. (2023), Model Evaluation for Extreme Risks.
3: Pace what?
- How should hazards be translated into covered activity for pacing interventions? Risks we would like to target, such as uncontrolled automation of AI R&D and bioweapon uplift, build up over various stages of AI research, development and deployment; capabilities will initially emerge at some point in training, and we may want to avoid such a point being reached, or we may care more about wider deployment (especially if capabilities have positive use cases we want to preserve). Research should compare candidate boundaries.
- Shevlane et al. (2023), Model Evaluation for Extreme Risks
- Hooker (2024), On the Limitations of Compute Thresholds as a Governance Strategy
- How can pacing thresholds be made specific and yet still cover distributed activity? A pacing intervention targeting a threshold could potentially be circumvented by distributing activities or artifacts such that each sits below the threshold. How can we design aggregation rules and methods to handle cumulative risk from activities that are divided across space, time, processes and entities, without hindering low-risk activities?.
- Seferis & Fist (2026), Detecting Compute Structuring in AI Governance Is Likely Feasible
- Rahman (2026), Does Distributed Training Undermine Compute Governance?
- What practical coverage is sufficient? If we consider the reach available through company control, infrastructure providers and national rules, including their supply-chain effects, can we estimate bounds on activities that would be effectively covered by an intervention and relevant activities that would be missed?
- Koopmanschap & Barten (2026), How to Catch a GPU. Maps issues with enforcement coverage as dangerous capabilities come to require progressively less compute
- Egan & Heim (2023), Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers
- How effective are different access restrictions once a dangerous capability has been released? How do factors such as access guardrails, alignment training, access to inference compute, ease-of-use, and tacit knowledge affect risk once diffusion has already occurred?
- Tamirisa et al. (2024), Tamper-Resistant Safeguards for Open-Weight LLMs. Identifies limitations in how reliable safeguards can be for open weight models.
- Ord (2025), Inference Scaling Reshapes AI Governance. Identifies inference scaling as a lever for released models.
- How can we permit exceptions to allow useful work, without being so permeable that it makes the rule useless? In the case of compute controls, it now seems technologically feasible, to some extent, to identify what uses a GPU is being put to. What other technical advances can allow interventions to be less blunt and more narrowly scoped?
- Gargiulo & Kulp (2026), Workload Identification with Physical Side Channels for AI Governance
- What are the tradeoffs between verification and invasiveness for different interventions? How can we push the frontier forward?
- Scher & Thiergart (2024), Mechanisms to Verify International Agreements About AI Development. Gives an overview of different verification mechanisms and what they require.
- Petrie et al. (2025), Flexible Hardware-Enabled Guarantees for AI Compute. Proposes verifying compliance without exposing sensitive information about AI development.
4: Pace how?
- How do the incentives of bound parties change across the lifecycle? To what extent can different actors reliably predict the behavior of other actors throughout the lifetime of a pacing intervention? How load-bearing are these predictions of behavior going to be for coordination of pacing interventions?
- Finke (2026), International Agreements to Limit Frontier AI: Objectives and Exit.
- Koremenos (2005), Contracting around International Uncertainty.
- Goldstein and Salib (2025), How to Stop an AI Arms Race.
- How does information sharing impact the credibility of coordinated pacing? Which information, in what granularity, through which channels, by which actors, matter most for credible coordinated pacing? How can this be kept compatible with national security considerations, commercial confidentiality, and cybersecurity? To what extent can this information sharing be kept robust to manipulation?
- Wasil et al. (2024), Verification Methods for International AI Agreements.
- Scher et al (2025), An International Agreement to Prevent the Premature Creation of Artificial Superintelligence.
- What evidence could legitimize a speedy initiation of pacing? What is likely to be the acceptable tolerance for unreliable evidence for different actors? Can scenarios and thresholds be specified in advance to a level of specificity that would garner coordinated buy-in to a rapid pacing onset? Can evidential and assessment processes be agreed on in advance?
- Karnofsky (2024), If-Then Commitments for AI Risk Reduction
- How could dry-run simulations inform and prepare for pacing interventions? What aspects of simulation design, delivery and follow-up affect their effectiveness and impact? Could simulations harm or misguide pacing interventions?
- Gruetzemacher et al. (2024), Strategic Insights from Simulation Gaming of AI Race Dynamics
- Bartels (2020), Building Better Games for National Security Policy Analysis
- How will our reliance on different sources of information about risks, and our methods for communication and coordination, change as AIs become more capable, autonomous, or integrated?
- Clymer et al. (2024), Safety Cases: How to Justify the Safety of Advanced AI Systems
- What does the space of exit scenarios look like? For example, small-scale interventions may allow for immediate release, whereas exits from ongoing large-scale interventions impacting multiple facets of the economy and society (e.g. export controls) may need to be staged. Other than staging, what other parameters of exits exist, and how could they be tuned?
- What is the relationship between initiation and exit triggers? Intuitively, the exit trigger should track whatever justified the intervention in the first place, but what could affect that link and what other factors should be considered?
- Cârlan et al. (2024), Dynamic Safety Cases for Frontier AI
- Karnofsky (2024), If-Then Commitments for AI Risk Reduction
5: Then what?
- What compensation schemes could make pacing interventions more desirable for actors who stand to lose financially from them? What are the precedents for such compensation, how could they be funded, and what secondary impacts might they have?
- Srivastav & Zaehringer (2024), The Economics of Coal Phaseouts
- Jobst Heitzig, Lessman & Zou (2018) Self-enforcing strategies to deter free-riding in the climate change mitigation game and other repeated public good games
- Which restrictions build what kinds of overhangs? How are these overhangs likely to play out if realized, and how dangerous might they be? Can overhangs be addressed through complementary policies? Are there pacing interventions which do not build up an overhang?
- Belrose (2023), AI Pause Will Likely Backfire. Looks at the negative case for a training pause, where a compute overhang leads to rapid progress.
- Which actors gain relative power under different interventions, and what are the expected consequences? What is the historical track record of uses and abuses of power when an activity comes under deliberate pacing intervention?
- Coe & Vaynman (2015), Collusion and the Nuclear Nonproliferation Regime. Looks at how nuclear nonproliferation entrenched superpower influence.
- Cassata & de Chadarevian (2025), Asilomar Across the Atlantic. Restrictions on recombinant-DNA research empowered certain scientific organizations.
- What safeguards can be deployed to guard against mission creep, where regulators or newly empowered authorities could gain power beyond what was intended and become hard to dislodge?
- Romano & Levin (2021), Sunsetting as an Adaptive Strategy
- Molloy (2021), Approach with Caution: Sunset Clauses as Safeguards of Democracy?
8.2. Appendix: longlist of pacing interventions
For concreteness, the following attempts to list the levers we have available to pace AI. Note that a lever’s inclusion here is not an argument in favor of acting on it.
One simple task for the pacing field is to have serious up-to-date research on each of the following levers, and to then model the dependencies and tensions between individual levers.
Compute → dangerous capabilities
- Cap training FLOPs per run (Heim & Koessler 2024; Calero-Forero 2026; Scher et al 2025)
- Require pre-registration and notice for planned training runs above a threshold (EO 14110; SB 53)
- Aggregation rules: make multi-cluster and distributed runs count toward the cap (Shavit 2023; Heim & Koessler 2024)
- Cap on total R&D compute per organization per year (Calero-Forero 2026)
- Cap on the RL share of total training compute (Irpan 2024)
- Minimum ratio of monitoring compute per inference compute (AI Futures Project 2026, Achiam 2026)
- Minimum ratio of safety spending per training compute (AI Futures Project 2026)
- Tax R&D compute above a threshold (Calero-Forero 2026)
- Regular reporting of each lab’s compute split into final runs, experiments, internal inference, and external inference (Epoch 2026; EO 14110).
Chips → compute → dangerous capabilities
- Export-control performance threshold for accelerators (BIS AC/S rule)
- Export controls on lithography, HBM and advanced packaging (BIS SME rule, Oct 2023; BIS rule 2024)
- Chip registry: serial numbers and owner of record above threshold, smuggling penalties (Sastry et al. 2024; Fist & Grunewald 2023)
- Location verification on accelerators (Brass & Aarne 2024; Chip Security Act, S.1705)
- Hardware-enabled mechanisms on new chips: offline licensing, metering, attestation (Aarne, Fist & Withers 2024; Kulp et al. 2024; FlexHEG 2025)
Datacenters (installed chips) → compute → dangerous capabilities
- Permit review threshold for datacenters (Sanders)
- Grid interconnect queue (Epoch 2024)
- Registry of datacenters with satellite-verified construction status (Epoch)
Data → dangerous capabilities
- Disclosure of synthetic-data share and RL environments used in frontier training (EU AI Act)
Algorithms → effective compute → dangerous capabilities
- Publication embargo on frontier algorithmic results (Bostrom 2017; Shevlane & Dafoe 2020)
- Total Research Transparency: mandatory disclosure of all frontier research (AI Futures Project 2026)
- Internal model transparency: mandatory sharing of all internal models with researchers at other labs (Tadepalli 2026).
- Structured access to code and checkpoints via vetted institutions (Shevlane 2022)
- Classification regime for capability-elicitation techniques (Shevlane & Dafoe 2020)
Talent → algorithms → dangerous capabilities
- Visa quota and processing time for frontier researchers (Zwetsloot et al. 2019, CSET)
- Vetting and cooling-off periods for staff with weight or R&D compute access (Nevo et al. 2024)
Capital → compute → dangerous capabilities
- Compute tax (dollar per FLOP on training runs above a threshold) (Calero-Forero 2026)
- Strict liability for catastrophic harms from frontier models (Weil 2024)
- Mandatory liability insurance, premiums priced by capability tier (Trout 2024)
- Investor-facing AI risk disclosure (SEC 2023)
AI R&D capabilities → algorithms → effective compute → dangerous capabilities
- Fraction of R&D compute consumed by autonomous agents (Stix et al. 2025; Charnock et al 2026)
- AI R&D speed-up trigger in safety frameworks (Anthropic RSP; DeepMind FSF; OpenAI Preparedness)
- Human review ratio for AI-written research code and experiment plans (Stix et al. 2025)
- Safety case required before internal deployment on the R&D stack (Clymer et al. 2024; Stix et al. 2025)
- Monitoring coverage of internal agent actions (Greenblatt et al. 2023; METR red-team of Anthropic’s internal monitoring, 2026)
Actors and cadence → race intensity → all other levers
- Licensing vs registration for frontier development (Anderljung et al. 2023; June 2026 EO)
- Minimum interval between frontier releases (FLI pause letter 2023)
- Pre-deployment testing window with government access (June 2026 EO; FRONTIER Act)
- Coordinated-pause trigger and duration across signatories (Alaga & Schuett 2023)
- Lead-margin reporting: months between top lab and next (Epoch; Karnofsky 2022)
Dangerous capabilities evals
- Pretraining data filtering for CBRN, cyber-offense and self-replication content (O’Brien et al. 2025)
- Verified unlearning of specified capabilities (Li et al. 2024, Feng et al 2025)
- Capability thresholds by domain (Koessler, Schuett & Anderljung 2024; the old Anthropic RSP)
Generality → dangerous capabilities
- Separate regulatory track for narrow AI, scientific models (Drexler 2019; EU AIA)
- Gating of long-horizon agentic post-training for general models (Chan et al. 2023; Kwa et al. 2025)
Legibility of model reasoning → control of dangerous capabilities
- Codebase / log audits for optimization pressure on chain-of-thought (Korbak et al. 2025; Baker et al. 2025)
- Disclosure and gating of latent reasoning architectures (Hao et al. 2024, Coconut; Korbak et al. 2025)
- Online weight updates in deployment off by default (Greenblatt et al. 2023)
- Persistent memory scope (Shavit et al. 2023, OpenAI)
- Canary-tagging or filtering of safety and eval content in training data (Berglund et al. 2023; Laine et al. 2024)
Weight security → blocking exfiltration
- Mandatory weight security level (“SL”) by capability (Nevo et al. 2024; Anthropic ASL-3)
- Two-person rule and hardware keys for weight access (Nevo et al. 2024)
- Insider-threat program coverage (Nevo et al. 2024)
- Weight retention policy (Anthropic 2025)
Inference → dangerous capabilities
- Per-query reasoning compute cap for the most capable models (Hooker 2024; Ord 2025)
- Token tax (Irwin 2026)
- Deployment tax by capability tier (Calero-Forero 2026)
- Liability allocation between deployer and developer for autonomous services (Weil 2024)
Autonomy → dangerous capabilities
- Agent permission tiers (Shavit et al. 2023, OpenAI; Chan et al. 2024)
- Spend limits per agent and per task (Shavit et al. 2023, OpenAI)
- Human approval for irreversible actions (Shavit et al. 2023, OpenAI; EU AI Act)
- Sub-agent spawn depth and maximum unattended run length (Kwa et al. 2025, METR)
- Agent identifiers and action-log retention (Chan et al. 2024)
- Kill-switch latency requirement (Orseau & Armstrong 2016; Shavit et al. 2023, OpenAI)
- Certification and fleet registration for embodied agents (EU Machinery Regulation 2023/1230)
Multi-agent → dangerous capabilities
- Instance-count reporting per deployment (Chan et al. 2024)
- Agent-to-agent communication logging with steganography checks (Motwani et al. 2024)
Safety research
- Minimum fraction of total compute reserved for safety research, above a total compute threshold (OpenAI 2023)
- Minimum safety headcount as a ratio of total research headcount (AI Lab Watch)
- External safety compute grants (minimum FLOP/year given to independent labs) (NAIRR)
Evaluation
- Third-party evaluator access depth (Casper et al. 2024)
- Elicitation budget per dangerous-capability eval (METR elicitation protocol; Barnett & Thiergart 2024)
- Sandbagging detection protocol (van der Weij et al. 2024)
Transparency
- Required “AI Assurance Level” for developers in the frontier tier (Brundage et al. 2026)
- Embedded auditors running regular audits with non-public access (Brundage et al. 2026)
- Audit scope including internal deployment, information security and safety decision-making (Brundage et al. 2026)
- Number of accredited audit providers and a standards body (Brundage et al. 2026)
- Incident reporting deadline (SB 53 §22757.13)
- Statutory whistleblower channel and protection (SB 53; Right to Warn letter 2024)
Governance response
- Indexing compute thresholds to measured algorithmic progress (Heim & Koessler 2024; Epoch 2025)
- Legislation with automatic clause triggers: e.g. once an eval result is shown, legal obligations come into force (Karnofsky 2024)
- Regulator capacity (CAISI; UK AISI; IFP 2026)
Coordination
- Treaty verifications: chip registry, datacenter inspections, interconnect bandwidth limits (Scher & Thiergart 2024; Baker et al. 2025)
- Training-run declarations exchanged between states (Shavit 2023; Baker et al. 2025)
- Verification R&D budget and a frontier-state incident hotline (Future Society 2026)
8.3. Appendix: Bibliography
Recent
-
Finke (2026), International Agreements to Limit Frontier AI: Objectives and Exit. arXiv
-
Larsen, Dean, Halstead, Lifland, Greenblatt & Kokotajlo (2026). AI 2040: Plan A. AI Futures Project
-
Lifland et al (2026). How to Pace the US Frontier. AI Futures Project.
-
Fist et al (2026). How Should the US Prepare for Increasingly Automated AI R&D?. Institute for Progress.
-
Koopmanschap, Barten (2026), How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements. arXiv
-
Institute for Progress (2026). Funding for CAISI. ifp.org
Foundations
-
Bostrom (2002). Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards. Journal of Evolution and Technology 9. nickbostrom.com
-
Shulman (2009). Arms Control and Intelligence Explosions. ECAP. intelligence.org
-
Marchant, Allenby & Herkert, eds. (2011). The Growing Gap Between Emerging Technologies and Legal-Ethical Oversight: The Pacing Problem. Springer. doi:10.1007/978-94-007-1356-7
-
Armstrong, Bostrom & Shulman (2016). Racing to the Precipice: A Model of Artificial Intelligence Development. AI & Society 31. doi:10.1007/s00146-015-0590-y
-
Christiano (2018). Takeoff Speeds. sideways-view.com
-
Maas (2018). Two Lessons from Nuclear Arms Control for the Responsible Governance of Military Artificial Intelligence. IOS.
-
Koessler, Schuett, Anderljung (2024). Risk thresholds for frontier AI. arXiv.
-
Dafoe (2018). AI Governance: A Research Agenda. GovAI / FHI. governance.ai
-
Kulveit, Douglas, Ammann, Turan, Krueger & Duvenaud (2025). Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development. arXiv:2501.16946
-
Barnett & Scher (2025). AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions. MIRI Technical Governance Team. arXiv:2505.04592
-
Favaro & Clark (2026). When AI Builds Itself. Anthropic Institute. anthropic.com
-
Bostrom (2017). Strategic Implications of Openness in AI Development. Global Policy 8(2). wiley.com, pdf
-
Drexler (2019). Reframing Superintelligence: Comprehensive AI Services as General Intelligence. FHI. ora.ox.ac.uk
-
Garfinkel & Dafoe (2019). How Does the Offense-Defense Balance Scale? Journal of Strategic Studies 42(6). tandfonline.com
-
Shevlane & Dafoe (2020). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse? arXiv:2001.00463
-
Leech et al. (2024). Shallow Review of Technical AI Safety: “Make AI Solve It”. shallowreview.ai
-
MacAskill & Moorhouse (2025). Preparing for the Intelligence Explosion. Forethought. forethought.org
-
Hobbhahn (2025). What’s the Short Timeline Plan? lesswrong.com
Economics
-
Aschenbrenner (2020). Existential Risk and Growth. GPI Working Paper 6-2020. leopoldaschenbrenner.github.io
-
Sandbrink, Hobbs, Swett, Dafoe & Sandberg (2022). Differential Technology Development: An Innovation Governance Consideration for Navigating Technology Risks. SSRN. ssrn.com
-
Jones (2024). The A.I. Dilemma: Growth versus Existential Risk. AER: Insights 6(4). nber.org (WP 31837)
-
Trammell & Aschenbrenner (2024). Existential Risk and Growth. GPI Working Paper 13-2024. philiptrammell.com
-
Salop & Scheffman (1983). Raising Rivals’ Costs. American Economic Review 73(2). repec.org
-
Acemoglu (2002). Directed Technical Change. Review of Economic Studies 69(4). mit.edu
-
Bostrom (2003). Astronomical Waste: The Opportunity Cost of Delayed Technological Development. Utilitas 15(3). nickbostrom.com
-
Bostrom (2005). The Fable of the Dragon-Tyrant. Journal of Medical Ethics 31(5). nickbostrom.com
-
Heitzig, Lessmann & Zou (2011). Self-Enforcing Strategies to Deter Free-Riding in the Climate Change Mitigation Game and Other Repeated Public Good Games. PNAS 108(38). pnas.org
-
Garicano, Lelarge & Van Reenen (2016). Firm Size Distortions and the Productivity Distribution: Evidence from France. American Economic Review 106(11). aeaweb.org
-
Johnson, Shriver & Goldberg (2023). Privacy and Market Concentration: Intended and Unintended Consequences of the GDPR. Management Science 69(10). informs.org
-
Merali (2024). Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Translation. arXiv:2409.02391
-
Srivastav & Zaehringer (2024). The Economics of Coal Phaseouts. arXiv:2406.14238
-
Weil (2024). Tort Law as a Tool for Mitigating Catastrophic Risk from Artificial Intelligence. SSRN
-
Google (2024). AI in Science. ai.google
-
Tomei, Jain & Franklin (2025). AI Governance through Markets. arXiv:2501.17755
-
Gundlach, Lynch, Mertens & Thompson (2025). The Price of Progress: Price Performance and the Future of AI. arXiv:2511.23455
-
Yan & Morck (2025). Who’s Afraid of Tariffs? The Geographic Distribution of Fear and Loss. NBER WP 34299. nber.org
-
Aubakirova, Atallah, Clark, Summerville & Midha (2026). State of AI: An Empirical 100 Trillion Token Study with OpenRouter. arXiv:2601.10088
-
Brynjolfsson, Collis, Eggers, Kazinnik & Nguyen (2026). What is Generative AI Worth? SSRN
-
Demirer, Musolff & Yang (2026). Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools. NBER WP 35275. nber.org
-
Forecasting Research Institute (2026). Experts Forecast Rapid AI Progress Could Bring Health and Wealth Without Happiness. forecastingresearch.substack.com
-
IntuitionLabs (2026). AI-Discovered Drugs in Clinical Trials. intuitionlabs.ai
-
Trout (2024). Insuring Uninsurable Risks from AI: Government as Insurer of Last Resort. arXiv:2409.06672
-
Irwin, Wu & Barez (2026). Position: Token Taxes Can Mitigate AI’s Economic Risks. arXiv:2603.04555
The pause debate
-
Grace (2022). Let’s Think About Slowing Down AI. AI Impacts / LessWrong. lesswrong.com
-
Belrose (2023). AI Pause Will Likely Backfire.
-
Buterin (2023). My Techno-optimizm. vitalik.eth.limo
-
Tallinn (2024). Priorities for AI Risk Reduction. jaan.online
-
Katzke & Futerman (2024). The Manhattan Trap: Why a Race to Artificial Superintelligence Is Self-Defeating. Convergence Analysis. arXiv:2501.14749
-
Larsen, Dean, Halstead, Lifland, Greenblatt & Kokotajlo (2026). AI 2040: Plan A. AI Futures Project. ai-2040.com
-
AI Impacts. Hardware Overhang. aiimpacts.org
-
AI Impacts (2023). Are There Examples of Overhang for Other Technologies? blog.aiimpacts.org
-
Miotti et al. (2024). A Narrow Path. ControlAI. narrowpath.co
-
Felstead (2025). Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks. arXiv:2511.08631
-
Felstead (2026). Can Frontier AI Labs Lawfully Agree to Pause? Lawfare. lawfaremedia.org
-
Employees of frontier AI companies (2026). Pacing the Frontier (open statement). pacingthefrontier.com
-
Karnofsky (2022). Racing through a Minefield: The AI Deployment Problem. Cold Takes. cold-takes.com
-
Future of Life Institute (2023). Pause Giant AI Experiments: An Open Letter. futureoflife.org
-
Alaga & Schuett (2023). Coordinated Pausing: An Evaluation-Based Coordination Scheme for Frontier AI Developers. arXiv:2310.00374
-
Cotra (2026). Total Research Transparency Would Be Nice. Planned Obsolescence. planned-obsolescence.org
International agreements
-
Ho, Barnhart, Trager, Bengio, Brundage, Casovan, Haas, Nemitz, Sastry, Weller, Zhang & Zhang (2023). International Institutions for Advanced AI. arXiv:2307.04699
-
Trager, Harack, Reuel, Carnegie, Heim, Ho, Kreps, Lall, Larter, Ó hÉigeartaigh, Staffell & Villalobos (2023). International Governance of Civilian AI: A Jurisdictional Certification Approach. arXiv:2308.15514
-
Hausenloy, Miotti & Dennis (2023). Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI. arXiv:2310.09217
-
Emery-Xu, Jordan & Trager (2025). International Governance of Advancing Artificial Intelligence. AI & Society 40. doi:10.1007/s00146-024-02050-7
-
Al Ramiah, Koopmanschap, Thorsteinson, Khan, Zhou, Noh, Meindertsma & Shafiq (2025). Toward a Global Regime for Compute Governance: Building the Pause Button. arXiv:2506.20530
-
Scher, Abecassis, Barnett & Abeyta (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. arXiv:2511.10783
-
Finke (2026). International Agreements to Limit Frontier AI: Objectives and Exit. TAIGR @ ICML 2026. arXiv:2607.16224
-
Koremenos (2005). Contracting around International Uncertainty. American Political Science Review 99(4). doi.org
-
Bartels (2020). Building Better Games for National Security Policy Analysis. RAND. rand.org
-
Gruetzemacher et al. (2024). Strategic Insights from Simulation Gaming of AI Race Dynamics. arXiv:2410.03092
-
Goldstein & Salib (2025). How to Stop an AI Arms Race. SSRN
Deterrence
- Hendrycks, Schmidt & Wang (2025). Superintelligence Strategy: Expert Version. arXiv:2503.05628
- Rehman, Mueller, Mazarr et al. (2025). Seeking Stability in the Competition for AI Advantage. RAND commentary. rand.org
- Abecassis (2025). Refining MAIM: Identifying Changes Required to Meet Conditions for Deterrence. MIRI. intelligence.org
- Arnold (2025). Superintelligence Deterrence Has an Observability Problem. AI Frontiers. ai-frontiers.org
- Hendrycks & Khoja (2025). AI Deterrence Is Our Best Option. AI Frontiers. ai-frontiers.org
- Delaney (2025). Crucial Considerations in ASI Deterrence. IAPS. iaps.ai
Verification
-
Brundage et al. (2020). Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims. arXiv:2004.07213
-
Baker (2023). Nuclear Arms Control Verification and Lessons for AI Treaties. arXiv:2304.04123
-
Scher & Thiergart (2024). Mechanisms to Verify International Agreements About AI Development. MIRI Technical Governance Team. arXiv:2506.15867
-
Wasil, Reed, Miller & Barnett (2024). Verification Methods for International AI Agreements. arXiv:2408.16074
-
Koopmanschap & Barten (2026). How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements. Existential Risk Observatory. arXiv:2607.22619
-
Choussat & Khoja (2026). An International AI Slowdown Is Ready Whenever Politicians Are. AI Frontiers. ai-frontiers.org
-
Baker, Kulp, Marks, Brundage & Heim (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. arXiv:2507.15916
-
The Future Society (2026). How To Make International AI Verification a Reality. thefuturesociety.org
Compute governance
-
Shavit (2023). What Does It Take to Catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv:2303.11341
-
Egan & Heim (2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv:2310.13625
-
Sastry, Heim, Belfield, Anderljung, Brundage, Hazell et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv:2402.08797
-
Heim, Fist, Egan, Huang, Zekany, Trager, Osborne & Zilberman (2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. Oxford Martin AIGI. oxfordmartin.ox.ac.uk
-
Aarne, Fist & Withers (2024). Secure, Governable Chips. CNAS. cnas.org
-
Kulp, Gonzales, Smith, Heim, Puri, Vermeer & Winkelman (2024). Hardware-Enabled Governance Mechanisms. RAND WR-A3056-1. rand.org
-
Petrie, Aarne, Ammann & Dalrymple (2024). Interim Report: Mechanisms for Flexible Hardware-Enabled Guarantees. Part I, arXiv:2506.15093
-
Brass & Aarne (2024). Location Verification for AI Chips. IAPS. iaps.ai
-
Heim & Koessler (2024). Training Compute Thresholds: Features and Functions in AI Regulation. arXiv:2405.10799
-
Hooker (2024). On the Limitations of Compute Thresholds as a Governance Strategy. arXiv:2407.05694
-
Ord (2025). Inference Scaling Reshapes AI Governance. arXiv
-
USITC (2023). Germanium and Gallium (Executive Briefing on Trade). usitc.gov
-
Ho, Besiroglu, Erdil et al. (2024). Algorithmic Progress in Language Models. Epoch AI. arXiv:2403.05812
-
Miller (2025). How US Export Controls Have (and Haven’t) Curbed Chinese AI. AI Frontiers. ai-frontiers.org
-
O’Gara, Kulp, Hodgkins, Petrie et al. (2025). Hardware-Enabled Mechanisms for Verifying Responsible AI Development. arXiv:2505.03742
-
Somala, Ho & Krier (2025). Three Challenges Facing Compute-Based AI Policies. Epoch AI. epochai.substack.com
-
Ansari (2026). Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification. arXiv:2604.04712
-
Fedasiuk & Torres (2026). The Lithography Loophole: How China Is Printing Its Way to Chip Self-Sufficiency. AEI. aei.org
-
Rahman (2026). Does Distributed Training Undermine Compute Governance? arXiv:2605.29359
-
Seferis & Fist (2026). Detecting Compute Structuring in AI Governance Is Likely Feasible. AAAI. aaai.org
-
Gargiulo & Kulp (2026). Workload Identification with Physical Side Channels for AI Governance. arXiv:2609.00309
-
Fist & Grunewald (2023). Preventing AI Chip Smuggling to China. CNAS. cnas.org
-
Irpan (2024). Late Takes on OpenAI o1. Sorta Insightful. alexirpan.com
-
Epoch AI (2024). Can AI Scaling Continue Through 2030? epoch.ai
-
Cottier & Owen (2025). How Many AI Models Will Exceed Compute Thresholds? Epoch AI. epoch.ai
-
Epoch AI. Data on Data Centers. epoch.ai
-
Denain & Wu (2026). Final Training Runs Account for a Minority of R&D Compute Spending. Epoch AI, Gradient Updates. epoch.ai
-
Ho (2026). Keeping Up with the GPTs. Epoch AI, Gradient Updates. epoch.ai
-
Calero-Forero (2026). How Should You Slow Down AI Progress, If It Becomes Necessary? lesswrong.com, substack
-
Achiam (2026). On monitoring-to-inference compute ratios. x.com
Law and regulation
- Zwetsloot, Dunham, Arnold & Huang (2019). Keeping Top AI Talent in the United States. CSET. cset.georgetown.edu
- European Union (2023). Machinery Regulation (EU) 2023/1230. eur-lex.europa.eu
- The White House (2023). Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. Federal Register. federalregister.gov
- Bureau of Industry and Security (2023). Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use. Federal Register. federalregister.gov
- Bureau of Industry and Security (2023). Export Controls on Semiconductor Manufacturing Items. Federal Register. federalregister.gov
- Bureau of Industry and Security (2024). Foreign-Produced Direct Product Rule Additions and Refinements to Controls for Advanced Computing. Federal Register. federalregister.gov
- SEC (2023). SEC Adopts Rules on Cybersecurity Risk Management, Strategy, Governance, and Incident Disclosure by Public Companies. Press release 2023-139. sec.gov
- Anderljung et al. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv:2307.03718
- European Union (2024). Artificial Intelligence Act. Article 14, Article 53, Code of Practice
- California Legislature (2025). SB 53: Transparency in Frontier Artificial Intelligence Act. leginfo.legislature.ca.gov
- U.S. Congress (2025). Chip Security Act, S.1705, 119th Congress. congress.gov
- Sanders & Ocasio-Cortez. Sanders, Ocasio-Cortez Announce AI Data Center Moratorium Act (press release). sanders.senate.gov
- NIST. Center for AI Standards and Innovation (CAISI). nist.gov
- UK AI Security Institute. aisi.gov.uk
- National Artificial Intelligence Research Resource Pilot. nairrpilot.org
- Foley Hoag (2026). Trump’s New AI Frontier: The Executive Order Regulating Frontier AI Models. foleyhoag.com
- Statt (2026). The FRONTIER Act: Federal AI Regulation in 2026. statt.com
Developer commitments
-
Shevlane, Farquhar, Garfinkel, Phuong, Whittlestone, Leung et al. (2023). Model Evaluation for Extreme Risks. arXiv:2305.15324
-
Clymer, Gabrieli, Krueger & Larsen (2024). Safety Cases: How to Justify the Safety of Advanced AI Systems. arXiv:2403.10462
-
Karnofsky (2024). If-Then Commitments for AI Risk Reduction. Carnegie Endowment. carnegieendowment.org
-
Cârlan, Gomez, Mathew, Krishna, King, Gebauer & Smith (2024). Dynamic Safety Cases for Frontier AI. arXiv:2412.17618
-
Christiano (2023). Thoughts on Responsible Scaling Policies and Regulation. Alignment Forum. alignmentforum.org
-
OpenAI (2025). Expanding on What We Missed with Sycophancy. openai.com
-
OpenAI (2025). How We Think About Safety and Alignment. openai.com
-
METR (2025). Common Elements of Frontier AI Safety Policies. metr.org
-
Anthropic (2026). Policy on the AI Exponential. anthropic.com
-
Anthropic (2026). Introducing Claude Fable 5 and Claude Mythos 5. anthropic.com
-
Anthropic (2026). Project Glasswing. anthropic.com
-
Anthropic (2026). Project Glasswing: Initial Update. anthropic.com
-
OpenAI (2026). Trusted Access for Cyber. openai.com
-
OpenAI (2026). Pacing Model Development for Cyber Capabilities. openai.com
-
OpenAI (2026). The AI Policy Window. openai.com
-
OpenAI (2026). GPT-6 Astra Deployment Safety Report: “Monitorability”. deploymentsafety.openai.com
-
OpenAI (2026). Research Acceleration: A View Inside OpenAI. openai.com
-
OpenAI (2023). Introducing Superalignment. openai.com
-
Anthropic. Responsible Scaling Policy (updates). anthropic.com
-
Google DeepMind (2024). Introducing the Frontier Safety Framework. deepmind.google
-
OpenAI. Preparedness Framework. openai.com
-
Right to Warn (2024). A Right to Warn about Advanced Artificial Intelligence (open letter). righttowarn.ai
-
Anthropic (2025). Activating AI Safety Level 3 Protections. anthropic.com
-
Anthropic (2025). Commitments on Model Deprecation and Preservation. anthropic.com
-
Stein-Perlman. AI Lab Watch. ailabwatch.org
Evaluations and forecasting
-
Brown et al. (2020). Language Models are Few-Shot Learners. NeurIPS. arXiv:2005.14165
-
Wei et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS. arXiv:2201.11903
-
Dell’Acqua et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. HBS Working Paper 24-013. SSRN
-
UK AI Security Institute (2024). Pre-Deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet. aisi.gov.uk
-
Pimpale, Højmark, Scheurer & Hobbhahn (2025). Forecasting Frontier Language Model Agent Capabilities. arXiv:2502.15850
-
Needham, Edkins, Pimpale, Bartsch & Hobbhahn (2025). Large Language Models Often Know When They Are Being Evaluated. arXiv:2505.23836
-
Bean et al. (2025). Measuring What Matters: Construct Validity in Large Language Model Benchmarks. arXiv:2511.04703
-
Wei & Heim (2025). Designing Incident Reporting Systems for Harms from General-Purpose AI. AAAI. arXiv:2511.05914
-
UK AI Security Institute (2025). More Compute, More Capability: Why AI Agent Evals Need to Account for Test-Time Compute. aisi.gov.uk
-
Epoch AI. AI Chip Production (data insight). epoch.ai
-
Epoch AI. Benchmarks: Epoch Capabilities Index. epoch.ai
-
Epoch AI (2026). An Update on AI’s Most Important Number. Gradient Updates. epoch.ai
-
Mengesha et al. (2026). A Pragmatic Classification Framework for AI Incident Monitoring. arXiv:2604.21412
-
Barrett et al. (2026). Lessons from External Review of DeepMind’s Scheming Inability Safety Case. arXiv:2604.21964
-
Guidelight AI Standards (2026). AI Control: An Assessment of Frontier Practices. guidelight.ai
-
METR (2026). Notes on Anthropic Researcher Uplift Estimates. metr.org
-
Casper et al. (2024). Black-Box Access Is Insufficient for Rigorous AI Audits. arXiv:2401.14446
-
Barnett & Thiergart (2024). Declare and Justify: Explicit Assumptions in AI Evaluations Are Necessary for Effective Regulation. arXiv:2411.12820
-
van der Weij et al. (2024). AI Sandbagging: Language Models Can Strategically Underperform on Evaluations. arXiv:2406.07358
-
METR. Guidelines for Capability Elicitation. Autonomy Evals Guide. metr.github.io
-
Brundage et al. (2026). Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies. arXiv:2601.11699
Safeguards and model security
-
Esvelt (2022). Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics. GCSP. gcsp.ch
-
Kirk et al. (2023). Understanding the Effects of RLHF on LLM Generalisation and Diversity. arXiv:2310.06452
-
Nevo et al. (2024). Securing AI Model Weights. RAND RR-A2849-1. rand.org
-
Tamirisa et al. (2024). Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv:2408.00761
-
Brent & McKelvey (2025). Contemporary AI Foundation Models Increase Biological Weapons Risk. arXiv:2506.13798
-
Chen, Joshi, Chen, Andriushchenko, Angell & He (2025). Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors. arXiv:2506.10949
-
Marshall et al. (2026). BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation. arXiv:2607.14479
-
Shevlane (2022). Structured Access: An Emerging Paradigm for Safe AI Deployment. arXiv:2201.05159
-
Li et al. (2024). The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning. arXiv:2403.03218
-
O’Brien et al. (2025). Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs. arXiv:2508.06601
-
Feng et al. (2025). Existing Large Language Model Unlearning Evaluations Are Inconclusive. arXiv:2506.00688
Agents, oversight and control
- Orseau & Armstrong (2016). Safely Interruptible Agents. UAI. intelligence.org
- Chan et al. (2023). Harms from Increasingly Agentic Algorithmic Systems. FAccT. arXiv:2302.10329
- Shavit et al. (2023). Practices for Governing Agentic AI Systems. OpenAI. openai.com
- Berglund et al. (2023). Taken Out of Context: On Measuring Situational Awareness in LLMs. arXiv:2309.00667
- Greenblatt, Shlegeris, Sachan & Roger (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv:2312.06942
- Chan et al. (2024). Visibility into AI Agents. FAccT. arXiv:2401.13138
- Motwani et al. (2024). Secret Collusion among AI Agents: Multi-Agent Deception via Steganography. NeurIPS. arXiv:2402.07510
- Laine et al. (2024). Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs. arXiv:2407.04694
- Hao et al. (2024). Training Large Language Models to Reason in a Continuous Latent Space. arXiv:2412.06769
- Baker et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. OpenAI. arXiv:2503.11926
- Kwa et al. (2025). Measuring AI Ability to Complete Long Tasks. METR. arXiv:2503.14499
- Stix et al. (2025). AI Behind Closed Doors: A Primer on the Governance of Internal Deployment. arXiv:2504.12170
- Korbak et al. (2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv:2507.11473
- Charnock et al. (2026). What Should Frontier AI Developers Disclose About Internal Deployments? arXiv:2604.23065
- METR (2026). Red-Teaming Anthropic’s Agent Monitoring. metr.org
Histories and precedents of restraint
- U.S. Naval Institute (1926). The Washington Treaties of 1922. Proceedings 52(5). usni.org
- GlobalSecurity.org. Treaty Cruisers. globalsecurity.org
- Haddon-Cave (2009). The Nimrod Review. HC 1025. gov.uk
- Leveson (2011). The Use of Safety Cases in Certification and Regulation. MIT. mit.edu
- Goitein & Patel (2015). What Went Wrong with the FISA Court. Brennan Center for Justice. brennancenter.org
- Coe & Vaynman (2015). Collusion and the Nuclear Nonproliferation Regime. Journal of Politics 77(4). andrewjcoe.com
- China State Council (2017). A New Generation Artificial Intelligence Development Plan (FLIA translation). flia.org
- Casey (2021). A Reckoning Looms for America’s 50-Year Financial Surveillance System. Cato Journal 41. cato.org
- Molloy (2021). Approach with Caution: Sunset Clauses as Safeguards of Democracy? northumbria.ac.uk
- Romano & Levin (2021). Sunsetting as an Adaptive Strategy. PNAS. ncbi.nlm.nih.gov
- Maas (2022). Paths Untaken: The History, Epistemology and Strategy of Technological Restraint, and Lessons for AI. Verfassungsblog. verfassungsblog.de
- Cassata & de Chadarevian (2025). Asilomar Across the Atlantic. ncbi.nlm.nih.gov
- Columbia Academic Commons. On the environmental costs of the slowed rollout of nuclear power. doi:10.7916/d8-qez9-6m49
- 31 U.S.C. §5324. Structuring Transactions to Evade Reporting Requirement Prohibited. law.cornell.edu
- Wikipedia. Year 2000 Problem. wikipedia.org
News and incidents
- CNBC (2025). Nvidia on Track to Hit Historic $5 Trillion Valuation amid AI Rally. cnbc.com
- The Legal Wire (2025). CAC Launches Special Campaign to Clear Up and Rectify the Abuse of AI Technology. thelegalwire.ai
- STAT News (2026). AI Ambient Scribes Bring Modest Time Savings in Clinical Documentation. statnews.com
- Axios (2026). Anthropic’s Revenue Growth. axios.com
- CNBC (2026). Micron Reaches a Trillion-Dollar Market Cap. cnbc.com
- The Guardian (2026). Anthropic Disables Advanced AI Models after US Government Order. theguardian.com
- Observer (2026). Anthropic Delays AI Model It Deems Too Powerful for Public Use. observer.co.uk
- Nature (2026). AI Researchers Reckon with the $1.5 Million ‘Academia Tax’. nature.com
- The Motley Fool (2026). Broadcom Is Less Than 5% from the $2 Trillion Club. fool.com
- Axios (2026). OpenAI’s GPT-5.6 Ban Lifted. axios.com
- Axios (2026). OpenAI Delays Astra Model over Cybersecurity Risks. axios.com
- METR (2026). OpenAI / Hugging Face Incident Investigation. metr.org
- SANS Institute (2026). “The Models Said No”: Inside the Hugging Face Post-Mortem. sans.org
- Toh (2026). The AI Chip Wars’ New Front: Control the Cloud, Not the Silicon. Forbes. forbes.com
- TechCrunch (2026). Group Funded by Andreessen Horowitz and Brockman Plans Data Center Ads to Sway Midterms. techcrunch.com
- Public First Action (2026). Public First Action and Defending Our Values PAC Launch First Ads Supporting Responsible AI Regulation. publicfirstaction.us
- CNBC (2026). OpenAI’s Astra Model Rated “Critical” for Cyber Risk. cnbc.com
- Yahoo Finance (2026). OpenAI Burning $12.3 Billion. yahoo.com
- The Verge (2026). Fable Won’t Answer Basic Biology Questions. theverge.com
- The Wall Street Journal (2026). Anthropic Researcher Quits over Out-of-Control AI Fears. wsj.com