Appendices

8.1. Appendix: All open questions

2: Why pace?

3: Pace what?

  • How should hazards be translated into covered activity for pacing interventions? Risks we would like to target, such as uncontrolled automation of AI R&D and bioweapon uplift, build up over various stages of AI research, development and deployment; capabilities will initially emerge at some point in training, and we may want to avoid such a point being reached, or we may care more about wider deployment (especially if capabilities have positive use cases we want to preserve). Research should compare candidate boundaries.
  • How can pacing thresholds be made specific and yet still cover distributed activity? A pacing intervention targeting a threshold could potentially be circumvented by distributing activities or artifacts such that each sits below the threshold. How can we design aggregation rules and methods to handle cumulative risk from activities that are divided across space, time, processes and entities, without hindering low-risk activities?.
  • What practical coverage is sufficient? If we consider the reach available through company control, infrastructure providers and national rules, including their supply-chain effects, can we estimate bounds on activities that would be effectively covered by an intervention and relevant activities that would be missed?
  • How effective are different access restrictions once a dangerous capability has been released? How do factors such as access guardrails, alignment training, access to inference compute, ease-of-use, and tacit knowledge affect risk once diffusion has already occurred?
  • How can we permit exceptions to allow useful work, without being so permeable that it makes the rule useless? In the case of compute controls, it now seems technologically feasible, to some extent, to identify what uses a GPU is being put to. What other technical advances can allow interventions to be less blunt and more narrowly scoped?
  • What are the tradeoffs between verification and invasiveness for different interventions? How can we push the frontier forward?

4: Pace how?

5: Then what?

8.2. Appendix: longlist of pacing interventions

For concreteness, the following attempts to list the levers we have available to pace AI. Note that a lever’s inclusion here is not an argument in favor of acting on it.

One simple task for the pacing field is to have serious up-to-date research on each of the following levers, and to then model the dependencies and tensions between individual levers.

Compute → dangerous capabilities

  1. Cap training FLOPs per run (Heim & Koessler 2024; Calero-Forero 2026; Scher et al 2025)
  2. Require pre-registration and notice for planned training runs above a threshold (EO 14110; SB 53)
  3. Aggregation rules: make multi-cluster and distributed runs count toward the cap (Shavit 2023; Heim & Koessler 2024)
  4. Cap on total R&D compute per organization per year (Calero-Forero 2026)
  5. Cap on the RL share of total training compute (Irpan 2024)
  6. Minimum ratio of monitoring compute per inference compute (AI Futures Project 2026, Achiam 2026)
  7. Minimum ratio of safety spending per training compute (AI Futures Project 2026)
  8. Tax R&D compute above a threshold (Calero-Forero 2026)
  9. Regular reporting of each lab’s compute split into final runs, experiments, internal inference, and external inference (Epoch 2026; EO 14110).

Chips → compute → dangerous capabilities

  1. Export-control performance threshold for accelerators (BIS AC/S rule)
  2. Export controls on lithography, HBM and advanced packaging (BIS SME rule, Oct 2023; BIS rule 2024)
  3. Chip registry: serial numbers and owner of record above threshold, smuggling penalties (Sastry et al. 2024; Fist & Grunewald 2023)
  4. Location verification on accelerators (Brass & Aarne 2024; Chip Security Act, S.1705)
  5. Hardware-enabled mechanisms on new chips: offline licensing, metering, attestation (Aarne, Fist & Withers 2024; Kulp et al. 2024; FlexHEG 2025)

Datacenters (installed chips) → compute → dangerous capabilities

  1. Permit review threshold for datacenters (Sanders)
  2. Grid interconnect queue (Epoch 2024)
  3. Registry of datacenters with satellite-verified construction status (Epoch)

Data → dangerous capabilities

  1. Disclosure of synthetic-data share and RL environments used in frontier training (EU AI Act)

Algorithms → effective compute → dangerous capabilities

  1. Publication embargo on frontier algorithmic results (Bostrom 2017; Shevlane & Dafoe 2020)
  2. Total Research Transparency: mandatory disclosure of all frontier research (AI Futures Project 2026)
  3. Internal model transparency: mandatory sharing of all internal models with researchers at other labs (Tadepalli 2026).
  4. Structured access to code and checkpoints via vetted institutions (Shevlane 2022)
  5. Classification regime for capability-elicitation techniques (Shevlane & Dafoe 2020)

Talent → algorithms → dangerous capabilities

  1. Visa quota and processing time for frontier researchers (Zwetsloot et al. 2019, CSET)
  2. Vetting and cooling-off periods for staff with weight or R&D compute access (Nevo et al. 2024)

Capital → compute → dangerous capabilities

  1. Compute tax (dollar per FLOP on training runs above a threshold) (Calero-Forero 2026)
  2. Strict liability for catastrophic harms from frontier models (Weil 2024)
  3. Mandatory liability insurance, premiums priced by capability tier (Trout 2024)
  4. Investor-facing AI risk disclosure (SEC 2023)

AI R&D capabilities → algorithms → effective compute → dangerous capabilities

  1. Fraction of R&D compute consumed by autonomous agents (Stix et al. 2025; Charnock et al 2026)
  2. AI R&D speed-up trigger in safety frameworks (Anthropic RSP; DeepMind FSF; OpenAI Preparedness)
  3. Human review ratio for AI-written research code and experiment plans (Stix et al. 2025)
  4. Safety case required before internal deployment on the R&D stack (Clymer et al. 2024; Stix et al. 2025)
  5. Monitoring coverage of internal agent actions (Greenblatt et al. 2023; METR red-team of Anthropic’s internal monitoring, 2026)

Actors and cadence → race intensity → all other levers

  1. Licensing vs registration for frontier development (Anderljung et al. 2023; June 2026 EO)
  2. Minimum interval between frontier releases (FLI pause letter 2023)
  3. Pre-deployment testing window with government access (June 2026 EO; FRONTIER Act)
  4. Coordinated-pause trigger and duration across signatories (Alaga & Schuett 2023)
  5. Lead-margin reporting: months between top lab and next (Epoch; Karnofsky 2022)

Dangerous capabilities evals

  1. Pretraining data filtering for CBRN, cyber-offense and self-replication content (O’Brien et al. 2025)
  2. Verified unlearning of specified capabilities (Li et al. 2024, Feng et al 2025)
  3. Capability thresholds by domain (Koessler, Schuett & Anderljung 2024; the old Anthropic RSP)

Generality → dangerous capabilities

  1. Separate regulatory track for narrow AI, scientific models (Drexler 2019; EU AIA)
  2. Gating of long-horizon agentic post-training for general models (Chan et al. 2023; Kwa et al. 2025)

Legibility of model reasoning → control of dangerous capabilities

  1. Codebase / log audits for optimization pressure on chain-of-thought (Korbak et al. 2025; Baker et al. 2025)
  2. Disclosure and gating of latent reasoning architectures (Hao et al. 2024, Coconut; Korbak et al. 2025)
  3. Online weight updates in deployment off by default (Greenblatt et al. 2023)
  4. Persistent memory scope (Shavit et al. 2023, OpenAI)
  5. Canary-tagging or filtering of safety and eval content in training data (Berglund et al. 2023; Laine et al. 2024)

Weight security → blocking exfiltration

  1. Mandatory weight security level (“SL”) by capability (Nevo et al. 2024; Anthropic ASL-3)
  2. Two-person rule and hardware keys for weight access (Nevo et al. 2024)
  3. Insider-threat program coverage (Nevo et al. 2024)
  4. Weight retention policy (Anthropic 2025)

Inference → dangerous capabilities

  1. Per-query reasoning compute cap for the most capable models (Hooker 2024; Ord 2025)
  2. Token tax (Irwin 2026)
  3. Deployment tax by capability tier (Calero-Forero 2026)
  4. Liability allocation between deployer and developer for autonomous services (Weil 2024)

Autonomy → dangerous capabilities

  1. Agent permission tiers (Shavit et al. 2023, OpenAI; Chan et al. 2024)
  2. Spend limits per agent and per task (Shavit et al. 2023, OpenAI)
  3. Human approval for irreversible actions (Shavit et al. 2023, OpenAI; EU AI Act)
  4. Sub-agent spawn depth and maximum unattended run length (Kwa et al. 2025, METR)
  5. Agent identifiers and action-log retention (Chan et al. 2024)
  6. Kill-switch latency requirement (Orseau & Armstrong 2016; Shavit et al. 2023, OpenAI)
  7. Certification and fleet registration for embodied agents (EU Machinery Regulation 2023/1230)

Multi-agent → dangerous capabilities

  1. Instance-count reporting per deployment (Chan et al. 2024)
  2. Agent-to-agent communication logging with steganography checks (Motwani et al. 2024)

Safety research

  1. Minimum fraction of total compute reserved for safety research, above a total compute threshold (OpenAI 2023)
  2. Minimum safety headcount as a ratio of total research headcount (AI Lab Watch)
  3. External safety compute grants (minimum FLOP/year given to independent labs) (NAIRR)

Evaluation

  1. Third-party evaluator access depth (Casper et al. 2024)
  2. Elicitation budget per dangerous-capability eval (METR elicitation protocol; Barnett & Thiergart 2024)
  3. Sandbagging detection protocol (van der Weij et al. 2024)

Transparency

  1. Required “AI Assurance Level” for developers in the frontier tier (Brundage et al. 2026)
  2. Embedded auditors running regular audits with non-public access (Brundage et al. 2026)
  3. Audit scope including internal deployment, information security and safety decision-making (Brundage et al. 2026)
  4. Number of accredited audit providers and a standards body (Brundage et al. 2026)
  5. Incident reporting deadline (SB 53 §22757.13)
  6. Statutory whistleblower channel and protection (SB 53; Right to Warn letter 2024)

Governance response

  1. Indexing compute thresholds to measured algorithmic progress (Heim & Koessler 2024; Epoch 2025)
  2. Legislation with automatic clause triggers: e.g. once an eval result is shown, legal obligations come into force (Karnofsky 2024)
  3. Regulator capacity (CAISI; UK AISI; IFP 2026)

Coordination

  1. Treaty verifications: chip registry, datacenter inspections, interconnect bandwidth limits (Scher & Thiergart 2024; Baker et al. 2025)
  2. Training-run declarations exchanged between states (Shavit 2023; Baker et al. 2025)
  3. Verification R&D budget and a frontier-state incident hotline (Future Society 2026)

8.3. Appendix: Bibliography

Recent

  • Finke (2026), International Agreements to Limit Frontier AI: Objectives and Exit. arXiv

  • Larsen, Dean, Halstead, Lifland, Greenblatt & Kokotajlo (2026). AI 2040: Plan A. AI Futures Project

  • Lifland et al (2026). How to Pace the US Frontier. AI Futures Project.

  • Fist et al (2026). How Should the US Prepare for Increasingly Automated AI R&D?. Institute for Progress.

  • Koopmanschap, Barten (2026), How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements. arXiv

  • Institute for Progress (2026). Funding for CAISI. ifp.org

Foundations

  • Bostrom (2002). Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards. Journal of Evolution and Technology 9. nickbostrom.com

  • Shulman (2009). Arms Control and Intelligence Explosions. ECAP. intelligence.org

  • Marchant, Allenby & Herkert, eds. (2011). The Growing Gap Between Emerging Technologies and Legal-Ethical Oversight: The Pacing Problem. Springer. doi:10.1007/978-94-007-1356-7

  • Armstrong, Bostrom & Shulman (2016). Racing to the Precipice: A Model of Artificial Intelligence Development. AI & Society 31. doi:10.1007/s00146-015-0590-y

  • Christiano (2018). Takeoff Speeds. sideways-view.com

  • Maas (2018). Two Lessons from Nuclear Arms Control for the Responsible Governance of Military Artificial Intelligence. IOS.

  • Koessler, Schuett, Anderljung (2024). Risk thresholds for frontier AI. arXiv.

  • Dafoe (2018). AI Governance: A Research Agenda. GovAI / FHI. governance.ai

  • Kulveit, Douglas, Ammann, Turan, Krueger & Duvenaud (2025). Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development. arXiv:2501.16946

  • Barnett & Scher (2025). AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions. MIRI Technical Governance Team. arXiv:2505.04592

  • Favaro & Clark (2026). When AI Builds Itself. Anthropic Institute. anthropic.com

  • Bostrom (2017). Strategic Implications of Openness in AI Development. Global Policy 8(2). wiley.com, pdf

  • Drexler (2019). Reframing Superintelligence: Comprehensive AI Services as General Intelligence. FHI. ora.ox.ac.uk

  • Garfinkel & Dafoe (2019). How Does the Offense-Defense Balance Scale? Journal of Strategic Studies 42(6). tandfonline.com

  • Shevlane & Dafoe (2020). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse? arXiv:2001.00463

  • Leech et al. (2024). Shallow Review of Technical AI Safety: “Make AI Solve It”. shallowreview.ai

  • MacAskill & Moorhouse (2025). Preparing for the Intelligence Explosion. Forethought. forethought.org

  • Hobbhahn (2025). What’s the Short Timeline Plan? lesswrong.com

Economics

  • Aschenbrenner (2020). Existential Risk and Growth. GPI Working Paper 6-2020. leopoldaschenbrenner.github.io

  • Sandbrink, Hobbs, Swett, Dafoe & Sandberg (2022). Differential Technology Development: An Innovation Governance Consideration for Navigating Technology Risks. SSRN. ssrn.com

  • Jones (2024). The A.I. Dilemma: Growth versus Existential Risk. AER: Insights 6(4). nber.org (WP 31837)

  • Trammell & Aschenbrenner (2024). Existential Risk and Growth. GPI Working Paper 13-2024. philiptrammell.com

  • Salop & Scheffman (1983). Raising Rivals’ Costs. American Economic Review 73(2). repec.org

  • Acemoglu (2002). Directed Technical Change. Review of Economic Studies 69(4). mit.edu

  • Bostrom (2003). Astronomical Waste: The Opportunity Cost of Delayed Technological Development. Utilitas 15(3). nickbostrom.com

  • Bostrom (2005). The Fable of the Dragon-Tyrant. Journal of Medical Ethics 31(5). nickbostrom.com

  • Heitzig, Lessmann & Zou (2011). Self-Enforcing Strategies to Deter Free-Riding in the Climate Change Mitigation Game and Other Repeated Public Good Games. PNAS 108(38). pnas.org

  • Garicano, Lelarge & Van Reenen (2016). Firm Size Distortions and the Productivity Distribution: Evidence from France. American Economic Review 106(11). aeaweb.org

  • Johnson, Shriver & Goldberg (2023). Privacy and Market Concentration: Intended and Unintended Consequences of the GDPR. Management Science 69(10). informs.org

  • Merali (2024). Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Translation. arXiv:2409.02391

  • Srivastav & Zaehringer (2024). The Economics of Coal Phaseouts. arXiv:2406.14238

  • Weil (2024). Tort Law as a Tool for Mitigating Catastrophic Risk from Artificial Intelligence. SSRN

  • Google (2024). AI in Science. ai.google

  • Tomei, Jain & Franklin (2025). AI Governance through Markets. arXiv:2501.17755

  • Gundlach, Lynch, Mertens & Thompson (2025). The Price of Progress: Price Performance and the Future of AI. arXiv:2511.23455

  • Yan & Morck (2025). Who’s Afraid of Tariffs? The Geographic Distribution of Fear and Loss. NBER WP 34299. nber.org

  • Aubakirova, Atallah, Clark, Summerville & Midha (2026). State of AI: An Empirical 100 Trillion Token Study with OpenRouter. arXiv:2601.10088

  • Brynjolfsson, Collis, Eggers, Kazinnik & Nguyen (2026). What is Generative AI Worth? SSRN

  • Demirer, Musolff & Yang (2026). Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools. NBER WP 35275. nber.org

  • Forecasting Research Institute (2026). Experts Forecast Rapid AI Progress Could Bring Health and Wealth Without Happiness. forecastingresearch.substack.com

  • IntuitionLabs (2026). AI-Discovered Drugs in Clinical Trials. intuitionlabs.ai

  • Trout (2024). Insuring Uninsurable Risks from AI: Government as Insurer of Last Resort. arXiv:2409.06672

  • Irwin, Wu & Barez (2026). Position: Token Taxes Can Mitigate AI’s Economic Risks. arXiv:2603.04555

The pause debate

  • Grace (2022). Let’s Think About Slowing Down AI. AI Impacts / LessWrong. lesswrong.com

  • Belrose (2023). AI Pause Will Likely Backfire.

  • Buterin (2023). My Techno-optimizm. vitalik.eth.limo

  • Tallinn (2024). Priorities for AI Risk Reduction. jaan.online

  • Katzke & Futerman (2024). The Manhattan Trap: Why a Race to Artificial Superintelligence Is Self-Defeating. Convergence Analysis. arXiv:2501.14749

  • Larsen, Dean, Halstead, Lifland, Greenblatt & Kokotajlo (2026). AI 2040: Plan A. AI Futures Project. ai-2040.com

  • AI Impacts. Hardware Overhang. aiimpacts.org

  • AI Impacts (2023). Are There Examples of Overhang for Other Technologies? blog.aiimpacts.org

  • Miotti et al. (2024). A Narrow Path. ControlAI. narrowpath.co

  • Felstead (2025). Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks. arXiv:2511.08631

  • Felstead (2026). Can Frontier AI Labs Lawfully Agree to Pause? Lawfare. lawfaremedia.org

  • Employees of frontier AI companies (2026). Pacing the Frontier (open statement). pacingthefrontier.com

  • Karnofsky (2022). Racing through a Minefield: The AI Deployment Problem. Cold Takes. cold-takes.com

  • Future of Life Institute (2023). Pause Giant AI Experiments: An Open Letter. futureoflife.org

  • Alaga & Schuett (2023). Coordinated Pausing: An Evaluation-Based Coordination Scheme for Frontier AI Developers. arXiv:2310.00374

  • Cotra (2026). Total Research Transparency Would Be Nice. Planned Obsolescence. planned-obsolescence.org

International agreements

  • Ho, Barnhart, Trager, Bengio, Brundage, Casovan, Haas, Nemitz, Sastry, Weller, Zhang & Zhang (2023). International Institutions for Advanced AI. arXiv:2307.04699

  • Trager, Harack, Reuel, Carnegie, Heim, Ho, Kreps, Lall, Larter, Ó hÉigeartaigh, Staffell & Villalobos (2023). International Governance of Civilian AI: A Jurisdictional Certification Approach. arXiv:2308.15514

  • Hausenloy, Miotti & Dennis (2023). Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI. arXiv:2310.09217

  • Emery-Xu, Jordan & Trager (2025). International Governance of Advancing Artificial Intelligence. AI & Society 40. doi:10.1007/s00146-024-02050-7

  • Al Ramiah, Koopmanschap, Thorsteinson, Khan, Zhou, Noh, Meindertsma & Shafiq (2025). Toward a Global Regime for Compute Governance: Building the Pause Button. arXiv:2506.20530

  • Scher, Abecassis, Barnett & Abeyta (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. arXiv:2511.10783

  • Finke (2026). International Agreements to Limit Frontier AI: Objectives and Exit. TAIGR @ ICML 2026. arXiv:2607.16224

  • Koremenos (2005). Contracting around International Uncertainty. American Political Science Review 99(4). doi.org

  • Bartels (2020). Building Better Games for National Security Policy Analysis. RAND. rand.org

  • Gruetzemacher et al. (2024). Strategic Insights from Simulation Gaming of AI Race Dynamics. arXiv:2410.03092

  • Goldstein & Salib (2025). How to Stop an AI Arms Race. SSRN

Deterrence

  • Hendrycks, Schmidt & Wang (2025). Superintelligence Strategy: Expert Version. arXiv:2503.05628
  • Rehman, Mueller, Mazarr et al. (2025). Seeking Stability in the Competition for AI Advantage. RAND commentary. rand.org
  • Abecassis (2025). Refining MAIM: Identifying Changes Required to Meet Conditions for Deterrence. MIRI. intelligence.org
  • Arnold (2025). Superintelligence Deterrence Has an Observability Problem. AI Frontiers. ai-frontiers.org
  • Hendrycks & Khoja (2025). AI Deterrence Is Our Best Option. AI Frontiers. ai-frontiers.org
  • Delaney (2025). Crucial Considerations in ASI Deterrence. IAPS. iaps.ai

Verification

  • Brundage et al. (2020). Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims. arXiv:2004.07213

  • Baker (2023). Nuclear Arms Control Verification and Lessons for AI Treaties. arXiv:2304.04123

  • Scher & Thiergart (2024). Mechanisms to Verify International Agreements About AI Development. MIRI Technical Governance Team. arXiv:2506.15867

  • Wasil, Reed, Miller & Barnett (2024). Verification Methods for International AI Agreements. arXiv:2408.16074

  • Koopmanschap & Barten (2026). How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements. Existential Risk Observatory. arXiv:2607.22619

  • Choussat & Khoja (2026). An International AI Slowdown Is Ready Whenever Politicians Are. AI Frontiers. ai-frontiers.org

  • Baker, Kulp, Marks, Brundage & Heim (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. arXiv:2507.15916

  • The Future Society (2026). How To Make International AI Verification a Reality. thefuturesociety.org

Compute governance

  • Shavit (2023). What Does It Take to Catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv:2303.11341

  • Egan & Heim (2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv:2310.13625

  • Sastry, Heim, Belfield, Anderljung, Brundage, Hazell et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv:2402.08797

  • Heim, Fist, Egan, Huang, Zekany, Trager, Osborne & Zilberman (2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. Oxford Martin AIGI. oxfordmartin.ox.ac.uk

  • Aarne, Fist & Withers (2024). Secure, Governable Chips. CNAS. cnas.org

  • Kulp, Gonzales, Smith, Heim, Puri, Vermeer & Winkelman (2024). Hardware-Enabled Governance Mechanisms. RAND WR-A3056-1. rand.org

  • Petrie, Aarne, Ammann & Dalrymple (2024). Interim Report: Mechanisms for Flexible Hardware-Enabled Guarantees. Part I, arXiv:2506.15093

  • Brass & Aarne (2024). Location Verification for AI Chips. IAPS. iaps.ai

  • Heim & Koessler (2024). Training Compute Thresholds: Features and Functions in AI Regulation. arXiv:2405.10799

  • Hooker (2024). On the Limitations of Compute Thresholds as a Governance Strategy. arXiv:2407.05694

  • Ord (2025). Inference Scaling Reshapes AI Governance. arXiv

  • USITC (2023). Germanium and Gallium (Executive Briefing on Trade). usitc.gov

  • Ho, Besiroglu, Erdil et al. (2024). Algorithmic Progress in Language Models. Epoch AI. arXiv:2403.05812

  • Miller (2025). How US Export Controls Have (and Haven’t) Curbed Chinese AI. AI Frontiers. ai-frontiers.org

  • O’Gara, Kulp, Hodgkins, Petrie et al. (2025). Hardware-Enabled Mechanisms for Verifying Responsible AI Development. arXiv:2505.03742

  • Somala, Ho & Krier (2025). Three Challenges Facing Compute-Based AI Policies. Epoch AI. epochai.substack.com

  • Ansari (2026). Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification. arXiv:2604.04712

  • Fedasiuk & Torres (2026). The Lithography Loophole: How China Is Printing Its Way to Chip Self-Sufficiency. AEI. aei.org

  • Rahman (2026). Does Distributed Training Undermine Compute Governance? arXiv:2605.29359

  • Seferis & Fist (2026). Detecting Compute Structuring in AI Governance Is Likely Feasible. AAAI. aaai.org

  • Gargiulo & Kulp (2026). Workload Identification with Physical Side Channels for AI Governance. arXiv:2609.00309

  • Fist & Grunewald (2023). Preventing AI Chip Smuggling to China. CNAS. cnas.org

  • Irpan (2024). Late Takes on OpenAI o1. Sorta Insightful. alexirpan.com

  • Epoch AI (2024). Can AI Scaling Continue Through 2030? epoch.ai

  • Cottier & Owen (2025). How Many AI Models Will Exceed Compute Thresholds? Epoch AI. epoch.ai

  • Epoch AI. Data on Data Centers. epoch.ai

  • Denain & Wu (2026). Final Training Runs Account for a Minority of R&D Compute Spending. Epoch AI, Gradient Updates. epoch.ai

  • Ho (2026). Keeping Up with the GPTs. Epoch AI, Gradient Updates. epoch.ai

  • Calero-Forero (2026). How Should You Slow Down AI Progress, If It Becomes Necessary? lesswrong.com, substack

  • Achiam (2026). On monitoring-to-inference compute ratios. x.com

Law and regulation

  • Zwetsloot, Dunham, Arnold & Huang (2019). Keeping Top AI Talent in the United States. CSET. cset.georgetown.edu
  • European Union (2023). Machinery Regulation (EU) 2023/1230. eur-lex.europa.eu
  • The White House (2023). Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. Federal Register. federalregister.gov
  • Bureau of Industry and Security (2023). Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use. Federal Register. federalregister.gov
  • Bureau of Industry and Security (2023). Export Controls on Semiconductor Manufacturing Items. Federal Register. federalregister.gov
  • Bureau of Industry and Security (2024). Foreign-Produced Direct Product Rule Additions and Refinements to Controls for Advanced Computing. Federal Register. federalregister.gov
  • SEC (2023). SEC Adopts Rules on Cybersecurity Risk Management, Strategy, Governance, and Incident Disclosure by Public Companies. Press release 2023-139. sec.gov
  • Anderljung et al. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv:2307.03718
  • European Union (2024). Artificial Intelligence Act. Article 14, Article 53, Code of Practice
  • California Legislature (2025). SB 53: Transparency in Frontier Artificial Intelligence Act. leginfo.legislature.ca.gov
  • U.S. Congress (2025). Chip Security Act, S.1705, 119th Congress. congress.gov
  • Sanders & Ocasio-Cortez. Sanders, Ocasio-Cortez Announce AI Data Center Moratorium Act (press release). sanders.senate.gov
  • NIST. Center for AI Standards and Innovation (CAISI). nist.gov
  • UK AI Security Institute. aisi.gov.uk
  • National Artificial Intelligence Research Resource Pilot. nairrpilot.org
  • Foley Hoag (2026). Trump’s New AI Frontier: The Executive Order Regulating Frontier AI Models. foleyhoag.com
  • Statt (2026). The FRONTIER Act: Federal AI Regulation in 2026. statt.com

Developer commitments

  • Shevlane, Farquhar, Garfinkel, Phuong, Whittlestone, Leung et al. (2023). Model Evaluation for Extreme Risks. arXiv:2305.15324

  • Clymer, Gabrieli, Krueger & Larsen (2024). Safety Cases: How to Justify the Safety of Advanced AI Systems. arXiv:2403.10462

  • Karnofsky (2024). If-Then Commitments for AI Risk Reduction. Carnegie Endowment. carnegieendowment.org

  • Cârlan, Gomez, Mathew, Krishna, King, Gebauer & Smith (2024). Dynamic Safety Cases for Frontier AI. arXiv:2412.17618

  • Christiano (2023). Thoughts on Responsible Scaling Policies and Regulation. Alignment Forum. alignmentforum.org

  • OpenAI (2025). Expanding on What We Missed with Sycophancy. openai.com

  • OpenAI (2025). How We Think About Safety and Alignment. openai.com

  • METR (2025). Common Elements of Frontier AI Safety Policies. metr.org

  • Anthropic (2026). Policy on the AI Exponential. anthropic.com

  • Anthropic (2026). Introducing Claude Fable 5 and Claude Mythos 5. anthropic.com

  • Anthropic (2026). Project Glasswing. anthropic.com

  • Anthropic (2026). Project Glasswing: Initial Update. anthropic.com

  • OpenAI (2026). Trusted Access for Cyber. openai.com

  • OpenAI (2026). Pacing Model Development for Cyber Capabilities. openai.com

  • OpenAI (2026). The AI Policy Window. openai.com

  • OpenAI (2026). GPT-6 Astra Deployment Safety Report: “Monitorability”. deploymentsafety.openai.com

  • OpenAI (2026). Research Acceleration: A View Inside OpenAI. openai.com

  • OpenAI (2023). Introducing Superalignment. openai.com

  • Anthropic. Responsible Scaling Policy (updates). anthropic.com

  • Google DeepMind (2024). Introducing the Frontier Safety Framework. deepmind.google

  • OpenAI. Preparedness Framework. openai.com

  • Right to Warn (2024). A Right to Warn about Advanced Artificial Intelligence (open letter). righttowarn.ai

  • Anthropic (2025). Activating AI Safety Level 3 Protections. anthropic.com

  • Anthropic (2025). Commitments on Model Deprecation and Preservation. anthropic.com

  • Stein-Perlman. AI Lab Watch. ailabwatch.org

Evaluations and forecasting

  • Brown et al. (2020). Language Models are Few-Shot Learners. NeurIPS. arXiv:2005.14165

  • Wei et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS. arXiv:2201.11903

  • Dell’Acqua et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. HBS Working Paper 24-013. SSRN

  • UK AI Security Institute (2024). Pre-Deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet. aisi.gov.uk

  • Pimpale, Højmark, Scheurer & Hobbhahn (2025). Forecasting Frontier Language Model Agent Capabilities. arXiv:2502.15850

  • Needham, Edkins, Pimpale, Bartsch & Hobbhahn (2025). Large Language Models Often Know When They Are Being Evaluated. arXiv:2505.23836

  • Bean et al. (2025). Measuring What Matters: Construct Validity in Large Language Model Benchmarks. arXiv:2511.04703

  • Wei & Heim (2025). Designing Incident Reporting Systems for Harms from General-Purpose AI. AAAI. arXiv:2511.05914

  • UK AI Security Institute (2025). More Compute, More Capability: Why AI Agent Evals Need to Account for Test-Time Compute. aisi.gov.uk

  • Epoch AI. AI Chip Production (data insight). epoch.ai

  • Epoch AI. Benchmarks: Epoch Capabilities Index. epoch.ai

  • Epoch AI (2026). An Update on AI’s Most Important Number. Gradient Updates. epoch.ai

  • Mengesha et al. (2026). A Pragmatic Classification Framework for AI Incident Monitoring. arXiv:2604.21412

  • Barrett et al. (2026). Lessons from External Review of DeepMind’s Scheming Inability Safety Case. arXiv:2604.21964

  • Guidelight AI Standards (2026). AI Control: An Assessment of Frontier Practices. guidelight.ai

  • METR (2026). Notes on Anthropic Researcher Uplift Estimates. metr.org

  • Casper et al. (2024). Black-Box Access Is Insufficient for Rigorous AI Audits. arXiv:2401.14446

  • Barnett & Thiergart (2024). Declare and Justify: Explicit Assumptions in AI Evaluations Are Necessary for Effective Regulation. arXiv:2411.12820

  • van der Weij et al. (2024). AI Sandbagging: Language Models Can Strategically Underperform on Evaluations. arXiv:2406.07358

  • METR. Guidelines for Capability Elicitation. Autonomy Evals Guide. metr.github.io

  • Brundage et al. (2026). Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies. arXiv:2601.11699

Safeguards and model security

  • Esvelt (2022). Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics. GCSP. gcsp.ch

  • Kirk et al. (2023). Understanding the Effects of RLHF on LLM Generalisation and Diversity. arXiv:2310.06452

  • Nevo et al. (2024). Securing AI Model Weights. RAND RR-A2849-1. rand.org

  • Tamirisa et al. (2024). Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv:2408.00761

  • Brent & McKelvey (2025). Contemporary AI Foundation Models Increase Biological Weapons Risk. arXiv:2506.13798

  • Chen, Joshi, Chen, Andriushchenko, Angell & He (2025). Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors. arXiv:2506.10949

  • Marshall et al. (2026). BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation. arXiv:2607.14479

  • Shevlane (2022). Structured Access: An Emerging Paradigm for Safe AI Deployment. arXiv:2201.05159

  • Li et al. (2024). The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning. arXiv:2403.03218

  • O’Brien et al. (2025). Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs. arXiv:2508.06601

  • Feng et al. (2025). Existing Large Language Model Unlearning Evaluations Are Inconclusive. arXiv:2506.00688

Agents, oversight and control

  • Orseau & Armstrong (2016). Safely Interruptible Agents. UAI. intelligence.org
  • Chan et al. (2023). Harms from Increasingly Agentic Algorithmic Systems. FAccT. arXiv:2302.10329
  • Shavit et al. (2023). Practices for Governing Agentic AI Systems. OpenAI. openai.com
  • Berglund et al. (2023). Taken Out of Context: On Measuring Situational Awareness in LLMs. arXiv:2309.00667
  • Greenblatt, Shlegeris, Sachan & Roger (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv:2312.06942
  • Chan et al. (2024). Visibility into AI Agents. FAccT. arXiv:2401.13138
  • Motwani et al. (2024). Secret Collusion among AI Agents: Multi-Agent Deception via Steganography. NeurIPS. arXiv:2402.07510
  • Laine et al. (2024). Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs. arXiv:2407.04694
  • Hao et al. (2024). Training Large Language Models to Reason in a Continuous Latent Space. arXiv:2412.06769
  • Baker et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. OpenAI. arXiv:2503.11926
  • Kwa et al. (2025). Measuring AI Ability to Complete Long Tasks. METR. arXiv:2503.14499
  • Stix et al. (2025). AI Behind Closed Doors: A Primer on the Governance of Internal Deployment. arXiv:2504.12170
  • Korbak et al. (2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv:2507.11473
  • Charnock et al. (2026). What Should Frontier AI Developers Disclose About Internal Deployments? arXiv:2604.23065
  • METR (2026). Red-Teaming Anthropic’s Agent Monitoring. metr.org

Histories and precedents of restraint

  • U.S. Naval Institute (1926). The Washington Treaties of 1922. Proceedings 52(5). usni.org
  • GlobalSecurity.org. Treaty Cruisers. globalsecurity.org
  • Haddon-Cave (2009). The Nimrod Review. HC 1025. gov.uk
  • Leveson (2011). The Use of Safety Cases in Certification and Regulation. MIT. mit.edu
  • Goitein & Patel (2015). What Went Wrong with the FISA Court. Brennan Center for Justice. brennancenter.org
  • Coe & Vaynman (2015). Collusion and the Nuclear Nonproliferation Regime. Journal of Politics 77(4). andrewjcoe.com
  • China State Council (2017). A New Generation Artificial Intelligence Development Plan (FLIA translation). flia.org
  • Casey (2021). A Reckoning Looms for America’s 50-Year Financial Surveillance System. Cato Journal 41. cato.org
  • Molloy (2021). Approach with Caution: Sunset Clauses as Safeguards of Democracy? northumbria.ac.uk
  • Romano & Levin (2021). Sunsetting as an Adaptive Strategy. PNAS. ncbi.nlm.nih.gov
  • Maas (2022). Paths Untaken: The History, Epistemology and Strategy of Technological Restraint, and Lessons for AI. Verfassungsblog. verfassungsblog.de
  • Cassata & de Chadarevian (2025). Asilomar Across the Atlantic. ncbi.nlm.nih.gov
  • Columbia Academic Commons. On the environmental costs of the slowed rollout of nuclear power. doi:10.7916/d8-qez9-6m49
  • 31 U.S.C. §5324. Structuring Transactions to Evade Reporting Requirement Prohibited. law.cornell.edu
  • Wikipedia. Year 2000 Problem. wikipedia.org

News and incidents

  • CNBC (2025). Nvidia on Track to Hit Historic $5 Trillion Valuation amid AI Rally. cnbc.com
  • The Legal Wire (2025). CAC Launches Special Campaign to Clear Up and Rectify the Abuse of AI Technology. thelegalwire.ai
  • STAT News (2026). AI Ambient Scribes Bring Modest Time Savings in Clinical Documentation. statnews.com
  • Axios (2026). Anthropic’s Revenue Growth. axios.com
  • CNBC (2026). Micron Reaches a Trillion-Dollar Market Cap. cnbc.com
  • The Guardian (2026). Anthropic Disables Advanced AI Models after US Government Order. theguardian.com
  • Observer (2026). Anthropic Delays AI Model It Deems Too Powerful for Public Use. observer.co.uk
  • Nature (2026). AI Researchers Reckon with the $1.5 Million ‘Academia Tax’. nature.com
  • The Motley Fool (2026). Broadcom Is Less Than 5% from the $2 Trillion Club. fool.com
  • Axios (2026). OpenAI’s GPT-5.6 Ban Lifted. axios.com
  • Axios (2026). OpenAI Delays Astra Model over Cybersecurity Risks. axios.com
  • METR (2026). OpenAI / Hugging Face Incident Investigation. metr.org
  • SANS Institute (2026). “The Models Said No”: Inside the Hugging Face Post-Mortem. sans.org
  • Toh (2026). The AI Chip Wars’ New Front: Control the Cloud, Not the Silicon. Forbes. forbes.com
  • TechCrunch (2026). Group Funded by Andreessen Horowitz and Brockman Plans Data Center Ads to Sway Midterms. techcrunch.com
  • Public First Action (2026). Public First Action and Defending Our Values PAC Launch First Ads Supporting Responsible AI Regulation. publicfirstaction.us
  • CNBC (2026). OpenAI’s Astra Model Rated “Critical” for Cyber Risk. cnbc.com
  • Yahoo Finance (2026). OpenAI Burning $12.3 Billion. yahoo.com
  • The Verge (2026). Fable Won’t Answer Basic Biology Questions. theverge.com
  • The Wall Street Journal (2026). Anthropic Researcher Quits over Out-of-Control AI Fears. wsj.com