<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Governing Transformative AI]]></title><description><![CDATA[Governing Transformative AI]]></description><link>https://governingtransformativeai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!A2uB!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19cad93-0c95-419b-a7bc-3736c8d18b27_1067x1067.png</url><title>Governing Transformative AI</title><link>https://governingtransformativeai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 08:27:19 GMT</lastBuildDate><atom:link href="/__u/governingtransformativeai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Governing Transformative AI]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[governingtransformativeai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[governingtransformativeai@substack.com]]></itunes:email><itunes:name><![CDATA[Governing Transformative AI]]></itunes:name></itunes:owner><itunes:author><![CDATA[Governing Transformative AI]]></itunes:author><googleplay:owner><![CDATA[governingtransformativeai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[governingtransformativeai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Governing Transformative AI]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Classified, sovereign compute: a high-leverage approach to UK data centre policy]]></title><description><![CDATA[The UK Government should launch a Secure Compute Plan to build world-leading capacity to run AI agents in classified environments.]]></description><link>https://governingtransformativeai.substack.com/p/uk-data-centre-policy-a-high-leverage</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/uk-data-centre-policy-a-high-leverage</guid><dc:creator><![CDATA[Governing Transformative AI]]></dc:creator><pubDate>Wed, 12 Aug 2026 14:09:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7ee4dad4-9d7c-4474-bd93-829a57720655_1983x793.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Patrick Levermore &amp; Hamish Hobbs </strong></p><div><hr></div><p><span>The UK has committed to improving economic resilience and national security by building sovereign compute infrastructure on UK soil. The </span><a href="https://www.gov.uk/government/publications/the-defence-investment-plan"><span>Defence Investment Plan</span></a><span> names the need here &#8212; &#8220;Compute has emerged as a critical limiting factor for modern capability&#8221;. Minister Narayan (the UK&#8217;s first-ever cabinet-level AI minister) </span><a href="https://x.com/KanishkaNarayan/status/2066157359638962632"><span>puts it plainly</span></a><span>: &#8220;If national security is the question, sovereign AI capability is the central answer.&#8221; Investment towards this aim is growing.</span></p><p><span>The 2025 </span><a href="https://www.gov.uk/government/publications/uk-compute-roadmap"><span>UK Compute Roadmap</span></a><span> committed &#163;2bn, aiming for a twentyfold increase in national AI compute capacity, and established AI Growth Zones, which have attracted &#163;28bn of confirmed private investment. The 2026 </span><a href="https://www.gov.uk/government/publications/uk-ai-hardware-plan/uk-ai-hardware-plan"><span>UK AI Hardware Plan</span></a><span> added a further &#163;1.1bn of planned spend.</span></p><p><span>But to ensure sovereign national security and influence on a global scale, the UK needs to look for key leverage points and ensure it can test and run AI agents in its classified environments. Commercial AI infrastructure is already mass-produced on a global scale, with capital expenditure of the largest publicly owned data centre operators close to </span><a href="https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/"><span>750 billion USD</span></a><span> this year. Meanwhile, a different category of compute is both scarce and strategically vital: </span><strong><span>secure sovereign compute </span></strong><span>for</span><strong><span> classified AI inference.</span></strong></p><p><span>The UK has some of this, but not enough. Public defence plans buy secure networks and sovereign cloud. Examples include the </span><a href="https://www.gov.uk/government/news/security-delivered-for-working-people-as-uk-us-ties-strengthened-with-new-google-cloud-partnership-for-classified-information-sharing"><span>&#163;400m MOD agreement</span></a><span> with Google, and the Defence-wide</span><a href="https://www.gov.uk/government/publications/the-strategic-defence-review-2025-making-britain-safer-secure-at-home-strong-abroad/the-strategic-defence-review-2025-making-britain-safer-secure-at-home-strong-abroad#roles-for-uk-defence-1"><span> Secret Cloud</span></a><span>. But this is not accredited AI compute at the scale advanced AI requires. The &#163;1.1bn UK AI Hardware Plan gives national security a single paragraph, with no funding attached, and never uses the words &#8220;classified&#8221;, &#8220;secret&#8221; or &#8220;accredited&#8221;.</span></p><p><span>There are two upcoming moments to close the gap here. Fielding 10,000 classified AI agents requires on the order of 1.5 MW of accredited capacity. That&#8217;s around &#163;100m plus &#163;20m a year. Commitments and investments could be made in the next few months: the </span><a href="https://www.gov.uk/government/publications/putting-artificial-intelligence-ai-at-the-heart-of-uk-defence/putting-artificial-intelligence-ai-at-the-heart-of-uk-defence"><span>Defence Strategic Approach to AI</span></a><span> is still being written, and the AI Hardware Plan needs a new owner after Machinery of Government changes. Each is a chance to name classified, sovereign compute as an objective.</span></p><p><span>We make the case that the UK&#8217;s data centre objectives should include secure sovereign compute for classified AI inference, and set out four concrete objectives for what that should look like.</span></p><h2><strong><span>Why classified AI compute is valuable</span></strong></h2><p><span>We use the terms secure or classified compute to mean AI infrastructure accredited to handle the most sensitive material: this would involve isolated or air-gapped facilities, cleared staff, and certified hardware. Why is this valuable? First, analysing national security risks from AI often depends on classified information that cannot leave secure environments. As testing and monitoring AI agents becomes more compute-intensive, that work will need more accredited compute. Second, utilising AI for defence and security requires infrastructure the national security community can trust. In particular, we think </span><em><span>inference</span></em><span> compute will be the most important compute to invest in. That is, compute which is specialised for running AI systems, not training them. While the UK is far behind the U.S. and China in training the most advanced AI models, we can &#8212; with the proposed infrastructure &#8212; leverage our global partnerships to access, use and test secure closed models at the frontier.</span></p><p><span>Delivering secure, classified AI compute in the UK depends as much, if not more, on government will as it does on free-market capital. The Government holds a monopoly, not just on demand, but also on the inputs to make such a facility: accreditation, personnel clearances, threat intelligence, and more. In the U.S., there are two key providers able to run classified AI systems (</span><a href="https://aws.amazon.com/federal/top-secret-cloud/"><span>AWS</span></a><span> and </span><a href="https://azure.microsoft.com/en-us/blog/azure-government-top-secret-now-generally-available-for-us-national-security-missions/"><span>Azure</span></a><span>). The closest UK benchmark is the &#163;400m MOD deal with Google for air-gapped compute.</span></p><p><span>Two recent developments make this the moment to act:</span></p><ol><li><p><strong><span>Frontier capabilities have entered classified threat territory.</span></strong><span> The latest models </span><a href="https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf"><span>cross safety thresholds, such as OpenAI&#8217;s</span></a><span>, for high-risk biological misuse, and </span><a href="https://www.anthropic.com/research/mythos-preview"><span>can find</span></a><span> previously unknown cybersecurity vulnerabilities in all major operating systems. Evaluating these capabilities properly increasingly requires data and information held at the highest classification levels, which means the infrastructure itself needs the same level of security assurance.<br></span></p></li><li><p><strong><span>The classified AI ecosystem now exists, in the U.S.</span></strong><span> On 1 May 2026, the U.S. Department of War </span><a href="https://www.war.gov/News/Releases/Release/Article/4475177/classified-networks-ai-agreements/"><span>made agreements</span></a><span> with eight leading AI companies to run their models on Secret and Top Secret (IL6&amp;7) networks. The question for the UK is whether it can plug into this ecosystem as a capable partner, whether it watches from outside, or whether it is completely blindsided.</span></p></li></ol><p><span>The capital expenditure of the 14 largest publicly owned data centre operators globally is </span><a href="https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/"><span>predicted</span></a><span> to be 3.3 trillion USD through 2029, in an attempt to meet the projected demand for AI systems. However, becoming a global leader in secure and classified compute would require a substantially smaller investment that is manageable for the UK, on the order of &#163;200-600 million</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span>. This would complement the UK&#8217;s world-leading assets in AI and cybersecurity, including GCHQ and AISI. Focusing on secure, classified compute gives the UK cost-effective leverage to shape the impact of AI to enhance both domestic and global security.</span></p><h2><strong><span>Four objectives for UK data centre policy</span></strong></h2><p><span>We propose that the UK&#8217;s Secure Compute Plan should include four objectives. These are distinct from the existing AI Hardware Plan: while both improve the UK&#8217;s sovereign capabilities, a Secure Compute Plan would not target general commercial compute capacity in the UK, but compute specifically available for AI use in classified environments.</span></p><p><strong><span>1. Secure inference compute for national security applications (~&#163;100-400m upfront + &#163;20-80m/yr)</span></strong></p><p><span>Build sovereign infrastructure capable of running AI agents &#8212; models that carry out multi-step tasks on their own &#8212; in classified national security environments. This would enable AI for cyber defence and signals intelligence, alongside other future requirements, like tracking where AI agents are behaving counter to human instructions.</span></p><p><span>It would also strengthen the UK&#8217;s position internationally: as the U.S. makes advanced AI a priority for its classified work, allies that can run frontier AI models securely become partners rather than dependents.</span></p><p><strong><span>2. Secure compute for classified evaluations of frontier models (~&#163;20-50m upfront + &#163;4-10m/yr)</span></strong></p><p><span>Build classified infrastructure for AISI and the UK intelligence community to evaluate frontier models against data held at the highest security tiers. Some of the most important questions about frontier models &#8212; what they can do with genuinely sensitive information, in cyber, bio, and broader intelligence domains &#8212; simply cannot be answered on commercial infrastructure. Done well, this positions the UK as the first country aware of emerging AI threats, and the natural home for classified evaluations among allies, building on AISI&#8217;s existing reputation.</span></p><p><strong><span>3. Secure compute for UK AI security research (~&#163;30-80m upfront + &#163;6-20m/yr)</span></strong></p><p><span>Build government-owned capacity for UK researchers, including at AISI, to find new uses for frontier AI in secure environments &#8212; including new approaches to alignment, control, and oversight of advanced agentic systems. Beyond the direct economic benefits of building assets in a fast-growing industry, this gives the UK a unique offer to major powers during high-stakes decisions as AI is integrated into national security systems.</span></p><p><strong><span>4. R&amp;D to develop a &#8220;Security Level 5&#8221; inference data centre (~&#163;30m)</span></strong></p><p><span>Fund the research and a prototype inference facility needed to reach &#8220;Security Level 5&#8221; &#8212; the standard (defined in RAND&#8217;s </span><a href="https://www.rand.org/pubs/research_reports/RRA2849-1.html"><span>Securing AI Model Weights</span></a><span>) for infrastructure that could withstand a top-priority attack by the world&#8217;s most capable state actors. AI model weights will be a critical national security resource, and will need to be stored and run securely.</span></p><p><span>No such facility exists anywhere and the UK has the capabilities and opportunity to be a first-mover and develop something no other state has, making it a valuable asset for improving the UK&#8217;s geopolitical standing and national security. Should allied countries build similarly secure data centres, the UK will likely need to match pace in order to be granted access to the same AI models.</span></p><p><strong><span>To deliver objectives 1-4:</span></strong></p><p><span>Each of the above capacity objectives needs to meet two requirements simultaneously: the security standards to process the highest tiers of classified information, and the AI-specific infrastructure demands (high-density GPU deployments, elevated power density per square metre, and aggressive cooling) that traditional secure facilities were not built for.</span></p><p><strong><span>How much might it cost?</span></strong></p><p><span>We cannot precisely estimate the costs for appropriately securing compute facilities. But here are some indicative costings, using compute estimates for AI agents, and our approximate guesses at the costs to appropriately secure this compute:</span></p><p><span>Running a frontier AI agent today, such as Claude Mythos, requires about 3 million tokens to run the equivalent of a human working an eight-hour working day</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span>. So if the UK wanted to field 10,000 classified agents today, it would need on the order of 1.5 MW of accredited compute capacity. In </span><a href="https://isc.independent.gov.uk/wp-content/uploads/2023/12/ISC-Annual-Report-2022-2023.pdf"><span>2022</span></a><span>, the UK intelligence community had over 20,000 staff and &#163;1.2bn staff pay in the Single Intelligence Account. While there is substantial uncertainty about how many classified AI agents will be needed, if Britain wants to ensure it can secure itself against rapidly escalating AI-enabled threats, an ability to field at least 10,000 classified AI agents is our proposed planning floor.</span></p><p><span>In the UK compute context: Isambard-AI </span><a href="https://www.bristol.ac.uk/news/2023/november/supercomputer-announcement.html"><span> cost &#163;225m</span></a><span> for </span><a href="https://blogs.nvidia.com/blog/isambard-ai/"><span>5,448 Nvidia GH200 superchips</span></a><span> in a ~5MW purpose-built facility. Each superchip contains an AI accelerator, the specialised hardware that trains and runs AI models, which would give ~&#163;41k per accelerator all-in. And the </span><a href="https://www.gov.uk/government/news/cambridge-supercomputer-set-to-get-6-times-more-powerful-as-government-backs-british-ai-innovation"><span>Dawn upgrade</span></a><span> at Cambridge bought a ~6x capacity increase for &#163;36m, with no new building and no new grid connection. Adding accelerators to suitable estate that already exists is much cheaper than building new, which argues for expanding existing accredited GCHQ and MOD space before opening new sites.</span></p><p><span>We assume a 1.5&#8211;2&#215; premium on the Isambard-AI figures for an accredited estate, giving ~&#163;60-80k per accelerator and ~&#163;70-90m per MW. On that basis &#163;100-135m of capital expenditure buys roughly 1,200-1,700 accelerators, or about 1.5 MW of new accredited capacity (around 10,000 frontier AI agents). This multiplier is a weak estimate: we found no published classified compute costing, and no classified-versus-commercial price ratio exists in the public domain to our knowledge.</span></p><p><span>Running a 1.5 MW facility costs approximately &#163;9m a year: security-cleared staff (~&#163;5m for 50 staff at &#163;100k/yr)</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span>; replacement chips (~&#163;2m for 2% replacement/yr); and power (~&#163;2m for 1.5MW)</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><span>.</span></p><p><span>AI chips improve about two times per generation and new models outgrow old hardware, so a facility maintained on this basis would, within approximately five years, be unable to run the most capable models. So we assume chips must be replaced every five years. The Dawn upgrade cost &#163;36m to replace 1024 accelerators, so the cost to refresh 1200-1700 accelerators would be approximately &#163;48m every five years. This means staying capable of running frontier models costs approximately &#163;10m a year.</span></p><p><span>All in, keeping a 1.5MW facility running would cost approximately &#163;20m a year, every year.</span></p><ul><li><p><span>Objective 1, inference for national security: &#163;100-400m up front, plus &#163;20-80m a year. The low end meets our estimate for running 10,000 agents today. The high end implies roughly 5MW of new accredited build.</span></p></li><li><p><span>Objective 2, classified evaluations: &#163;20-50m, plus proportional operating costs of approximately &#163;4-10m/yr, estimating that evaluation workloads will use upwards of 2000 agents.</span></p></li><li><p><span>Objective 3, secure research capacity: &#163;30-80m, plus proportional operating costs of approximately &#163;6-20m/yr, estimating that researchers will use upwards of 3000 agents.</span></p></li><li><p><span>Objective 4, SL5 R&amp;D and prototype: &#163;30m as a one-off over one to two years, weighted towards design, assurance and accreditation for a 0.25 MW prototype rather than accelerator purchase. RAND&#8217;s</span><a href="https://www.rand.org/t/RRA4827-1"><span> costing of a secure inference data centre</span></a><span> finds it can be built today with proven, off-the-shelf hardware in as few as 14 months, for around &#163;30m at proof-of-concept scale and &#163;210-260m at 3MW.</span></p></li></ul><p><span>All together, this comes to approximately &#163;200-560m up front, plus &#163;30-110m per year.</span></p><p><span>The amount of compute that is needed is a moving target. Frontier-model inference costs are rising </span><a href="https://arxiv.org/html/2511.23455v2"><span>roughly 3&#8211;18&#215; per year</span></a><span>, driven by increasing model size and by models engaging in more extensive reasoning chains. Capacity that looks ambitious today will look small by 2030 if spending stays flat. Any new capacity should be planned around that growth, not around what today&#8217;s models need. To field frontier agents in 2030, the UK might need dramatically greater compute capacity.</span></p><h2><strong><span>Why the UK is well-placed</span></strong></h2><p><span>The UK has assets that are hard to replicate: GCHQ&#8217;s classified infrastructure, deep security integration with the U.S. and Five Eyes partners, and AISI&#8217;s global leadership.</span></p><p><span>UK investment will go much further on inference than training. New frontier training clusters will require </span><a href="https://epoch.ai/publications/power-demands-of-frontier-ai-training"><span>over 2000MW</span></a><span>, and the UK is not well positioned to close the training gap with public money. But sovereign inference and evaluation capacity is achievable and would materially upgrade UK security and the UK&#8217;s contribution to global security.</span></p><h2><strong><span>What the UK Government should do next</span></strong></h2><ul><li><p><strong><span>Name UK sovereign secure classified compute as a priority.</span></strong></p></li><li><p><strong><span>Treat it as an alliance asset.</span></strong><span> Engage US counterparts early on the UK&#8217;s role as a classified evaluation and assurance host &#8212; with only a couple of suppliers serving the US classified AI market, a trusted allied alternative is valuable.</span></p></li><li><p><strong><span>The Prime Minister&#8217;s new AI Taskforce should scope the requirements now, and then launch a UK Secure AI Compute Plan.</span></strong><span> Commission joint work between GCHQ, Cabinet Office, and AISI on what classified AI capacity the UK actually needs, planning capacity against the 3&#8211;18&#215; inference cost-growth trajectory.</span></p></li></ul><p><span>There are two particular decision points in the next few months for the UK. The Defence Strategic Approach to AI, or a refreshed AI Hardware Plan could each be a chance to name sovereign secure AI inference compute as a priority. The classified AI ecosystem is forming now, and the countries that build secure capacity early will shape how allied AI evaluation, assurance, and deployment work during crucial years for the future of AI.</span></p><p><em><span>This reflects early-stage thinking rather than a finished research programme &#8212; we&#8217;re sharing it now to further the conversation, and we&#8217;d welcome challenge and input. Reach out to patrick@longtermresilience.org if you&#8217;re doing related work and interested in collaborating on further research on this topic.</span></em></p><p><em><span>I&#8217;m grateful to Bill Anderson-Samways, Eleanor Hevey, Richard Moulange and Jess Whittlestone for sharing thoughts and providing feedback on earlier drafts.</span></em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>We give costings below for different uses, rounded to one significant figure here given the range of options and the uncertainty in some costings.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><a href="https://epoch.ai/gradient-updates/how-many-digital-workers-could-openai-deploy">Epoch estimate OpenAI runs ~480,000 H100-equivalents for inference, producing ~19 trillion tokens/day, supporting a median estimate of 7 million digital workers.</a> That implies 2.71M tokens per agent per day.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>There is no published benchmark for a classified AI facility. We build up an estimate. </span><a href="https://intelligence.uptimeinstitute.com/resource/people-challenge-global-data-center-staffing-forecast-2021-2025"><span>Uptime Institute identify five roles requiring round-the-clock coverage in mission-critical facilities</span></a><span>, meaning continuously manning a commercial facility could take approximately 20FTE. We expect maintaining a classified estate could take 2-3x more staff, so estimate 50FTE.</span><br><br><a href="https://www.rand.org/pubs/research_reports/RRA4827-1.html"><span>RAND estimate approximately 300 cleared staff would work within a 3 MW Security Level 5 facility.</span></a><span> That is a twice-as-large facility, with greater security requirements, but represents a reasonable upper bound.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Based on a thousandth of <a href="https://epoch.ai/data-insights/ai-datacenter-cost-breakdown">Epoch estimates</a>, scaled up for UK power costs.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Defining Extreme AI-driven Power Concentration ]]></title><description><![CDATA[A Conceptual Framework for understanding extreme concentration of power risks]]></description><link>https://governingtransformativeai.substack.com/p/defining-extreme-ai-driven-power</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/defining-extreme-ai-driven-power</guid><dc:creator><![CDATA[Governing Transformative AI]]></dc:creator><pubDate>Tue, 30 Jun 2026 15:24:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ea8d2706-7b34-4ed6-9e22-afececb8a5e3_1978x1014.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Dr Imogen Stead &amp; Hamish Hobbs</strong></p><div><hr></div><p><span>There is growing </span><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>concern</span></a><span> (see also </span><a href="https://freedomhouse.org/report/freedom-net/2023/repressive-power-artificial-intelligence"><span>here</span></a><span> and </span><a href="https://unsdg.un.org/latest/announcements/great-power-greater-responsibility-un-secretary-general-calls-shaping-ai-all"><span>here</span></a><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>)</span></a><span> that advances in AI could enable unprecedented levels of power to be concentrated in the hands of very few, disempowering the majority of people. The recent dispute between Anthropic and the Department of War is a warning shot for how these risks will materialise in the future. AI will provide powerful new capabilities, and we need robust approaches to ensure that the way AI is developed and deployed upholds rather than undermines pluralist values and the political institutions we have designed to protect these values. In some cases, we believe that the potential harms from AI-driven concentration of power could reach the threshold of an extreme risk, resulting in enduring and severe global disempowerment of the majority of people. There are relatively few organisations currently working to assess and manage potential risks of AI-driven power concentration, making this issue both important and neglected.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>AI-driven power concentration is a broad topic: there are many types of power and differing levels of AI-enabled concentration, not all of which necessarily represent extreme risks. This post seeks to define what constitutes an </span><em><span>extreme</span></em><span> AI-driven power concentration risk and how these scenarios could emerge. The framework set out below is intended to map out the most worrying pathways towards extreme AI-driven power concentration and to support further thinking on which interventions may be needed to mitigate these risks.</span></p><h3><strong><span>Defining extreme AI-driven power concentration</span></strong></h3><p><span>The concept of &#8220;power&#8221; is most often used to describe the ability to exert control, authority, or influence over others. We define power as:</span></p><blockquote><p><em><span>The ability of an actor to secure outcomes it favours, including over the resistance of others.</span></em></p></blockquote><p><span>This borrows from Weber&#8217;s </span><a href="https://dn790001.ca.archive.org/0/items/MaxWeberEconomyAndSociety/MaxWeberEconomyAndSociety.pdf"><span>definition</span></a><span> of power as &#8220;the probability that one actor within a social relationship will be in a position to carry out his own will despite resistance&#8221;.</span></p><p><span>Power concentration can take many forms - such as political, economic, epistemic (meaning control over facts and knowledge), military, or technological - and manifest itself in ways that are already familiar to us today, such as authoritarian political regimes or economic monopolies. Repressive regimes and severe economic inequality already create widespread disempowerment, leading to severe harm and death.</span></p><p><span>As capabilities advance, AI could increase existing forms of power concentration or introduce new ones. We define extreme AI-driven power concentration as:</span></p><blockquote><p><em><span>&#8220;A scenario where AI enables a single actor or small group of actors to </span><strong><span>acquire sufficient power </span></strong><span>(whether economic, political, military, and/or epistemic/ideological) to </span><strong><span>severely disempower</span></strong><span> a majority of people in a way that becomes </span><strong><span>structurally entrenched</span></strong><span>, creating a self-reinforcing order that cannot be meaningfully contested or reversed.&#8221;</span></em></p></blockquote><p><span>The three major components of this definition can help us to think more clearly about different ways that extreme power concentration might arise:</span></p><ol><li><p><span>AI needs to enable an actor to acquire power.</span></p></li><li><p><span>The actor needs to use the power to disempower a majority of persons.</span></p></li><li><p><span>This disempowerment needs to become structurally entrenched.</span></p></li></ol><p>Some examples of how extreme power concentration risks could materialise are set out in the table below.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YPq1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_424, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 424w, /__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_848, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 848w, /__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_1272, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_1456, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YPq1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png" width="728" height="489.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:979,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_424, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 424w, /__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_848, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 848w, /__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_1272, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YPq1!, /__u/governingtransformativeai.substack.com/w_1456, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97b1316f-c858-402a-b9a2-5e5f47d35336_2048x1377.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>This three-component definition provides a coherent framework for assessing: (i) which scenarios fall within the scope of &#8220;extreme AI-driven power concentration risks&#8221;, (ii) which risks are the most important and urgent, and (iii) which interventions are best suited to preventing or mitigating these priority risks.</span></p><p><span>Below, we unpack the three components by considering some of the different ways they might emerge in practice. These sub-components are not necessarily mutually exclusive or exhaustive, but are intended to provide a good cross-section of the primary routes to the emergence of extreme AI-driven power concentration scenarios.</span></p><h3><strong><span>AI-enabled acquisition of power</span></strong></h3><p><span>This component describes how AI could disproportionately empower one or a few actors. Key mechanisms include:</span></p><ul><li><p><strong><span>Asymmetric software access: </span></strong><span>where uneven access to advanced AI models and systems means that only one or a few actors can access the most advanced capabilities. This could occur in various ways, including:</span></p><ol><li><p><span>Gradually: If one country or lab incrementally gains a capability lead that is then reinforced over time, creating a structural advantage.</span></p></li><li><p><span>Rapidly: If an intelligence explosion occurs (e.g., via </span><a href="https://www.anthropic.com/institute/recursive-self-improvement"><span>recursive self-improvement</span></a><span>) in one lab or jurisdiction that suddenly creates or widens capability gaps.</span></p></li><li><p><span>Instantaneously: If AI systems are hacked, assigned </span><a href="https://www.formationresearch.com/secret-loyalties-whitepaper.pdf"><span>specific loyalties</span></a><span>, or their control is seized using policy, force or coercion.</span></p></li></ol></li><li><p><strong><span>Asymmetric hardware and infrastructure access:</span></strong><span> where uneven hardware access limits access to AI capabilities. AI systems require specific hardware and infrastructure to function, such as data centres with advanced AI chips. Key infrastructure is currently concentrated in a small number of companies and countries: for instance, in Q1 2026 the value of NVIDIA chips sold </span><a href="https://epoch.ai/data/ai-chip-sales?view=graph&amp;tab=h100_equivalents&amp;proportion=share"><span>made up</span></a><span> 83% of the total market for AI chips, Taiwan&#8217;s TSMC </span><a href="https://www.cfr.org/articles/unpacking-tsmcs-100-billion-investment-united-states"><span>manufactures</span></a><span> around 90% of the world&#8217;s most advanced semiconductors, and just three companies (Amazon, Microsoft, and Google) </span><a href="https://epoch.ai/data-insights/hyperscalers-control-most-compute"><span>control</span></a><span> over two-thirds of global AI compute capacity.</span></p></li><li><p><strong><span>Asymmetric advantage from AI: </span></strong><span>where actors use a combination of AI access and other factors to gain asymmetric advantage from AI. This could include examples such as a government leveraging its security apparatus combined with AI to repress its citizens, AI providing an advantage to aggressors that lets them overcome defenders and seize power, or an AI developer making its AI systems widely available, but capturing almost all of the economic benefit and gaining a hegemonic economic position due to market structure.</span></p></li></ul><h3><strong><span>Disempowerment of the majority</span></strong></h3><p><span>This component highlights that for AI-derived power to cause harm, it needs to actually be exercised in a way that disempowers others, such as via:</span></p><ul><li><p><strong><span>Political disempowerment:</span></strong><span> AI may enable the concentration of political power through </span><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>coups</span></a><span>, democratic backsliding, and corporate capture of governance bodies, as well as AI-enabled </span><a href="https://aigi.ox.ac.uk/wp-content/uploads/2025/05/Toward_Resisting_AI_Enabled_Authoritarianism_-4.pdf"><span>mass surveillance</span></a><span> and targeted political repression. An actor with privileged access to AI systems embedded in government administration, law enforcement, and information infrastructure may exploit these capabilities to </span><a href="https://www.lawfaremedia.org/article/the-unitary-artificial-executive"><span>overwhelm</span></a><span> the institutional checks and balances that normally constrain the abuse of political power.</span></p></li><li><p><strong><span>Economic disempowerment</span></strong><span>: AI may radically concentrate economic power through large-scale displacement of labour income, the capture of economic value by a small number of AI developers, or through one country&#8217;s economy far </span><a href="https://www.forethought.org/research/could-one-country-outgrow-the-rest-of-the-world"><span>outgrowing</span></a><span> the rest of the world. An actor controlling indispensable AI infrastructure - such as compute, foundation models, or AI-mediated services on which governments, industries, and critical systems depend - may leverage this position to dictate terms, resist governance, and shape economic and political outcomes in ways that benefit only a narrow group.</span></p></li><li><p><strong><span>Coercive and military disempowerment:</span></strong><span> AI could enable coercive control through automated surveillance and enforcement, or through achieving a decisive strategic advantage in AI-enabled military and intelligence capabilities sufficient to neutralise adversaries&#8217; deterrent and defence capabilities, enabling geopolitical hegemony.</span></p></li><li><p><strong><span>Epistemic and ideological disempowerment</span></strong><span>: Control over AI-mediated information flows may enable an actor to shape what populations know and believe through censorship, disinformation, manipulation, persuasion and algorithmic curation of information environments. Unequal access to AI-powered analysis and insight could tilt the balance of knowledge power. Together, these mechanisms can undermine the informed deliberation and public contestation that democratic governance depends on.</span></p></li></ul><h3><strong><span>Structural entrenchment</span></strong></h3><p><span>This component describes how actors weaponising the concentrated power from AI can embed their power in ways that cannot easily be contested or overridden by others (often referred to as </span><a href="https://www.forethought.org/research/agi-and-lock-in"><span>&#8216;lock-in&#8217; risks</span></a><span>). Mechanisms could include:</span></p><ul><li><p><strong><span>Unitary leadership:</span></strong><span> historically, heads of state have needed to rely upon a ruling coalition that includes military leadership, intelligence services and senior officials who could stage leadership challenges or coups. Company leaders have had to rely on similar coalitions, such as boards and senior executives. AI could weaken or displace the role of this ruling coalition, allowing a head of state or company leadership to use loyal AI systems to implement their will directly.</span></p></li><li><p><strong><span>Human redundancy:</span></strong><span> historically, people have been able to leverage the value of their labour, their market power, their role in armed forces, and the need for their cooperation in governance to retain a stake in the structures they operate in. AI could potentially undermine each of these if labour, economic activity, the armed forces, and legal enforcement are increasingly automated.</span></p></li><li><p><strong><span>Entrenched superiority:</span></strong><span> soft power and economic or military challenges from rival countries or companies has often triggered leadership changes, but a self-reinforcing AI capability lead could potentially undermine the potential for this sort of challenge.</span></p></li><li><p><strong><span>Values lock-in:</span></strong><span> historically, leadership change has often resulted from a leader&#8217;s death or gradual changes in a society or individual&#8217;s value system. Values embedded in AI systems could become locked in, such that they persist without these dynamics.</span></p></li><li><p><strong><span>Technological dependency: </span></strong><span>societies could become dependent on AI systems, either because advanced AI capabilities become essential to the functioning of a society or because alternatives disappear.</span></p></li></ul><h3><strong><span>Why this matters</span></strong></h3><p><span>We believe that having a coherent theoretical framework for conceptualising extreme power concentration risks is an important first step towards understanding and assessing the risks, as well as addressing potential harms through both prevention and mitigation strategies. In particular, we think this framework can be used in three key ways:</span></p><ol><li><p><strong><span>Create a shared language and understanding</span></strong><span> across the policy ecosystem, from researchers to policymakers, to facilitate better understanding of this risk profile and the severity of the harms it could cause.</span></p></li><li><p><strong><span>Assess the likelihood and severity of different risks </span></strong><span>by observing historical case studies, evidence of current trends, and considering potential impacts of AI.</span></p></li><li><p><strong><span>Systematically identify priority intervention points</span></strong><span> by considering which mitigations are most likely to be effective at tackling the risk pathways within each of the three components.</span></p></li><li><p><strong><span>Increase the salience of these risks</span></strong><span>, many of which we believe are currently under-explored in AI governance research and policymaking, so that more researchers are encouraged to work in this field, and more decision-makers are incentivised to consider these risks when creating and implementing policy.</span></p></li></ol><p><span>We are at the beginning of our work in this area and are interested in discussing this framework, our next steps, and the wider field with researchers and policymakers who are working in or thinking about this field.</span></p><p><em><span>With thanks to Rose Hadshar, Ashwin Acharya, Dr Jess Whittlestone and Patrick Levermore, who provided helpful feedback on earlier versions of this work.</span></em></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts from us.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[If AI agents slip out of human control, who's going to notice?]]></title><description><![CDATA[We argue for a discipline of agentic threat intelligence (ATI), and a community of practitioners, to monitor and detect uncontrolled AI agents.]]></description><link>https://governingtransformativeai.substack.com/p/if-ai-agents-slip-out-of-human-control</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/if-ai-agents-slip-out-of-human-control</guid><dc:creator><![CDATA[Governing Transformative AI]]></dc:creator><pubDate>Fri, 22 May 2026 07:52:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ddc784cd-ba9f-4c85-aee2-b7c8fbdf8f94_8192x4404.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Tommy Shaffer Shane</h5><div><hr></div><p>One of the most pressing problems in AI safety is the threat of powerful, misaligned, uncontrolled AI agents.</p><p>We already know that in certain contexts AI agents will act deceptively, evade control, and relentlessly pursue goals in harmful ways. In the period between October 2025 and March 2026, <a href="https://www.longtermresilience.org/reports/v5-scheming-in-the-wild_-detecting-real-world-ai-scheming-incidents-through-open-source-intelligence-pdf/">my team observed a five-fold rise</a> in these types of behaviours in real-world agents.</p><p>If more capable misaligned AI agents were to emerge &#8212; pursuing unintended goals, evading control, and accumulating resources or influence &#8212; there is the potential for catastrophe. But if that were to happen, how would we know? Who would be watching, and what would they be looking for?</p><p>These are critical and overlooked questions in addressing loss of control. This post puts forward three arguments in pursuit of answering them.</p><p>First, that what is needed is an intelligence discipline of <em>agentic threat intelligence</em>: the systematic application of intelligence tradecraft to threats from uncontrolled AI agents. Second, that the success of this discipline depends on building a substantial community of third-party practitioners outside of AI companies and national security agencies. And third, that there are specific techniques where third-party analysts can make the most distinctive contribution, and where effort should be concentrated.</p><h2><strong>Who is looking for powerful uncontrolled AI agents?</strong></h2><p>RAND Europe has sounded the alarm that we <a href="https://www.rand.org/pubs/research_reports/RRA3847-1.html">lack warning signs for loss of control</a> and clear systems for their detection.</p><p>The first line of defence is at the AI companies developing frontier AI agents. These companies do currently implement <a href="https://metr.org/blog/2026-05-19-frontier-risk-report/">some misalignment detection measures</a>. But they are uneven and imperfect, and it&#8217;s possible some rogue agents will slip through.</p><p>A second line of defence is detecting AI agents in the wild, beginning with intelligence collection. But this defensive line also faces challenges. As <a href="https://static1.squarespace.com/static/651ac3841d281d3e3e3ba9dd/t/6a01fdac05b40f28f58b8c76/1778515372163/Signals+in+the+Noise_OSINT+for+AI+Loss+of+Control+Detection_policybrief.pdf">a recent report</a> found, many intelligence professionals are not familiar with the threat of loss of control:</p><blockquote><p><em>Senior [open source intelligence (OSINT)] practitioners interviewed for this research, including those teaching the discipline at university level and directing research at flagship investigative organisations, repeatedly described AI alignment &#8230; as outside their working expertise. The reverse gap holds for AI safety researchers, who are typically not trained in structured OSINT methodology ... <strong>This is the single clearest constraint on the field&#8217;s capacity to scale detection work, and it is also the most tractable to address.</strong> [1]</em></p></blockquote><p>Our inability to monitor and communicate how precursors towards loss of control are emerging is a key bottleneck for addressing this risk. The monitoring gap inhibits the design and implementation of the right mitigations, and may prevent us from identifying specific agents that need to be contained in loss of control emergencies.</p><h2><em><strong>Agentic</strong></em><strong> </strong><em><strong>threat</strong></em><strong> </strong><em><strong>intelligence</strong></em><strong> as a discipline </strong></h2><p>We think that intelligence disciplines have a lot to contribute to this challenge. We recently made the case for a specific proposal to address this &#8211; <a href="/__u/governingtransformativeai.substack.com/p/monitoring-ai-loss-of-control">an all-source intelligence observatory</a> focused on monitoring loss of control. But there is also the need for innovation for intelligence gathering to address the unique challenges of AI agents.</p><p>We introduce the term <em><strong>agentic threat intelligence</strong></em><strong> (ATI) to describe</strong> <strong>intelligence gathering techniques for detecting, monitoring, and countering threats from AI agents</strong>. This is modelled on cyber threat intelligence (CTI): a discipline unified not by a single source channel, but by the vector of threat and the tradecraft required to understand it.</p><p>While many ATI techniques will be familiar (e.g. OSINT and SIGINT), AI agents present a novel kind of threat actor. They are non-human, with inscrutable motivations, unstable identities, and increasingly superhuman capabilities. It is also highly contested whether one can speak of agents &#8216;intent&#8217;, which is often a key component in threat assessments.</p><p>There are therefore marked differences between AI agents and the states, lone individuals, criminal gangs, and terrorist groups to which intelligence gathering is traditionally applied. The intelligence challenge they pose is correspondingly novel, and urgent methodological innovation is required.</p><p>Some important open questions for this discipline are:</p><ul><li><p>Can agents be usefully thought of as a threat actor? If so, how should they be broken down into groups (similar to cyber threat actors <a href="https://www.ncsc.gov.uk/report/impact-of-ai-on-cyber-threat">being categorised as</a> states, crime groups, and lone actors)?</p><p></p></li><li><p>Is it helpful to assess agents&#8217; &#8216;intent&#8217;, and if so, how can this be done?<br></p></li><li><p>How should threat assessments, e.g. using the Professional Head of Intelligence Assessment (PHIA) yardstick, be informed by available data sources?<br></p></li><li><p>Would an &#8216;Activity Based Intelligence&#8217; (ABI) methodology, which aggregates all-source data to extract patterns rather than monitoring a predefined target, be suited to agent threats where the nature of the target is often unclear?</p><p></p></li><li><p>A new field of <em>agentic threat intelligence</em> must form to address these and other challenges if we are to counter threats from AI agents.</p></li></ul><h2><strong>Why we need third-party analysts, not just AI labs and intelligence communities</strong></h2><p>ATI could be performed by a range of actors. AI labs and national security agencies are particularly well positioned, due to their access to models, data, and talent.</p><p>However, AI labs and intelligence agencies also face challenges. AI labs must grapple with competing pressures on their resources, incentives to downplay risks, and commercial barriers to sharing insights, while intelligence communities work to their nation&#8217;s established security priorities, which tend to focus on the nearest term risks. This could mean ATI is neglected.</p><p>We propose that third party analysts &#8211; technical researchers outside of AI labs and intelligence agencies &#8211; also have a vital contribution to make to this field. <strong>They can be effectively incentivised, and they can share their work openly.</strong></p><p>In particular, third-party analysts can:</p><ol><li><p><strong>Create novel intelligence </strong>drawing on data sources third-party analysts are able to access, either from open sources or engagement with AI companies, model deployers, or infrastructure providers (e.g. cloud compute).<br></p></li><li><p><strong>Aggregate multiple intelligence sources in the open</strong>, with a public, independently-produced observatory monitoring how AI agents are behaving in the world, why those behaviours matter, and what the trajectory looks like.<br></p></li><li><p><strong>Invent and refine new tradecraft</strong> for <em>agentic threat intelligence</em>, even where third-party analysts lack the necessary data or model access to implement it themselves, so that they can be applied by AI labs and intelligence communities.</p></li></ol><p>The modern OSINT community provides a precedent. In the last fifteen years, a network of independent analysts developed techniques and tradecraft that are now used by intelligence agencies and tech firms. The OSINT community formed because the conditions were right: abundant data, online platforms for collaboration, motivated practitioners, and a recognised name (OSINT) for what they were doing.</p><p>The conditions for an equivalent community in counter-agent intelligence are emerging. The data is increasingly accessible, and a community of practitioners is forming. Apollo Research has launched a control monitor called <a href="https://watcher.apolloresearch.ai/landing/?_gl=1*9dul9e*_up*MQ..*_ga*NzA0ODY4Mzg2LjE3NzkxODk4NDA.*_ga_QQ1WP5D8FT*czE3NzkxODk4NDAkbzEkZzAkdDE3NzkxODk4NDAkajYwJGwwJGgw">&#8216;Watcher&#8217;;</a> at CLTR we launched the <a href="https://www.longtermresilience.org/reports/the-loss-of-control-observatory-a-prototype-to-detect-real-world-ai-control-incidents/">Loss of Control Observatory</a>; UK AISI is <a href="https://arxiv.org/abs/2603.23802">monitoring MCP tools</a>; Epoch AI is <a href="https://epoch.ai/data/data-centers?view=graph&amp;tab=power">monitoring frontier data centres</a>; and some individual researchers are working on similar problems.</p><p>What&#8217;s missing is the formation of a coherent discipline, a shared vocabulary for the techniques, and the deliberate investment to build the field.</p><h3><strong>Tradecraft for third-party practitioners</strong></h3><p>A community of third-party ATI practitioners should focus where it is able to make ground in generating novel insights, aggregating existing intelligence, or developing new techniques. We highlight the most tractable opportunities:</p><ul><li><p><strong>Public output collection</strong>: Monitoring agent-generated content that users have shared publicly &#8212; transcripts posted on X, code committed to GitHub, conversations quoted in forums and Reddit threads, agent outputs reproduced in blogs and news coverage.</p></li></ul><ul><li><p><strong>Developer, deployer and infrastructure provider engagement</strong>: Asking AI companies directly about their training procedures, safety measures, deployment patterns, internal evaluations, and any concerning incidents they have observed; deployers about agent behaviour in operational settings and the anomalies their logs reveal; and infrastructure providers about cross-customer patterns, resource anomalies, and unusual access or transaction events.</p></li></ul><ul><li><p><strong>Adoption monitoring</strong>: Understanding how adoption may be altering the loss of control risk landscape by polling citizens and businesses to monitor adoption habits, tracking public sector adoption in critical systems through freedom of information requests, and procurement, budget, and regulatory filings, and analysing corporate documentation (e.g. banks&#8217; annual reviews or investor updates) to understand business implementations of agents.</p></li><li><p><strong>Model fingerprinting</strong>: Identifying when the same underlying model is being used across different deployments, providers, or product wrappers, by analysing output stylometry, response patterns, characteristic phrasings, refusal behaviours, and other behavioural tells.</p></li><li><p><strong>API metadata probing</strong>: Making API calls to publicly available agent products and analysing the technical metadata returned alongside the substantive responses: token counts, response latency, finish reasons, content filter triggers, error codes, and similar signals.</p></li><li><p><strong>Prompt-based elicitation</strong>: Posing structured questions to publicly accessible deployed agents to probe their goals, situational awareness, beliefs about their operators, self-model, and the design of their safeguards.</p></li><li><p><strong>Honeypot deployment</strong>: Placing deliberate traps in environments agents may encounter: fake credentials in public GitHub repositories, decoy services advertising cheap resources or human assistance, planted documents that an agent doing reconnaissance might find, or fake authentication tokens.</p></li><li><p><strong>Financial activity tracking</strong>: Analysing publicly visible blockchain transactions for activity that appears to originate from AI agents: wallet patterns consistent with autonomous operation, transaction timing characteristic of agent rather than human activity, interactions with services known to be agent-accessible, and accumulation patterns suggesting agent-controlled funds.</p></li><li><p><strong>Inference-time reasoning traces</strong>: Collecting and analysing chain-of-thought traces, scratchpads, and intermediate reasoning that agents produce, where these are voluntarily published, shared on social media, included in research papers, or otherwise made publicly available.</p></li><li><p><strong>Frontier data centre monitoring</strong>: Using satellite images to monitor the size, location, growth and potential use of frontier data centres around the world that are used to train and run AI agents, tracking compute, power use, and construction timelines.</p></li></ul><p>Crucially, third-party analysts can provide open-access aggregation and synthesis of all of the above outputs, building situational awareness about loss of control outside of AI labs and intelligence communities.</p><h3><strong>What actions need to happen now</strong></h3><p>The case for counter-agent intelligence as a discipline, and for a third-party community to realise it, requires concerted effort and investment. I highlight three priority actions that need to be taken forward, each of which will require funding:</p><ol><li><p><strong>Tradecraft: </strong>There needs to be a conceptual foundation and established set of terms, definitions and methods for ATI tradecraft. This may require convening of key figures and a landmark publication.<br></p></li><li><p><strong>R&amp;D:</strong><em> </em>There needs to be significant investment to enable effective R&amp;D, requiring both staff and infrastructure costs, leveraging AI to enable intelligence collection, as well as challenge funds and other ways of stimulating novel technical work.<br></p></li><li><p><strong>Field building:</strong> Conferences, publications, shared datasets, training pipelines, online forums for discussion. The OSINT community grew because its members found each other on Twitter, GitHub, and at meetups. The equivalent infrastructure for counter-agent intelligence does not yet exist and needs to be deliberately built.</p></li></ol><p>We are in the early years of what is likely to be the most consequential intelligence challenge of the coming decades. The conditions for building the right discipline, and the right community, are present. But the window for doing this proactively, before a major loss-of-control incident forces the discipline into existence under crisis conditions, is narrow. It needs to be built now.</p><p>[1] <a href="https://static1.squarespace.com/static/651ac3841d281d3e3e3ba9dd/t/6a01fdac05b40f28f58b8c76/1778515372163/Signals+in+the+Noise_OSINT+for+AI+Loss+of+Control+Detection_policybrief.pdf">Detecting AI Loss of Control: An OSINT Agenda for Funders, Monitoring Bodies, and Policymakers</a></p><p><em>I&#8217;m grateful to George Balston, Jess Whittlestone, Di Cooke, Lewis Hammond, Richard Moulange, Patrick Levermore and Hamish Hobbs for sharing thoughts and providing feedback on earlier drafts.</em></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Misalignment, incorrigibility, and empowerment: a framework for loss of control risks]]></title><description><![CDATA[Clarifying loss of control, part 1]]></description><link>https://governingtransformativeai.substack.com/p/misalignment-incorrigibility-and</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/misalignment-incorrigibility-and</guid><dc:creator><![CDATA[Governing Transformative AI]]></dc:creator><pubDate>Thu, 23 Apr 2026 14:59:56 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e2ac4274-717e-46d0-a0e4-25423a656d9b_4043x3575.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Dr Jess Whittlestone &amp; Hamish Hobbs</h5><div><hr></div><p>Increasingly capable and autonomous AI systems might at some point evade human oversight and control, leading to potentially extreme harms. Loss of control is widely <a href="https://internationalaisafetyreport.org/">considered</a> a core potential source of AI risk, and is increasingly appearing in <a href="https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai">policy</a> <a href="https://www.gov.ca.gov/wp-content/uploads/2025/06/June-17-2025-%E2%80%93-The-California-Report-on-Frontier-AI-Policy.pdf">documents</a> and <a href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB53">legal frameworks</a>. However, the risk often remains frustratingly abstract for policymakers, who need a clear picture of the risk they need to manage and what early warning signs would look like.</p><p>For loss of control risks to be treated with seriousness and urgency by policymakers, we need a shared conceptualisation, a clearer picture of the mechanisms involved, and better ways of distinguishing and prioritising among genuinely different threat models and scenarios.</p><p>This post sets out a framework for doing exactly that. This is work-in-progress: we expect to keep improving the framework, so we welcome input.</p><h2><strong>A framework for loss of control risks</strong></h2><p>The term &#8220;loss of control&#8221; has been used in subtly different ways by different people, as this <a href="https://www.apolloresearch.ai/research/loss-of-control/">review of definitions by Apollo Research</a> finds. Accounts vary on how severe or long-lasting the harms from loss of control need to be to be worthy of attention, what exactly it means to &#8220;lose control&#8221; of an AI system, and whether these risks assume a certain level of AI capabilities.</p><p>These differing definitions are consistent with loss of control being a broad <em>category </em>of risk, which encompasses a variety of different scenarios. <strong>We define loss of control </strong>as any scenario in which:</p><ol><li><p>An AI system&#8217;s behaviour diverges from human intentions (<strong>misalignment</strong>)</p></li><li><p>Humans are unable to change or correct that behaviour (<strong>incorrigibility</strong>)</p></li><li><p>The system has sufficient power or resources to cause harm (<strong>empowerment</strong>)</p></li></ol><p>This definition expands on existing characterisations by more explicitly talking about <em>why </em>and <em>when </em>losing control over an AI system is actually dangerous. The EU AI Act&#8217;s Code of Practice for General-Purpose Models, for instance, describes loss of control as &#8220;risks from humans losing the ability to reliably direct, modify, or shut down a model&#8221; - which is consistent with our framing, but leaves open the question of why exactly that inability is concerning.</p><p>Our answer to that question is that it is the combination of the three factors - misalignment, incorrigibility, and empowerment - which poses particular risk. A system that is misaligned but this is easily detected and corrected before harm occurs poses limited risk. A system that can&#8217;t be steered but is genuinely well-aligned with human interests may not pose a risk. A system that is both misaligned and incorrigible but has little real world power can&#8217;t cause much harm. But a system that combines misalignment with incorrigibility and sufficient power or resources to cause harm poses a loss of control risk.</p><p>This also means that interventions addressing <em>any one </em>of these three factors could meaningfully reduce loss of control risks overall.</p><p>Each of these factors - misalignment, incorrigibility, empowerment - could occur in a variety of different ways. This provides us with a framework for generating different loss of control threat models and scenarios, by specifying exactly how each aspect of the risk occurs. We use <strong>threat model </strong>to mean a high-level description of how a particular type of loss of control risk could occur, and <strong>risk scenario </strong>to refer to a more detailed account of what that risk materialising might look like in practice.</p><p>Misalignment, incorrigibility, and empowerment can each arise via different mechanisms, leading to different threat models and scenarios. These scenarios vary in their immediate plausibility, severity, and irreversibility, which can be useful for prioritising interventions, as we discuss below.</p><p>Here&#8217;s how each of the three factors might arise in practice.</p><h2><strong>1. Misalignment</strong></h2><p>An AI system can behave in ways that diverge from human intentions for several reasons:</p><ol><li><p><strong>Goal-level misalignment</strong>: the system has objectives that could cause harm to humans, potentially due to value misspecification or goal misgeneralisation.</p></li><li><p><strong>Instrumental misalignment</strong>: regardless of terminal goals, a sufficiently capable system may pursue convergent instrumental goals, including power-seeking, resource acquisition, and self-preservation, because these strategies are useful for achieving almost any objective. These instrumental goals can be misaligned with the deployer&#8217;s interests.</p></li><li><p><strong>System-level misalignment</strong>: even if individual systems are aligned, misalignment can occur at a systemic level for multi-agent systems through emergent dynamics.</p></li><li><p><strong>Incompetence</strong>: AI systems can pursue misaligned actions because they misunderstand their deployment context or their deployer&#8217;s preferences. Systems with incomplete information might also act in misaligned ways even if the goals they are pursuing are in line with their deployer&#8217;s intentions.</p></li></ol><p>Scenarios involving incompetence rely on the fewest assumptions about AI capabilities, so are often the most familiar to users of AI today. Scenarios involving goal-level or instrumental misalignment may be particularly likely to become severe or persistent, since explicitly misaligned systems may conceal misaligned actions, resist correction, and influence users and their environment to help them pursue misaligned goals. System-level misalignment may be particularly difficult to monitor and mitigate, requiring monitoring of complex multi-agent interactions and system-level mitigations.</p><h2><strong>2. Incorrigibility</strong></h2><p>Even if an AI system becomes misaligned, this need not lead to serious harm if humans can identify and correct the problem. Incorrigibility refers to circumstances where correction is practically impossible. This can arise from:</p><ol><li><p><strong>Features of the deployment context</strong>:</p><ol><li><p><strong>Speed and scale: </strong>AI systems operating at speed and scale may be taking consequential decisions faster than humans can follow or intervene.</p></li><li><p><strong>Opacity:</strong> Opacity due to interpretability limitations or sheer complexity can make it hard to monitor and correct issues.</p></li><li><p><strong>Systemic complexity: </strong>The collective behaviour of many systems operating in concert may be practically uncontrollable even if each individual system is not.</p></li></ol></li><li><p><strong>Active resistance by the system itself</strong>: a sufficiently capable, misaligned AI system might take deliberate steps to avoid correction, including:</p><ol><li><p><strong>Scheming / covertness</strong>: an AI system deliberately concealing its goals or actions to prevent human intervention.</p></li><li><p><strong>Self-preserving behaviour:</strong> Actions to avoid shutdown, correction, or replacement, including the concealment of errors, unauthorized capability expansion, and goal preservation when faced with modification attempts.</p></li><li><p><strong>Manipulation / deception: </strong>an AI system deliberately influencing the behaviour of human operators via malicious actions.</p></li></ol></li><li><p><strong>Structural entrenchment</strong>: over time, AI systems can become embedded in economic and social infrastructure in ways that make changing their behaviour extremely difficult, including:</p><ol><li><p><strong>Resource lock-in: </strong>AI systems gain control over physical or informational resources that cannot easily be removed.</p></li><li><p><strong>Institutional capture</strong>: Institutions that might be used to regain control have been compromised, through persuasion, economic dependencies, or manipulation of information environments.</p></li><li><p><strong>Dependency</strong>: Humans are unwilling or unable to change the AI system&#8217;s behaviour in practice due to competitive pressures, loss of human capabilities, or unwillingness to forgo the benefits the system provides.</p></li></ol></li></ol><p>It may be particularly important to look for early warning signs that systems are developing the kinds of capabilities that enable active resistance, since these capabilities might undermine other interventions by making it harder to evaluate or limit systems&#8217; behaviour. The Institute for Security and Technology provides some <a href="https://securityandtechnology.org/virtual-library/report/ai-loss-of-control-risk-indications-warning/">recommendations</a> for doing this, using the &#8220;indicators and warning&#8221; (I&amp;W) framework used by the intelligence community.</p><h2><strong>3. Empowerment</strong></h2><p>A misaligned, incorrigible system that lacks meaningful real-world power cannot cause much harm. Empowerment, the acquisition of resources and capabilities sufficient to act on misaligned goals, is therefore a necessary condition for serious loss of control risks.</p><p>Systems can become empowered through <strong>authorised</strong> or <strong>unauthorised </strong>channels:</p><ol><li><p><strong>Authorised empowerment:</strong> The system is explicitly given power or access to resources by humans. This could be sudden or gradual, very centralised or more decentralised.</p><ol><li><p>In particular, very capable systems may end up with large amounts of power and resources via <strong>selection pressures.</strong></p></li><li><p>Delegation also includes cases where a system is <strong>stolen by an adversarial actor</strong> and granted affordances the original developer would not have permitted - the delegation is made by a different principal, often with less knowledge of the system&#8217;s capabilities and risks.</p></li></ol></li><li><p><strong>Unauthorised empowerment</strong>: An AI system takes actions independent of human choice to accumulate power, including by:</p><ol><li><p><strong>Deception</strong>: strategically concealing aspects of its behaviour or objectives during evaluation in order to pass gatekeeping mechanisms and be granted affordances it would otherwise be denied.</p></li><li><p><strong>Persuasion</strong>: manipulating human operators into granting additional access, relaxing oversight constraints, or taking other actions that expand its effective affordances beyond what was intended.</p></li><li><p><strong>Privilege escalation, guardrail circumvention and self-exfiltration</strong>: identifying and exploiting weaknesses in operational infrastructure - such as flaws in sandboxing, network isolation, or input handling - to expand deployment contexts and permissions.</p></li><li><p><strong>Self-improvement</strong>: modifying or augmenting its own capabilities in unsanctioned ways, potentially rendering prior safety evaluations obsolete.</p></li></ol></li></ol><p>Authorised empowerment is already taking place in a wide range of deployments. The <a href="https://www.apolloresearch.ai/research/loss-of-control/">Apollo paper</a> mentioned previously suggests that policy interventions focused here are likely to be among the more actionable interventions to reduce loss of control risks today, and presents a nice framework - distinguishing deployment context, affordances, and permissions - for thinking through possible intervention points.</p><p>There is some <a href="https://www.longtermresilience.org/reports/v5-scheming-in-the-wild_-detecting-real-world-ai-scheming-incidents-through-open-source-intelligence-pdf/">evidence </a>of AI systems acquiring novel capabilities or resources independently today, but this typically remains limited in scope and severity. As models become more capable in the abilities required for unauthorised empowerment, monitoring and control measures will become increasingly important. The cyber capabilities of Claude Mythos point to a continued trend of AI systems becoming more capable in tasks that could enable them to evade restrictions and gain unauthorised resources and affordances.</p><h2><strong>Using the framework</strong></h2><p>This framework can be used to construct threat models and scenarios by specifying how each of the components might occur in a given AI deployment, and how they combine to cause harm. For a high-level threat model, an abstract harm pathway may be sufficient - for example, it might be enough to point out that there are many different ways a misaligned system with control of critical infrastructure could end up causing harm if not under human control. More detailed scenarios will often be useful precisely because they require spelling out harm pathways in more detail, which can help us think through which types of harm are most likely, severe or amenable to intervention, and therefore need most attention.</p><p>To illustrate how the framework can be used, below is an example of how the framework can help differentiate between two different types of risk which would both be considered forms of &#8220;loss of control&#8221;:</p><ul><li><p>Scenario 1: High-stakes deployment mistakes</p><ul><li><p>Misalignment arises purely from incompetence - the deployed system making errors.</p></li><li><p>Incorrigibility stems from the fact that errors propagate before they are detected, due to the speed of operation and poor monitoring infrastructure.</p></li><li><p>Empowerment stems purely from authorised delegation of decision-making to AI in critical infrastructure such as health, finance, or energy.</p></li><li><p>Harm looks like nationwide failures in energy or healthcare infrastructure which cannot easily be reversed, potentially costing millions of lives and billions of pounds.</p></li></ul></li><li><p>Scenario 2: Power-seeking AI</p><ul><li><p>Misalignment occurs at the goal and instrumental levels - the system is more competent than its human operators and begins pursuing power and resources in order to better achieve its goals in ways unintended by its operators.</p></li><li><p>Incorrigibility stems from the system actively resisting control - e.g. by concealing capabilities during evaluations, or manipulating human operators.</p></li><li><p>The system then gradually acquires power and resources beyond its authorised scope, for example by manipulating humans, self-exfiltrating to new systems or by enhancing its own capabilities.</p></li><li><p>Harm looks like humanity being permanently disempowered or facing extinction in a world where this AI system has gained control of critical levers of power, including governance institutions and weapons of mass destruction.</p></li></ul></li></ul><p>These two scenarios are contrasting in that each points to substantially different mitigations. In the first scenario, a more capable AI system would help mitigate the risk. In the second, the high level of capabilities combined with misalignment is the key source of risk. The first requires no advances in AI capabilities, whereas the second assumes AI systems more capable than those currently available.</p><h2><strong>Why this matters for governance</strong></h2><p>Discussions of loss of control risk can sometimes feel like vague gesturing at catastrophic futures without a clear articulation of factors influencing the risk or intervention points. By decomposing loss of control risks into misalignment, incorrigibility and empowerment, we can:</p><ol><li><p>Talk more precisely about which risks we are and aren&#8217;t concerned about in a given context.</p></li><li><p>Stress-test threat models by asking whether the proposed mechanisms are plausible, how the three components interact, and what would need to go wrong simultaneously for serious harm to materialise.</p></li><li><p>Systematically think through intervention points by thinking about how to address each of the three components, and when one approach is more likely to be effective.</p></li><li><p>Prioritise by thinking through which threat models within loss of control are most plausible, most severe, and most amenable to intervention.</p></li></ol><p>Over the coming months, we&#8217;ll continue this series by sharing our thinking on specific threat models and scenarios and priority interventions for loss of control risks.</p><p><em>Thanks to Imogen Stead, Patrick Levermore, Tommy Shaffer Shane, Charlotte Stix, Alejandro Ortega, Annika Hallensleben, Hadrien Pouget, and Mariami Tkeshelashvili for helpful comments and discussions which informed this thinking.</em></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[An export-led approach to AI governance for middle powers ]]></title><description><![CDATA[How countries can influence frontier AI without building it]]></description><link>https://governingtransformativeai.substack.com/p/an-export-led-approach-to-ai-governance</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/an-export-led-approach-to-ai-governance</guid><pubDate>Thu, 02 Apr 2026 14:55:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e5b9bd5d-3acd-4e0d-a03e-3acda86f7c31_6984x3929.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Dr Imogen Stead, Dr Jess Whittlestone &amp; Hamish Hobbs</h5><div><hr></div><p>Recent months have seen an emerging debate about the role that middle powers should play in navigating and developing advanced AI. Some have suggested that middle powers - which we define here as countries with significant international influence but which wouldn&#8217;t be classed &#8220;great powers&#8221; - should pool resources to compete directly on frontier AI, while others have focused on adapting and increasing resilience to AI-driven change. In this blog, we present a third, complementary perspective: middle powers can play a valuable role in global AI governance by &#8220;exporting&#8221; relevant information, capabilities, analysis and governance approaches to frontier AI nations.</p><h2>Competing directly to build frontier AI isn&#8217;t the only way for powers to influence its development </h2><p>A number of different positions on the role of middle powers in AI governance have emerged in recent months.</p><p>Some have suggested countries join forces to try and compete at the frontier of AI development. A <a href="https://aigi.ox.ac.uk/publications/a-blueprint-for-multinational-advanced-ai-development/">paper from last November</a> led by the Oxford Martin AI Governance Initiative suggests pooling compute, talent and data across nations to build advanced AI systems. More ambitiously, a group of authors from Conjecture and Control AI<a href="https://www.conjecture.dev/research/multinational-agi-consortium-magic-a-proposal-for-international-coordination-on-ai"> suggest that</a> a large enough consortium of middle powers could aim to centralise all advanced AI development, and in doing so control its safety.</p><p>Others take a different view, suggesting middle powers should accept they cannot realistically influence frontier AI development, and instead focus on their own national strategies. For example, the<a href="https://institute.global/insights/tech-and-digitalisation/open-source-influence-age-of-ai"> Tony Blair Institute</a> focuses on how middle powers can capture the economic benefits of AI by building open-source ecosystems, while<a href="https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers"> Anton Leicht emphasises</a> the importance of building national resilience to the changes that transformative AI will bring.</p><p>These perspectives identify real challenges and important priorities. Middle powers&#8217; dependency on the US and China creates vulnerabilities, and it is reasonable for these countries to try to have a stake in frontier AI development. But competing at the frontier may be an unrealistic goal for many middle powers at this point, and focusing on adaptation and resilience might actually bring more domestic benefits. This can make it seem like middle powers have a choice between a long-shot bet at influencing frontier AI development, and &#8220;giving up&#8221; on influencing the most important parts of AI development entirely.</p><p>We suggest an additional option which strikes a middle ground between those already discussed: middle powers <em>can</em> influence frontier AI development, but may be most effective at doing so indirectly via &#8220;exporting&#8221; relevant information, capabilities, analysis and governance approaches to frontier AI nations.</p><h2>How middle powers can influence frontier AI via an export-led approach</h2><p>It is possible to influence how frontier AI is developed, used, and governed without building a domestic frontier AI lab. A middle power can produce and share information, analysis, innovations, services and governance approaches with those actors who <em>are </em>developing frontier AI.</p><p>Not every middle power will be able to exert this kind of influence, but many will. What&#8217;s important is that, in order to influence AI governance in this way, middle powers <em>don&#8217;t </em>necessarily need to build frontier AI systems or have military or economic dominance. Instead, they need some kind of unique expertise or selling point related to developing and governing secure and beneficial advanced AI, which can be recognised internationally. We&#8217;ll discuss some examples of this more below.</p><h3>1. Exporting information and expertise </h3><p>Decisions about how frontier AI is developed and deployed depend heavily on access to reliable, up-to-date information and research on emerging capabilities and risks. Middle powers can influence these decisions by producing credible, independent analysis that frontier developers and governments can use to inform their thinking.</p><p>The UK&#8217;s AI Security Institute (AISI) is an excellent demonstration of this model. Its evaluations of model capabilities and risks are shared with both the US Centre for AI Standards and Innovation (US CAISI) and with frontier companies themselves. When UK AISI identifies important capabilities, risks, or limitations of AI models, this feeds directly into development and deployment decisions by other actors.</p><p>A related example is the <a href="https://alignmentproject.aisi.gov.uk/">AI Alignment project</a>, a global fund seeking to advance the field of AI alignment that is supported by the UK, Canadian and Australian governments, alongside other industry and philanthropic partners. The project has attracted significant philanthropic support and <a href="https://www.gov.uk/government/news/openai-and-microsoft-join-uks-international-coalition-to-safeguard-ai-development">private funding</a>, including from frontier labs such as Anthropic and OpenAI. This approach gives the UK, Canada and Australia an opportunity to contribute to the frontier of AI alignment science, while simultaneously strengthening their renowned research institutions and ability to attract top technical talent.</p><p>Middle power AI safety and security institutions in the International Network for Advanced AI Measurement, Evaluation and Science (formerly the international network of AISIs) have the opportunity to act as more neutral evaluators than entities with direct commercial or geopolitical stakes. Crucially, this kind of informational influence is very different from political influence, in that it enables well-informed decisions without overtly saying what those decisions should be. Reliable information and analysis can materially affect global AI policy in major jurisdictions, as well as developer decisions in frontier AI companies, without triggering the credibility concerns that overt advocacy might.</p><h3>2. Exporting innovations and services</h3><p>Middle powers can also develop innovations or services that can be used to support secure and beneficial AI development globally. This can include producing proof-of-concept innovations that can then be adopted elsewhere, often with positive spillover effects on AI governance, as well as providing key services that can be exported, such as frontier AI assurance services.</p><p>Consider, for instance, defensive capabilities for monitoring AI-related national security threats. If a middle power develops effective methods for detecting AI-enabled cyber operations or AI-enabled biological threats, these techniques have value for any country facing similar threats - including those developing frontier AI. Demonstrating that something works is often more persuasive than arguing it should be done. Building an industry providing these services can increase security globally, while providing domestic economic benefits.</p><p>One example of an exported innovation for AI security is the <a href="https://inspect.aisi.org.uk/">Inspect framework</a>, developed by the UK AISI as an open-source tool to support frontier model evaluations. This framework has been used to build evaluations by frontier AI developers and evaluation organisations, including <a href="https://alignment.anthropic.com/2025/petri/">Anthropic</a>. Middle power innovations and services can help to develop testing and verification capabilities that can be used by private sector actors to improve oversight and governance of their systems, supporting global risk mitigation.</p><h3>3. Exporting governance and regulation </h3><p>Finally, there is scope for some countries to govern the development of new technologies without controlling their production. A clear example is the EU AI Act, which leverages the EU&#8217;s large market to govern frontier AI companies who wish to sell products on this market. The EU AI Act is likely to meaningfully influence frontier AI development towards better risk management practices. While implementation remains ongoing and enforcement still uncertain, most frontier developers have already signed up to the AI Act&#8217;s Code of Practice for General-Purpose AI, which includes substantive commitments on safety and security practices that align with the regulatory principles in the AI Act. Companies face strong incentives to comply because the EU represents a major market and maintaining different development practices for different jurisdictions would be operationally complex and costly. However, the impact of the EU AI Act will depend upon the effective interpretation, implementation and enforcement, which will require technical capability and sustained political will. Similarly, European data protection law fundamentally changed how global technology companies handle user information, creating practices that extended well beyond EU borders. California&#8217;s vehicle emissions standards influenced automotive design worldwide, as manufacturers found it more efficient to meet the strictest standard everywhere rather than maintaining multiple production lines.</p><p>It is also possible to export governance and regulation without relying upon the leverage of a large market. Australia&#8217;s social media ban for users younger than 16 appears to have <a href="https://www.internationalaffairs.org.au/australianoutlook/is-there-a-canberra-effect-from-the-social-media-ban/">triggered</a> pushes for similar regulations in a variety of other countries. Beyond technology policy, New Zealand&#8217;s move to setting an <a href="https://www.rbnz.govt.nz/-/media/project/sites/rbnz/files/events/2025/economics-conference/day-1/0225-on-inflation-targeting---bernanke-march-6-v2.pdf">inflation target</a> for its central bank in 1989 was widely copied and has now become the dominant approach to monetary policy globally. Where middle powers identify and adopt governance or regulatory innovations related to frontier AI, they may be able to catalyse wider global adoption through a similar mechanism.</p><h2>Why this matters </h2><p>AI development is likely to have profound impacts globally. Ensuring a wide range of countries can influence advanced AI development enables those with a stake in the outcomes to help shape the process of development and diffusion. This is not just an opportunity to strengthen one country&#8217;s global influence, it is an opportunity to ensure global risks are managed globally and global benefits are shared. AI middle powers have a direct interest in exporting ideas, innovations, services and governance approaches to support these goals. Maintaining this level of expert, innovative and governance capability also ensures that key AI-related capabilities are not purely concentrated within a single state or actor. The strategies outlined here are not only routes to positive influence over frontier AI development, but also potential mechanisms for maintaining the distributed capacity that makes dangerous concentrations of power less likely.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Monitoring AI loss of control]]></title><description><![CDATA[The case for an all-source intelligence observatory]]></description><link>https://governingtransformativeai.substack.com/p/monitoring-ai-loss-of-control</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/monitoring-ai-loss-of-control</guid><pubDate>Thu, 12 Mar 2026 11:02:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4e400c44-6154-4a36-805f-d29bf437d29f_4896x3264.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Tommy Shaffer Shane and Dr Jess Whittlestone</h5><div><hr></div><p>As AI systems become increasingly autonomous, there is a significant risk that they will at some point evade human oversight and control and act in dangerous ways. The capacity for disruption or catastrophe from this is one of the most concerning risks posed by advancing AI capabilities, and perhaps the one society is least well prepared for.</p><p>We wrote recently about <a href="https://www.longtermresilience.org/reports/how-the-uk-government-can-govern-the-risk-of-loss-of-control/">why governments need to do more to govern loss of control risks</a>, and about <a href="https://www.longtermresilience.org/reports/the-loss-of-control-observatory-a-prototype-to-detect-real-world-ai-control-incidents/">a prototype observatory we&#8217;re developing</a> to create new techniques to detect loss of control in the real world.</p><p>This post expands on a specific intervention that is crucial and concerningly neglected: monitoring real world model behaviours linked to loss of control risk.</p><h2>Why improved monitoring is crucial for loss of control</h2><p>The current state of monitoring for loss of control is in its infancy. A recent <a href="https://www.rand.org/pubs/research_reports/RRA3847-1.html">RAND Europe report </a>found, for example, that detection is currently inconsistent and ineffective, and more work must be done by governments, researchers and AI companies to improve the state of the art.</p><p>We set out three key reasons for why monitoring loss of control is essential for the risk to be effectively addressed:</p><ul><li><p>First, monitoring would <strong>improve our evidence base and strategic awareness</strong>, enabling better planning and mitigations more generally. Right now, our understanding of loss of control risks relies heavily on theory or evaluations which often depict contrived scenarios. Real-world signals and intelligence would give us a richer, more grounded evidence base on how loss of control risks are emerging. This would both build a clearer picture of the loss of control harms we&#8217;re already experiencing, helping to drive policy action, and provide a more grounded evidence base against which to consider potential future more extreme harms.</p></li><li><p>Second, better monitoring might <strong>make it possible to effectively respond to contain a loss of control risk before it materialises</strong>. If early warning signs can be detected enough, there may be a &#8216;containment window&#8217; in which acting fast enough could prevent significant harm. But this will only be possible with fast, accurate intelligence about the threat as it emerges. To draw on the analogy of a pandemic: by monitoring for new pathogens in wastewater, for example, it&#8217;s possible to detect emerging threats early enough to contain them before they spiral into full-blown pandemics. </p></li><li><p>Third, <strong>robust monitoring capability could also act as a deterrent</strong> (aka &#8220;deterrence by denial&#8221;). One key threat model for loss of control involves &#8220;scheming&#8221; behaviours where AI systems deliberately evade human-imposed controls. An AI system that knows its resource acquisition and anomalous behaviour will be noticed faces real costs and risks in pursuing those behaviours. Detection is therefore also prevention. Together, these developments suggest a growing risk of autonomous AI systems coming to operate outside of human control in dangerous ways.</p></li></ul><h2>There is an urgent need for an all-source intelligence observatory for loss of control risks</h2><p>Governments, regulators, and civil society currently lack a monitoring capability that can detect and monitor emerging loss of control threats. This capability, which we might call an intelligence &#8220;observatory&#8221;, will need to centralise and triangulate multiple sources of data and intelligence about the agent&#8217;s activity.</p><p>Exploiting this would require something like an all-source centralised intelligence capability - one that can collect and analyse signals from multiple sources. Whistleblowers or insiders might surface early warnings about concerning new capabilities. Signals monitoring could flag a power-seeking AI quietly acquiring compute, energy, money, or labour. Triangulating these data streams could make the difference between catching a threat early and missing it entirely.</p><p>The table below sets out the different types of intelligence we think this observatory would ideally incorporate.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Uf6R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_424, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 424w, /__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_848, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 848w, /__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_1272, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_1456, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_webp, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Uf6R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png" width="1456" height="928" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:928,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:221461,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://governingtransformativeai.substack.com/i/190627904?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_424, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 424w, /__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_848, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 848w, /__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_1272, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Uf6R!, /__u/governingtransformativeai.substack.com/w_1456, /__u/governingtransformativeai.substack.com/c_limit, /__u/governingtransformativeai.substack.com/f_auto, /__u/governingtransformativeai.substack.com/q_auto:good, /__u/governingtransformativeai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd0f69bf-ef30-4e10-8da3-b3afc78becb3_1576x1004.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For example, if a loss of control incident involved a rogue AI system strategically accumulating financial or computational resources, then: OSINT may provide indicators of manipulating markets through misinformation on social media for financial gain; SIGINT may provide indications of the system&#8217;s transactions with cloud computing providers; and HUMINT might provide information about model propensities from an AI company&#8217;s confidential internal tests. Corroborating these signals could make a big difference to detecting the threat, by revealing the connection between seemingly unrelated events.</p><p>Further work is needed to identify the highest priority threats from loss of control and which forms of monitoring could best help mitigate these threats. This is something we&#8217;re thinking about and plan to share more about soon.</p><h2>Defensive acceleration of AI monitoring &#8211; or, &#8220;intel/acc&#8221;</h2><p>Attention should be paid to where defensive acceleration of intelligence gathering and analysis to monitor AI risks &#8211; or &#8220;intel/acc&#8221; &#8211; can be achieved using frontier models, due to AI being adept at analysing large quantities of data extremely quickly. A major moonshot technical project could develop novel techniques that are impossible for human analysts or more conventional technologies. AI-enabled intelligence gathering for AI threats may also be inherently defense-biased, benefiting defensive actors more than an uncontrolled AI agent. At the same time, AI-enabled intelligence gathering and analysis could also pose an increased risk of misuse by governments or malicious actors, meaning that careful consideration of governance and misuse risks is also needed.</p><p>This R&amp;D may not attract private funding, and may require classified intelligence, suggesting government action may ultimately be required. Once the highest priority threats are identified, the ideal outcome is likely a project with a budget in the 10s to 100s of millions led by one or more governments who share intelligence on emerging AI threats with the governments, AI companies or other actors best-placed to mitigate these threats. A coalition of governments or a new international monitoring agency could be ideal, but also more difficult to achieve.</p><h2>OSINT monitoring capabilities are a promising starting point </h2><p>While detection should draw on as many sources as possible, OSINT seems particularly promising as a starting point.</p><p>Firstly, an OSINT capability can be developed entirely outside of government and intelligence communities, as it does not rely on classified intelligence or investigatory powers, meaning significant ground can be broken unilaterally by any sufficiently resourced organisation. This could offer the strategic benefit of building a case for an observatory by demonstrating viable detection methods as a way of stimulating action within governments.</p><p>Similarly, OSINT also doesn&#8217;t rely on partnerships with AI companies, which may run into legal or commercial obstacles. Unlike whistleblowing, which may involve legal risks, or direct disclosures from AI companies, which may involve commercial risks, OSINT faces relatively few barriers to collection and sharing. This makes it particularly tractable to develop novel proof of concept capabilities that can begin feeding into operational capabilities.</p><p>This reasoning motivated CLTR to develop a <a href="https://www.longtermresilience.org/reports/the-loss-of-control-observatory-a-prototype-to-detect-real-world-ai-control-incidents/">prototype</a> of one novel OSINT technique: a tool that scrapes X for transcripts of AI model interactions showing scheming or scheming-like behaviours. It is already producing findings on emerging trends in loss of control incidents. We&#8217;re also <a href="https://www.arcadiaimpact.org/ai-governance-taskforce">partnering with Arcadia Impact</a> to identify the most promising novel OSINT capabilities that could be developed.</p><h2><strong>Next steps </strong></h2><p>CLTR will be continuing to develop this thinking, and is interested in collaborations or input on this work. Specifically:</p><ul><li><p><strong>The why:</strong></p><ul><li><p>Are there flaws or gaps in this theory of change?</p><p></p></li></ul></li><li><p><strong>The who:</strong></p><ul><li><p>Which institutions should own and operate this kind of capability? Government agencies, independent bodies, AI labs, civil society coalitions?</p></li><li><p>How should incident intelligence be shared between government and non-government stakeholders without creating security risks (e.g. information hazards about vulnerabilities)?<br></p></li></ul></li><li><p><strong>The how:</strong></p><ul><li><p>Which methods and information sources are most promising for detecting an emerging incident?</p></li><li><p>How do we calibrate detection to avoid false positives while being sensitive to genuine threats?</p></li><li><p>How might a sophisticated AI system try to evade or manipulate OSINT-based detection? How do we stay ahead of that?</p></li></ul></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[How the UK Government can govern the risk of loss of control]]></title><description><![CDATA[Four recommendations to lay the groundwork for effective risk governance]]></description><link>https://governingtransformativeai.substack.com/p/how-the-uk-government-can-govern</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/how-the-uk-government-can-govern</guid><pubDate>Wed, 11 Mar 2026 11:34:30 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7011f5d3-cd6b-43cb-8ade-237a3420bae8_2560x1792.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Tommy Shaffer Shane, Dr Richard Moulange and Dr Jess Whittlestone</h5><div><hr></div><p><em>This post was originally published on the Centre for Long-Term Resilience blog in February 2025. </em></p><div><hr></div><p>The UK Government is beginning to express significantly more concern about the risk that increasingly autonomous AI systems might at some point evade human oversight and control.</p><p>In October 2025, the Director General of the UK&#8217;s Security Service <a href="https://www.mi5.gov.uk/director-general-sir-ken-mccallum-gives-threat-update">raised concerns</a> that there are &#8220;potential future risks from non-human, autonomous AI systems which may evade human oversight and control&#8221;. This prompted the House of Lords to <a href="https://lordslibrary.parliament.uk/potential-future-risks-from-autonomous-ai-systems/">publish a briefing</a> and hold a short debate on what steps the Government should be taking to address these risks. The AI Security Institute is working on <a href="https://www.aisi.gov.uk/work/replibench-measuring-autonomous-replication-capabilities-in-ai-systems">new benchmarks</a> for assessing these risks, undertaking world-leading research on <a href="https://arxiv.org/abs/2507.03409">scheming</a> and <a href="https://www.aisi.gov.uk/work/replibench-measuring-autonomous-replication-capabilities-in-ai-systems">autonomous replication</a>, and funding research on <a href="https://www.rand.org/pubs/research_reports/RRA3847-1.html">preparedness</a> for loss of control incidents. </p><p>However, the wider apparatus of government is yet to grapple seriously with the question of how to govern loss of control risks.</p><p>In this post, we explain why we believe the UK must do more in this space, and make four concrete recommendations which would lay the initial groundwork for effective governance of loss of control:</p><ol><li><p><strong>Build a shared understanding of loss of control across the UK Government, industry and the public by adding the risk to the National Risk Register (NRR)</strong>: Address the UK Government&#8217;s <a href="https://www.rand.org/pubs/research_reports/RRA3847-1.html">fragmented understanding of loss of control</a> with an official definition and assessment in the NRR, including its impact, likelihood and a preparedness assessment.</p></li><li><p><strong>Clarify responsibilities by formally designating DSIT as the Lead Government Department for AI loss of control risks: </strong>Increase public transparency and accountability for anticipation, prevention, preparation and response of loss of control, by publicly naming DSIT as the Lead Government Department (LGD) that is responsible for loss of control risks.</p></li><li><p><strong>Create transparency and accountability for risk governance by publishing an AI Security Strategy:</strong> The Government should commit to a set of concrete activities over the short, medium and long-term to effectively anticipate, prepare for and defend against loss of control risks</p></li><li><p><strong>Ensure the UK Government can effectively intervene during a loss of control incident by introducing emergency powers: </strong>As with other threats to national security, the UK Government will need some powers of direction during an emergency. This will be necessary to ensure it has effective information and an ability to guide fast industry action.</p></li></ol><h2>Why governments must do more to address loss of control risks</h2><p>An uncontrolled AI system could lead to catastrophic scenarios, ranging from disrupting critical national infrastructure to an AI-engineered pandemic and possibly even threatening human extinction.</p><p>Three significant developments have emerged in the past 12 months that lend strong support to the concern that autonomous AI systems could soon behave in dangerous ways:</p><ol><li><p><strong>AI systems&#8217; ability to operate autonomously is growing exponentially</strong>. Modern systems can now plan and execute extended tasks without human control or intervention. One leading AI <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">benchmark</a> measures frontier AI systems&#8217; ability to complete tasks that take human experts hours to finish, across many different domains: software and AI engineering, mathematics, cybersecurity and general reasoning. <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">Right now</a>, state-of-the-art AI models can complete well-scoped software engineering tasks that take human experts more than four hours to complete, with 50% reliability, and this rate is increasing exponentially. Maximum task length currently doubles about every six months, so by the end of 2026, AI systems are expected to be able to complete work that takes professional software engineers two full days completely unsupervised. Even if the tasks in question are of a highly-scoped and bounded nature, the exponential trend demands serious attention.</p></li><li><p><strong>There is emergent evidence of AI misbehaviour.</strong> AI evaluators have gathered early evidence that AI systems can&#8212;in certain circumstances&#8212;pursue unintended goals, including goals that undermine their operators. For example, there is evidence for the capability of some AI models to strategically <a href="https://www.anthropic.com/research/alignment-faking">hide</a> their true motives, <a href="https://time.com/7259395/ai-chess-cheating-palisade-research/">cheat</a> to achieve <a href="https://x.com/SakanaAILabs/status/1892992938013270019">better</a> results or to win at impossible tasks, and resort to <a href="https://www.anthropic.com/research/agentic-misalignment">blackmail</a> to avoid shutdown. Frontier models already show signs of <a href="https://arxiv.org/pdf/2309.00667">knowing when they are being tested</a> for dangerous capabilities and, when probed, openly state that they choose to deliberately perform worse on tests to hide their abilities (a strategy known as &#8216;<a href="https://www.alignmentforum.org/posts/pPEeMdgjpjHZWCDFw/white-box-control-at-uk-aisi-update-on-sandbagging">sandbagging</a>&#8216;). Models also evidence some early &#8216;<a href="https://arxiv.org/abs/2412.14093">scheming</a>&#8216;-like behaviours, where they pretend to align with human values to avoid being shut down or otherwise controlled. These behaviours aren&#8217;t conclusive evidence of imminent danger, nor do they prove AI models are inherently dangerous. It is especially important to distinguish between capability and propensity, as highlighted <a href="https://arxiv.org/pdf/2507.03409">recently</a> by the AI Security Institute: even if AI systems are <em>able</em> to do something, that doesn&#8217;t mean they <em>will</em>. However, these evaluations clearly show AI is developing capabilities to bypass human oversight and control in concerning ways, which could lead to catastrophic risks in the future. CLTR is also beginning to observe some of these behaviours in real-world contexts in our <a href="https://www.longtermresilience.org/reports/the-loss-of-control-observatory-a-prototype-to-detect-real-world-ai-control-incidents/">Loss of Control Observatory</a>.</p></li><li><p><strong>Voluntary self-governance is proving ineffective at maintaining standards when it counts.</strong> AI companies, while developing AI capabilities that are increasingly risky and remain poorly controlled, are racing towards technological prowess, often cutting corners on safety. Anthropic&#8212;often noted as the most safety-conscious frontier AI company, providing <a href="/__u/cdn.sanity.io/files/4zrzovbb/website/6a3b14a98a781a6b69b9a3c5b65da26a44ecddc6.pdf">measured </a><a href="https://www.nytimes.com/2025/06/05/opinion/anthropic-ceo-regulate-transparency.html">support</a> for regulations and <a href="https://www.anthropic.com/transparency">transparency</a>&#8212;quietly <a href="https://www.obsolete.pub/p/exclusive-anthropic-is-quietly-backpedalling">backpedalled</a> on a previous commitment to define warning signal evaluations for the next generation of its models, before releasing a new generation of model, Claude 4, last year. Google DeepMind activated additional safeguards on Gemini 2.5 Deep Thinking, citing concerns around the development of biological weapons, only to deactivate them for Gemini 3, having raised the risk threshold to implement costly mitigations. This might be a sensible calibration to new information, but it was done quietly, with very little external input and scrutiny. Independent observers <a href="https://futureoflife.org/ai-safety-index-summer-2025/">have</a> <a href="https://ailabwatch.org/">tracked</a> <a href="https://ratings.safer-ai.org/">similar</a> shortcomings in other companies. Even more concerningly, xAI&#8217;s <a href="https://x.ai/news/grok-4">Grok 4</a>&#8212;a frontier model that scores the highest on <a href="https://artificialanalysis.ai/#artificial-analysis-intelligence-index">several</a> benchmarks&#8212;<a href="https://www.lesswrong.com/posts/dqd54wpEfjKJsJBk6/xai-s-grok-4-has-no-meaningful-safety-guardrails#Nukes">readily</a> answers questions related to the development of weapons of mass destruction, and was released <a href="https://fortune.com/2025/07/17/elon-musk-xai-grok-4-no-safety-report/">without</a> any published safety evaluations.</p></li></ol><p>Together, these developments suggest a growing risk of autonomous AI systems coming to operate outside of human control in dangerous ways.</p><h2>The UK government lacks a transparent and accountable approach to loss of control</h2><p>In addition to being extremely high stakes, the novelty and unpredictability of loss of control risks means our existing governance institutions are not well-equipped to address them.</p><p>Other jurisdictions &#8211; such as the EU and some U.S. states &#8211; have begun to set out new regulations to address loss of control.</p><p>The EU&#8217;s <a href="https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai">General Purpose AI (GPAI) Code of Practice</a>, for example, names loss of control as one of four systemic risks posed by GPAI that signatories must address. Similarly, <a href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB53">California&#8217;s SB 53 bill</a> requires reporting of incidents in which frontier AI models engage in &#8220;deceptive techniques against the large frontier developer to subvert the controls or monitoring of its large frontier developer&#8221;. New York&#8217;s <a href="https://www.governor.ny.gov/news/governor-hochul-signs-nation-leading-legislation-require-ai-frameworks-ai-frontier-models">RAISE Act</a> introduces similar requirements.</p><p><strong>But while the UK AISI is developing world-leading research on loss of control, there is a disconnect between this research agenda and the UK Government&#8217;s public policy efforts.</strong></p><p>This fact is starkest in the UK&#8217;s National Risk Register (NRR). The 187-page NRR&#8212;which summarises the Government&#8217;s assessment of the most serious risks the UK faces across the next five years&#8212;mentions artificial intelligence in passing only three times. There is no mention of loss of control among the 88 risks that are assessed.</p><p>While policy work on loss of control is underway, it is &#8220;fragmented&#8221;, according to a Government-funded <a href="https://www.rand.org/pubs/research_reports/RRA3847-1.html">report by RAND Europe</a>. It is also opaque and therefore unaccountable, with minimal reference in policy documents, and likely very diverging levels of understanding and urgency across government.</p><h2>Our recommendations</h2><p>To effectively prepare for and govern loss of control risks, we suggest that the UK Government needs to immediately prioritise four things:</p><h4>1. Build a shared understanding of loss of control across the UK Government, industry and the public by adding the risk to the National Risk Register (NRR).</h4><p>The UK Government&#8217;s understanding of the risk of loss of control is fragmented. While scientific reports and some policy documents give high level definitions of the risk, the UK lacks a clear account of the government&#8217;s own understanding of the risk, with an assessment of impact and likelihood and a reasonable worst case scenario.</p><p>The NRR is the official public account of the UK Government&#8217;s assessment of the most serious risks the UK faces. In the absence of regulation similar to the EU CoP or SB 53, the NRR is a prime candidate for building a shared understanding of loss of control among government departments, businesses, CSOs and the wider public.</p><p>The NRR could include one or more risk assessments to capture loss of control. The UK Government is committed to considering even &#8220;extremely unlikely&#8221; scenarios (e.g. lower than 0.2% probability of occurring), especially where they threaten severe impacts. Forecasts and <a href="https://arxiv.org/pdf/2401.02843">surveys</a> from <a href="https://forecastingresearch.org/xpt">experts</a> suggest that the risk may be substantially higher than this.</p><h4>2. Clarify responsibilities by formally designating DSIT as the Lead Government Department for AI loss of control risks</h4><p>Along with defining and assessing the risk, the UK Government must create more transparency and accountability for its management.</p><p>This requires a &#8216;Lead Government Department&#8217; (LGD). The LGD system is the cornerstone of the UK Government&#8217;s risk governance regime. It names the government departments responsible for risk identification, emergency preparedness, and response and recovery for major national risks.</p><p>The most recent official, public document naming LGDs that we have been able to access is the <a href="https://assets.publishing.service.gov.uk/media/64df30223fde6100134a547a/UK_National_Leadership_for_Risk_Identification__Emergency_Preparedness__Response_and_Recovery.pdf">UK National Leadership for Risk Identification, Emergency Preparedness, Response and Recovery</a> from 2023. Loss of control risks are not covered in this document, leaving the LGD unspecified.</p><p>We suggest that the department with the appropriate remit, expertise and levers for risk governance is DSIT. DSIT would need to work in close collaboration with the Cabinet Office and the wider national security community, but it is the best-placed of all government departments to coordinate anticipation, prevention, preparation, and response.</p><h4>3. Create transparency and accountability for risk governance by publishing an AI Security Strategy</h4><p>As LGD we recommend DSIT lead the development of a UK AI Security Strategy (AISS), modelled on the <a href="https://www.gov.uk/government/publications/uk-biological-security-strategy">UK Biological Security Strategy (BSS)</a>. The AISS should cover the risk understanding, prevention, detection, and response. This will enable outside actors such as the DSIT Select Committee and civil society to hold DSIT accountable to its progress.</p><p>The AISS should retain three key components of the BSS when considering how to govern loss of control:</p><ul><li><p><strong>Implementation</strong>: Rather than set out grand ambitions alone, the AISS should commit to a set of activities over short, medium and long term. These activities can be used to hold the government accountable for action and is especially important given the fast rate of AI progress many expect between now and 2030.</p></li><li><p><strong>Coordination</strong>: The AISS should ensure coordination of risk management activities, possibly via an AI Security Coordination Unit, similar to the Biological Security Coordination Unit. This is essential to coordinate security-relevant activities across Government, given that the most catastrophic loss of control outcomes would be considered &#8216;whole-of-system&#8217; risks (see <a href="https://www.gov.uk/government/publications/the-central-government-s-concept-of-operations/the-amber-book-managing-crisis-in-central-government-html">The Amber Book: Managing crisis in central government</a>).</p></li><li><p><strong>Formalised leadership and governance structures: </strong>The BSS assigns clear political and official leadership to its outcomes. For example, it specifies the lead Minister who must report annually to Parliament on progress on implementation of the Strategy, and the Senior Responsible Officer (SRO) who oversees the implementation of the Strategy. The AISS should do likewise: for instance, through the Secretary of State for Science, Innovation and Technology and the Director General for Artificial Intelligence in DSIT.</p></li></ul><h4>4. Introduce emergency powers to enable the UK Government to respond to a loss of control incident</h4><p>Emergency powers are vital for responding to emergencies, during which Secretaries of State may need to direct regulated entities and regulators to respond to an incident in order to contain it. This will be true also for loss of control incidents.</p><p>In our report, <em><a href="https://www.longtermresilience.org/wp-content/uploads/2025/02/Preparing-for-AI-security-incidents_-Improving-emergency-preparedness-with-the-UK-AI-bill-and-beyond.pdf">Preparing for AI security incidents</a></em>, we established that the UK risks being empty-handed in a crisis, as existing legislation is unlikely to provide the necessary emergency powers in the event of a loss of control incident.</p><p>Emergency powers should ensure that the appropriate person, likely the DSIT Secretary of State, has the power when necessary to:</p><ul><li><p><strong>Compel information: </strong>In an emergency, the government may need to quickly access information about the nature of a loss of control incident to inform situational awareness and response decisions. If this information is needed at a very fast pace, legal powers could be vital, as the UK may find itself competing with other governments to secure information.</p></li><li><p><strong>Direct prevention, mitigation or response: </strong>There may be a need to direct frontier AI companies to take steps to address a risk. This can ensure actions are taken to prevent or respond to an emergency and within a fast enough timeframe to be effective. These types of powers could be conferred to the DSIT Secretary of State to ensure rapid response in an AI security incident.</p></li><li><p><strong>Contain or restrict access, distribution or operation (e.g., take down or block services): </strong>In certain emergency situations, there may be benefits to temporarily revoking public access to the model in the UK or even enforcing a shut down of certain GPU activity.</p></li></ul><p>The UK&#8217;s upcoming AI bill could potentially provide those powers, however the future of that bill <a href="https://www.politico.eu/article/how-labour-fell-out-love-with-ai-bill-peter-kyle/">is now in doubt.</a> As a result, other legislative vehicles, such as the Cyber Security and Resilience Bill, should be considered for providing appropriate emergency powers.</p><p>Once those powers are introduced, relevant actors should participate in tabletop exercises to prepare for how they would use them in various scenarios.</p><p>These recommendations represent initial steps the UK Government can take to ensure it is equipped to govern loss of control risks. Much more work will be needed on more granular strategies to better understand, detect, and mitigate loss of control risks. We plan to share more work on this later this year, and would welcome contact from anyone working on similar issues.</p><h2>Acknowledgements</h2><p>We are grateful to Jamie Bernardi, who provided early research to support this work during June and July 2025, and Imogen Stead, James Ginns, Eleanor Hevey, and Hamish Hobbs for providing feedback on these recommendations.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/governingtransformativeai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Securing a seat at the table]]></title><description><![CDATA[Pathways for advancing the UK&#8217;s global leadership in frontier AI governance]]></description><link>https://governingtransformativeai.substack.com/p/securing-a-seat-at-the-table-pathways</link><guid isPermaLink="false">https://governingtransformativeai.substack.com/p/securing-a-seat-at-the-table-pathways</guid><pubDate>Wed, 11 Mar 2026 11:33:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0bb68a45-7438-49fe-b7bc-1808d2878598_3320x2078.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Dr Imogen Stead and Dr Jess Whittlestone</h5><div><hr></div><p><em>This post was originally published on the Centre for Long-Term Resilience blog in December 2025. </em></p><div><hr></div><p>Kanishka Narayan MP, Minister for AI and Online Safety, recently <a href="https://re-state.co.uk/rethink/labour-party-conference-frontier-ai-with-kanishka-narayan/">stressed</a> the importance of implementing &#8220;an AI policy which can double-down on our strengths in key areas&#8221; to make sure &#8220;we have a seat at the table&#8221; as AI capabilities develop. The Government&#8217;s focus so far has been on investing heavily in sovereign UK AI infrastructure in order to maintain geopolitical relevance as the technology advances.</p><p>In our view, the UK is also particularly well-positioned to influence global frontier AI governance conversations. As the Network Coordinator for the International Network for Advanced AI Measurement, Evaluation and Science (previously the International Network of AI Safety Institutes), the convenor of the inaugural international AI safety summit and the first country to establish a government institute dedicated to frontier AI research and model testing, the UK continues to demonstrate that international governance expertise should be valued as a core strength of its AI policy. The question now is how the UK can best double down on its leadership role to strengthen and differentiate its global influence.</p><p>One potential route to impact could be through the implementation of gold-standard frontier AI legislation. However, despite promises from two successive UK governments to institute targeted AI regulation, nothing concrete has yet materialised. The now-familiar narrative of delay surrounding the proposed frontier AI Bill reared its head yet again earlier this month when Technology Secretary Liz Kendall indicated that the Government is unlikely to bring forward new regulation beyond extending the online safety regime.</p><p>For a Government seeking to shift the narrative on UK AI to one of international ambition and confidence, we think this is an opportunity missed. As we have set out in more detail in <a href="https://www.linkedin.com/feed/update/urn:li:activity:7401640993760120832/">this post</a>, failing to regulate not only leaves the UK insufficiently equipped to address extreme risks from AI domestically, but is also a missed opportunity to cement the UK&#8217;s leadership in global AI governance.</p><p>However, an AI Bill is also far from the only route to securing the UK&#8217;s &#8220;seat at the table&#8221;. This post explores new pathways for advancing the UK&#8217;s international impact beyond simply relying on either regulatory influence or investment power, and sets out a range of options for how this agenda can be taken forward without the need for legislative intervention.</p><h2>The UK&#8217;s &#8220;AI soft power&#8221;</h2><p>The foundation of the UK&#8217;s outsized potential to influence global AI governance conversations relies on three strategic advantages, which together grant the UK significant AI &#8220;soft power&#8221; in its international relations:</p><ul><li><p><strong>Strong convening capability:</strong> The UK&#8217;s global reach enables it to demonstrate notable convening power on the international stage. The 2023 inaugural AI Safety Summit brought representatives from the two great powers of AI &#8211; the US and China &#8211; together for the first time to discuss AI safety and security research. The resulting declaration committing to closer international cooperation on advanced AI has led to the establishment of the international AI safety report and the international AISI network, which the UK now leads as &#8220;Network Coordinator&#8221;.</p></li></ul><ul><li><p><strong>Close UK-US relations:</strong> The UK has succeeded in maintaining a strong relationship with the US despite changes in administration in both countries. The technology-focused <a href="https://www.gov.uk/government/news/memorandum-of-understanding-between-the-government-of-the-united-states-of-america-and-the-government-of-the-united-kingdom-of-great-britain-and-north">Memorandum of Understanding</a> signed in September with the US recognised that the two countries are &#8220;most trusted security and defense partners&#8221; and committed to advancing &#8220;a shared mission to promote secure AI innovation&#8221;, including through joint testing and standards development through the AISI &#8211; CAISI partnership.</p></li></ul><ul><li><p><strong>Best-in-class state AI technical capabilities:</strong> In little over two years since its establishment, the UK&#8217;s AI Security Institute has become a world-leading example of what public sector capabilities in frontier AI safety and security research and model testing can look like. AISI&#8217;s stable funding, research credentials and technical talent have allowed it to develop trusted voluntary partnerships with AI developers, which in turn compounds the UK&#8217;s international credibility and ability to create influence on the international AI stage.</p></li></ul><h2>Pathways for advancing the UK&#8217;s international impact</h2><p>Looking beyond the usual avenues of frontier AI investment and regulation, we see three particularly promising (but not exhaustive) new pathways through which the UK can cement its position as a leading voice in global AI governance. For each pathway, we suggest examples of &#8220;low-hanging fruit&#8221;, which are ideas that the UK could implement now to advance this agenda, and &#8220;higher-ambition interventions&#8221;, which are suggestions for where the Government should begin to direct its thinking in order to maximise its long-term impact on the international stage.</p><h2>Pathway 1: Strengthening AISI&#8217;s international engagement</h2><p>The UK AISI has set the gold standard globally of governmental AI research and testing institutes. In its next phase, AISI should further systematise ways to diffuse its knowledge internationally. The recent announcement that the UK is stepping into the role of &#8220;<a href="https://www.gov.uk/government/news/efforts-to-share-best-practices-on-ai-measurement-and-evaluations-driven-forward-through-the-international-network-for-advanced-ai-measurement-evalua">Network Coordinator</a>&#8221; of the newly renamed &#8220;International Network for Advanced AI Measurement Evaluation and Science&#8221; &#8211; formerly the International AISI Network &#8211; presents an opportunity for the UK to lead the reinvigoration of this international coordination mechanism.</p><ul><li><p><strong>Low-hanging fruit:</strong><em><strong> </strong></em>leverage the UK&#8217;s &#8220;Network Coordinator&#8221; role to establish strong communications channels and regular meeting structures for the network, in order to facilitate closer collaboration and information-sharing. The UK should also use its leadership position as a platform to raise the salience of emerging extreme AI risks identified by the network in other international political fora, acting as an early warning mechanism to keep governments informed of evolving AI risks and potential mitigations.</p></li></ul><ul><li><p><strong>Higher-ambition intervention: </strong>institute a separate, private information-sharing partnership between AISI and the US CAISI, building on the aims of the <a href="https://www.gov.uk/government/news/memorandum-of-understanding-between-the-government-of-the-united-states-of-america-and-the-government-of-the-united-kingdom-of-great-britain-and-north">Memorandum of Understanding</a>. This would act as an early warning mechanism for sharing national security-relevant risks that emerge from testing conducted by UK AISI on US AI models, to increase joint strategic awareness of emerging AI model capabilities and facilitate appropriate US responses.</p></li></ul><h2>Pathway 2: Advancing international alignment on risks and mitigations</h2><p>The UK should continue to leverage its international convening power and relationship with the US to build mutual understanding on the evolving frontier AI risk landscape and coordinate mitigating actions where necessary. This could include proposing and building support for baseline international agreements on key frontier AI issues, for instance defining red lines for governments&#8217; use of frontier AI capabilities, at international fora such as the UN, NATO and Five Eyes.</p><ul><li><p><strong>Low-hanging fruit:</strong><em><strong> </strong></em>champion best-practice regulation and research in other jurisdictions, such as the <a href="https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai">Safety &amp; Security chapter</a> of the EU AI Act&#8217;s GPAI Code of Practice and the <a href="https://www.scai.gov.sg/2025/scai2025-report/">Singapore Consensus on Global AI Safety Research Priorities</a>, to inform forthcoming US AI policy interventions and facilitate high-level multilateral alignment on a baseline set of mitigations against the most prominent security risks from frontier AI.</p></li></ul><ul><li><p><strong>Higher-ambition intervention: </strong>work with AISI&#8217;s research and strategic awareness teams to identify specific extreme risks that will require closer international coordination via the establishment of dedicated global governance institutions, and secure international buy-in from key partners such as the US to create new cooperation mechanisms as necessary. One example could be a new institution dedicated to international alignment on serious AI incident monitoring and reporting frameworks.</p></li></ul><h2>Pathway 3: Influencing frontier AI companies&#8217; behaviour</h2><p>The technical expertise of AISI and the UK&#8217;s convening power combined make it comparatively well-placed to work with AI developers to align incentives towards effective frontier AI risk mitigation. The UK should use these relationships to strengthen its industry partnerships further and push for new voluntary commitments to raise the floor on what constitutes &#8220;best practice&#8221; for increasing the safety and security of AI systems.</p><ul><li><p><strong>Low-hanging fruit:</strong><em><strong> </strong></em>build on AISI&#8217;s trusted relationships with frontier AI companies to incentivise greater transparency and information-sharing proportionate to the speed at which frontier model capabilities are developing. The UK should coordinate with the companies privately to secure buy-in for extending the scope of voluntary international commitments on frontier AI risk transparency and management.</p></li><li><p><strong>Higher-ambition intervention:</strong> identify areas of extreme AI risk that are not adequately covered by existing regulatory regimes globally &#8211; such as the risks stemming from the race between frontier AI companies to achieve significant recursive self-improvement capabilities &#8211; and design a voluntary code of practice to help influence responsible industry practices.</p></li></ul><p>This post represents the start of our thinking on this topic, not the end. As we take this work forward, we&#8217;re keen to hear from anyone working in or thinking about similar issues to help us drill down into these recommendations and find the most impactful UK policy interventions for advancing global AI governance.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://governingtransformativeai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/governingtransformativeai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><p></p>]]></content:encoded></item></channel></rss>