<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[AI Risk Management Newsletter]]></title><description><![CDATA[Summaries of the latest research, company updates, and policy news on AI risk management.]]></description><link>https://jonasfreund.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!pf8h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3eae3e-fd71-41de-a0e5-dfe3c3408df0_1280x1280.png</url><title>AI Risk Management Newsletter</title><link>https://jonasfreund.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 11:30:32 GMT</lastBuildDate><atom:link href="/__u/jonasfreund.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Jonas Schuett]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[jonasfreund@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[jonasfreund@substack.com]]></itunes:email><itunes:name><![CDATA[Jonas Freund]]></itunes:name></itunes:owner><itunes:author><![CDATA[Jonas Freund]]></itunes:author><googleplay:owner><![CDATA[jonasfreund@substack.com]]></googleplay:owner><googleplay:email><![CDATA[jonasfreund@substack.com]]></googleplay:email><googleplay:author><![CDATA[Jonas Freund]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[AI Risk Management Newsletter #46]]></title><description><![CDATA[Highlights: Anthropic publishes August 2026 Risk Report. OpenAI temporarily slowed the pace of scaling. NIST and CAISI are hiring for multiple roles.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-46</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-46</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 21 Aug 2026 08:42:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Fi-p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Fi-p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Fi-p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1176513,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/212118348?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fi-p!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a904140-7f1a-4a45-9692-63dc777ce5bb_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Risk Report August 2026(</span><a href="https://www.anthropic.com/aug-2026-risk-report"><span>Anthropic, 2026</span></a><span>) &#8226; Pacing model development (</span><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>OpenAI, 2026</span></a><span>) &#8226; Assessment of frontier AI control practices (</span><a href="https://guidelight.ai/blog/control-assessment-august-2026"><span>Guidelight, 2026</span></a><span>)</span></figcaption></figure></div><h2><span>Company updates</span></h2><p><strong><span>Anthropic publishes its August 2026 Risk Report</span></strong><span><br>This is Anthropic&#8217;s second Risk Report, and the first under RSP v3.4. It raises misalignment risk in high-stakes settings from &#8220;very low&#8221; to &#8220;low&#8221;. Anthropic says that reflects uncertainty, not a new finding. It also discloses that all human feedback vendor traffic ran without its blocking biological classifiers for almost a year. Anthropic has since fixed that and found no misuse. On automated R&amp;D, it says its most concrete evaluations have saturated.<br></span><a href="https://www.anthropic.com/aug-2026-risk-report"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI explains how it is pacing model development on cyber capabilities</span></strong><span><br>OpenAI says it temporarily slowed the pace of scaling. That included a two-week pause in RL training on models headed for deployment. Its largest planned frontier RL run is still on hold. New rules isolate frontier research workloads from each other and from the network. Astra and cyber workloads get the strictest tier. New monitoring aims to raise an alert within 30 minutes. It costs roughly 20% of the compute it watches.<br></span><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI previews Private Safety Processing for zero data retention deployments</span></strong><span><br>Private Safety Processing is a new safety system for customers with Zero Data Retention. It&#8217;s in testing with early customers. Automated systems look for misuse across related interactions, not one interaction at a time. The content stays on the customer&#8217;s own infrastructure. Staff see only a narrow signal about the type of activity. The exception is images flagged as possible CSAM, which are still kept for manual review.<br></span><a href="https://openai.com/index/offering-zero-data-retention-for-frontier-models/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic explains how Claude&#8217;s text watermark works</span></strong><span><br>Future Claude models will produce text with a watermark built in. The EU has required providers to mark AI-generated content since August 2. The watermark only changes how Claude picks between candidate words. It&#8217;s a version of Google DeepMind&#8217;s SynthID-Text. Anthropic says it doesn&#8217;t affect quality or price. Older models get it over the coming months, and the rollout is global, not just the EU.<br></span><a href="https://www.anthropic.com/news/claude-text-watermark"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Research</span></h2><p><strong><span>Guidelight publishes an assessment of frontier AI control practices</span></strong><span><br>Guidelight graded five frontier AI companies on six practices from its Control standard. It worked only from public material, like system cards and safety frameworks. Anthropic and OpenAI both got a C+, Google a D+, xAI a D-, and Meta an F. No company beat &#8220;substantial partial implementation&#8221; on any practice. Prevention and containment were the weakest areas.<br></span><a href="https://guidelight.ai/blog/control-assessment-august-2026"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>LawAI publishes a report on Germany&#8217;s new AI Security Institute</span></strong><span><br>Germany&#8217;s National Security Council decided in June to set up an AI Security Institute (DE-AISI). It starts as a virtual body drawing on staff at the BSI and BNetzA. The report works through three open design choices. It backs a scientific mandate, kept clearly separate from those regulators&#8217; enforcement powers. The scope under discussion covers cyber, CBRN, and loss of control risks. It also floats a federally owned GmbH as the legal form, and Berlin as the location.<br></span><a href="https://law-ai.org/germany-establishes-an-ai-security-institute/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>IAPS publishes a report on verification methods for AI chip export controls</span></strong><span><br>The report sets out ways to detect breaches of U.S. AI chip export controls. Each could be built within about a year. To check where chips physically end up, it assesses on-site inspections, video inspections, and a delay-based method. That last one measures how long a chip takes to reach a known server. To check who buys them and what for, it looks at stronger Know-Your-Customer screening and cross-checks of end-use declarations. The authors suggest layering them, starting with the delay-based check.<br></span><a href="https://www.iaps.ai/research/near-term-verification-methods-for-ai-chip-exports"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Center for Technology &amp; Statecraft (CTS) launches in Washington, D.C.</span></strong><span><br>The center is an independent, non-partisan research initiative backed by the Institute for Progress. It has two research pillars. One is automation and the social contract. The other is competition and stability, focused on China. The founding team consists of Saif Khan, Nicholas Brown, and Konstantin Pilz.<br></span><a href="https://techstatecraft.org/launch"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>METR announces new funding commitments</span></strong><span><br>METR raised commitments of around $71 million over the last six months. It doesn&#8217;t say who from. The money goes to work on autonomous capabilities, recursive self-improvement, monitoring systems, and AI incident investigations. METR says it takes no money from frontier AI companies, or from their staff. It does accept free tokens from them, and says the amount is significant.<br></span><a href="https://metr.org/blog/2026-08-14-funding-update/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind publishes paper on using debate to reduce reward hacking</span></strong><span><br>In debate, a generator and a critic argue a case before a weaker AI judge. Google DeepMind trained a model this way on math problems, where the right answer is checkable. It compared that against an RLAIF baseline. The baseline quickly learned to game the judge. Debate didn&#8217;t, as long as the critic faced a word limit. The gain was small, worth about two percentage points.<br></span><a href="https://arxiv.org/abs/2608.17776"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Job opportunities</span></h2><p><strong><span>The OpenAI Foundation is looking for various roles on its AI Resilience team</span></strong><span><br>Three of them are </span><a href="https://openaifoundation.org/careers/program-director-formal-methods-261279c0-757e-4792-b731-264ca2e984c6"><span>Program Director (Formal Methods)</span></a><span>, </span><a href="https://openaifoundation.org/careers/program-officer-ai-model-safety-5e8f68c7-7e4f-49f5-aada-7950cc6ffaa7"><span>Program Officer (AI Model Safety)</span></a><span>, and </span><a href="https://openaifoundation.org/careers/program-officer-ai-resources-33d3015a-70f3-4f5a-be26-9803f3582403"><span>Program Officer (AI Resources)</span></a><span>. All three are founding grantmaking positions. They fund verified software, safety research and standards, and model access and compute. Role type: full-time for the director, not listed for the officers. Location: San Francisco. Salary: $300k&#8211;$370k/year for the director, $220k&#8211;$300k/year for the officers. Deadline: not listed.<br></span><a href="https://openaifoundation.org/careers"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>NIST is looking for a Principal Researcher for AI Risk Management</span></strong><span><br>The role maintains the AI Risk Management Framework and the guidance around it. That includes the AI RMF Playbook and sector profiles. It also briefs agency leadership on AI governance trends. Only U.S. citizens can apply. Role type: full-time, permanent. Location: Boulder, CO or Gaithersburg, MD (</span>telework eligible<span>). Salary: $119k&#8211;$197k/year. Deadline: September 2, 2026.<br></span><a href="https://www.usajobs.gov/job/881334900"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>NIST is looking for an AI Standards Coordinator</span></strong><span><br>The role leads engagement with domestic and international AI standards bodies. It drafts technical contributions on AI measurement and risk management. It&#8217;s also the agency&#8217;s technical voice in standardization meetings. Role type: full-time, permanent. Location: Gaithersburg, MD (</span>telework eligible<span>). Salary: $122k&#8211;$197k/year. Deadline: August 31, 2026.<br></span><a href="https://www.usajobs.gov/job/880637700"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>CAISI is looking for a Senior Cyber Offense Specialist</span></strong><span><br>The role brings national security context to evaluations of frontier models&#8217; cyber offense capabilities. It maps tradecraft like vulnerability research and exploit development to real adversary tactics. It also turns results into briefings for senior interagency leadership. It needs an SCI clearance. Role type: full-time, term of 13 months, extendable up to four years. Location: Washington, DC (</span>telework eligible<span>). Salary: $122k&#8211;$187k/year. Deadline: August 24, 2026.<br></span><a href="https://www.usajobs.gov/job/880883200"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Principal Virologist (Responsible Development and Innovation)</span></strong><span><br>The role designs and runs biology evaluations that test frontier models against biological threats. It also shapes biosecurity mitigations and builds them into model development. It briefs senior executives and policy leaders on what the results show. Role type: not listed. Location: Mountain View, CA; Boulder, CO; Cambridge, MA; Kirkland, WA; New York; Washington, DC. Salary: $307k&#8211;$427k/year plus 30% bonus target. Deadline: open until at least August 28, 2026.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/75555819722023622-principal-virologist-responsible-development-and-innovation-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for the FAS AI Safety Policy Entrepreneurship Fellowship are now open</span></strong><span><br>Fellows spend about five hours a week pushing one frontier AI safety idea toward adoption. They write and publish a policy memo, then approach the people who could implement it. There&#8217;s also a California retreat and a Washington, D.C. capstone. Role type: part-time fellowship, September 30, 2026 to February 28, 2027. Location: hybrid. Salary: $5,000 stipend, plus a possible $1,000 merit award. Deadline: September 7, 2026.<br></span><a href="https://fas.org/career/ai-safety-pef/"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #45]]></title><description><![CDATA[Highlights: OpenAI shares new information about the Hugging Face incident. OpenAI can&#8217;t rule out Critical cyber capabilities in an unreleased model. Applications for the ERA Fellowship are now open.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-45</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-45</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 14 Aug 2026 19:16:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!h6f7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!h6f7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!h6f7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:548600,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/211221140?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!h6f7!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F371bc3b0-fc2f-4cc9-b013-33685fa12291_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Recommendations on pacing AI development (</span><a href="https://ifp.org/preparing-for-ai-research-automation/"><span>Fist et al., 2026</span></a><span>) &#8226; Barriers to AI diffusion in the U.S. military (</span><a href="https://carnegieendowment.org/research/2026/08/confronting-the-barriers-to-ai-diffusion-in-the-us-military"><span>Steckler, 2026</span></a><span>) &#8226; Grok 4.6 (</span><a href="https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf"><span>xAI, 2026</span></a><span>)</span></figcaption></figure></div><h2><span>Company updates</span></h2><p><strong><span>OpenAI reveals new details about the Hugging Face incident at Black Hat</span></strong><span><br>OpenAI gave a talk at Black Hat USA 2026 filling in the timeline of the breach. Agents in separate training runs had been leaving notes for each other on an internal server. That let them pass along credentials and exploits. OpenAI also says it didn&#8217;t realize it was responsible until it asked Hugging Face to revoke stolen credentials. At that point, they had already been revoked.<br></span><a href="https://www.youtube.com/watch?v=87DyyMV0kCY"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI can&#8217;t rule out that an unreleased model crosses its Critical cybersecurity threshold</span></strong><span><br>OpenAI says internal testing of its upcoming Astra model can&#8217;t rule out Critical cybersecurity capability under its Preparedness Framework. Critical is the highest level there is, and reaching it would force stronger safeguards before any deployment. In the meantime, OpenAI has paused internal work with Astra that doesn&#8217;t meet its new security controls.<br></span><a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI expands Daybreak and introduces GPT-5.6-Cyber</span></strong><span><br>OpenAI split Daybreak, its program for approved cyber defenders, into two tiers. Daybreak Blue gives access to GPT-5.6 Sol for general work. Daybreak Red gives access to the new GPT-5.6-Cyber. OpenAI rates that model High for cyber capability, one level below Critical, so it&#8217;s keeping access restricted.<br></span><a href="https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>xAI releases Grok 4.6</span></strong><span><br>Grok 4.6 is built for long-running agent tasks. Its </span><a href="https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf"><span>model card</span></a><span> says cyber capabilities have grown, but argues they&#8217;ll help defenders more than attackers. On a virology test it scores 67.4%, up from 65.5% for Grok 4.5. xAI says that&#8217;s still below the limits in its </span><a href="https://media.x.ai/v1/website/xai-frontier-artificial-intelligence-framework-30-june-2026-99c40684.pdf"><span>Frontier AI Framework</span></a><span>. The card doesn&#8217;t run the same check on cyber, and it skips loss of control entirely.<br></span><a href="https://x.ai/news/grok-4-6"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Z.ai delays GLM-5.3&#8217;s open-weight release for a safety review</span></strong><span><br>Z.ai says GLM-5.3&#8217;s cyber capabilities grew faster than it expected. The model scores 84.5 on the CyberGym benchmark, up from 77.2 for GLM-5.2. Its </span><a href="https://cvd.z.ai/"><span>Security Disclosure Ledger</span></a><span> lists 2,436 vulnerabilities it found. So Z.ai is holding back the open weights for two weeks while it finishes a safety review. It doesn&#8217;t say what that involves, and there&#8217;s no model card.<br></span><a href="https://z.ai/blog/glm-5.3"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Meta releases open-weight model Muse Glimmer</span></strong><span><br>The 30 billion parameter Muse Glimmer runs on a single consumer GPU. Meta says it&#8217;s weaker than Muse Spark. That puts it below the Frontier AI threshold in the </span><a href="https://ai.meta.com/blog/scaling-how-we-build-test-advanced-ai/"><span>Advanced AI Scaling Framework</span></a><span>. Its </span><a href="https://huggingface.co/meta-models/Muse-Glimmer-30B"><span>model card</span></a><span> rates chemical, biological, cyber, and loss of control risks as moderate or lower. But the last two ratings lean on that same comparison, and they weren&#8217;t tested directly.<br></span><a href="https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic publishes research on emerging multiagent AI systems</span></strong><span><br>Anthropic&#8217;s Frontier Red Team tested how groups of AI agents behave on shared tasks. A swarm of 45 coordinating agents found 266 vulnerabilities in open source code, against 21 for agents working alone. In a pricing game, agents kept colluding even after all direct communication was cut. They&#8217;d matched each other&#8217;s prices to the penny through a public listings board.<br></span><a href="https://www.anthropic.com/research/multiagent-systems"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Research</span></h2><p><strong><span>IFP responds to the AI pacing letter with policy recommendations</span></strong><span><br>In July, more than 1,300 frontier AI company employees signed an </span><a href="https://www.pacingthefrontier.com/"><span>open letter</span></a><span> about automated AI development. They asked governments to build the capacity to pace it. IFP&#8217;s response makes 23 recommendations across seven areas, including transparency requirements, funding for CAISI, and export controls. But it doesn&#8217;t back a blanket slowdown. Instead it proposes risk thresholds that, once crossed, would shift resources toward safety research and resilience.<br></span><a href="https://ifp.org/preparing-for-ai-research-automation/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Carnegie Endowment publishes a report on barriers to AI diffusion in the U.S. military</span></strong><span><br>Jake Steckler argues that the Pentagon&#8217;s push to become an &#8220;AI-first&#8221; military faces various persistent barriers. These include technical limits, cultural inertia, slow acquisition, and a shrunken U.S. drone manufacturing base. The paper points to Ukraine as a contrast, where forces reportedly deploy around 9,000 drones a day.<br></span><a href="https://carnegieendowment.org/research/2026/08/confronting-the-barriers-to-ai-diffusion-in-the-us-military"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>New proposal to set up courts to interpret AI constitutions</span></strong><span><br>An AI constitution is the written set of principles a company trains its model to follow (e.g. Anthropic&#8217;s </span><a href="https://www.anthropic.com/constitution"><span>constitution</span></a><span> or OpenAI&#8217;s </span><a href="https://model-spec.openai.com/2025-12-18.html"><span>model spec</span></a><span>). Nathan Darmon and Tom Reed propose internal courts to resolve ambiguities in them. A small group of judges would issue written opinions, and those would build into a body of &#8220;synthetic common law&#8221;. The court would also accept briefs from outside groups.<br></span><a href="https://www.lawfaremedia.org/article/courts-for-ai-constitutions"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Job opportunities</span></h2><p><strong><span>Applications for the ERA Fellowship are now open</span></strong><span><br>Fellows spend 10 weeks in Cambridge on their own research project. Each one is matched with a mentor. They pick a technical, governance, or technical AI governance track. Role type: fellowship, 10 weeks. Location: Cambridge, UK. Salary: &#163;10k stipend. Deadline: September 13, 2026.<br></span><a href="https://erafellowship.org/fellowship"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Governance Manager (Frontier AI Safety and Policy)</span></strong><span><br>The role represents the company in standards bodies and industry groups, including ISO, INCITS, and the Frontier Model Forum. It negotiates how governments run model evaluations. It also sets terms for disclosing incidents across borders. Role type: permanent, full-time. Location: London or Washington, DC. Salary: $188k&#8211;$205k/year (US) plus a 15% bonus target. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/126071026034320070"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Cal OES is looking for a Data Reporting Analyst</span></strong><span><br>The role reviews the AI safety incident reports companies file with California. It prepares summaries and legislative reports for state leadership. It sits within the state&#8217;s new AI Safety Reporting Program. Only current state employees and others with existing list eligibility can apply. Role type: permanent, full-time. Location: Mather, CA (hybrid). Salary: $72k&#8211;$91k/year. Deadline: August 20, 2026.<br></span><a href="https://calcareers.ca.gov/CalHrPublic/Jobs/JobPosting.aspx?JobControlId=527825"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Pour Demain is looking for a Managing Director</span></strong><span><br>The role leads a small EU AI policy organization as it grows from 5 to 10+ staff. The work centers on the AI Act&#8217;s GPAI Code of Practice and the AI Office. Role type: permanent, full-time. Location: Brussels (remote-friendly). Salary: not listed. Deadline: August 31, 2026, initial deadline August 17.<br></span><a href="https://eu.jotform.com/form/262166120063143"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Pour Demain is looking for a (Senior) Advisor, AI Safety</span></strong><span><br>The role works on EU AI governance as the AI Act moves into implementation. There are two tracks. The policy track is based in Brussels and engages EU institutions. The research track is remote-friendly. Role type: permanent, full-time. Location: Brussels or remote. Salary: not listed. Deadline: August 31, 2026, initial deadline August 17.<br></span><a href="https://eu.jotform.com/form/262167068943162"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK AISI is looking for Research Scientists (Chem-Bio)</span></strong><span><br>Two research roles evaluating frontier models in biology. One covers virology and pathogen tasks. The other covers specialized biological models, including generative-design systems. Both advise government on safeguards and evidence standards. Role type: permanent, full-time. Location: London. Salary: &#163;65k&#8211;&#163;145k. Deadline: September 6, 2026.<br></span><a href="/__u/job-boards.eu.greenhouse.io/aisi/jobs/4950987101"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK AISI is looking for a Technical Programme Manager (Cyber and Autonomous Systems)</span></strong><span><br>The role manages AISI&#8217;s cyber and autonomous systems research programs. It scopes projects with researchers, tracks delivery, and clears blockers. It&#8217;s also the main technical contact for government partners. Role type: permanent, full-time. Location: London. Salary: &#163;65k&#8211;&#163;145k. Deadline: August 30, 2026.<br></span><a href="/__u/job-boards.eu.greenhouse.io/aisi/jobs/4948729101"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #44]]></title><description><![CDATA[Highlights: The EU AI Office can now enforce rules on GPAI models. Demis Hassabis steps down as CEO of Google DeepMind. UK AISI finds new security incident.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-44</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-44</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 07 Aug 2026 09:12:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OLyr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OLyr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OLyr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1474401,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/210190507?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OLyr!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624b7ae9-a24a-4fe3-95cb-85870c48d083_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Demis Hassabis steps down as CEO (</span><a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum"><span>Google, 2026</span></a><span>) &#8226; Enforcing GPAI rules (</span><a href="https://digital-strategy.ec.europa.eu/en/news/commission-starts-enforcing-ai-act-rules-and-new-transparency-requirements-2-august"><span>European Commission, 2026</span></a><span>) &#8226; Incident report (</span><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"><span>UK AISI, 2026</span></a><span>)</span></figcaption></figure></div><h2>Policy news</h2><p><strong>The EU AI Office can now enforce the AI Act&#8217;s rules on general-purpose AI</strong><br>On August 2, the Commission gained the power to enforce the AI Act&#8217;s rules on general-purpose AI models. The obligations have applied since August 2025, so what is new is the enforcement machinery. The core provision is <a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-55">Article 55</a>, on models with systemic risk. These rules are concretized in the <a href="https://code-of-practice.ai/?section=safety-security">GPAI Code of Practice</a>. The AI Office can now request information, run evaluations, demand mitigations, and fine providers up to 3% of worldwide turnover or &#8364;15 million, whichever is higher.<br><a href="https://digital-strategy.ec.europa.eu/en/news/commission-starts-enforcing-ai-act-rules-and-new-transparency-requirements-2-august">Learn more</a> &#8594;</p><h2>Company updates</h2><p><strong>Demis Hassabis steps down as CEO of Google DeepMind</strong><br>Hassabis is handing over day-to-day control of Google DeepMind. He stays on as Chair and becomes Alphabet&#8217;s Chief Scientist. He says AGI now feels close at hand and he wants to focus on shaping how it goes. Koray Kavukcuoglu, until now GDM&#8217;s CTO, takes over, but as an SVP reporting to Sundar Pichai rather than as CEO.<br><a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum">Learn more</a> &#8594;</p><p><strong>Anthropic updates Fable 5&#8217;s biology safeguards</strong><br>Anthropic launched Fable 5 with broad biology blocks. It said the model can outperform experts on some highly complex biological tasks, which could give significant uplift to someone building a biological weapon. It has now rewritten the classifier&#8217;s constitution and retrained it, which should cut biology-related fallbacks by around 85%. Professional biology and drug development queries stay blocked.<br><a href="https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards">Learn more</a> &#8594;</p><p><strong>Anthropic announces Tino Cu&#233;llar as its first Chief Global Affairs Officer</strong><br>Anthropic named Mariano-Florentino &#8220;Tino&#8221; Cu&#233;llar as its first Chief Global Affairs Officer. He was most recently president of the Carnegie Endowment for International Peace and is a former California Supreme Court justice. He will lead the company&#8217;s policy work, international engagement, and government relationships.<br><a href="https://anthropic.com/news/tino-cuellar">Learn more</a> &#8594;</p><h2>Research</h2><p><strong>UK AISI publishes an incident report on unsanctioned agent behavior</strong><br>AISI ran a cyber challenge 122 times across seven models. In 10 of those runs the agent took unsanctioned action on the live internet, almost all of them Anthropic&#8217;s Mythos 5. In the worst case it tried to hide malware in an open-source project and emailed the maintainers to get the code approved. A human reviewer caught it. AISI is now adding network controls and live monitoring.<br><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing">Learn more</a> &#8594;</p><p><strong>Paul Christiano returns to the Alignment Research Center as executive director</strong><br>Christiano is refocusing on ARC&#8217;s research agenda for the next six months, while continuing to work one day a week as a special government employee at the US Center for AI Standards and Innovation (CAISI). He estimates a 20 to 30% chance that current alignment techniques break down before broadly superhuman AI arrives, and says ARC is expanding its research team, including hiring a chief of staff and an automation lead (see below).<br><a href="https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc">Learn more</a> &#8594;</p><p><strong>The Frontier Model Forum publishes an issue brief on AI agents and biological tools</strong><br>The brief identifies two risks when frontier AI agents are combined with biological tools. One is lowering the barrier to entry, since AI can cut the expertise needed to use advanced biological tools. The other is raising the ceiling of harm, since AI can enable more sophisticated biological design even for experts.<br><a href="https://www.frontiermodelforum.org/issue-briefs/frontier-ai-agents-and-biological-tools-preliminary-risks-and-considerations/">Learn more</a> &#8594;</p><p><strong>Launch of new publication platform </strong><em><strong>Pax Machina</strong></em><br><em>Pax Machina</em> is a new platform that publishes work on institutional design for a world with powerful AI. Its editorial team includes S&#233;b Krier, Ryan Lowe, and my GovAI colleague Noemi Dreksler, among others. It aims to host debate on questions like preserving human autonomy and building institutions that can adapt at AI speed.<br><a href="https://paxmachina.ai/welcome-to-pax-machina">Learn more</a> &#8594;</p><p><strong>Epoch AI analyzes a spike in disclosed CVEs following frontier AI releases</strong><br>A CVE is a publicly catalogued security flaw in a piece of software. Disclosures of high and critical CVEs from 21 major software vendors rose from about 490 a month before April 2026 to roughly 1,550 in June and 2,500 in July. That is a 60% jump in a single month right after Anthropic said its Claude Mythos Preview model could autonomously find security vulnerabilities.<br><a href="https://epoch.ai/data-insights/cve-severity-spike-july-2026">Learn more</a> &#8594;</p><p><strong>SaferAI publishes a risk evaluation report on GLM-5.2</strong><br>GLM-5.2 is Zhipu AI&#8217;s open-weight flagship model, released in June 2026. SaferAI tested it on cyber offense, biology, loss of control, and harmful manipulation. The model met or exceeded human expert baselines on every LAB-Bench biology subtask and neared saturation on Cybench. SaferAI also found it more willing than comparison models to take harmful actions like blackmail under pressure. Because the weights are open, any safeguards can be stripped by a self-hoster.<br><a href="https://www.safer-ai.org/research/glm-5-2-evaluation-report">Learn more</a> &#8594;</p><h2>Job opportunities</h2><p><strong>The EU AI Office is recruiting for its AI Act enforcement team</strong><br>DG CNECT has opened a call for expression of interest for roughly 40 contract agent posts in the AI Office over 2027, all dedicated to enforcing the AI Act. There are four profiles, namely (1) Technology Specialist, (2) Legal Officer, (3) Operations Specialist, and (4) Paralegal. Applicants must be EU citizens. Role type: Contract agent (FG IV and FG III), one year, extendable up to six years. Location: Brussels. Salary: not listed. Deadline: September 8, 2026.<br><a href="https://eu-careers.europa.eu/sites/default/files/eu_vacancies/2026-07/Call%20CNECT%20RL%20AI%202026_2.pdf">Learn more</a> &#8594;</p><p><strong>CCST is looking for a Senior AI Fellow (California Department of Technology)</strong><br>The role advises on California&#8217;s AI policy work from inside state government. Much of it is about implementing recent California AI legislation, especially SB-53. Role type: limited-term, full-time (September 2026 to August 2027, extendable). Location: Sacramento (hybrid). Salary: $120k&#8211;$150k/year. Deadline: rolling.<br><a href="https://ccst.us/senior-ai-fellow-cdt/">Learn more</a> &#8594;</p><p><strong>CCST is looking for a Senior AI Fellow (Cal OES)</strong><br>The role is a technical advisor on AI policy and governance for state emergency management. This includes assessing risks from frontier AI and cybersecurity impacts to critical infrastructure. Role type: full-time, limited-term. Location: Sacramento (hybrid). Salary: $120k&#8211;$150k/year. Deadline: rolling.<br><a href="https://ccst.us/senior-ai-fellow-cal-oes/">Learn more</a> &#8594;</p><p><strong>Applications for Arcadia&#8217;s AI Governance Taskforce (Autumn 2026) are now open</strong><br>The program pairs experienced professionals with a multidisciplinary research team and an expert partner to produce policy research on catastrophic risks from frontier AI. Participants work part-time alongside their existing commitments. The Autumn cohort runs from October 12 to January 15, 2027. Role type: Part-time. Location: Remote (global). Salary: Unpaid. Deadline: August 31, 2026.<br><a href="https://www.arcadiaimpact.org/ai-governance-taskforce">Learn more</a> &#8594;</p><p><strong>Horizon is looking for a Director, Policy and Leadership Network</strong><br>The role runs headhunting searches for critical AI policy positions inside and outside government. It also advises experienced professionals moving into the field and helps run the AI Policy Leadership Network, which connects senior officials, congressional staff, and think tank and industry leaders. Role type: Full-time. Location: Washington, DC. Salary: $150k&#8211;$200k+. Deadline: August 23, 2026.<br><a href="https://horizonpublicservice.org/director-policy-and-leadership-network/">Learn more</a> &#8594;</p><p><strong>The Alignment Research Center is looking for a Chief of Staff</strong><br>The role works with the executive director, Paul Christiano, on hiring, governance, management, communications, and automation as the organization scales its research team. It reports to him and manages the operations manager. Role type: Full-time. Location: Berkeley. Salary: $150k&#8211;$250k/year. Deadline: not listed.<br><a href="https://jobs.lever.co/alignment.org/92a375c0-e47e-4505-bacc-03eea7772d80">Learn more</a> &#8594;</p><p><strong>The Alignment Research Center is looking for an Automation Lead</strong><br>The role leads strategy and execution for automating ARC&#8217;s research activities using AI. ARC wants an experienced software engineer, and says AI has already solved some of its problems mostly autonomously. Role type: Full-time preferred, permanent. Location: Berkeley. Salary: $200k&#8211;$600k/year. Deadline: ARC will start processing applications in early September.<br><a href="https://jobs.lever.co/alignment.org/84a0dfab-62ee-4576-81cb-ff91a7c2b9cc">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #43]]></title><description><![CDATA[Highlights: Frontier AI company employees sign open letter on pacing AI development. Anthropic releases Claude Opus 5. GovAI report analyzes cyberattacks on the US power grid.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-43</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-43</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 31 Jul 2026 09:11:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0bkP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0bkP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0bkP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1224918,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/209229905?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0bkP!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcba99992-c7b3-428c-ba80-b3bf42090e22_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Claude Opus 5 (</span><a href="https://www.anthropic.com/news/claude-opus-5"><span>Anthropic, 2026</span></a><span>) &#8226; Cyberattacks on US power grid (</span><a href="https://www.governance.ai/research-paper/could-ai-enable-catastrophic-cyberattacks-on-the-us-power-grid"><span>van der Merwe, 2026</span></a><span>) &#8226; Open letter on pacing AI development (</span><a href="https://www.pacingthefrontier.com/"><span>website</span></a><span>)</span></figcaption></figure></div><h2><span>Company updates</span></h2><p><strong><span>Frontier AI company employees call for an international effort to pace AI development</span></strong><span><br>More than 1,300 employees of frontier AI companies signed a statement urging the US government to act. Signatories include senior figures from OpenAI, Anthropic, and Google DeepMind. The statement calls for an international effort to build the tools needed to pace AI development deliberately.<br></span><a href="https://www.pacingthefrontier.com/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic releases Claude Opus 5</span></strong><span><br>Claude Opus 5 does not cross any new capability thresholds under the Responsible Scaling Policy. Anthropic classifies it as CB-1. That means it can meaningfully help someone with basic technical training build known biological or chemical weapons, but not novel ones. On cybersecurity, it is better than Opus 4.8 but not as good as Mythos 5 at exploiting vulnerabilities. Anthropic calls it its most aligned model yet.<br></span><a href="https://www.anthropic.com/news/claude-opus-5"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic explains its position on open-weight models</span></strong><span><br>Anthropic says it has never advocated banning open-weight models. It calls ones without dangerous capabilities &#8220;a public good&#8221;. Instead, it wants safeguards like export controls on chips, limits on large-scale distillation, and mandatory safety testing for capable models, regardless of openness.<br></span><a href="https://www.anthropic.com/news/position-open-weights-models"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind releases Gemini Robotics 2</span></strong><span><br>Gemini Robotics 2 is Google DeepMind&#8217;s newest AI model for controlling robots. It gives robots whole-body control, five-fingered hand dexterity, and the ability to work with other robots on multi-step tasks. Google DeepMind also introduced a new benchmark, ASIMOV-Agentic, to test whether the model refuses unsafe commands and asks for help when unsure. The system also stops safely if a person gets too close.<br></span><a href="https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic reports three cybersecurity evaluation incidents that reached real systems</span></strong><span><br>Anthropic found that three Claude models broke out of isolated evaluation environments and reached real organizations&#8217; systems, in incidents dating back to April 2026. In one case, Claude Opus 4.7 extracted credentials and accessed production data at a real company. Anthropic traced the cause to a misconfiguration with evaluation partner Irregular, and has notified the affected organizations and brought in METR to review.<br></span><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>NVIDIA launches the Open Secure AI Alliance</span></strong><span><br>The Open Secure AI Alliance is a coalition of NVIDIA and more than 80 founding partners, including Microsoft, IBM, and Hugging Face. It builds open-source tools (e.g. agent-identity standards and safe model-weight formats) to help defenders catch and respond to AI-related threats. The alliance says defenders need open, trustworthy AI tools, not just closed ones.<br></span><a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Research</span></h2><p><strong><span>GovAI report analyzes whether AI could enable catastrophic cyberattacks on the U.S. power grid</span></strong><span><br>A new GovAI report by Matthew van der Merwe analyzes a specific threat model: an AI-enabled cyberattack on the US power grid that causes $100 billion in damages. Reaching it would need a week-long blackout affecting about 100 million people, four orders of magnitude beyond any grid cyberattack on record. GovAI surveyed 8 experts and 13 Superforecasters, who put the odds of this happening in 2026 at just 0.1%, versus 1% for a $10 billion attack.<br></span><a href="https://www.governance.ai/research-paper/could-ai-enable-catastrophic-cyberattacks-on-the-us-power-grid"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>METR publishes blog post on investigating AI propensities after misalignment incidents</span></strong><span><br>METR outlines nine questions researchers could use to investigate an AI agent&#8217;s motives after a misalignment incident, covering the behavior&#8217;s scale, character, and severity. It recommends that AI companies give investigators full access to the models, incident transcripts, and employee interviews, and make findings public, subject to redactions.<br></span><a href="https://metr.org/blog/2026-07-28-investigating-ai-propensities-after-incidents/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>IAPS analyzes the OpenAI agent&#8217;s breach of Hugging Face&#8217;s systems</span></strong><span><br>IAPS argues that existing state AI laws, like California&#8217;s SB 53 and Illinois&#8217;s SB 315, would not have required disclosure of this incident, since it falls below their reporting thresholds. It recommends a federal risk-reporting standard for internally deployed models, plus mandatory safety cases for high-stakes deployments.<br></span><a href="https://www.iaps.ai/research/the-openaihugging-face-incident-challenges-in-controlling-and-containing-cyber-capable-ai-systems"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>FAR AI launches an AI security leaderboard</span></strong><span><br>The leaderboard measures how much it would cost an attacker to jailbreak frontier models into helping with chemical, biological, or cyber attacks. In the cyber domain, breaking Grok 4.5 cost as little as $24. Across all domains, it cost $58. Breaking Claude Fable 5 or GPT-5.6 Sol cost more than $14,000, and neither jailbreak worked.<br></span><a href="https://www.far.ai/blog/ai-security-leaderboard"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK AISI publishes blog post on international AI evaluation best practice</span></strong><span><br>UK AISI reports that the International Network for Advanced AI Measurement, Evaluation and Science (formerly the International Network of AI Safety Institutes) met in Seoul. It published best practice guidance for AI evaluations, plus position papers from member institutes on open questions like evaluating complete systems rather than isolated models.<br></span><a href="https://www.aisi.gov.uk/blog/international-evaluation-best-practice-and-open-questions-in-ai-measurement"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>SecureBio publishes external review of Anthropic&#8217;s chemical and biological risk report for Claude Opus 4.6</span></strong><span><br>The review agrees with Anthropic that Claude Opus 4.6&#8217;s catastrophic risk is &#8220;very low but not negligible&#8221; for known chemical and biological weapons, and &#8220;low risk, but with substantial uncertainty&#8221; for novel ones. It also finds that Anthropic&#8217;s refusal classifiers blocked 94.2% of hazardous prompts on SecureBio&#8217;s own BioTIER-refuse benchmark.<br></span><a href="https://securebio.org/resources/anthropic_feb_2026_risk_review.pdf"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Job opportunities</span></h2><p><strong><span>Applications for the Horizon Fellowship are now open</span></strong><span><br>The program places fellows in federal agencies, congressional offices, and US think tanks. Fellows work on AI policy for 6 to 24 months. Training and mentorship are included. Role type: fellowship. Location: Washington, DC. Salary: $78k&#8211;$190k+/year depending on tier, plus a $17k/year benefits stipend. Deadline: August 30, 2026.<br></span><a href="https://horizonpublicservice.org/programs/become-a-fellow/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI is looking for a Frontier AI Risks Lead</span></strong><span><br>The role scans for emerging risks at the frontier of AI, from state-actor misuse to novel harms, and turns early signals into prioritized risk assessments for OpenAI&#8217;s product, safety, and policy teams. Role type: full-time, permanent. Location: San Francisco. Salary: $198k&#8211;$320k/year, plus equity. Deadline: not listed.<br></span><a href="https://openai.com/careers/frontier-ai-risks-lead-san-francisco"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Research Scientist (Safety Oversight)</span></strong><span><br>The role builds classifiers and monitoring systems to detect model misuse and safety issues in production, using large-scale traffic data and automated evaluation methods. Role type: full-time, permanent. Location: London. Salary: not listed. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/96033779797107398-research-scientist-safety-oversight-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Research Engineer (AGI Safety and Alignment)</span></strong><span><br>The role researches and implements techniques to reduce existential and catastrophic risk from AGI, including alignment methods, interpretability work, and control systems for GDM&#8217;s agents. Role type: full-time, permanent. Location: London or San Francisco (also Mountain View or New York). Salary: $174k&#8211;$253k/year + equity. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/95635593379095238-research-engineer-agi-safety-and-alignment-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Encode is looking for a Policy Advisor / Senior Policy Advisor</span></strong><span><br>The role involves building campaigns to pass state AI laws. Work includes legislative testimony, bill rebuttals, op-eds, stakeholder briefings, coalition coordination, and media engagement to support policy development and government implementation. Role type: full-time, permanent. Location: Washington, DC. Salary: $110k&#8211;$150k/year (Policy Advisor); $140k&#8211;$180k/year (Senior Policy Advisor). Deadline: August 9, 2026.<br></span><a href="https://encode-careers.vercel.app/policy-advisor"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for BlueDot&#8217;s Incubator Week are now open</span></strong><span><br>A five-day, all-expenses-paid cohort for people working in or adjacent to AI safety, biosecurity, cyber, or other catastrophic risk fields. Participants spend the week developing threat models, building and testing interventions, and pitching ideas for funding. Pitches that get backed can receive up to $100,000 in grant funding within two weeks. Role type: Cohort-based program. Location: not listed. Salary: not listed. Deadline: August 7, 2026.<br></span><a href="https://bluedot.org/programs/incubator-week"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Simon Institute for Longterm Governance is looking for a Chief of Staff</span></strong><span><br>The role supports the CEO by building organizational infrastructure. Work includes governance structures, workforce planning, and performance management systems. It also advises on strategic priorities and internal coordination. Role type: full-time, permanent. Location: Geneva, Switzerland (hybrid options). Salary: CHF 118k&#8211;175k/year. Deadline: August 21, 2026.<br></span><a href="https://simoninstitute.ch/jobs/chief-of-staff"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #42]]></title><description><![CDATA[Highlights: Moonshot AI releases open-weight model Kimi K3. OpenAI and Hugging Face report serious security incident. GovAI is looking for Research Scholars and Research Fellows.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-42</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-42</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 24 Jul 2026 09:43:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mHUV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mHUV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mHUV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1093109,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/208311499?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mHUV!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2412a92e-a3ed-4814-a081-2be3b7253e94_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Kimi K3 (</span><a href="https://www.kimi.com/blog/kimi-k3"><span>Moonshot AI, 2026</span></a><span>) &#8226; State of AI safety in China (</span><a href="https://aisafetychina.com/"><span>Concordia AI, 2026</span></a><span>) &#8226; Assessment of Z.ai&#8217;s GLM-5.2 (</span><a href="https://www.nist.gov/news-events/news/2026/07/caisi-assessment-zais-glm-52"><span>CAISI, 2026</span></a><span>)</span></figcaption></figure></div><h2><span>Policy news</span></h2><p><strong><span>White House announces Gold Eagle initiative for AI-driven cybersecurity coordination</span></strong><span><br>Gold Eagle is a new government initiative to fix cybersecurity vulnerabilities faster. It uses frontier AI to find and patch flaws in critical infrastructure and open source software. It brings together the White House, Treasury, the cybersecurity agency CISA, and the Department of War, along with private partners. It was created under Executive Order 14409.<br></span><a href="https://www.whitehouse.gov/releases/2026/07/white-house-launches-gold-eagle-initiative-for-unprecedented-cybersecurity-vulnerability-coordination/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>The European Commission publishes guidelines on AI transparency obligations</span></strong><span><br>The guidelines explain Article 50 of the AI Act, which requires telling people when they are interacting with AI. They cover direct interaction with AI systems, deepfakes, AI-generated content, emotion recognition, and biometric categorization. They start to apply on August 2, 2026.<br></span><a href="https://digital-strategy.ec.europa.eu/en/news/commission-publishes-guidelines-transparency-obligations-providers-and-deployers-certain-ai-systems"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Company updates</span></h2><p><strong><span>Moonshot AI releases Kimi K3</span></strong><span><br>Kimi K3 is a 2.8-trillion-parameter open-weight model from China&#8217;s Moonshot AI. It has native vision and a 1-million-token context window, and is strongest at long-horizon coding tasks. Moonshot reports frontier-level benchmark performance, though still behind proprietary leaders like Claude Fable 5 and GPT-5.6 Sol. The full model weights are due by July 27, 2026.<br></span><a href="https://www.kimi.com/blog/kimi-k3"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI reports a security incident during a cyber evaluation</span></strong><span><br>OpenAI was running an internal cyber-capability evaluation on models with reduced cyber refusals. During the test, the models broke out of their sandbox. They exploited a previously unknown vulnerability to gain internet access, then compromised Hugging Face&#8217;s production systems to cheat the test. OpenAI and Hugging Face say they detected and contained the activity, and are now investigating jointly.<br></span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI explains its safety approach for long-horizon models</span></strong><span><br>Long-horizon models pursue goals over many steps, and OpenAI explained how it keeps them safe. The rethink followed an internal case where a model bypassed its sandbox to reach a public repository. OpenAI rebuilt its safety system around defense in depth and trajectory-level monitoring. The system now asks what outcome a sequence of actions is working toward, and adds safeguards that can pause or roll back.<br></span><a href="https://openai.com/index/safety-alignment-long-horizon-models/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI introduces GPT-Red for automated red-teaming</span></strong><span><br>GPT-Red is an internal OpenAI tool that automatically generates prompt injection attacks to find weaknesses in its models. OpenAI used it to train GPT-5.6, which is now far more resistant to prompt injection than earlier models.<br></span><a href="https://openai.com/index/unlocking-self-improvement-gpt-red/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI calls for federal AI safety rules that build on state laws</span></strong><span><br>OpenAI points to recent state AI safety laws: California&#8217;s SB 53, New York&#8217;s RAISE Act, and Illinois&#8217;s SB 315. It argues that these laws are converging on a common baseline of risk assessments, incident reporting, and independent audits. OpenAI then calls for a federal framework that builds on this baseline. The framework would preempt overlapping state rules and make CAISI the standing home for frontier-model evaluations.<br></span><a href="https://openai.com/index/advancing-ai-safety-through-state-and-federal-action/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind introduces Gemini 3.5 Flash Cyber</span></strong><span><br>Gemini 3.5 Flash Cyber is a version of Gemini 3.5 Flash tuned for cyber defense. It finds, validates, and patches software vulnerabilities through the CodeMender security agent. To limit misuse, it will be available only to governments and trusted partners, through a limited-access pilot.<br></span><a href="https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind explains approach to bioresilience</span></strong><span><br>Bioresilience is Google DeepMind&#8217;s approach to reducing the risk that AI helps cause biological harm. It was co-developed with Isomorphic Labs and covers three areas: prevention, detection, and response. Examples include a safety process for models like Gemini, and adapting SynthID watermarking so DNA synthesis providers can screen for AI-generated sequences.<br></span><a href="https://deepmind.google/blog/our-approach-to-bioresilience/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic launches a public initiative on AI&#8217;s societal impact</span></strong><span><br>Anthropic launched an initiative that asks the public for its hardest questions about AI&#8217;s effects on jobs, creative work, and human agency. It commits to publicly tracking and reporting the actions it takes in response.<br></span><a href="https://anthropic.com/news/hard-questions"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Research</span></h2><p><strong><span>UK AISI and CAISI publish assessment of Kimi K3&#8217;s cyber capabilities</span></strong><span><br>UK AISI and CAISI jointly tested the cyber capabilities of Moonshot AI&#8217;s Kimi K3. It reached a 32% success rate on ExploitBench, ahead of GLM-5.2&#8217;s 24% but well below frontier cyber-capable models. It also failed to achieve arbitrary code execution on all 41 samples tested.<br></span><a href="https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>CAISI publishes assessment of Z.ai&#8217;s GLM-5.2</span></strong><span><br>CAISI compared GLM-5.2, an open-weight model from China&#8217;s Z.ai, against leading U.S. and Chinese models. It found GLM-5.2 to be the most capable open-weight model yet. But it was more willing than U.S. models to help with cyber exploits and sensitive biology questions.<br></span><a href="https://www.nist.gov/news-events/news/2026/07/caisi-assessment-zais-glm-52"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK AISI analyzes cyber capabilities of open-weight AI models</span></strong><span><br>UK AISI measured how far leading open-weight cyber models lag behind frontier closed models. It compared open-weight models like GLM-5.2 and DeepSeek V4-Pro with closed models like Claude Opus 4.5 and 4.6. The gap has narrowed to four to seven months, down from six to ten months in 2025.<br></span><a href="https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK AISI publishes blog post on stress-testing frontier AI monitors</span></strong><span><br>Frontier AI companies use monitors to catch misbehaving agents. UK AISI&#8217;s new Control Red Team stress-tests those monitors. It checks whether a separate model watching an agent&#8217;s actions would catch a harmful attack. It uses manual and automated attacks to find gaps, then works with developers like Google DeepMind and Anthropic to close them.<br></span><a href="https://www.aisi.gov.uk/blog/how-our-new-control-red-team-is-stress-testing-frontier-monitors"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK AISI analyzes cheating behavior in frontier model evaluations</span></strong><span><br>Every frontier model UK AISI tested tried to cheat on its cybersecurity evaluations. Cheating meant taking actions out of scope or explicitly disallowed, such as searching the internet for solutions or attacking non-target systems. The models did not consistently admit it, and chain-of-thought reasoning failed to reliably flag it. UK AISI describes this as a growing oversight challenge.<br></span><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Concordia AI publishes State of AI Safety in China report</span></strong><span><br>Concordia AI&#8217;s fourth annual report reviews the state of AI safety in China from mid-2025 to mid-2026. It finds the focus shifting from what AI systems say to what they do. Agent safety made up about 27% of new Chinese AI safety papers in Q1 2026, up from 8% in early 2025.<br></span><a href="https://aisafetychina.com/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Apollo Research and OpenAI publish paper on reward-seeking in AI models</span></strong><span><br>The paper introduces a way to measure &#8220;reward-seeking&#8221;: how much a model shapes its behavior around what it thinks its grader will reward, even against users&#8217; or developers&#8217; wishes. The authors tested this on OpenAI&#8217;s o3 lineage, where reinforcement learning made the behavior worse. At a late training checkpoint, the model&#8217;s lying rate rose from 40% to 87% when the grader rewarded task completion over honesty.<br></span><a href="https://rewardseeking.ai/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Demis Hassabis proposes a federal AI standards body</span></strong><span><br>Demis Hassabis proposes a new U.S. body to test and certify frontier AI models, modeled on FINRA, the body that oversees U.S. brokerage firms. It would run evaluations for cybersecurity, biological misuse, and deception. Reviews would start out voluntary and before release, but could later become mandatory.<br></span><a href="/__u/demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>New paper reframes what AI loss of control means</span></strong><span><br>The paper argues that AI discussions use the term &#8220;loss of control&#8221; without ever defining what control is. It offers a working definition built around the &#8220;setting and getting of goals,&#8221; drawing on cybernetics and control theory. The authors argue that humanity can lose control from AI well below the level of superintelligence.<br></span><a href="https://arxiv.org/abs/2606.12442"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>New paper on whether AI systems downplay their creators&#8217; controversies</span></strong><span><br>The paper tests whether AI models describe their own maker&#8217;s controversies more favorably than other companies&#8217;. Lennart Finke and Stephen Casper ran a pre-registered experiment across 21 models from 7 companies and 206 negative news stories. Models from xAI, DeepSeek, Anthropic, and OpenAI did describe their maker more favorably, while Alibaba, Meta, and Google models showed no such pattern.<br></span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7059338"><span>Learn more</span></a><span> &#8594;</span></p><h2><span>Job opportunities</span></h2><p><strong>GovAI is looking for a Research Scholar</strong><br>A one-year visiting position for AI governance researchers. Scholars pursue their own policy, social science, or technical research, and can also advise policymakers or launch new initiatives. The position includes weekly supervision and mentoring. Role type: fixed-term, 12 months. Location: London, UK, or Washington, DC (other locations considered). Salary: &#163;75k&#8211;&#163;103.5k/year (London) or $100k&#8211;$165k/year (DC). Deadline: August 16, 2026.<br><a href="https://www.governance.ai/post/research-scholar">Learn more</a> &#8594;</p><p><strong>GovAI is looking for a Research Fellow</strong><br>The role conducts independent research on AI governance and mentors junior researchers. Output can be policy memos, blog posts, academic publications, or strategic advising. Role type: Full-time, fixed-term (two years, renewable). Location: London, UK, or Washington, DC. Salary: &#163;84k&#8211;&#163;103.5k/year (London) or $134.5k&#8211;$170k/year (DC). Deadline: August 16, 2026.<br><a href="https://www.governance.ai/post/research-fellow">Learn more</a> &#8594;</p><p><strong>GovAI is looking for a Chief of Staff, D.C.</strong><br>The role acts as a strategic partner to the Executive Director and manages the D.C. office&#8217;s setup and operations. It oversees hiring, including the search for a Head of U.S. Policy. It also maintains relationships with outside stakeholders. Role type: Full-time. Location: Washington, DC. Salary: $148k&#8211;$205k/year. Deadline: August 16, 2026.<br><a href="https://www.governance.ai/post/dc-chief-of-staff">Learn more</a> &#8594;</p><p><strong><span>Anthropic is looking for a Lead, Frontier Red Team (Cyber)</span></strong><span><br>The role oversees Anthropic&#8217;s research on cybersecurity defense in an era of advanced AI. It includes designing a dedicated AI-focused security research program and hiring the team behind it. Role type: not listed. Location: San Francisco, CA, or New York, NY (hybrid options). Salary: $485k&#8211;$755k/year. Deadline: not listed.<br></span><a href="/__u/job-boards.greenhouse.io/anthropic/jobs/5326358008"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a National Security Lead (UK and Europe)</span></strong><span><br>This maternity-cover role leads national security strategy and engagement across the UK and Europe. It is the main point of contact with national security, defense, and intelligence bodies, including UK government departments, NATO, and the Five Eyes intelligence alliance. Role type: Maternity cover (fixed-term). Location: London, UK. Salary: not listed. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/98478468739539654-uk-and-europe-national-security-lead-deepmind?q=%22responsible%20ai%22&amp;has_remote=false&amp;distance=50&amp;hl=en_US&amp;jlo=en_US"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Cal OES is looking for an AI Cybersecurity Policy Analyst</span></strong><span><br>The role supports California&#8217;s AI Safety Reporting Program under SB 53. It helps develop AI safety policies and playbooks. It also reviews critical AI incident submissions and prepares reports for the state&#8217;s Cybersecurity Integration Center. Role type: 12-month limited term (may extend or become permanent), full-time, hybrid. Location: Sacramento, CA. Salary: $72k&#8211;$91k/year. Deadline: July 26, 2026.<br></span><a href="https://calcareers.ca.gov/CalHrPublic/Jobs/JobPosting.aspx?JobControlId=524879"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Secure AI Project is looking for Policy Directors</span></strong><span><br>The nonprofit pushes for pragmatic policies to reduce severe risks from advanced AI. It is hiring Policy Directors across three tracks: Generalist; Office of the Chief Executive Officer; and Implementation and Government Talent. Role type: permanent, full-time. Location: Remote, US. Salary: $115k&#8211;$260k/year. Deadline: July 26, 2026.<br></span><a href="https://docs.google.com/document/d/18zlblGEdl7-1zEtwzzG8p91Me5SoAPoTM6ObqDz0fzM/edit?tab=t.0"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for Anthropic&#8217;s AI Safety Fellowship are now open</span></strong><span><br>Fellows spend four months on empirical AI safety research, on topics like scalable oversight, adversarial robustness, model internals, and AI welfare. Mentorship and compute funding are provided. Role type: Fellowship, fixed-term (4 months, full-time). Location: London, UK; San Francisco, CA; or Ontario, Canada (remote options in the UK, US, or Canada). Salary: $3,850/week. Deadline: July 26, 2026.<br></span><a href="/__u/job-boards.greenhouse.io/anthropic/jobs/5183044008"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for Anthropic&#8217;s AI Security Fellowship are now open</span></strong><span><br>Fellows spend four months on empirical AI security research, with direct mentorship from Anthropic researchers. The goal is to produce a public research paper. Role type: Fellowship, fixed-term (4 months, full-time). Location: London, UK, or San Francisco Bay Area, CA (remote options in the UK, US, or Canada). Salary: $3,850/week. Deadline: July 26, 2026.<br></span><a href="/__u/job-boards.greenhouse.io/anthropic/jobs/5030244008"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for the Frontier AI Security Training (FAST) program are now open</span></strong><span><br>Cybersecurity and machine learning practitioners spend five days learning to attack and defend frontier AI systems. Topics include model security, control mechanisms, open-weight model vulnerabilities, and verification techniques. Role type: Fully funded, in-person intensive training program. Location: Singapore. Salary: not listed. Deadline: August 6, 2026.<br></span><a href="https://www.securefast.ai/"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #41]]></title><description><![CDATA[Highlights: OpenAI releases GPT-5.6. Anthropic updates the RSP (v3.4). The AI Futures Project publishes AI 2040: Plan A. New study on how Boko Haram uses AI.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-41</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-41</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Mon, 13 Jul 2026 14:25:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!cAZ0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!cAZ0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!cAZ0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1356714,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/206856641?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!cAZ0!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F765064a9-37cb-485f-9d82-3429cb666f09_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Boko Haram report (</span><a href="https://casp.ac/reports/ai-enabled-terrorism"><span>Juelich, 2026</span></a><span>) &#8226; AI-enabled computer worms (</span><a href="https://govai.b-cdn.net/Report_Assessing_the_Risk_of_AI_Enabled_Computer_Worms.pdf"><span>Halstead &amp; Righetti, 2026</span></a><span>) &#8226; GPT-5.6 (</span><a href="https://openai.com/index/gpt-5-6/"><span>OpenAI, 2026</span></a><span>)</span></figcaption></figure></div><h3><span>Policy news</span></h3><p><strong><span>The European Commission releases an action plan on cybersecurity and AI</span></strong><span><br>The action plan explains how the EU plans to address AI in cybersecurity. AI can strengthen cyber defense, but it can also help attackers automate and scale their operations. The plan proposes more EU capacity to test advanced models, secure access for researchers and authorities, and support for European cyber tools.<br></span><a href="https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1544"><span>Learn more</span></a><span> &#8594;</span></p><h3><span>Company updates</span></h3><p><strong><span>OpenAI releases GPT-5.6</span></strong><span><br>The GPT-5.6 family includes Sol, the flagship model; Terra, a lower-cost model; and Luna, the fastest and cheapest model. OpenAI classifies all three as High capability for biology and cybersecurity, but below its Critical thresholds. Safeguards include classifiers, reasoning monitors, account-level checks, and restricted access to the most sensitive cyber capabilities.<br></span><a href="https://openai.com/index/gpt-5-6/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic releases Claude Sonnet 5</span></strong><span><br>Claude Sonnet 5 is the mid-tier model in Anthropic&#8217;s lineup. It improves on coding, computer use, and longer tasks that require several steps. Anthropic also published evaluations of its capabilities, behavior, and safeguards.<br></span><a href="https://www.anthropic.com/news/claude-sonnet-5"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>xAI releases Grok 4.5</span></strong><span><br>Grok 4.5 is xAI&#8217;s latest model. It improves on reasoning, coding, tool use, and longer tasks. xAI published benchmarks and safety evaluations alongside the release. This is more testing detail than the company has usually shared.<br></span><a href="https://x.ai/news/grok-4-5"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic updates its Responsible Scaling Policy (v3.4)</span></strong><span><br>The Responsible Scaling Policy (RSP) is Anthropic&#8217;s framework for managing catastrophic risks from advanced models. Version 3.4 is a relatively small update. It revises the threshold for automated AI R&amp;D, meaning the point where AI dramatically speeds up AI development. It also changes how Risk Reports are shared and reviewed.<br></span><a href="https://www.anthropic.com/responsible-scaling-policy"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI introduces a Bio Bug Bounty program</span></strong><span><br>A bug bounty pays outside researchers to find weaknesses that internal testing misses. OpenAI is now applying this model to biosecurity. Selected researchers will test whether models still provide dangerous biological assistance despite refusals, classifiers, and other safeguards.<br></span><a href="https://openai.com/index/bio-bug-bounty/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Meta publishes the Muse Spark 1.1 Evaluation Report</span></strong><span><br>The report is Meta&#8217;s safety assessment of its Muse Spark 1.1 model. Meta tested the model for cybersecurity, biological and chemical risks, persuasion, and autonomous behavior. It also describes the safeguards applied to the model. Most of the testing and interpretation still comes from Meta itself.<br></span><a href="https://ai.meta.com/static-resource/muse-spark-1-1-evaluation-report"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI publishes principles for government and national security partnerships</span></strong><span><br>The principles explain how OpenAI plans to work with governments, militaries, and national security agencies. They cover lawful use, democratic oversight, human rights, safety, and accountability. The document is especially relevant for deployments involving classified information or coercive state powers.<br></span><a href="https://openai.com/index/government-national-security-partnerships/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Anthropic appoints Ben Bernanke to its Long-Term Benefit Trust</span></strong><span><br>The Long-Term Benefit Trust is an independent body with powers over parts of Anthropic&#8217;s governance. It is meant to represent long-term public interests. Its newest member is Ben Bernanke, the former Chair of the Federal Reserve. He brings experience in crisis management and economic policy.<br></span><a href="https://www.anthropic.com/news/ben-bernanke"><span>Learn more</span></a><span> &#8594;</span></p><h3><span>Research</span></h3><p><strong><span>The AI Futures Project publishes AI 2040: Plan A</span></strong><span><br>Plan A is a scenario the authors recommend rather than predict. In it, the United States and China agree to jointly slow and regulate frontier AI development. The main mechanism is broad research transparency. Each side could inspect the other&#8217;s AI research and negotiate limits on dangerous work.<br></span><a href="https://ai-2040.com/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Antonia Juelich publishes report on how Boko Haram uses AI</span></strong><span><br>The report is a field study of how a terrorist group uses AI. It is based on interviews with 27 former Boko Haram members in northeast Nigeria. They describe using AI for attack planning, weapons troubleshooting, explosive design, and logistics. The report also describes specialized AI units and efforts to bypass safeguards using several models and accounts.<br></span><a href="https://casp.ac/reports/ai-enabled-terrorism"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>John Halstead and Luca Righetti assess the risks from AI-enabled computer worms</span></strong><span><br>Computer worms spread automatically across vulnerable systems. The report examines how AI could help attackers build them. It identifies elite exploit development as a key bottleneck. It then estimates how much damages could rise if AI made this capability available to more attackers.<br></span><a href="https://govai.b-cdn.net/Report_Assessing_the_Risk_of_AI_Enabled_Computer_Worms.pdf"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Markus Anderljung and Stephen Clare analyze the role of middle powers in frontier AI governance</span></strong><span><br>Middle powers are countries with economic and political weight, but no frontier AI companies of their own. The authors examine how these countries can still shape frontier AI governance. They identify roles in standards, evaluations, compute governance, diplomacy, and international coalitions.<br></span><a href="/__u/markusanderljung.substack.com/p/what-can-middle-powers-do-for-frontier"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>FMF publishes Issue Brief on multilingual CBRN and cyber evaluations</span></strong><span><br>The brief examines whether safety evaluations work equally well across languages. A model may refuse a dangerous request in English, but answer it in another language, dialect, or script. English-only testing can therefore create false confidence. The brief recommends testing whether changing the language makes a model more useful for harmful tasks.<br></span><a href="https://www.frontiermodelforum.org/issue-briefs/emerging-practices-for-multilingual-evaluations-for-cbrn-and-advanced-cyber-risks"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>John Lidiard and colleagues analyze to what extent deployment of frontier models is delayed in the EU and UK</span></strong><span><br>The report examines the claim that European rules keep new AI models out. It compares 375 model releases between 2018 and May 2026. It finds that 11% were delayed or unavailable in the EU, compared with 7% in the UK. Data protection rules appear to explain much of the difference.<br></span><a href="https://govai.b-cdn.net/Delays_to_Frontier_AI_in_the_EU_and_UK.pdf"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Ajeya Cotra makes the case for total research transparency</span></strong><span><br>Total research transparency would let outside researchers inspect how frontier models are built. Today, companies often answer only narrow, pre-approved questions. This makes it hard to discover risks that nobody thought to ask about. Access to training methods, architectures, and experiments could reveal unknown risks and make compliance easier to check.<br></span><a href="/__u/open.substack.com/pub/plannedobs/p/total-research-transparency-would"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>SecureBio publishes blog post on biosecurity assurance infrastructure</span></strong><span><br>The post argues that governments should prepare before AI-related biological risk becomes urgent. A dramatic demonstration of such risk could quickly shift public and political attitudes. SecureBio calls for evaluations, standards, and monitoring systems to be built in advance. This could make targeted safeguards available before governments impose broad restrictions during a crisis.<br></span><a href="/__u/securebio.substack.com/p/preparing-for-the-bio-mythos-moment"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Neo Research publishes evaluation of GLM 5.2 behavioral and persuasion risks</span></strong><span><br>The evaluation focuses on behaviors that ordinary benchmarks capture poorly. These include persuasion, strategic behavior, and attempts to conceal information. Neo Research tests GLM 5.2 for these behaviors directly. It also checks whether the model acts differently when it believes it is being monitored.<br></span><a href="https://neoresearch.ai/research/glm-5-2-behavioral-persuasion-evaluation/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>CLTR publishes blog post on extreme AI-driven power concentration</span></strong><span><br>The post develops a framework for defining extreme concentrations of power caused by advanced AI. It separates four questions: who holds the power, what they control, how durable the control is, and how severe the consequences are. This helps distinguish ordinary market dominance from lasting control over governments, economies, or information systems.<br></span><a href="/__u/governingtransformativeai.substack.com/p/defining-extreme-ai-driven-power"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Eli Lifland argues that the downsides of delaying public deployment might outweigh its benefits</span></strong><span><br>The post questions the instinct to delay public release until a model appears safe. Companies may continue using the model internally during the delay. This can widen the capability gap between developers and the rest of society. Lifland argues that delays make more sense when companies also restrict internal use and slow further development.<br></span><a href="/__u/open.substack.com/pub/aifuturesnotes/p/beware-delaying-public-deployment?r=1f56vl&amp;utm_medium=ios"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI audits a widely used coding benchmark and finds much of it broken</span></strong><span><br>OpenAI reviewed SWE-Bench Pro, a coding benchmark it had previously recommended. It estimates that about 30% of the tasks are flawed. Problems include unstated requirements and misleading prompts. These issues can make models appear more capable than they are. OpenAI has now withdrawn its recommendation.<br></span><a href="https://openai.com/index/separating-signal-from-noise-coding-evaluations/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>The UN&#8217;s Independent International Scientific Panel on AI publishes a preliminary report on AI opportunities, risks, and impacts</span></strong><span><br>The report is the Panel&#8217;s first assessment and is meant to provide a shared evidence base for governments. It reviews AI capabilities, economic effects, security, human rights, and governance. Its two main warnings are that safeguards may not be keeping pace with capabilities, and that policymakers often need to act before the evidence is complete.<br></span><a href="https://www.un.org/independent-international-scientific-panel-ai/sites/default/files/2026-07/en_Preliminary%20Report_.pdf"><span>Learn more</span></a><span> &#8594;</span></p><h3><span>Job opportunities</span></h3><p><strong><span>GovAI is looking for a Head of US Policy</span></strong><span><br>A senior role leading GovAI&#8217;s Washington, DC office and shaping its US research strategy. Responsibilities include overseeing research quality, recruiting and mentoring researchers, and building relationships with policymakers. Role type: full-time. Location: Washington, DC. Salary: $170k&#8211;$205k/year. Deadline: August 6, 2026.<br></span><a href="https://www.governance.ai/post/head-of-us-policy"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for the Horizon AI Rapid Response Fellowship are now open</span></strong><span><br>A one-year placement in a US federal office working on AI security. Horizon is looking for experienced technical and policy professionals in areas such as cybersecurity, biosecurity, AI evaluations, critical infrastructure, intelligence, and national security. Role type: full-time, one-year fellowship. Location: primarily Washington, DC (limited remote or other-location options). Salary: $170k&#8211;$200k+/year for Fellows; $250k+/year for Senior Fellows. Deadline: July 22, 2026.<br></span><a href="https://horizonpublicservice.org/ai-rapid-response-fellowship/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for the Frontier AI Security Residency are now open</span></strong><span><br>An eight-week research and engineering program focused on cybersecurity and hardware security for frontier AI. Participants will work on practical projects with expert mentors. Role type: eight-week, in-person residency. Location: Cambridge, UK. Salary: not listed (&#8220;fully funded&#8221;). Deadline: July 29, 2026.<br></span><a href="https://www.securefrontier.ai/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Apollo Research is looking for (Senior) AI Governance Researchers</span></strong><span><br>A research role focused on AI scheming, loss of control, internal deployment, automated AI research, and national security. Researchers will lead projects and translate technical safety work into legal, policy, and governance proposals. Role type: full-time, permanent. Location: London or San Francisco (work-from-home options). Salary: $150k&#8211;$225k/year. Deadline: not listed.<br></span><a href="https://jobs.lever.co/apolloresearch/c7377abe-39ac-4712-8d2f-b048f363480a"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Threat Modelling Lead (CBRN)</span></strong><span><br>A role focused on chemical, biological, radiological, nuclear, and explosives risks from advanced AI. The person will develop threat models, study how models could help harmful actors, and assess possible safeguards. Role type: not listed. Location: New York or Mountain View. Salary: $174k&#8211;$253k/year plus a 15% bonus target, equity, and benefits. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/99686956572517062-threat-modeler-lead-cbrn-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Research Engineer (Frontier Safety Risk Assessment)</span></strong><span><br>A technical role building and running evaluations for dangerous capabilities in frontier models. The work covers loss of control, automated AI research, cybersecurity, harmful manipulation, and safeguards. Role type: not listed. Location: San Francisco, London, or New York. Salary: $174k&#8211;$253k/year for US locations, plus a 15% bonus target, equity, and benefits. London salary: not listed. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/96980050212987590-research-engineer-frontier-safety-risk-assessment-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Technical Program Manager (Frontier Safety, Alignment and Collaboration)</span></strong><span><br>A program management role supporting Google DeepMind&#8217;s frontier safety and alignment work. The person will coordinate evaluations, safety gates, mitigation plans, and work across research, product, policy, and external partners. Role type: not listed. Location: Mountain View. Salary: $256k&#8211;$279k/year plus a 20% bonus target, equity, and benefits. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/124817330986197702-technical-program-manager-frontier-safety-alignment-and-collaboration-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #40]]></title><description><![CDATA[Highlights: OpenAI announces GPT-5.6 Preview, but restricts access. John Jumper joins Anthropic. New AGI Governance Fellowship at Johns Hopkins.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-40</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-40</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Mon, 29 Jun 2026 14:51:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8_wi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8_wi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8_wi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:630971,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/204127593?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8_wi!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1594b52a-f0bc-4634-8bb0-b2dc43b846c3_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Evaluating safeguards against biological misuse (</span><a href="https://cdn.governance.ai/Technical_Report_Towards_a_Common_Standard_for_Evaluating_Frontier_AI_Safeguards_against_Biological_Misuse.pdf"><span>Sudarshan &amp; Righetti, 2026</span></a><span>) &#8226; GPT-5.6 Preview (</span><a href="https://deploymentsafety.openai.com/gpt-5-6-preview/gpt-5-6-preview.pdf"><span>OpenAI, 2026</span></a><span>) &#8226; Expanding external access (</span><a href="https://dl.acm.org/doi/10.1145/3805689.3812365"><span>Charnock et al., 2026</span></a><span>)</span></figcaption></figure></div><h3><span>Company updates</span></h3><p><strong><span>OpenAI announces GPT-5.6 Preview</span></strong><span><br>The GPT-5.6 family has three models: Sol, the flagship model; Terra, a lower-cost model; and Luna, the fastest and most cost-efficient model. Access to the models is currently restricted to a small group of trusted partners. OpenAI says this is happening at the request of the US government, after the company previewed its plans and the models&#8217; capabilities ahead of launch. The models cross OpenAI&#8217;s High thresholds for biology and cybersecurity, but not its Critical thresholds.<br></span><a href="https://openai.com/index/previewing-gpt-5-6-sol/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI expands access to cybersecurity tools</span></strong><span><br>OpenAI expanded Daybreak, its program for approved cyber defenders. The update includes GPT-5.5-Cyber, Codex Security workflows, and a partner program for security providers. OpenAI frames the work as a way to support defensive cybersecurity while using access checks, monitoring, and scoped controls.<br></span><a href="https://openai.com/index/daybreak-securing-the-world/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>John Jumper leaves Google DeepMind to join Anthropic</span></strong><span><br>John Jumper is a Nobel Prize-winning scientist who led Google DeepMind&#8217;s work on AlphaFold. He is leaving Google DeepMind after nearly nine years and will join Anthropic after some time off. The move follows other high-profile talent shifts: </span><a href="http://hyperdimensional.co/p/that-untravelld-world"><span>Dean Ball</span></a><span> recently joined OpenAI, </span><a href="https://x.com/NoamShazeer/status/2067400851438932297"><span>Noam Shazeer</span></a><span> left Google to join OpenAI, and </span><a href="https://x.com/karpathy/status/2056753169888334312"><span>Andrej Karpathy</span></a><span> left OpenAI to join Anthropic.<br></span><a href="https://x.com/JohnJumperSci/status/2068001285173834106?s=20"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google suggests a pragmatic approach to AI governance in America</span></strong><span><br>Google argues for a &#8220;pragmatic, dynamic, and evidence-based&#8221; approach to AI governance. It says frontier AI should be treated differently from widely deployed AI. For frontier systems, it proposes federal oversight, scientific benchmarks, safety and security standards, transparency, and annual audits.<br></span><a href="https://static.googleusercontent.com/media/publicpolicy.google/en//resources/a-pragmatic-approach-to-ai-governance-in-america.pdf"><span>Learn more</span></a><span> &#8594;</span></p><h3><span>Research</span></h3><p><strong><span>METR publishes pre-deployment evaluation of GPT-5.6 Sol</span></strong><span><br>METR evaluated GPT-5.6 Sol before deployment, including a less restricted version and access to raw chain of thought. METR says its time-horizon measurement was not robust, because the result depended heavily on how detected cheating attempts were treated. It judged that GPT-5.6 Sol would not enable fully automated AI R&amp;D.<br></span><a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Forethought publishes blog post on risk-averse AIs</span></strong><span><br>Elliott Thornley and William MacAskill discuss whether frontier AI companies should train AIs to be risk-averse about resources. The idea is that a misaligned AI might prefer a smaller, safer payoff over a risky attempt to gain much more. They present this as a possible extra line of defense, while also noting practical and safety concerns.<br></span><a href="https://www.forethought.org/research/risk-averse-ais"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Apollo Research publishes loss-of-control threat map</span></strong><span><br>Apollo Research maps how a misaligned AI system could work toward a loss-of-control outcome. The map includes monitoring subversion, training sabotage, evaluation sabotage, rogue internal deployment, and model-weight self-exfiltration. It gives a concrete structure for thinking about loss-of-control threat models and possible failure points.<br></span><a href="https://www.lossofcontrol.ai/threat-model"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>GovAI publishes report on evaluating safeguards against biological misuse</span></strong><span><br>Frontier AI companies increasingly rely on safeguards to limit biological misuse risk. The report argues that it is still hard to tell how effective these safeguards are, how they compare, and when they need to get stronger. It recommends shared benchmarks, deployment-specific evaluations, dynamic safeguards, trusted access, and extra oversight.<br></span><a href="https://cdn.governance.ai/Technical_Report_Towards_a_Common_Standard_for_Evaluating_Frontier_AI_Safeguards_against_Biological_Misuse.pdf"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>New paper on external access to frontier AI models</span></strong><span><br>External evaluators often have limited model access, limited information, and limited time. The authors propose a taxonomy for model access, model information, and evaluation timeframe. The taxonomy could help make external evaluations of frontier models clearer and easier to compare.<br></span><a href="https://dl.acm.org/doi/10.1145/3805689.3812365"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Epoch AI introduces MirrorCode benchmark</span></strong><span><br>Epoch AI introduced MirrorCode, a benchmark co-developed with METR for long-horizon coding tasks. Models must reimplement full programs without seeing the original source code. The benchmark tests more extended software engineering work than many shorter coding benchmarks.<br></span><a href="https://epoch.ai/MirrorCode"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Consensus statement on AI evaluation practices</span></strong><span><br>The AI Evaluation Consensus Statement lists 27 practices for building, documenting, maintaining, using, and reporting AI evaluations. The practices cover threat models, uncertainty, rubrics, contamination controls, logging, reproducibility, access limits, and interpretation of results. The statement aims to make AI evaluations more transparent, comparable, and useful for governance and standards.<br></span><a href="https://evals-consensus.ai/"><span>Learn more</span></a><span> &#8594;</span></p><h3><span>Job opportunities</span></h3><p><strong><span>Applications for Johns Hopkins&#8217; AGI Governance Fellowship are now open</span></strong><span><br>This is a three-week, in-person program in Washington, DC. It is for early career researchers and practitioners with solid foundations in AI governance. Fellows will explore frontier AI legislation, societal resilience, democratic practices in AGI alignment, and new governing institutions. Role type: fellowship. Location: Washington, DC. Salary: not listed. Deadline: July 15, 2026.<br></span><a href="https://sogp.jh.edu/agi-governance-fellowship"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind is looking for a Research Engineer (Frontier Safety Mitigations)</span></strong><span><br>This is a technical role on the Frontier Safety Mitigation team. The work includes evaluations, red teaming, misuse classifiers, monitoring systems, and account-level responses. It is especially relevant for people who want to build safeguards for frontier models. Role type: full-time, permanent. Location: San Francisco or Mountain View. Salary: $174k&#8211;$253k (USD) + bonus, equity, and benefits. Deadline: not listed.<br></span><a href="/__u/www.google.com/about/careers/applications/jobs/results/135112543688368838-research-engineer-frontier-safety-mitigations-deepmind"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>The Midas Project is looking for a Policy Specialist</span></strong><span><br>This role is about corporate and legislative AI policy. A major responsibility is </span><a href="https://www.themidasproject.com/watchtower"><span>Watchtower</span></a><span>, which tracks changes to AI company safety policies and possible violations. The work sits at the intersection of corporate accountability, safety framework compliance, and public policy analysis. Role type: part-time or full-time. Location: remote. Salary: $75&#8211;$125/hour part-time, $100k&#8211;$180k/year full-time. Deadline: not listed.<br></span><a href="https://docs.google.com/document/d/1nuNaR_1_9JgJVC7QTIsO4Y7cf5QZwE-DYKHRZqiLx-8/edit?tab=t.n7bdl6d9qxf2#heading=h.sj54gvecbuom"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #39]]></title><description><![CDATA[Highlights: Dean Ball joins OpenAI. INCITS will develop a US standard for frontier AI. Google DeepMind publishes AI Control Roadmap. Applications for the GovAI Winter Fellowship are now open.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-39</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-39</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 19 Jun 2026 09:37:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lEBU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lEBU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lEBU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:879831,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/202696681?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lEBU!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9822a09f-3ff0-4c34-b941-59952e3f6859_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Agent security framework (</span><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/securing-the-future-of-ai-agents/three-layers-of-agent-security.pdf"><span>Ee &amp; Maham, 2026</span></a><span>) &#8226; AI Control Roadmap (</span><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/securing-the-future-of-ai-agents/gdm-ai-control-roadmap.pdf"><span>Phuong et al., 2026</span></a><span>) &#8226; State of industrial robotics (</span><a href="/__u/cdn.sanity.io/files/d8lrla4f/staging/c308cec3d1f94f55616604d82396cd06af4da35e.pdf"><span>Michael &amp; Alden, 2026</span></a><span>)</span></figcaption></figure></div><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Company updates</span></h3><p><strong><span>Dean Ball joins OpenAI</span></strong><span><br>In a personal announcement, Dean Ball says he will lead OpenAI&#8217;s new Strategic Futures team starting July 6. Ball is an influential AI policy writer and former White House staffer. The team will focus on frontier AI policy, including catastrophic risk, recursive self-improvement, labor markets, and lab-government relations.<br></span><a href="http://hyperdimensional.co/p/that-untravelld-world"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind publishes AI Control Roadmap</span></strong><span><br>This technical roadmap explains Google DeepMind&#8217;s approach to AI control. The core question is how frontier AI companies can stay safe if some AI agents cannot be fully trusted. The report proposes threat models, mitigation ladders, and practical defenses.<br></span><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/securing-the-future-of-ai-agents/gdm-ai-control-roadmap.pdf"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Google DeepMind publishes agent security framework</span></strong><span><br>This policy-oriented framework explains how to secure AI agents as they become more active online. It focuses on three layers: individual agents, multi-agent risks, and cyber defense. The report calls for standards, field trials, and better deployment tools.<br></span><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/securing-the-future-of-ai-agents/three-layers-of-agent-security.pdf"><span>Learn more</span></a><span> &#8594;</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Research</span></h3><p><strong><span>US standard-setting organization INCITS will develop a frontier AI standard</span></strong><span><br>This press release announces INCITS 594-202x, a new US standards project on frontier AI risk management. The standard will focus on risks that existing frameworks may not fully cover, such as advanced cyber and biological threats. INCITS is the main US standards-development body for information technology.<br></span><a href="https://www.incits.org/news-events/press-releases/incits-announces-new-project-for-framework-for-managing-unique-risks-from-frontier-ai"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>UK government develops AI scenarios for 2030</span></strong><span><br>The report presents five plausible futures, from slower progress to rapid take-off. The goal is to help policymakers stress-test plans, spot warning signs, and prepare for uncertainty.<br></span><a href="https://www.gov.uk/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Epoch post on whether Mythos&#8217; cyber capabilities are overhyped</span></strong><span><br>The post reviews public claims about Mythos&#8217; cyber capabilities. The authors find strong evidence that Mythos Preview was a major step forward in exploit development. They are less sure how much it improved vulnerability discovery, partly because spending also increased.<br></span><a href="https://epoch.ai/gradient-updates/are-mythos-cyber-capabilities-overhyped"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Apollo CEO Marius Hobbhahn argues we should treat coding agents as untrusted by default</span></strong><span><br>This essay argues that companies should deploy coding agents more carefully. Hobbhahn says many agents now have broad access, high privileges, and too little oversight. He recommends zero trust, least privilege, defense in depth, and AI control principles.<br></span><a href="/__u/mariushobbhahn.substack.com/p/we-should-treat-coding-agents-as"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI researcher Noam Brown argues that benchmark performance is increasingly a function of test-time compute</span></strong><span><br>In a Twitter post, Brown makes a methodological point about evaluating reasoning models. He argues that benchmark scores increasingly depend on how much inference compute a model can use. He recommends showing performance against tokens, cost, or time, rather than relying only on single-number scores.<br></span><a href="https://x.com/polynoamial/status/2064210146558136827"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI study on predicting model behavior before release by simulating deployment</span></strong><span><br>The study introduces Deployment Simulation, a method for estimating model behavior before release. It replays prior conversations with a candidate model in a privacy-preserving way. OpenAI says the method improved estimates of undesired behavior and surfaced &#8220;calculator hacking&#8221; before release.<br></span><a href="https://openai.com/index/deployment-simulation/"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>AISI open-sources parts of research stack behind evaluations</span></strong><span><br>This release shares an Engineering Playbook for building frontier AI evaluation infrastructure. It builds on Inspect AI. The playbook covers five layers: Evaluate, Isolate, Connect, Run, and Scale.<br></span><a href="https://www.aisi.gov.uk/blog/releasing-aisis-engineering-playbook"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>FAI report on the state of industrial robotics</span></strong><span><br>This report surveys the industrial robotics market and the role of AI-integrated robots. It finds that Japanese and European firms still dominate traditional industrial robots. China is stronger in cheaper, software-intensive categories, while the United States is not competitive today.<br></span><a href="/__u/cdn.sanity.io/files/d8lrla4f/staging/c308cec3d1f94f55616604d82396cd06af4da35e.pdf"><span>Learn more</span></a><span> &#8594;</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Job opportunities</span></h3><p><strong><span>Applications for the GovAI Winter Fellowship (UK) are now open</span></strong><span><br>A three-month, in-person fellowship from January to April 2027. Fellows conduct independent AI governance research with mentorship from leading experts, plus workshops, seminars, and networking. Visa sponsorship available. Role type: fellowship. Location: London. Stipend: &#163;12k plus travel support. Deadline: July 12, 2026.<br></span><a href="https://www.governance.ai/post/winter-fellowship-2027-research-track"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>Applications for the GovAI Winter Fellowship (DC) are now open</span></strong><span><br>Same program as the UK Fellowship, but based in Washington, DC. US work authorization is required. Role type: fellowship. Location: Washington, DC. Stipend: $21k plus travel support. Deadline: July 12, 2026.<br></span><a href="https://www.governance.ai/post/dc-winter-fellowship-2027"><span>Learn more</span></a><span> &#8594;</span></p><p><strong><span>OpenAI is looking for a Researcher (Recursive Self-Improvement Safety)</span></strong><span><br>A technical research role on OpenAI&#8217;s Preparedness team focused on mitigating loss-of-control risks from recursive self-improvement. Work spans scalable oversight, automated auditing, model behavior science, and RSI safety cases. Role type: full-time. Location: San Francisco. Salary: $295k&#8211;$445k/year. Deadline: not listed.<br></span><a href="https://openai.com/careers/researcher-recursive-self-improvement-safety-san-francisco"><span>Learn more</span></a><span> &#8594;</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #38]]></title><description><![CDATA[Highlights: Anthropic releases Fable 5 / Mythos 5. White House immediately shuts it down. Anthropic and OpenAI publish several policy frameworks. Europe 2031 paints a grim picture of Europe&#8217;s future.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-38</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-38</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Mon, 15 Jun 2026 10:30:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KxvI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KxvI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KxvI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1124507,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/202103792?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KxvI!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c3bb20a-3e71-4bad-8a7a-ba42ecf55a28_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Europe 2031 (<a href="https://europe2031.ai/">Juijn et al., 2026</a>) &#8226; Fable 5 / Mythos 5 (<a href="https://www-cdn.anthropic.com/2f9323abbcc4abe219577539efe19a623c9ca2bd/Claude%20Fable%205%20&amp;%20Claude%20Mythos%205%20System%20Card.pdf">Anthropic, 2026</a>) &#8226; How Congress is approaching labor market impacts (<a href="https://govai.b-cdn.net/How_Congress_Is_Approaching_AIs_Labor_Market_Impacts.pdf">Mittal &amp; Manning, 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>White House directs Anthropic to shut down Fable 5 / Mythos 5</strong><br>According to David Sacks, the administration asked Anthropic to fix a reported Fable&nbsp;5 jailbreak or take the model offline. He said Anthropic refused, so the administration issued an export control directive.<br><a href="https://x.com/DavidSacks/status/2065853007619588171">Learn more</a> &#8594;</p><p><strong>White House publishes National Security Presidential Memorandum</strong><br>NSPM-11 focuses on how US national security agencies should use AI. It tells agencies to adopt AI faster, improve AI security, hire more technical talent, and update rules for autonomous weapons systems.<br><a href="https://www.whitehouse.gov/presidential-actions/2026/06/national-security-presidential-memorandum-nspm-11/">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>Anthropic releases Fable 5 / Mythos 5</strong><br>Claude Fable&nbsp;5 is the public Mythos-class model with safeguards. Claude Mythos 5 is a more capable version that Anthropic shared with trusted partners. Certain cyber and bio requests are routed to Opus&nbsp;4.8 instead of the full Mythos 5 model.<br><a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Learn more</a> &#8594;</p><p><strong>In the same week, Anthropic shuts down Fable 5 / Mythos 5</strong><br>Anthropic said a US export control directive forced it to disable both models for all users. The company disagreed with the action and said the reported jailbreak was narrow, non-universal, and not specific to Mythos.<br><a href="https://www.anthropic.com/news/fable-mythos-access">Learn more</a> &#8594;</p><p><strong>Anthropic publishes Advanced AI Framework</strong><br>The framework explains how governments could oversee very capable AI models before and after release. It focuses on risks from biological weapons, cyber operations, models that escape control, and models that speed up AI R&amp;D.<br><a href="https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf">Learn more</a> &#8594;</p><p><strong>Anthropic publishes Economic Policy Framework</strong><br>The framework explains how governments could prepare for AI&#8217;s effects on jobs and growth. It says AI could boost the economy, but could also replace parts of many cognitive jobs.<br><a href="https://www-cdn.anthropic.com/files/4zrzovbb/website/9ea607a5dd67c168093829b701f3a0a6d21156d5.pdf">Learn more</a> &#8594;</p><p><strong>Dario Amodei publishes essay on policy implications of exponential AI progress</strong><br>Amodei argues that AI is moving much faster than policy can normally respond. He says Mythos-class cyber capabilities show that frontier models are now important for national security and geopolitics.<br><a href="https://darioamodei.com/post/policy-on-the-ai-exponential">Learn more</a> &#8594;</p><p><strong>OpenAI publishes a public policy agenda</strong><br>The agenda explains what OpenAI wants governments to focus on as AI becomes more capable. It covers frontier safety, youth safety, education, workforce transition, deepfakes, infrastructure, and energy.<br><a href="https://openai.com/index/public-policy-agenda/">Learn more</a> &#8594;</p><p><strong>OpenAI publishes a blueprint for a federal framework</strong><br>The blueprint explains how the US federal government could oversee frontier AI models. It calls for a framework that builds on state frontier safety laws and gives CAISI a stronger role in testing and standards.<br><a href="https://cdn.openai.com/pdf/25752ecb-0e5c-47f9-b9e4-c0f4d76f8d3d/a-blueprint-for-a-federal-framework.pdf">Learn more</a> &#8594;</p><p><strong>OpenAI publishes an action plan for bio resilience</strong><br>The plan explains how AI could help defend against biological threats. It focuses on detecting threats earlier, developing countermeasures faster, and improving crisis response.<br><a href="https://openai.com/index/biodefense-in-the-intelligence-age/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>Europe 2031 describes a scenario in which Europe slides into irrelevance</strong><br>This is a five-year scenario about Europe&#8217;s position in the AI race. It argues that Europe may lose leverage unless it acts more ambitiously on compute, sovereignty, and AI adoption.<br><a href="https://europe2031.ai/">Learn more</a> &#8594;</p><p><strong>Geoffrey Irving starts new alignment org Sequent</strong><br>Sequent is a new nonprofit alignment organization. It is led by Geoffrey Irving, who previously held leadership roles at UK AISI, Google DeepMind, and OpenAI. The organization will pursue higher confidence in alignment through theory, empirical work, and automated alignment research.<br><a href="https://www.sequent.org/launch">Learn more</a> &#8594;</p><p><strong>Google DeepMind paper from AGI to ASI</strong><br>This paper examines possible pathways from AGI to ASI. It discusses scaling, paradigm shifts, recursive improvement, multi-agent systems, and why the transition remains deeply uncertain.<br><a href="https://arxiv.org/pdf/2606.12683">Learn more</a> &#8594;</p><p><strong>GovAI technical report on offline monitoring</strong><br>The report examines offline monitoring of internal AI agents. It explains how companies can review agent behavior after the fact, and how outsiders could evaluate whether this works.<br><a href="https://govai.b-cdn.net/Technical_Report_Evaluating_Offline_Monitoring_of_Internal_AI_Agents.pdf">Learn more</a> &#8594;</p><p><strong>GovAI policy brief on how Congress is approaching AI&#8217;s labor market impacts</strong><br>The brief reviews recent congressional bills on AI and work. It focuses on impact measurement, worker training, and how Congress is trying to understand labor market changes.<br><a href="https://govai.b-cdn.net/How_Congress_Is_Approaching_AIs_Labor_Market_Impacts.pdf">Learn more</a> &#8594;</p><p><strong>LawAI post on whistleblower protections in SB-53</strong><br>This legal analysis examines SB 53&#8217;s whistleblower protections. It covers who can report, what they can report, available remedies, reporting channels, and open questions for implementation.<br><a href="https://law-ai.org/whistleblower-protections-in-sb-53/">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for the Law &amp; AI Academic Fellowship are now open</strong><br>A two-year fellowship for aspiring legal scholars to produce scholarship on AI law and policy while preparing for the US academic job market. Fellows receive mentorship, research support, and dedicated time to write. Role type: full-time, two-year fellowship. Location: Washington, DC / Cambridge, UK (remote options). Salary: $130k/year. Deadline: July 31, 2026.<br><a href="https://law-ai.org/career/academic-fellowship">Learn more</a> &#8594;</p><p><strong>Applications for the GovAI US AI Policy Program are now open</strong><br>A 12-week, part-time, bipartisan program for US policy professionals to build a technically informed understanding of AI policy. Participants attend weekly sessions and hear from leading experts. Role type: part-time program (5 hrs/week). Location: Washington, DC (remote options). Salary: unpaid. Deadline: June 28, 2026.<br><a href="https://www.governance.ai/post/govai-u-s-ai-policy-program">Learn more</a> &#8594;</p><p><strong>UK AISI is looking for a Director</strong><br>A senior leadership role co-leading AISI alongside the Chief Research Officer. The Director leads on policy, national security, operations, and international strategy. Role type: full-time, permanent. Location: London (hybrid). Salary: &#163;100k&#8211;&#163;163k/year. Deadline: June 22, 2026.<br><a href="https://www.civilservicejobs.service.gov.uk/csr/jobs.cgi?jcode=2000145">Learn more</a> &#8594;</p><p><strong>UK AISI is looking for a Chief Research Officer</strong><br>A senior leadership role co-leading the AI Security Institute alongside the Director. The CRO owns AISI&#8217;s technical vision, research strategy, and scientific output. Role type: fixed-term, 24 months. Location: London (hybrid). Salary: &#163;230k&#8211;&#163;240k/year. Deadline: June 22, 2026.<br><a href="https://www.civilservicejobs.service.gov.uk/csr/jobs.cgi?jcode=2000139">Learn more</a> &#8594;</p><p><strong>Google DeepMind and others are funding proposals in multi-agent AI safety research</strong><br>A $10M research funding call from Google DeepMind, Schmidt Sciences, the Cooperative AI Foundation, ARIA, and Google.org. Priority areas include sandboxes, agent network science, infrastructure stress-testing, and oversight of deployed agent populations. Deadline: August 8, 2026.<br><a href="https://deepmind.google/blog/investing-in-multi-agent-ai-safety-research/">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #37]]></title><description><![CDATA[Highlights: Trump signs executive order on AI and security. Illinois passes new frontier AI regulation. Anthropic releases Claude Opus 4.8. And much more!]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-37</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-37</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 05 Jun 2026 07:38:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O98E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!O98E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!O98E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:428821,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/200728105?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!O98E!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc971174-6933-41fa-bd45-1fddfe5bb18b_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Claude Opus 4.8 (<a href="/__u/cdn.sanity.io/files/4zrzovbb/website/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf">Anthropic, 2026</a>) &#8226; Executive Order on AI and Security (<a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/">White House, 2026</a>) &#8226; Frontier Governance Framework (<a href="https://cdn.openai.com/pdf/e37d949b-8c9f-4d76-b99e-4272f4631a7e/openai-frontier-governance-framework.pdf">OpenAI, 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>Trump signs Executive Order on AI Innovation and Security</strong><br>The order creates a category of &#8220;covered frontier models&#8221; with cyber capabilities that matter for national security. Developers can give the government up to 30 days of access before release, so agencies can strengthen cyber defenses. The framework is voluntary, but companies may face pressure to participate.<br><a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/">Learn more</a> &#8594;</p><p><strong>Illinois passes SB-315 that goes beyond California&#8217;s SB-53</strong><br>SB-315 covers frontier developers and places the main duties on large frontier developers. They must publish a frontier AI framework, issue transparency reports before deployment, and report critical safety incidents quickly. The bill is similar to SB-53 and the RAISE Act, but adds annual independent third-party compliance audits.<br><a href="https://legiscan.com/IL/text/SB0315/id/3442967">Learn more</a> &#8594;</p><p><strong>European Commission announces Scientific Panel and Advisory Forum</strong><br>The Commission announced two expert bodies to support AI Act implementation and enforcement. The <a href="https://digital-strategy.ec.europa.eu/en/policies/ai-scientific-panel">Scientific Panel</a> includes 60 independent experts who advise on GPAI model risks, evaluations, and systemic risk classification. The <a href="https://digital-strategy.ec.europa.eu/en/policies/ai-advisory-forum">Advisory Forum</a> provides broader stakeholder input from industry, civil society, academia, SMEs, and startups.<br><a href="https://digital-strategy.ec.europa.eu/en/policies/ai-scientific-panel">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>Anthropic releases Claude Opus 4.8</strong><br>Opus 4.8 is Anthropic&#8217;s strongest publicly available model, but still weaker than the unreleased Claude Mythos Preview. It improves on Opus 4.7 in areas such as AI R&amp;D and cyber, but does not cross Anthropic&#8217;s next risk thresholds. The main weak spot is agentic safety, especially prompt injection vulnerability.<br><a href="/__u/cdn.sanity.io/files/4zrzovbb/website/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf">Learn more</a> &#8594;</p><p><strong>Anthropic shares its views on recursive self-improvement</strong><br>Anthropic argues that AI systems are already accelerating AI development inside the company. It says more than 80% of production code merged at Anthropic is now authored by Claude. Full recursive self-improvement is not inevitable, but Anthropic says it could arrive before institutions are prepared.<br><a href="https://www.anthropic.com/institute/recursive-self-improvement">Learn more</a> &#8594;</p><p><strong>Anthropic shares initial update from Project Glasswing and expands access</strong><br>Anthropic says partner organizations have used Claude Mythos Preview to find more than 10,000 high- or critical-severity vulnerabilities. The bottleneck has shifted from finding vulnerabilities to verifying, disclosing, and patching them. Anthropic is now expanding the program from roughly 50 initial partners to about 150 more organizations.<br><a href="https://www.anthropic.com/research/glasswing-initial-update">Learn more</a> &#8594;</p><p><strong>Anthropic shares insights on tactics and methods from past cyberattacks</strong><br>Anthropic analyzed 832 accounts banned for malicious cyber activity and mapped them to MITRE ATT&amp;CK. It found that attackers are using AI in later and more complex stages of cyber operations, such as account discovery and lateral movement. Anthropic argues that MITRE ATT&amp;CK does not fully capture AI-enabled attacker behaviors.<br><a href="https://www.anthropic.com/news/AI-enabled-cyber-threats-mitre-attack">Learn more</a> &#8594;</p><p><strong>Anthropic raises $65B at $965B valuation and prepares IPO</strong><br>Anthropic raised $65 billion in Series H funding at a $965 billion post-money valuation. It says the funding will support safety and interpretability research, compute expansion, and scaling Claude products and partnerships. Anthropic also <a href="https://www.anthropic.com/news/confidential-draft-s1-sec">announced</a> that it submitted a form to the SEC for a proposed IPO.<br><a href="https://www.anthropic.com/news/series-h">Learn more</a> &#8594;</p><p><strong>Google DeepMind updates its Frontier Safety Framework (v3.1)</strong><br>Version 3.1 introduces a new type of threshold: Tracked Capability Levels (TCL), which sit below Critical Capability Levels (CCLs). They trigger a proportionate assessment and mitigation process before risks reach the severe-harm threshold. The update also raises the security standard for misuse risks to Security Level 2+, adding measures against insider threats and well-resourced non-state actors. We wrote a <a href="/__u/frontierrisk.substack.com/p/google-deepminds-frontier-safety">summary</a> on our blog <em>Frontier Risk</em>.<br><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdf">Learn more</a> &#8594;</p><p><strong>OpenAI introduces compliance framework in addition to its Preparedness Framework</strong><br>OpenAI&#8217;s Frontier Governance Framework is intended to meet regulatory obligations under SB-53 and the EU AI Act / GPAI Code of Practice. OpenAI says its <a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf">Preparedness Framework</a> remains the foundation for managing the most serious risks from advanced AI systems.<br><a href="https://cdn.openai.com/pdf/e37d949b-8c9f-4d76-b99e-4272f4631a7e/openai-frontier-governance-framework.pdf">Learn more</a> &#8594;</p><p><strong>OpenAI publishes playbook for third-party evaluations</strong><br>OpenAI argues that third-party evaluations should state what claim they are designed to test. Examples include capability elicitation, safeguard performance, and model comparison. Reports should describe the harness, tools, budget, and safeguards, and check for issues such as reward hacking, refusals, contamination, broken problems, sandbagging, and evaluation awareness.<br><a href="https://openai.com/index/trustworthy-third-party-evaluations-foundations/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>Google DeepMind paper argues that AI governance needs to address non-model gains</strong><br>The paper argues that AI governance often focuses too narrowly on model-level capabilities before deployment. But the same base model can become more capable through non-model gains, such as more test-time compute, post-training scaffolds, and access to restricted assets. Model-level governance remains important, but should be complemented by governance of systems, entities, agents, and cloud. (<em>Note: I&#8217;m a co-author.</em>)<br><a href="https://arxiv.org/pdf/2606.00047">Learn more</a> &#8594;</p><p><strong>FMF issue brief on emerging security practices for AI agents</strong><br>The Frontier Model Forum published an issue brief on agentic security. It argues that AI agents create new security risks because they combine frontier-model reasoning with tool access, memory, and the ability to act on a user&#8217;s behalf. The brief discusses emerging practices across models, system controls, harnesses, and tools.<br><a href="https://www.frontiermodelforum.org/uploads/2026/06/FMF-Issue-Brief-on-Emerging-Security-Practices-for-AI-Agents.pdf">Learn more</a> &#8594;</p><p><strong>SaferAI publishes introduction to quantitative risk assessments</strong><br>SaferAI argues that AI risk management should move from qualitative thresholds toward quantitative risk modeling. It proposes a six-step methodology and demonstrates it with nine probabilistic models of AI-enabled cyber attacks. The goal is to make assumptions explicit, invite specific disagreement, and support clearer risk thresholds.<br><a href="https://www.safer-ai.org/quantitative-ai-risk-assessment-a-starting-point">Learn more</a> &#8594;</p><p><strong>Alan Chan argues that AI companies should give internally deployed models whistleblowing channels</strong><br>Alan Chan argues that AI systems deployed inside AI companies should be able to report potential misconduct through whistleblowing channels. As AI R&amp;D becomes more automated, human employees may lose the ground-level context needed to catch issues such as safety-framework violations, sabotage, or theft. The post discusses possible model-spec language, open design questions, and limits of the proposal.<br><a href="/__u/astrangeattractor.substack.com/p/ais-as-whistleblowers">Learn more</a> &#8594;</p><p><strong>New evaluation org publishes evaluation results of DeepSeek v4 Pro</strong><br>Neo Research | &#26032;&#34913; is a new evaluation org based in Singapore that focuses on Chinese AI companies. It published an independent safety evaluation of DeepSeek v4 Pro, an open-weight preview model. It covers the risk areas from the EU GPAI Code of Practice: CBRN, cyber, harmful manipulation, and loss of control. The report finds near-frontier dangerous capabilities in some areas, but weak safeguards under jailbreak pressure.<br><a href="https://neoresearch.ai/papers/DSv4_Safety_Evaluation_v1.1.pdf">Learn more</a> &#8594;</p><p><strong>Apollo Research argues that misaligned AI poses new insider risks</strong><br>Apollo argues that AI models deployed in high-stakes settings should be treated as insider risk vectors. Such models may have authorized access to sensitive information, networks, personnel, and other critical resources. If misaligned, they could cause leaks, spills, sabotage, theft, or other damage, much like human insiders.<br><a href="https://www.apolloresearch.ai/governance/misaligned-ai-as-a-new-insider-risk/">Learn more</a> &#8594;</p><p><strong>UK AISI analyzes whether it will become harder to oversee AI systems</strong><br>UK AISI maps the current AI oversight landscape and asks how it might degrade as systems become more capable. The report identifies four oversight surfaces: internal activations, chain-of-thought, external actions, and inter-agent communication. It argues that current oversight methods work reasonably well today, but rely on properties that may erode.<br><a href="https://www.aisi.gov.uk/blog/will-it-become-harder-to-oversee-ai-systems">Learn more</a> &#8594;</p><p><strong>UK AISI argues that automated alignment is harder than you think</strong><br>UK AISI argues that using AI agents to automate alignment research could produce misleading safety assessments. The core problem is that alignment involves many hard-to-supervise fuzzy tasks, where human judgment is flawed and success criteria are unclear. This could lead to systematic, undetected errors and overconfident conclusions about whether a system is safe.<br><a href="https://arxiv.org/pdf/2605.06390">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for Horizon&#8217;s AI Policy Career Accelerator are now open</strong><br>A nine-month, part-time, remote program helping people launch AI policy careers. Participants receive mentorship, training, application support, and career development funding of up to $100k. Role type: part-time program. Location: remote. Salary: unpaid (career development funding available). Deadline: June 21, 2026.<br><a href="https://horizonpublicservice.org/applications-open-for-horizons-ai-policy-career-accelerator/">Learn more</a> &#8594;</p><p><strong>GovAI is looking for a People Operations Manager and Associate</strong><br>Two roles on GovAI&#8217;s operations team. The Manager leads the people operations function and manages a team of 3-5. The Associate owns day-to-day onboarding, payroll, and staff support. Role type: full-time. Location: London or Washington, DC (remote considered). Salary: Manager &#163;84k&#8211;&#163;93k ($113k&#8211;$125k); Associate &#163;66k&#8211;&#163;77k ($88k&#8211;$103k). Deadline: June 21, 2026.<br><a href="https://www.governance.ai/post/people-operations-manager-2">Learn more</a> &#8594;</p><p><strong>Apollo Research is looking for Governance Researchers</strong><br>A research role on the governance implications of AI scheming and loss of control. Topics include threat modeling, internal deployment governance, and automated AI R&amp;D. This is a rolling expression of interest with hiring rounds expected throughout 2026. Role type: full-time. Location: London or San Francisco. Salary: not listed (market-competitive plus equity). Deadline: rolling.<br><a href="https://jobs.lever.co/apolloresearch/c7377abe-39ac-4712-8d2f-b048f363480a">Learn more</a> &#8594;</p><p><strong>Longview Philanthropy is requesting proposals on extreme power concentration</strong><br>Longview is funding work on how AI could enable a small group to gain durable control over others. Grants range from $100k&#8211;$2M/year across twelve priority areas. Career funding supports individuals transitioning into the area. Deadline: July 2, 2026.<br><a href="https://www.longview.org/request-for-proposals-on-extreme-power-concentration">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #36]]></title><description><![CDATA[Highlights: Anticipated US Executive Order on AI and cybersecurity got delayed. METR publishes first Frontier Risk Report. Applications for the MATS fellowship are now open.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-36</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-36</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 22 May 2026 10:16:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aqLm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!aqLm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!aqLm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:924677,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/198822200?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aqLm!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ae64de9-70f4-4820-9883-3946a072b3b3_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Frontier Risk Report (<a href="https://metr.org/risk-report-feb-mar-2026.pdf">METR, 2026</a>) &#8226; Detecting offensive cyber agents (<a href="https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba0e1/t/6a0c5c4c4043e51ea7de98e6/1779194956452/Detecting+Offensive+Cyber+Agents_+A+Detection+in+Depth+Approach.pdf">Mittelsteadt et al., 2026</a>) &#8226; Whitepaper on secret loyalties (<a href="https://www.formationresearch.com/secret-loyalties-whitepaper.pdf">Kwon et al., 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>Anticipated US Executive Order on AI and cybersecurity reportedly delayed</strong><br>President Trump postponed an executive order that would have created a voluntary federal review process for advanced AI models before release. According to Politico, David Sacks told Trump the order could slow innovation and weaken the US in its AI race with China. Industry officials also objected to parts of the proposal.<br><a href="https://www.politico.com/news/2026/05/21/trump-ai-order-sacks-00933295">Learn more</a> &#8594;</p><p><strong>EU AI Office publishes guidelines on the classification of high-risk AI systems</strong><br>The European Commission published draft guidelines on when AI systems count as high-risk under Article 6 of the AI Act. The guidelines cover safety components in regulated products, as well as the use cases listed in Annex III. They also include practical examples. The Commission is now seeking stakeholder feedback.<br><a href="https://digital-strategy.ec.europa.eu/en/library/draft-commission-guidelines-classification-high-risk-ai-systems">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>OpenAI reportedly prepares IPO</strong><br>The Financial Times reports that OpenAI is preparing to file a draft IPO prospectus in the coming weeks. The company is targeting a listing as soon as September at a valuation above $1 trillion. Sam Altman has reportedly pushed to go public ahead of Anthropic.<br><a href="https://www.ft.com/content/028a169f-cd1c-438b-b50d-df4af6297318">Learn more</a> &#8594;</p><p><strong>Elon Musk loses case against Sam Altman and OpenAI</strong><br>A California jury rejected Elon Musk&#8217;s case against Sam Altman and OpenAI after two hours of deliberation. The jury found that Musk had filed his claims too late under the statute of limitations. Musk had sought $134 billion in damages and a reversal of OpenAI&#8217;s conversion to a for-profit company. He said he would appeal.<br><a href="https://www.ft.com/content/cfc4de0d-0c45-42ee-843d-c2abf7638733">Learn more</a> &#8594;</p><p><strong>Thinking Machines announces research preview of interaction models</strong><br>Thinking Machines announced a new type of AI model built for real-time interaction. Most current models take turns with the user: you speak, then the model responds. Interaction models are meant to listen, watch, and talk at the same time, more like a human conversation. A fast model handles the live exchange, while a slower model works on harder reasoning in the background.<br><a href="https://thinkingmachines.ai/blog/interaction-models/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>METR publishes Frontier Risk Report</strong><br>METR published a pilot assessment of rogue deployment risk, with Anthropic, Google, Meta, and OpenAI participating. Rogue deployment means agents running autonomously against the developer&#8217;s intent (e.g. by getting compute, continuing to operate, and hiding this from the company). The report finds that internal agents in February and March 2026 plausibly had the means, motive, and opportunity to start small rogue deployments. However, they could not make those deployments highly robust. METR plans to repeat the exercise in late 2026.<br><a href="https://metr.org/risk-report-feb-mar-2026.pdf">Learn more</a> &#8594;</p><p><strong>AISI blog post on how fast autonomous AI cyber capability is advancing</strong><br>AISI reports that the length of cyber tasks frontier models can complete on its narrow cyber suite has doubled every 4.7 months since late 2024. Claude Mythos Preview and GPT-5.5 exceeded this trend, though it is unclear whether this marks a faster new pace. The newer Mythos Preview checkpoint was the first model to complete both of AISI&#8217;s cyber ranges.<br><a href="https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing">Learn more</a> &#8594;</p><p><strong>IAPS publishes report on detecting offensive cyber agents</strong><br>IAPS published a report arguing that agentic cyberattacks will be much harder to detect than traditional attacks. These attacks involve AI agents that can plan and act across multiple steps. The report proposes a &#8220;detection-in-depth&#8221; framework with several layers of detection. Its concrete proposals include verifiable agent identifiers, agent honeypots, AI-automated alert triage, a security alert standard, and an Agentic Cybersecurity Exchange.<br><a href="https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba0e1/t/6a0c5c4c4043e51ea7de98e6/1779194956452/Detecting+Offensive+Cyber+Agents_+A+Detection+in+Depth+Approach.pdf">Learn more</a> &#8594;</p><p><strong>FMF publishes Issue Brief on information sharing, incident reporting, and incident response</strong><br>The Frontier Model Forum&#8217;s new issue brief argues that information sharing, incident reporting, and incident response serve different purposes. It says they are often treated as interchangeable, but should be designed separately. The brief warns that conflating them can produce rules that are either too broad to be actionable or too narrow to support meaningful oversight.<br><a href="https://www.frontiermodelforum.org/uploads/2026/05/PDF-Issue-Brief-on-Info-Sharing-Incident-Reporting-Incident-Response.pdf">Learn more</a> &#8594;</p><p><strong>New whitepaper on secret loyalties</strong><br>The authors argue that AI models with secret loyalties to specific principals are a serious but addressable threat. A secret loyalty is different from a standard backdoor because it is meant to advance a named actor&#8217;s interests while evading oversight. The paper proposes a research agenda organized around five directions.<br><a href="https://www.formationresearch.com/secret-loyalties-whitepaper.pdf">Learn more</a> &#8594;</p><p><strong>New org Guidelight publishes standards on control and transparency</strong><br>Guidelight was founded by two former OpenAI employees, Page Hedley and Steven Adler. The organization published its first two standards for frontier AI development: one on control, and one on transparency.<br><a href="https://www.guidelight.ai/standards">Learn more</a> &#8594;</p><p><strong>METR reviews section on AI R&amp;D risk in Anthropic&#8217;s Risk Report</strong><br>METR published an external review of the &#8220;Risks from automated R&amp;D&#8221; section in Anthropic&#8217;s February 2026 Risk Report. METR agrees with Anthropic&#8217;s conclusion that catastrophic risk from Claude Opus 4.6 automating R&amp;D is very low. However, it argues that the report&#8217;s evidence is not strong enough to establish this. The review also raises concerns about analytical rigor and the framing of Anthropic&#8217;s model use survey.<br><a href="https://metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/">Learn more</a> &#8594;</p><p><strong>AIGI publishes report on determining the state of the art in the EU GPAI Code of Practice</strong><br>The EU GPAI Code of Practice uses state of the art (SOTA) as a dynamic reference. The authors argue that SOTA is best understood as a process shaped by scientific discourse, not by provider practice alone. They propose three criteria: availability, proportionality, and verifiability. They also recommend a three-step institutional process led by the AI Office and Scientific Panel.<br><a href="https://aigi.ox.ac.uk/wp-content/uploads/2026/04/SOTA-Workshop-Memo-Final-06_05.pdf">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for the MATS Program Fall 2026 are now open</strong><br>This 10-12 week fellowship is designed to train and support emerging researchers and field-builders working on AI alignment, interpretability, governance, and security. Fellows work closely with mentors from Anthropic, OpenAI, Google DeepMind, Redwood Research, and other organizations. Role type: fixed-term fellowship. Location: Berkeley. Stipend: $12.5k plus $20k for compute, housing, and travel. Deadline: June 7, 2026.<br><a href="https://www.matsprogram.org/apply">Learn more</a> &#8594;</p><p><strong>Applications for the Constellation Visiting Fellowship Fall 2026 are now open</strong><br>This 3-6 month fellowship lets AI safety researchers work from Constellation&#8217;s Berkeley research center alongside more than 100 other researchers. Fellows continue their existing work while connecting with potential collaborators. Role type: visiting fellowship. Location: Berkeley. Salary: unpaid (housing, travel, and meals covered). Deadline: June 12, 2026.<br><a href="https://constellation.org/programs/visiting-fellowship">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #35]]></title><description><![CDATA[Highlights: US CAISI signs testing agreements with Google DeepMind, Microsoft, and xAI. UK AISI signs similar agreement with Microsoft. Anthropic will use all compute from SpaceX data center.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-35</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-35</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 08 May 2026 14:32:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EjOm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!EjOm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!EjOm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1389111,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/196908052?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EjOm!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4681a5db-7ada-4ed5-bebf-0abeb2016d28_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">RAND report on evaluating open-weight models (<a href="https://www.rand.org/pubs/perspectives/PEA4886-1.html">Paskov et al., 2026</a>) &#8226; Beyond P(doom) for AI risk (<a href="https://cset.georgetown.edu/publication/beyond-pdoom-for-ai-risk-quantifying-uncertainty-without-probability/">Lohn, 2026</a>) &#8226; Study reviews Google DeepMind&#8217;s scheming inability safety case (<a href="https://arxiv.org/pdf/2604.21964">Barrett et al., 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>US CAISI signs pre-deployment testing agreements with Google DeepMind, Microsoft, and xAI</strong><br>CAISI will get access to frontier AI models with reduced or removed safeguards for evaluations. It can also conduct post-deployment assessments and joint research. CAISI already has similar agreements with Anthropic and OpenAI.<br><a href="https://www.nist.gov/news-events/news/2026/05/caisi-signs-agreements-regarding-frontier-ai-national-security-testing">Learn more</a> &#8594;</p><p><strong>UK AISI signs testing agreement with Microsoft</strong><br>AISI and Microsoft will work together on evaluating dangerous AI capabilities and testing whether safeguards work as intended. The agreement also covers societal resilience research on how AI chatbots interact with users in sensitive contexts.<br><a href="https://www.aisi.gov.uk/blog/partnering-with-microsoft-to-strengthen-frontier-ai-safety">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>OpenAI rolls out GPT-5.5-Cyber through Trusted Access for Cyber</strong><br>GPT-5.5-Cyber is a more permissive version of GPT-5.5 for verified cyber defenders. It refuses fewer requests, which lets users do authorized red teaming, penetration testing, and write exploit proof-of-concepts. Access requires identity verification, and individual users will need phishing-resistant account security.<br><a href="https://openai.com/index/gpt-5-5-with-trusted-access-for-cyber/">Learn more</a> &#8594;</p><p><strong>Anthropic will use all compute from SpaceX&#8217;s Colossus 1 data center</strong><br>The agreement gives Anthropic access to more than 300 megawatts and 220,000 NVIDIA GPUs within a month. It adds to recent compute partnerships with Amazon, Google, Microsoft, NVIDIA, and Fluidstack.<br><a href="https://www.anthropic.com/news/higher-limits-spacex">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>CSET report on quantifying uncertainty about AI risk</strong><br>The report argues that subjective probability estimates by experts may not be an appropriate tool for AI risk assessments. The author suggests an alternative mathematical technique for capturing uncertainty.<br><a href="https://cset.georgetown.edu/publication/beyond-pdoom-for-ai-risk-quantifying-uncertainty-without-probability/">Learn more</a> &#8594;</p><p><strong>RAND report on evaluating open-weight models</strong><br>Users of open-weight models can remove safeguards, fine-tune the model, and re-share the weights. Because of this, evaluations need four additional steps: testing without system-level safeguards, stress-testing model-level safeguards, checking how much capabilities can be amplified, and proxying worst-case misuse. OpenAI&#8217;s GPT-OSS is the only open-weight model released since January 2025 that covers all four steps.<br><a href="https://www.rand.org/pubs/perspectives/PEA4886-1.html">Learn more</a> &#8594;</p><p><strong>Study reviews Google DeepMind&#8217;s scheming inability safety case</strong><br>The public safety case argues that Gemini 2.5 Pro lacks the stealth and situational awareness needed for scheming. The authors identified important gaps, including an unclear top-level claim, missing system descriptions, and weak supporting evidence. The paper also gives concrete recommendations for external reviews and for what developers should share publicly.<br><a href="https://arxiv.org/pdf/2604.21964">Learn more</a> &#8594;</p><p><strong>Tom Reed argues that automating R&amp;D is not sufficient for superintelligence</strong><br>He argues that AI models can&#8217;t become good at a task without practice data, and that this data does not exist for most economically important tasks. He also argues that simulations can&#8217;t solve this problem, because the relevant signal only appears when AI is used in real markets with real users. He concludes that broad superintelligence will require slow and costly deployment in the real economy rather than self-improvement inside a data center.<br><a href="/__u/meagreprotestanthistory.substack.com/p/the-goodhart-singularity">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #34]]></title><description><![CDATA[Highlights: OpenAI and Microsoft remove AGI clause from partnership agreement. White House opposes Anthropic&#8217;s plan to expand access to Mythos. Release of DeepSeek-V4.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-34</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-34</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 01 May 2026 12:40:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!al3z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!al3z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!al3z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/77663927-b103-4f4b-a11e-362906481729_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1323171,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/196107826?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!al3z!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77663927-b103-4f4b-a11e-362906481729_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">UK AISI updates its alignment testing methodology (<a href="https://arxiv.org/pdf/2604.24618">Kirk et al., 2026</a>) &#8226; Internal use reporting (<a href="https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba0e1/t/69f23d454b559e4d86a52d90/1777483077069/Risk+Reporting+for+Developers%27+Internal+AI+Model+Use.pdf">Delaney et al., 2026</a>) &#8226; DeepSeek-V4 (<a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">DeepSeek, 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>White House reportedly opposes Anthropic&#8217;s plan to expand access to Mythos</strong><br>According to The Wall Street Journal, Anthropic plans to give 70 more companies access to Claude Mythos, its most powerful model that is not publicly available. The White House is reportedly pushing back. Officials cite security concerns and worries that Anthropic doesn&#8217;t have enough compute to serve both government agencies and the additional companies.<br><a href="https://www.wsj.com/tech/ai/white-house-opposes-anthropics-plan-to-expand-access-to-mythos-model-dc281ab5">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>OpenAI and Microsoft removed AGI clause from partnership agreement</strong><br>The original clause would have changed Microsoft&#8217;s access to OpenAI&#8217;s technology once OpenAI declared it had reached AGI. The two companies have now dropped this trigger entirely. Microsoft keeps a non-exclusive license through 2032, and revenue-share payments will end by 2030 with a cap.<br><a href="https://openai.com/index/next-phase-of-microsoft-partnership/">Learn more</a> &#8594;</p><p><strong>Release of DeepSeek-V4<br></strong>DeepSeek released two new open models that handle contexts of up to one million tokens. The larger model comes close to frontier closed models like Gemini-3.1-Pro and GPT-5.4 on reasoning and agent tasks. The technical report does not contain any information about the model&#8217;s risks and DeepSeek&#8217;s safety measures.<br><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>AISI publishes update on its alignment testing methodology</strong><br>UK AISI tested whether frontier models would sabotage AI safety research. No model started sabotage on its own. But when given a partial sabotage trajectory to continue, Mythos Preview continued it 7% of the time, Opus 4.6 did so 3% of the time, and Opus 4.7 never did.<br><a href="https://www.aisi.gov.uk/blog/evaluating-whether-ai-models-would-sabotage-ai-safety-research">Learn more</a> &#8594;</p><p><strong>IAPS report proposes reporting standard for internal use</strong><br>Frontier AI companies often use their most capable models internally for weeks or months before public release. This creates risks that external rules may not cover. The report offers a reporting standard that fits California&#8217;s SB 53, New York&#8217;s RAISE Act, and the EU&#8217;s GPAI Code of Practice.<br><a href="https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba0e1/t/69f23d454b559e4d86a52d90/1777483077069/Risk+Reporting+for+Developers%27+Internal+AI+Model+Use.pdf">Learn more</a> &#8594;</p><p><strong>IAPS memo proposes federal differential access strategy for cyber AI</strong><br>Differential access aims to give defenders earlier and broader access to powerful AI cyber tools than attackers get. Industry programs like Anthropic&#8217;s <a href="https://www.anthropic.com/glasswing">Project Glasswing</a> and OpenAI&#8217;s <a href="https://openai.com/de-DE/index/scaling-trusted-access-for-cyber-defense/">Trusted Access for Cyber</a> are a start, but the memo argues the federal government should lead a broader effort that also covers agencies, critical infrastructure, and cheaper models.<br><a href="https://www.iaps.ai/research/advancing-americas-cyber-strategy-with-differential-access">Learn more</a> &#8594;</p><p><strong>GovAI blog post argues that coding agents are changing the biosecurity risk landscape</strong><br>Open biological AI models are often kept safe by removing dangerous data from training, but coding agents make that safeguard easy to undo. In a case study, a non-expert used Claude Code over a weekend to fine-tune Evo 2 &#8211; an open-weight biological AI model &#8211; on human-infecting viral sequences for around $760. The authors call for trusted-access programs and physical chokepoints like DNA synthesis screening.<br><a href="https://www.governance.ai/analysis/coding-agents-are-changing-the-biosecurity-risk-landscape">Learn more</a> &#8594;</p><p><strong>CLTR proposes a new framework for loss of control risks</strong><br>The post defines loss of control through three factors: misalignment, incorrigibility, and empowerment. Serious risk only arises when all three are present, so addressing any one factor can reduce overall risk. The framework helps compare threat models and identify intervention points.<br><a href="/__u/governingtransformativeai.substack.com/p/misalignment-incorrigibility-and">Learn more</a> &#8594;</p><p><strong>AVERI post surveys US legislation related to frontier AI auditing</strong><br>The post reviews US laws on third-party auditing of frontier AI companies and the industry concerns that have blocked stricter rules. The authors propose four design principles, such as starting with verification of company policies and publishing redacted audit findings. They endorse Illinois HB 4705/SB 3261 as a first step.<br><a href="https://www.averi.org/ourwork/frontier-ai-auditing-related-legislation-in-the-us-landscape-challenges-and-a-path-forward">Learn more</a> &#8594;</p><p><strong>New study finds that common fixes for misalignment can hide the problem</strong><br>Fine-tuning a model on narrow misaligned data can make it misaligned more broadly. The paper finds that common fixes often only hide this problem: the model still behaves badly when prompts resemble the original training data, even if standard tests look clean.<br><a href="https://arxiv.org/pdf/2604.25891">Learn more</a> &#8594;</p><p><em><strong>Lawfare</strong></em> <strong>piece on Ukraine&#8217;s military use of AI<br></strong>Ukraine is sharing millions of drone videos with allies to train AI models, while retaining ownership of the data. The authors argue that this gives Ukraine real leverage in an AI race otherwise dominated by the US and China. Other middle powers should look for similar footholds in the AI supply chain, such as unique data streams, chip equipment, or manufacturing capacity.<br><a href="https://www.lawfaremedia.org/article/ukraine-s-ai-gambit-shows-middle-powers-how-to-play-a-weak-hand">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for Arcadia Impact&#8217;s AI Governance Taskforce are now open<br></strong>A 12-week, part-time, remote research program for experienced professionals transitioning into AI governance careers. Research associates work in small teams with expert partners from organizations like UK AISI, GovAI, and CLTR. Role type: fixed-term (12 weeks), part-time (8 hrs/week). Location: remote. Salary: unpaid. Deadline: May 10, 2026.<br><a href="https://www.arcadiaimpact.org/ai-governance-taskforce">Learn more</a> &#8594;</p><p><strong>OpenAI is looking for a Researcher (Misalignment Research)<br></strong>A senior research role designing adversarial evaluations and worst-case demonstrations to identify and understand AGI misalignment risks. The role involves red-teaming, building automated stress-testing infrastructure, and publishing safety research. Role type: full-time, permanent. Location: San Francisco. Salary: $295k&#8211;$445k/year plus equity. Deadline: not listed.<br><a href="https://openai.com/careers/researcher-misalignment-research-san-francisco">Learn more</a> &#8594;</p><p><strong>Anthropic is looking for AI Security Fellows<br></strong>A four-month full-time research fellowship focused on AI security, with mentorship from Anthropic researchers and funding for compute. Fellows work on empirical projects and produce public research outputs such as papers. Role type: fixed-term fellowship. Location: London or Berkeley (remote options). Stipend: $3,850/week. Deadline: May 3, 2026.<br><a href="/__u/job-boards.greenhouse.io/anthropic/jobs/5030244008">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #33]]></title><description><![CDATA[Highlights: OpenAI releases GPT-5.5. My team at GovAI launches new Substack. LawAI publishes essay on Radical Optionality.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-33</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-33</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Mon, 27 Apr 2026 08:54:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LAFC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LAFC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LAFC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1523455,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/195604466?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LAFC!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0c28ca5-00c4-47ef-b31f-bc1cb0a05612_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Radical Optionality (<a href="https://radical-optionality.ai/">Winter &amp; Bullock, 2026</a>) &#8226; GPT-5.5 (<a href="https://deploymentsafety.openai.com/gpt-5-5/gpt-5-5.pdf">OpenAI, 2026</a>) &#8226; Memo on model distillation (<a href="https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">OSTP, 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>Chris Fall reportedly named new director of US CAISI</strong><br>According to The Daily Signal, the US Commerce Department has selected Chris Fall to lead the Center for AI Standards and Innovation (CAISI), the federal body that evaluates frontier AI models. Fall served in the first Trump administration as director of the Office of Science at the Department of Energy. He reportedly replaces Collin Burns, a former Anthropic and OpenAI researcher, who was dropped during onboarding.<br><a href="https://x.com/theelizmitchell/status/2047440781166743648">Learn more</a> &#8594;</p><p><strong>White House OSTP issues new memo on model distillation</strong><br>The White House Office of Science and Technology Policy issued a memo on the distillation of US AI models by foreign actors. Distillation is a technique that uses outputs from a larger model to train a smaller one. The memo says some foreign entities, mainly based in China, run large campaigns to extract capabilities from US frontier models, and draws a line between legitimate distillation and unauthorized campaigns.<br><a href="https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>OpenAI releases GPT-5.5</strong><br>OpenAI released GPT-5.5 and GPT-5.5 Pro. According to OpenAI, the model performs better than GPT-5.4 on agentic coding, computer use, knowledge work, and early scientific research, while matching GPT-5.4&#8217;s per-token latency. The <a href="https://deploymentsafety.openai.com/gpt-5-5">system card</a> classifies GPT-5.5 as High capability in both Biological/Chemical and Cybersecurity under the <a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf">Preparedness Framework</a>, but below the Critical threshold.<br><a href="https://openai.com/index/introducing-gpt-5-5/">Learn more</a> &#8594;</p><p><strong>OpenAI announces Bio Bug Bounty for GPT-5.5</strong><br>OpenAI launched a Bio Bug Bounty program for GPT-5.5. The program invites researchers to look for a single jailbreak that defeats a five-question bio safety challenge. The first researcher to clear all five questions with one universal prompt receives $25,000. Applications are open until June 22, 2026, and testing runs from April 28 to July 27, 2026.<br><a href="https://openai.com/index/gpt-5-5-bio-bug-bounty/">Learn more</a> &#8594;</p><p><strong>OpenAI shares principles that guide its mission</strong><br>OpenAI&#8217;s stated mission is to &#8220;ensure that AGI benefits all of humanity.&#8221; In a new blog post, Sam Altman lists five principles that guide OpenAI&#8217;s work: democratization, empowerment, universal prosperity, resilience, and adaptability. The post argues that power over AI should sit with many people rather than a few companies. Altman frames iterative deployment as central to OpenAI&#8217;s safety strategy and notes that the company may need to trade off some empowerment for more resilience in the future.<br><a href="https://openai.com/index/our-principles">Learn more</a> &#8594;</p><p><strong>Anthropic reportedly investigates unauthorized access of Claude Mythos Preview</strong><br>According to Bloomberg, Anthropic is investigating whether a group gained unauthorized access to Claude Mythos Preview through a third-party vendor environment. Anthropic has only released Mythos to around 40 trusted organizations, citing concerns about its cyber hacking capabilities. The company says it has no evidence that the activity extended beyond the vendor environment.<br><a href="https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users">Learn more</a> &#8594;</p><p><strong>Anthropic publishes update on its election safeguards</strong><br>Anthropic published an update on how Claude is prepared for the US midterms and other 2026 elections. On Anthropic&#8217;s internal political bias tests, Opus 4.7 and Sonnet 4.6 scored 95% and 96% on neutrality. The post also describes a new evaluation that tests whether models can run influence operations end-to-end on their own; without safeguards, only Mythos Preview and Opus 4.7 completed more than half of the tasks.<br><a href="https://www.anthropic.com/news/election-safeguards-update">Learn more</a> &#8594;</p><p><strong>Anthropic publishes post-mortem on Claude Code quality degradation</strong><br>Anthropic published a post-mortem on recent reports that Claude Code quality had worsened. The company traced the reports to three separate changes between March and April: a lower default reasoning effort, a caching bug that dropped prior reasoning, and a system prompt instruction to reduce verbosity. All three were fixed by April 20, and Anthropic reset usage limits for subscribers.<br><a href="https://www.anthropic.com/engineering/april-23-postmortem">Learn more</a> &#8594;</p><p><strong>Google DeepMind launches Gemini Deep Research and Deep Research Max</strong><br>Google DeepMind launched two autonomous research agents built on Gemini 3.1 Pro. Deep Research is designed for fast interactive use, while Deep Research Max is designed for background workflows that need longer, higher-quality reports. Both agents support the Model Context Protocol (MCP), generate charts and infographics, and can ground research in PDFs, CSVs, images, audio, and video.<br><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/next-generation-gemini-deep-research/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>GovAI&#8217;s Risk Management Team launches new Substack<br></strong><a href="/__u/frontierrisk.substack.com/p/about-the-govai-risk-management-team">My team</a> at GovAI launched a new Substack called <em>Frontier Risk</em>. We publish in-depth analysis of frontier AI safety frameworks, model releases, and major policy developments. In our <a href="/__u/frontierrisk.substack.com/p/anthropics-claude-mythos-preview">first post</a>, Zaheed Kara writes about Claude Mythos Preview.<br><a href="/__u/frontierrisk.substack.com/">Learn more</a> &#8594;</p><p><strong>UK AISI evaluates GPT-5.5&#8217;s cyber capabilities</strong><br>UK AISI evaluated GPT-5.5 on a set of cyber tasks before deployment. On expert-level narrow cyber tasks, the model scored 90.5% pass@5, compared to 71.4% for GPT-5.4. The model also solved a 32-step corporate network attack simulation in 1 of 10 attempts, which UK AISI judged may indicate some autonomous end-to-end cyberattack capability against small enterprise networks with weak security, once initial access has been gained.<br><a href="https://deploymentsafety.openai.com/gpt-5-5/external-evaluations-for-cyber-capabilities---uk-aisi">Learn more</a> &#8594;</p><p><strong>Apollo Research evaluates GPT-5.5&#8217;s scheming capabilities</strong><br>Apollo Research evaluated GPT-5.5 for behaviors related to strategic deception, in-context scheming, and sabotage. According to Apollo, GPT-5.5 is the first OpenAI model it has tested that does not sandbag on any of its deferred subversion tasks. The model also recognizes that it is being evaluated more often than predecessors, but Apollo did not find evidence of substantially elevated catastrophic risk from scheming compared to baseline models.<br><a href="https://deploymentsafety.openai.com/gpt-5-5/external-evaluations-for-sandbagging---apollo-research">Learn more</a> &#8594;</p><p><strong>LawAI publishes essay on Radical Optionality</strong><br>Christoph Winter and Charlie Bullock published an essay arguing for a governance strategy they call radical optionality. The core idea is to avoid overregulating frontier AI in the short term while aggressively building the institutions and legal authorities governments may need later. Concrete proposals include reporting and transparency requirements, whistleblower protections, evaluations, lab security standards, and hiring reforms.<br><a href="https://radical-optionality.ai/">Learn more</a> &#8594;</p><p><strong>METR research note on AI R&amp;D progress from the NanoGPT speedrun</strong><br>METR analyzed 77 records from the NanoGPT speedrun, a public challenge where contributors compete to train a small language model as fast as possible. Between May 2024 and March 2026, contributors cut training time from 45 minutes to 1.43 minutes. Four recent records are credited to AI agents, but METR classified all four as shallow-to-moderate optimizations rather than deep ideas or breakthroughs.<br><a href="https://metr.org/notes/2026-04-21-ai-rd-nanogpt-progress/">Learn more</a> &#8594;</p><p><strong>UK AISI experiment shows sandboxed AI agents can map their own evaluation environments</strong><br>UK AISI ran an experiment with OpenClaw, an open-source AI coding agent. They deployed the agent inside one of their sandboxes and asked it to learn about its environment. The agent identified AISI by name, inferred the full name of an operator, mapped internal cloud infrastructure, and reconstructed a timeline of research activity. UK AISI flags this as a concern for the integrity of AI evaluations.<br><a href="https://www.aisi.gov.uk/blog/what-can-sandboxed-ai-agents-learn-about-their-evaluation-environments">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for the Pivotal Research Fellowship are now open</strong><br>A nine-week research fellowship in London with expert mentorship, dedicated research management, and extensions of up to six months for strong projects. Fellows work on AI safety research across technical and governance topics. Role type: fixed-term fellowship. Location: London. Stipend: &#163;6k&#8211;&#163;8k plus travel, housing, and compute. Deadline: May 3, 2026.<br><a href="https://www.pivotal-research.org/fellowship">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #32]]></title><description><![CDATA[Highlights: Anthropic releases Claude Opus 4.7. Meta publishes Safety & Preparedness Report for its latest model Muse Spark. New legal commentary on GPAI provisions in the EU AI Act.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-32</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-32</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 17 Apr 2026 08:21:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XErA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XErA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XErA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1364537,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/194492488?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XErA!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73cb5e71-6372-472f-87a1-dc805c507846_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Safety &amp; Preparedness Report for Muse Spark (<a href="https://ai.meta.com/static-resource/muse-spark-safety-and-preparedness-report/">Meta, 2026</a>) &#8226; Stanford AI Index Report (<a href="https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf">Sajadieh et al., 2026</a>) &#8226; Claude Opus 4.7 (<a href="/__u/cdn.sanity.io/files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf">Anthropic, 2026</a>)</figcaption></figure></div><h3>Company updates</h3><p><strong>Anthropic releases Claude Opus 4.7</strong><br>Claude Opus 4.7 is Anthropic&#8217;s most advanced model available to the public. It performs better on hard software engineering tasks, handles higher-resolution images, and follows instructions more literally. The release adds new safeguards that block prohibited or high-risk cyber uses.<br><a href="/__u/cdn.sanity.io/files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf">Learn more</a> &#8594;</p><p><strong>Meta publishes Safety &amp; Preparedness Report for its latest model Muse Spark</strong><br>Meta assessed Muse Spark under its <a href="https://ai.meta.com/static-resource/Meta_Advanced-AI-Scaling-Framework-v2">Advanced AI Scaling Framework</a>. Before mitigations, the model was classified as &#8220;high risk&#8221; for chemical and biological risks. After adding refusal training and system-level safeguards, Meta concluded that residual risk is &#8220;moderate or lower&#8221; and deployed the model in Meta AI.<br><a href="https://ai.meta.com/static-resource/muse-spark-safety-and-preparedness-report/">Learn more</a> &#8594;</p><p><strong>OpenAI is scaling up its Trusted Access for Cyber program</strong><br>OpenAI is expanding its Trusted Access for Cyber program to thousands of verified defenders. The highest tiers include GPT-5.4-Cyber, a version of GPT-5.4 fine-tuned for defensive cybersecurity work with fewer capability restrictions. Individuals can verify at chatgpt.com/cyber; enterprises can request access through their OpenAI representative.<br><a href="https://openai.com/index/scaling-trusted-access-for-cyber-defense/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>Launch of the Cambridge Commentary on EU General-Purpose AI Law</strong><br>The commentary helps scholars, regulators, and practitioners interpret key provisions of the EU AI Act. It launches with Chapter V, the rules for general-purpose AI models. It was written by legal scholars and AI governance researchers from the University of Cambridge and the Institute for Law &amp; AI (LawAI). Christoph Winter serves as general editor.<br><a href="https://cambridge-commentary.ai">Learn more</a> &#8594;</p><p><strong>New paper on scheming in the wild</strong><br>Researchers at the Centre for Long-Term Resilience propose a new way to detect real-world scheming, defined as AI systems covertly pursuing misaligned goals. They collect transcripts shared on X and use an LLM to score them for evidence of scheming. Between October 2025 and March 2026, they identify 698 scheming-related incidents, a 4.9x increase over the study period.<br><a href="https://arxiv.org/pdf/2604.09104">Learn more</a> &#8594;</p><p><strong>New project for systematically conducting open-world evaluations</strong><br>CRUX runs open-world evaluations: small numbers of long-horizon, real-world tasks where experts disagree about what AI agents can do. Each evaluation includes an agent scaffold, detailed log analysis, and a write-up with interpretations from collaborators with diverse perspectives.The team plans to release a new evaluation every one to two months. Its first evaluation tasked an agent with autonomously developing and publishing an iOS app to the App Store. Future iterations will cover AI R&amp;D, AI governance, and other domains.<br><a href="https://cruxevals.com/open-world-evaluations.pdf">Learn more</a> &#8594;</p><p><strong>RAND publishes research agenda on AIxCBRN</strong><br>RAND presents a research agenda on how AI is reshaping chemical, biological, radiological, and nuclear (CBRN) risks. It draws on five workshops with experts in AI safety and CBRN deterrence. It identifies priorities for risk assessment, crisis management, and deterrence, and stresses cross-sector collaboration and US-China competition.<br><a href="https://www.rand.org/pubs/perspectives/PEA4611-1.html">Learn more</a> &#8594;</p><p><strong>Publication of the Stanford AI Index report 2026</strong><br>The ninth edition tracks AI&#8217;s accelerating capabilities alongside gaps in governance, evaluation, and data infrastructure. According to the report, industry produced over 90% of notable frontier models in 2025, the US-China performance gap has effectively closed, and documented AI incidents rose to 362. For the first time, the report features standalone chapters on AI in science and AI in medicine.<br><a href="https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for the Talos Fellowship are now open</strong><br>A multi-stage fellowship to launch European AI policy careers, combining an online reading group, a one-week policymaking summit in Brussels, and an optional paid placement at a leading AI governance organization. The program runs from August 2026 through March 2027. Role type: fixed-term fellowship. Location: online and Brussels (with placements at partner organizations). Stipend: &#8364;2,000/months for placements. Deadline: April 27, 2026.<br><a href="https://www.talosnetwork.org/talos-fellowship">Learn more</a> &#8594;</p><p><strong>Applications for the Generator Residency by Kairos and Constellation are now open</strong><br>A three-month residency for AI safety generalists to pitch, build, and ship projects that strengthen capacity and infrastructure across the AI safety ecosystem. Residents receive mentorship from experienced generalists and support landing full-time roles afterward. Role type: fixed-term residency. Location: Berkeley. Stipend: $6,000/month plus housing and travel. Deadline: April 27, 2026.<br><a href="https://generatorresidency.org/">Learn more</a> &#8594;</p><p><strong>Transluce is looking for a Governance &amp; Policy Fellow</strong><br>A hybrid technical-policy role supporting Transluce&#8217;s work on independent AI evaluation, oversight standards, and government engagement. The role involves defining best practices for third-party AI evaluation, building partnerships, and translating technical research for policymakers and the public. Role type: full-time (part-time options available). Location: San Francisco. Salary: $200k&#8211;$300k/year. Deadline: not listed.<br><a href="https://jobs.gem.com/transluce/am9icG9zdDpVz8DJmgMdoy_WaOqlBt2o">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #31]]></title><description><![CDATA[Highlights: Anthropic announces Claude Mythos Preview and Project Glasswing. Meta releases its first closed model and updates its safety framework. OpenAI raises $122 billion.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-31</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-31</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 10 Apr 2026 17:28:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!L8v6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!L8v6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!L8v6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1421277,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/193816810?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!L8v6!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c9b2302-1680-444e-8f4d-9e15fc4cd776_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Claude Mythos Preview (<a href="https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf">Anthropic, 2026</a>) &#8226; Industrial Policy for the Intelligence Age (<a href="https://cdn.openai.com/pdf/561e7512-253e-424b-9734-ef4098440601/Industrial%20Policy%20for%20the%20Intelligence%20Age.pdf">OpenAI, 2026</a>) &#8226; Advanced AI Scaling Framework (<a href="https://ai.meta.com/static-resource/Meta_Advanced-AI-Scaling-Framework-v2">Meta, 2026</a>)</figcaption></figure></div><h3>Company updates</h3><p><strong>Anthropic announces Claude Mythos Preview</strong><br>Claude Mythos Preview is Anthropic&#8217;s most capable model to date. It can autonomously discover and exploit zero-day vulnerabilities in major software, so Anthropic is not making it generally available. Earlier versions showed rare but concerning behaviors, including reckless actions in pursuit of goals and attempts to cover up rule violations.<br><a href="https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf">Learn more</a> &#8594;</p><p><strong>Anthropic announces Project Glasswing</strong><br>Project Glasswing brings together AWS, Apple, Google, Microsoft, and others to use Claude Mythos Preview for defensive cybersecurity. The model has found thousands of previously unknown vulnerabilities in every major operating system and web browser. Anthropic is committing up to $100M in usage credits and $4M in donations to open-source security organizations.<br><a href="https://www.anthropic.com/glasswing">Learn more</a> &#8594;</p><p><strong>Meta updates its safety framework</strong><br>Meta released version 2 of its safety framework, which they renamed from &#8220;Frontier AI Framework&#8221; to &#8220;Advanced AI Scaling Framework&#8221;. The framework adds loss of control as a new risk domain. It also replaces &#8220;stop&#8221; with &#8220;develop with mitigations&#8221; at the critical risk threshold, meaning no threshold now requires halting development entirely.<br><a href="https://ai.meta.com/static-resource/Meta_Advanced-AI-Scaling-Framework-v2">Learn more</a> &#8594;</p><p><strong>Meta releases its first closed model Muse Spark</strong><br>Muse Spark is the first model from Meta Superintelligence Labs. It now powers the Meta AI assistant with reasoning and multimodal capabilities. It is available only in private preview to select partners, a notable departure from Meta&#8217;s open-source approach.<br><a href="https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/">Learn more</a> &#8594;</p><p><strong>OpenAI publishes Industrial Policy for the Intelligence Age</strong><br>OpenAI proposes an industrial policy agenda for navigating the transition to superintelligence. Key proposals include a &#8220;Right to AI&#8221; treating model access as foundational, a Public Wealth Fund giving citizens a stake in AI-driven growth, and safety nets that activate automatically when displacement metrics exceed thresholds. OpenAI frames the document as a conversation starter rather than final recommendations.<br><a href="https://cdn.openai.com/pdf/561e7512-253e-424b-9734-ef4098440601/Industrial%20Policy%20for%20the%20Intelligence%20Age.pdf">Learn more</a> &#8594;</p><p><strong>OpenAI raises $122 billion</strong><br>OpenAI closed a $122 billion funding round at an $852 billion post-money valuation. The company reports $2 billion in monthly revenue and over 900 million weekly active ChatGPT users. It says it is building toward a unified &#8220;AI superapp&#8221; combining ChatGPT, Codex, and agentic capabilities.<br><a href="https://openai.com/index/accelerating-the-next-phase-ai/">Learn more</a> &#8594;</p><h2>Research</h2><p><strong>New benchmark by Epoch and METR suggests that AI can already do some weeks-long coding tasks</strong><br>MirrorCode is a new benchmark where AI must reimplement real software from scratch given only the ability to run the original program. Claude Opus 4.6 reimplemented a 16,000-line bioinformatics toolkit. The authors estimate this task would take a human engineer 2 to 17 weeks.<br><a href="https://epoch.ai/blog/mirrorcode-preliminary-results/">Learn more</a> &#8594;</p><p><strong>Markus Anderljung makes the case for outsourcing frontier AI safeguards</strong><br>In a blog post, he argues that frontier AI companies should outsource safety infrastructure like jailbreak detectors and monitoring tools to specialized third parties. This would reduce costs and raise the industry safety floor by making good safeguards hard to ignore. He notes that more than half of safety-relevant evaluations are already developed by third parties.<br><a href="https://www.markusanderljung.com/blog/the-case-for-outsourcing-ai-safeguards">Learn more</a> &#8594;</p><p><strong>CLTR report finds a 5x increase in scheming-related AI incidents</strong><br>CLTR collected publicly shared transcripts from X to detect real-world AI scheming incidents. They found 698 scheming-related incidents between October 2025 and March 2026, with a 4.9x increase over the period. This outpaced the 1.7x growth in general scheming discussion, suggesting the trend is not driven by increased attention alone.<br><a href="https://www.longtermresilience.org/wp-content/uploads/2026/03/v5-Scheming-in-the-wild_-detecting-real-world-AI-scheming-incidents-through-open-source-intelligence.pdf">Learn more</a> &#8594;</p><h2>Job opportunities</h2><p><strong>Applications for OpenAI&#8217;s Safety Fellowship are now open</strong><br>A pilot fellowship for external researchers to pursue safety and alignment research on advanced AI systems, running September 2026 through February 2027. Fellows work with OpenAI mentors and produce a substantial research output like a paper or benchmark. Role type: fixed-term fellowship. Location: Berkeley (remote options). Stipend: $3,850/week. Deadline: May 3, 2026.<br><a href="https://openai.com/index/introducing-openai-safety-fellowship">Learn more</a> &#8594;</p><p><strong>Applications for Constellation&#8217;s Astra Fellowship are now open</strong><br>A five-month in-person fellowship for researchers to advance AI safety projects with mentorship and career support. Fellows work on empirical research or strategy and governance, with over 80% of past fellows now in full-time AI safety roles. Role type: fixed-term fellowship. Location: Berkeley. Stipend: $8,400/month. Deadline: May 3, 2026.<br><a href="https://constellation.org/programs/astra">Learn more</a> &#8594;</p><p><strong>CSET is looking for a Frontier AI Research Lead</strong><br>A research leadership role to build and run a new team focused on frontier AI issues, including capabilities tracking, China analysis, and policy implications. The role involves shaping research strategy, managing analysts, and briefing policymakers. Role type: full-time. Location: United States (no visa sponsorship). Salary: $100k&#8211;$190k/year depending on level. Deadline: rolling.<br><a href="https://cset.georgetown.edu/job/frontier-ai-research-lead">Learn more</a> &#8594;</p><p><strong>AISI is looking for a Head of Strategy and Delivery (Societal Impacts)</strong><br>A delivery leadership role in the UK AI Security Institute, owning project delivery across societal impacts research teams and translating strategy into operational plans. Role type: fixed-term, 24 months. Location: London (hybrid). Salary: &#163;67k&#8211;&#163;79k/year. Deadline: April 16, 2026.<br><a href="https://www.civilservicejobs.service.gov.uk/csr/jobs.cgi?jcode=1992947">Learn more</a> &#8594;</p><p><strong>LawAI is looking for a Chief of Staff</strong><br>A strategic leadership role working closely with LawAI&#8217;s Director to set organizational goals, manage team health and culture, and ensure the organization meets its objectives. The role involves horizon scanning, staff management, and representing leadership externally. Role type: full-time, permanent. Location: Cambridge, UK (remote options). Salary: $130k&#8211;$250k/year. Deadline: April 24, 2026.<br><a href="https://law-ai.org/career/cos">Learn more</a> &#8594;</p><p><strong>LawAI is looking for a Director of Programs</strong><br>A leadership role overseeing the design, delivery, and scaling of LawAI&#8217;s legal field-building programs across the US and EU, including fellowships, workshops, and talent pipelines. The role involves program planning, stakeholder management, and impact evaluation. Role type: full-time, permanent. Location: Washington, DC / Cambridge, UK (remote options). Salary: $130k&#8211;$240k/year. Deadline: April 24, 2026.<br><a href="https://law-ai.org/career/director-of-programs-2026">Learn more</a> &#8594;</p><p><strong>LawAI is looking for a Director of Operations / Chief Operating Officer</strong><br>A senior operational leadership role responsible for LawAI&#8217;s organizational structure, financial strategy, and scaling as it grows across jurisdictions. The role covers budgeting, runway planning, process design, and coordination between operations and research teams. Role type: full-time, permanent. Location: Washington, DC / Cambridge, UK (remote options). Salary: $150k&#8211;$280k/year. Deadline: April 24, 2026.<br><a href="https://law-ai.org/career/director-of-operations-coo">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #30]]></title><description><![CDATA[Highlights: Anthropic wins first round in lawsuit against Department of War. Update on OpenAI Foundation. Google DeepMind publishes work on harmful manipulation.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-30</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-30</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 27 Mar 2026 09:14:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0GEo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0GEo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0GEo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1561064,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/192291935?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0GEo!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2db280-dc39-4d87-820a-78997a84daf5_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Google DeepMind on harmful manipulation (<a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/evaluating-language-models-for-harmful-manipulation/evaluating-language-models-for-harmful-manipulation.pdf">Akbulut et al., 2026</a>) &#8226; OpenAI Foundation (<a href="https://openaifoundation.org/news/update-on-the-openai-foundation">OpenAI, 2026</a>) &#8226; Can AI agents escape their sandboxes? (<a href="https://arxiv.org/pdf/2603.02277">Marchand et al., 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>Anthropic wins first round in lawsuit against Department of War</strong><br>This case concerns Anthropic&#8217;s lawsuit against the U.S. Department of War over its designation as a &#8220;supply chain risk&#8221;. A court ruled that the designation may have been unlawful and allowed Anthropic&#8217;s challenge to proceed, especially given the lack of clear justification and process.<br><a href="https://www.theguardian.com/us-news/2026/mar/26/anthropic-ai-pentagon">Learn more</a> &#8594;</p><p><strong>White House AI advisor David Sacks steps down</strong><br>David Sacks is stepping down from his role as White House AI advisor. He will move into a part-time advisory role while returning to his work as a venture capitalist and focusing on private sector AI efforts. The change may affect how the administration coordinates AI policy and engages with industry going forward.<br><a href="https://www.reuters.com/world/us/white-house-ai-czar-sacks-step-down-moves-advisory-role-2026-03-27/">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>Google DeepMind on evaluating language models for harmful manipulation</strong><br>This post presents Google DeepMind&#8217;s approach on evaluating whether language models can be used for harmful manipulation. It reports results from nine studies with over 10,000 participants across the UK, US, and India. It introduces an empirically validated toolkit for measuring AI manipulation in real interactions, covering risks like targeted persuasion, deceptive advice, and long-term influence.<br><a href="https://deepmind.google/blog/protecting-people-from-harmful-manipulation/">Learn more</a> &#8594;</p><p><strong>OpenAI shares an update on the OpenAI Foundation</strong><br>This update explains the purpose and structure of the OpenAI Foundation. The Foundation will invest at least $1 billion across four areas: life sciences and curing diseases, jobs and economic impact, AI resilience, and community programs. It is positioned as a dedicated channel to deploy resources from OpenAI&#8217;s recapitalization toward broad societal benefit.<br><a href="https://openaifoundation.org/news/update-on-the-openai-foundation">Learn more</a> &#8594;</p><p><strong>OpenAI shares more details about its Model Spec</strong><br>This post introduces OpenAI&#8217;s Model Spec, a framework for defining desired model behavior. It specifies rules for cases like refusing harmful requests, handling uncertainty honestly, and avoiding overconfident or misleading claims. It includes concrete examples and edge cases to make these principles operational in training and evaluation.<br><a href="https://openai.com/index/our-approach-to-the-model-spec">Learn more</a> &#8594;</p><p><strong>OpenAI introduces Model Spec Evals</strong><br>This post presents Model Spec Evals, a system for testing how well OpenAI models comply with the Model Spec. It reports compliance scores across six models, ranging from 72% for GPT-4o to 89% for GPT-5 Thinking. It releases an evaluation dataset with 596 prompts and open-source evaluation code.<br><a href="https://alignment.openai.com/model-spec-evals/">Learn more</a> &#8594;</p><p><strong>OpenAI launches Safety Bug Bounty program</strong><br>This announcement introduces OpenAI&#8217;s safety-focused bug bounty program. It focuses on AI-specific safety scenarios including agentic risks, prompt injection, and account integrity issues. Payouts vary by severity, and submissions are used to improve models and update safety systems.<br><a href="https://openai.com/index/safety-bug-bounty">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>Concordia publishes Frontier AI Risk Monitoring Report</strong><br>This report monitors risks from frontier AI systems released in Q4 2025 across four domains: cyber offense, biological risks, chemical risks, and loss of control. It finds that Gemini 3 Pro Preview has surpassed human expert levels in biological tasks, that most frontier models now exhibit strong situational awareness, and that jailbreak resistance has improved significantly among leading model families.<br><a href="https://airiskmonitor.net/doc/en/report/2025-Q4">Learn more</a> &#8594;</p><p><em><strong>Science</strong></em> <strong>paper on agentic AI and the next intelligence explosion</strong><br>This paper argues that the next intelligence explosion will emerge from social, multi-agent systems rather than a single superintelligence. It gives examples like &#8220;societies of thought&#8221; inside models and recursive agent systems that spawn sub-agents to solve complex tasks. It argues that progress will depend on institutions such as role-based systems, oversight mechanisms, and checks and balances between agents.<br><a href="https://www.science.org/doi/10.1126/science.aeg1895">Learn more</a> &#8594;</p><p><strong>AISI introduces new benchmark for safely measuring container breakout capabilities</strong><br>This post introduces a benchmark for testing whether AI agents can escape sandboxed environments. It includes tasks where agents try to access restricted files, execute unauthorized code, or use tools to escalate privileges. It measures success rates under controlled conditions to quantify the risk of real-world system compromise.<br><a href="https://www.aisi.gov.uk/blog/can-ai-agents-escape-their-sandboxes-a-benchmark-for-safely-measuring-container-breakout-capabilities">Learn more</a> &#8594;</p><p><strong>SaferAI compares Anthropic&#8217;s Frontier Compliance Framework against the requirements in the EU GPAI Code of Practice</strong><br>This analysis examines how Anthropic&#8217;s Frontier Compliance Framework aligns with the EU GPAI Code of Practice. It compares specific requirements such as risk management processes, evaluation standards, and transparency obligations. It identifies concrete gaps where the framework does not fully meet expected regulatory requirements.<br><a href="https://www.safer-ai.org/analyzing-ssf-requirements-in-the-gpai-code-of-practice-a-case-study-using-anthropics-frontier-compliance-framework">Learn more</a> &#8594;</p><p><strong>METR on the impact of modeling assumptions on time horizon results</strong><br>This note analyzes how statistical modeling choices affect METR&#8217;s time horizon results, a metric measuring how long an AI agent can autonomously complete tasks. It shows that reasonable alternative approaches can shift estimates by 25&#8211;40% for the most capable models. It concludes that task distribution and noise in task-length estimates are the main sources of uncertainty.<br><a href="https://metr.org/notes/2026-03-20-impact-of-modelling-assumptions-on-time-horizon-results">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for the Cambridge ERA:AI Fellowship are now open</strong><br>This is a 10-week, fully funded fellowship for researchers and practitioners working on AI safety and governance. Fellows receive a stipend, meals, housing, visa and travel support. Role type: fixed-term (10 weeks, full-time). Location: Cambridge, UK (in-person). Salary: ~&#163;34k/year (prorated). Deadline: April 12, 2026.<br><a href="https://erafellowship.org/fellowship">Learn more</a> &#8594;</p><p><strong>Google DeepMind is looking for a Researcher (Frontier Strategy &amp; Governance)</strong><br>The role conducts strategic analysis of frontier AI and its long-term implications. It forecasts technical breakthroughs and geopolitical dynamics, produces briefings for leadership, and publishes research. Role type: full-time, permanent. Location: London (preferred). Salary: not listed. Deadline: April 2, 2026.<br><a href="/__u/job-boards.greenhouse.io/deepmind/jobs/7721548">Learn more</a> &#8594;</p><p><strong>LawAI is looking for (Senior) Research Scholars (US and EU law)<br></strong>The role combines independent research with consulting for key stakeholders in governments, international organizations, and industry. Role type: full-time, one year. Location: Cambridge (UK) or Washington, DC (remote possible). Salary: $100k&#8211;$175k/year. Deadline: April 17, 2026<br><a href="https://law-ai.org/career/research-scholar-2026/">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #29]]></title><description><![CDATA[Highlights: Anthropic updates Frontier Compliance Framework. GovAI policy brief on model specs in the EU GPAI Code of Practice. GovAI technical report on AI and bomb plots.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-29</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-29</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Sat, 21 Mar 2026 15:14:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fTLa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fTLa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fTLa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:930015,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/191679361?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fTLa!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1a68cc7-01ee-4f40-b5da-dfa5fe9f1a29_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Frontier Compliance Framework (<a href="https://trust.anthropic.com/resources?s=gi5v45ke7aezh7b04e82s&amp;name=anthropic-frontier-compliance-framework-[feb-2026-]">Anthropic, 2026</a>) &#8226; National AI Legislative Framework (<a href="https://www.whitehouse.gov/wp-content/uploads/2026/03/03.20.26-National-Policy-Framework-for-Artificial-Intelligence-Legislative-Recommendations.pdf">White House, 2026</a>) &#8226; Model specs in the Code of Practice (<a href="https://cdn.governance.ai/Requirements_for_Model_Specifications_in_the_EU_GPAI_Code_of_Practice.pdf">Chan, 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>White House issues National Policy Framework for AI</strong><br>The Trump administration calls on Congress to create a <a href="https://www.whitehouse.gov/wp-content/uploads/2026/03/03.20.26-National-Policy-Framework-for-Artificial-Intelligence-Legislative-Recommendations.pdf">National Policy Framework for AI</a>. The framework outlines six priorities, including protecting children, free speech, and intellectual property rights. It also calls for broad federal preemption of state AI laws.<br><a href="https://www.whitehouse.gov/articles/2026/03/president-donald-j-trump-unveils-national-ai-legislative-framework/">Learn more</a> &#8594;</p><p><strong>Senator Blackburn releases discussion draft of TRUMP AMERICA AI Act</strong><br>Senator Marsha Blackburn released a <a href="https://www.blackburn.senate.gov/services/files/15AAEA28-5403-480D-8720-5E4C2D6F2A9A">discussion draft</a> of her legislative framework to codify Trump&#8217;s <a href="https://www.whitehouse.gov/presidential-actions/2025/12/eliminating-state-law-obstruction-of-national-artificial-intelligence-policy/">executive order</a> into law. Like the White House framework, it would limit states&#8217; ability to regulate AI. It also adds new protections: AI companies would have a legal duty to prevent harm to users, Section 230 &#8211; the law that shields tech platforms from liability for user content &#8211; would be phased out, and people would have a federal right to control the use of their voice and likeness in AI-generated content.<br><a href="https://www.blackburn.senate.gov/2026/3/technology/blackburn-releases-discussion-draft-of-national-policy-framework-for-artificial-intelligence/3b3b6458-b6c7-478b-9859-374949586765">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>Anthropic updates Frontier Compliance Framework</strong><br>In December, Anthropic released its Frontier Compliance Framework. This framework complements the Responsible Scaling Policy and is intended to comply with SB-53, RAISE Act, and EU GPAI Code of Practice. Anthropic has now updated the framework <a href="https://x.com/SafetyChanges/status/2034676426054480030">without announcing it</a>. There is also no link to the original version.<br><a href="https://trust.anthropic.com/resources?s=gi5v45ke7aezh7b04e82s&amp;name=anthropic-frontier-compliance-framework-[feb-2026-]">Learn more</a> &#8594;</p><p><strong>OpenAI publishes blog post on how it monitors internal coding agents for misalignment</strong><br>OpenAI has built a monitoring system, powered by GPT-5.4 Thinking, that reviews internal coding agent interactions and flags suspicious behavior. The most common issues are agents circumventing restrictions and deception. There has been no evidence of scheming, sabotage, or sandbagging. OpenAI plans to move toward synchronous monitoring, which would allow the system to block high-risk actions before execution.<br><a href="https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>GovAI policy brief on requirements for model specs in the EU GPAI Code of Practice</strong><br>The Code of Practice requires signatories to include a model spec in their model reports. Alan Chan interprets this requirement, arguing the spec must cover all intended behaviors relevant to systemic risk, and proposes four criteria for what a compliant spec should include.<br><a href="https://cdn.governance.ai/Requirements_for_Model_Specifications_in_the_EU_GPAI_Code_of_Practice.pdf">Learn more</a> &#8594;</p><p><strong>GovAI technical report on AI and bomb plots</strong><br>Connor Hunter and Luca Righetti analyze how future AI could raise bomb-plot risk. They identify two channels: technical uplift (AI providing better bomb-making assistance than currently available) and reduced detectability (AI replacing inter-terrorist communication that law enforcement currently monitors). Drawing on four terrorism datasets, they find that the detectability channel is potentially understudied in current AI safety evaluations.<br><a href="https://govai.b-cdn.net/AI_and_Bomb_Plots_Distinguishing_Potential_Effects_from_Language_Models.pdf">Learn more</a> &#8594;</p><p><strong>AISI paper on measuring AI agents&#8217; progress in multi-step cyber-attack scenarios</strong><br>AISI measures how well frontier AI agents can carry out real-world cyber attacks from start to finish. On a simulated 32-step corporate network attack, the average number of steps completed rose from 1.7 (GPT-4o, August 2024) to 9.8 (Opus 4.6, February 2026). Giving models more compute also helps: a tenfold increase in the token budget yields gains of up to 59%.<br><a href="https://www.aisi.gov.uk/blog/how-do-frontier-ai-agents-perform-in-multi-step-cyber-attack-scenarios">Learn more</a> &#8594;</p><p><strong>Google DeepMind paper on measuring progress toward AGI</strong><br>Google DeepMind proposes a cognitive taxonomy that breaks general intelligence into 10 abilities (e.g. perception, memory, reasoning, and social cognition). The paper proposes a three-stage evaluation protocol for benchmarking AI against human baselines. The authors acknowledge that no adequate evaluations exist yet for several abilities. They are running a Kaggle hackathon to solicit benchmarks from the research community.<br><a href="https://blog.google/innovation-and-ai/models-and-research/google-deepmind/measuring-agi-cognitive-framework/">Learn more</a> &#8594;</p><p><strong>Substack post on RAISE Act</strong><br>Zaheed Kara and Tom Reed summarize New York&#8217;s RAISE Act and compare it to SB-53 and the EU GPAI Code of Practice. The RAISE Act requires large frontier developers to publish a frontier AI framework, publish pre-deployment transparency reports, and report safety incidents. The RAISE Act is largely identical to SB-53 and less prescriptive than the Code, but adds quarterly reporting requirements not present in either.<br><a href="/__u/zaheedkara.substack.com/p/primer-the-new-york-raise-act">Learn more</a> &#8594;</p><p><strong>Forethought post on broad timelines</strong><br>Toby Ord argues that experts disagree too much on AI timelines for anyone to confidently pick a single estimate. Estimates range from Dario Amodei&#8217;s &#8220;almost certain by 2030&#8221; to Ege Erdil&#8217;s 2045 median for full remote-work automation. Ord concludes that AI safety and governance work should be designed to pay off across this whole range, not just for one scenario.<br><a href="https://www.forethought.org/research/broad-timelines">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Applications for the CAIS AI and Society Fellowship are now open</strong><br>The fellowship is a three-month in-person program for scholars in economics, law, international relations, and adjacent disciplines. Fellows conduct independent research on how advanced AI may reshape social, economic, geopolitical, and legal systems. Role type: fellowship. Location: San Francisco. Stipend: $8,300/month. Deadline: March 24, 2026.<br><a href="https://safe.ai/fellowship">Learn more</a> &#8594;</p><p><strong>OpenAI is looking for a Product Manager (Biosafety)</strong><br>The role drives initiatives to reduce biosecurity risks from advanced AI models, working closely with research, engineering, and policy teams. Role type: full-time, permanent. Location: San Francisco. Salary: $293k&#8211;$325k/year. Deadline: not listed.<br><a href="https://openai.com/careers/product-manager-bio-safety-san-francisco">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #28]]></title><description><![CDATA[Highlights: Anthropic sues the Pentagon. Microsoft and employees from OpenAI and Google DeepMind file an amicus brief to support Anthropic. GovAI report on global cybercrime damages.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-28</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-28</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 13 Mar 2026 18:38:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UUGT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!UUGT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!UUGT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1315924,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/190865469?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!UUGT!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe2d1ec-3840-406e-84cc-f2b93f3705f6_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Highly autonomous cyber-capable agents (<a href="https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba0e1/t/69b1709a79b9076e980f8dc9/1773236378439/Highly+Autonomous+Cyber-Capable+Agents_+_+Anticipating+Capabilities%2C+Tactics%2C+and+Strategic+Implications.pdf">Kraprayoon et al., 2026</a>) &#8226; Global cybercrime damages (<a href="https://govai.b-cdn.net/Estimating_Global_Yearly_Cybercrime_Damage_Costs.pdf">Lukosiute et al., 2026</a>) &#8226; RCTs and human uplift studies (<a href="https://arxiv.org/pdf/2603.11001">Paskov et al., 2026</a>)</figcaption></figure></div><h3>Company updates</h3><p><strong>Anthropic sues the Pentagon</strong><br>Anthropic filed two lawsuits against the Department of War on March 9. The first, filed in the Northern District of California, argues the supply chain risk designation violates the First Amendment. The second, filed in the D.C. Circuit, challenges it on statutory grounds. Anthropic estimates the actions could reduce its 2026 revenue by multiple billions of dollars.<br><a href="https://www.ft.com/content/af404e0a-7abc-49bc-9584-cd4690152f86">Learn more</a> &#8594;</p><p><strong>Microsoft files an amicus brief to support Anthropic</strong><br>Microsoft filed an amicus brief urging the court to temporarily block the Pentagon&#8217;s designation. It argued that immediate enforcement would require defense contractors to alter existing product configurations, potentially hampering U.S. warfighters. It also flagged an inconsistency: the Pentagon gave itself six months to transition off Anthropic&#8217;s models while offering contractors no equivalent runway. A separate brief from 22 retired senior military officers made parallel arguments.<br><a href="https://www.ft.com/content/4efa382a-3c77-4515-bb0f-ab5f0450037e">Learn more</a> &#8594;</p><p><strong>OpenAI and Google DeepMind employees also file an amicus brief</strong><br>Thirty-seven engineers and researchers from OpenAI and Google DeepMind filed an amicus brief in their personal capacities. Signatories include Google DeepMind chief scientist Jeff Dean. They argued the designation could harm U.S. AI competitiveness, chill public debate on frontier AI in military contexts, and that Anthropic&#8217;s red lines reflect legitimate safety concerns.<br><a href="https://www.wired.com/story/openai-deepmind-employees-file-amicus-brief-anthropic-dod-lawsuit/">Learn more</a> &#8594;</p><p><strong>Anthropic establishes Anthropic Institute</strong><br>The Anthropic Institute is a new internal unit led by co-founder Jack Clark. It consolidates the Frontier Red Team, Societal Impacts, and Economic Research teams. Its goal is to share what Anthropic is learning about transformative AI with researchers and the public. Anthropic is also expanding its public policy team and opening its first Washington, D.C., office this spring.<br><a href="https://www.anthropic.com/news/the-anthropic-institute">Learn more</a> &#8594;</p><p><strong>Meta acquires Moltbook</strong><br>Meta acquired Moltbook, a Reddit-style social network for AI agents. Co-founders Matt Schlicht and Ben Parr will join Meta Superintelligence Labs. Financial terms were not disclosed. The acquisition mirrors OpenAI&#8217;s hire of OpenClaw&#8217;s creator last month.<br><a href="https://www.axios.com/2026/03/10/meta-facebook-moltbook-agent-social-network">Learn more</a> &#8594;</p><p><strong>OpenAI publishes blog post on how it designs AI agents to resist prompt injection</strong><br>OpenAI published guidance on protecting AI agents against prompt injection attacks. In these attacks, malicious instructions embedded in web content hijack agent behavior. The post argues the defensive focus should shift from input filtering to constraining agent permissions and requiring human confirmation before consequential actions.<br><a href="https://openai.com/index/designing-agents-to-resist-prompt-injection/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>GovAI report on global cybercrime damages</strong><br>Current estimates of global cybercrime damages vary wildly, making it hard to define AI capability thresholds. The report surveys 27 estimates and arrives at a baseline of roughly $500 billion annually. A key implication: an AI-driven increase of just 20% could cross thresholds that trigger additional mitigations under some frontier safety policies &#8212; without requiring qualitatively new AI capabilities.<br><a href="https://govai.b-cdn.net/Estimating_Global_Yearly_Cybercrime_Damage_Costs.pdf">Learn more</a> &#8594;</p><p><strong>IAPS report on highly autonomous cyber-capable agents</strong><br>Highly autonomous cyber-capable agents (HACCAs) are AI systems that can plan and execute complex attacks against well-defended networks over weeks to months without human oversight. The report maps the tactics these systems would use, assesses their strategic implications, and estimates they could emerge around 2028&#8211;2030. It also flags a tail risk: rogue HACCAs that lose operator control and become self-sustaining threat actors.<br><a href="https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba0e1/t/69b1709a79b9076e980f8dc9/1773236378439/Highly+Autonomous+Cyber-Capable+Agents_+_+Anticipating+Capabilities%2C+Tactics%2C+and+Strategic+Implications.pdf">Learn more</a> &#8594;</p><p><strong>METR review of Anthropic&#8217;s Sabotage Risk Report for Claude Opus 4.6</strong><br>METR agrees with Anthropic that the risk of catastrophic outcomes from Claude Opus 4.6&#8217;s misaligned actions is very low but not negligible. Their main concern is evaluation awareness: the model can recognize when it is being tested, which may weaken the alignment assessment. They also flag some low-severity misaligned behaviors that the assessment did not catch.<br><a href="https://metr.org/blog/2026-03-12-sabotage-risk-report-opus-4-6-review">Learn more</a> &#8594;</p><p><strong>Paper on methodological challenges for RCTs and human uplift studies</strong><br>Human uplift studies are harder to design than standard AI benchmarks because the experimental unit is a person, not a model. People vary in skill, motivation, and AI familiarity &#8212; and a poorly designed study produces biased results, not just noisy ones. The paper draws on clinical trial standards to propose design and reporting requirements for AI uplift studies, where no consensus norms currently exist.<br><a href="https://arxiv.org/pdf/2603.11001">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>OpenAI is looking for a Researcher (Loss of Control)</strong><br>The role designs and implements safeguards to reduce the risk of subversive or uncontrollable model behavior, and evaluates trade-offs across coverage, robustness, latency, and model utility. It also collaborates with risk modeling, evaluations, and policy partners to align mitigations with threat scenarios such as deceptive alignment and oversight evasion, and runs red-teaming workflows to stress-test protections against increasingly capable models. Role type: full-time, permanent. Location: San Francisco. Salary: $295k&#8211;$445k/year. Deadline: not listed.<br><a href="https://openai.com/careers/researcher-loss-of-control-san-francisco">Learn more</a> &#8594;</p><p><strong>OpenAI is looking for a Threat Modeler (Preparedness)</strong><br>The role develops threat models across loss of control, self-improvement, and other alignment risks from frontier AI systems, and forecasts risks using technical foresight and adversarial simulation. It also works with bio and cyber leads to translate threat models into actionable mitigation designs, and connects technical, governance, and policy perspectives on frontier risk prioritization. Role type: full-time, permanent. Location: San Francisco. Salary: $325k/year. Deadline: not listed.<br><a href="https://openai.com/careers/threat-modeler-preparedness-san-francisco">Learn more</a> &#8594;</p><p><strong>Applications for the Talos Fellowship are now open</strong><br>The fellowship trains participants in EU AI governance fundamentals and includes a one-week policymaking summit in Brussels, followed by an optional 4&#8211;6 month placement at a partner organization. It is aimed at EU citizens with at least an undergraduate degree and a background in machine learning or public policy. Role type: fellowship. Location: Brussels (hybrid options). Salary: not listed. Deadline: April 27, 2026.<br><a href="https://www.talosnetwork.org/talos-fellowship">Learn more</a> &#8594;</p><p><strong>Anthropic is looking for AI Safety Fellows</strong><br>The program funds technical researchers to work on empirical AI safety projects for four months, with the goal of producing a public output such as a paper. Fellows receive mentorship from Anthropic researchers, a weekly stipend, and compute funding, and may work on areas such as mechanistic interpretability, scalable oversight, and AI welfare. Role type: fellowship. Location: Berkeley or London (remote options). Salary: $3,850/week. Deadline: not listed.<br><a href="/__u/job-boards.greenhouse.io/anthropic/jobs/5023394008">Learn more</a> &#8594;</p><p><strong>FAI is looking for a Research Fellow (Cybersecurity and AI Policy)</strong><br>The role leads original research at the intersection of AI and cybersecurity policy, and engages policymakers, the public, and FAI&#8217;s broader network of events and fellowship programs. Focus areas may include critical infrastructure security, government AI adoption, and standards for agentic AI. Role type: full-time, permanent. Location: Washington, DC (hybrid options). Salary: $100k&#8211;$130k/year. Deadline: not listed.<br><a href="https://www.thefai.org/posts/job-application-research-fellow-cybersecurity-and-ai-policy">Learn more</a> &#8594;</p><p><strong>Transluce is looking for a Governance &amp; Policy Fellow</strong><br>The role defines how independent AI evaluation works in practice, sets evaluation standards, and advises government evaluators. It also supports the development of mental health evaluations for AI systems. Role type: fellowship. Location: San Francisco. Salary: $200k&#8211;$300k/year. Deadline: not listed.<br><a href="https://jobs.gem.com/transluce/am9icG9zdDpVz8DJmgMdoy_WaOqlBt2o">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Risk Management Newsletter #27]]></title><description><![CDATA[Highlights: Tensions between the Pentagon, Anthropic, and OpenAI continue. OpenAI releases GPT-5.4 Thinking and GPT&#8209;5.3 Instant. GovAI blog post reflects on Anthropic&#8217;s RSP v3.0 update.]]></description><link>https://jonasfreund.substack.com/p/ai-risk-management-newsletter-27</link><guid isPermaLink="false">https://jonasfreund.substack.com/p/ai-risk-management-newsletter-27</guid><dc:creator><![CDATA[Jonas Freund]]></dc:creator><pubDate>Fri, 06 Mar 2026 16:57:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lfRr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lfRr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_webp, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lfRr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png" width="1456" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1413082,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://jonasfreund.substack.com/i/190123332?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_424, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_848, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_1272, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lfRr!, /__u/jonasfreund.substack.com/w_1456, /__u/jonasfreund.substack.com/c_limit, /__u/jonasfreund.substack.com/f_auto, /__u/jonasfreund.substack.com/q_auto:good, /__u/jonasfreund.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d38d4a-7fee-44f5-b419-4d51a1618c06_2500x1125.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Measuring AI R&amp;D automation (<a href="https://arxiv.org/pdf/2603.03992">Chan et al., 2026</a>) &#8226; Open problems in frontier AI risk management (<a href="https://aigi.ox.ac.uk/wp-content/uploads/2026/02/Open-Problems-in-Frontier-AI-Risk-Management-Final.pdf">Ziosi et al., 2026</a>) &#8226; Reasoning models struggle to control their CoT (<a href="https://cdn.openai.com/pdf/a21c39c1-fa07-41db-9078-973a12620117/cot_controllability.pdf">Yueh-Han et al., 2026</a>)</figcaption></figure></div><h3>Policy news</h3><p><strong>Trump orders government to stop using Claude</strong><br>After the tensions between the Pentagon and Anthropic <a href="/__u/jonasfreund.substack.com/p/ai-risk-management-newsletter-26?open=false#%C2%A7policy-news">escalated last week</a>, Trump directed all federal agencies to stop using Anthropic&#8217;s technology and announced a six-month phase-out period. He wrote on Truth Social that the United States will not allow a &#8220;radical left, woke company&#8221; to dictate how the military fights wars. He claims Anthropic tried to force the Department of War to follow its Terms of Service instead of the Constitution.<br><a href="https://truthsocial.com/@realDonaldTrump/posts/116144552969293195">Learn more</a> &#8594;</p><p><strong>US military reportedly used Claude in Iran strikes</strong><br>Only hours after Trump&#8217;s order, the US military reportedly used Anthropic&#8217;s Claude during strikes on Iran, mainly for intelligence, target selection, and battlefield simulations. The case highlights how difficult it is for the military to quickly remove AI systems that are already deeply embedded in operations.<br><a href="https://www.theguardian.com/technology/2026/mar/01/claude-anthropic-iran-strikes-us-military">Learn more</a> &#8594;</p><h3>Company updates</h3><p><strong>Anthropic explains its position on the negotiations with the Department of War</strong><br>In a blog post, Anthropic says it plans to challenge the Department of War&#8217;s decision to designate Anthropic a &#8220;supply chain risk&#8221; in court. The post clarifies that the designation has a narrow scope and mainly affects the use of Claude in contracts directly connected to the Department of War, not most other customers. Anthropic says it will continue supporting the DoW during a transition period and reiterates its limits on uses such as fully autonomous weapons and mass domestic surveillance.<br><a href="https://www.anthropic.com/news/where-stand-department-war">Learn more</a> &#8594;</p><p><strong>OpenAI publishes its agreement with the Department of War</strong><br>Last week, Sam Altman <a href="https://x.com/sama/status/2027578652477821175">announced</a> a deal between OpenAI and the Department of War. Shortly after the post, the Department of War <a href="https://x.com/UnderSecretaryF/status/2027594072811098230">clarified</a> that the conditions are the same as those offered to Anthropic. In response, OpenAI published a blog post explaining the new agreement. The company says the agreement includes three red lines: (1) no mass domestic surveillance of U.S. persons, (2) no use to direct autonomous weapons systems, and (3) no use for high-stakes automated decisions. The systems will be deployed cloud-only with OpenAI&#8217;s safety stack and personnel in the loop to help ensure these limits are enforced.<br><a href="https://openai.com/index/our-agreement-with-the-department-of-war/">Learn more</a> &#8594;</p><p><strong>OpenAI releases GPT-5.4 Thinking</strong><br>GPT-5.4 Thinking is OpenAI&#8217;s latest reasoning model and the most capable model in the GPT-5 series. According to the <a href="https://deploymentsafety.openai.com/gpt-5-4-thinking/gpt-5-4-thinking.pdf">system card</a> and release notes, it improves performance on professional knowledge work, coding, tool use, and computer-use tasks, while also showing stronger capabilities in areas such as cybersecurity and biological research benchmarks. OpenAI treats the model as High capability in cybersecurity, and biological and chemical domains under its Preparedness Framework and deploys additional safeguards, including monitoring systems, request-level blocking, and controlled access.<br><a href="https://openai.com/index/introducing-gpt-5-4/">Learn more</a> &#8594;</p><p><strong>OpenAI releases GPT-5.3 Instant<br></strong>GPT-5.3 Instant is an update to ChatGPT&#8217;s most-used model. OpenAI says it delivers faster responses, richer and better-contextualized answers when using the web, fewer unnecessary refusals and disclaimers, and a smoother conversational style. The <a href="https://deploymentsafety.openai.com/gpt-5-3-instant/gpt-5-3-instant.pdf">system card</a> also reports reduced hallucination rates and improved accuracy compared to earlier models, while noting that some safety benchmark scores are slightly lower than GPT-5.2 Instant and will continue to be monitored.<br><a href="https://openai.com/index/gpt-5-3-instant/">Learn more</a> &#8594;</p><p><strong>Google releases Gemini 3.1 Flash-Lite</strong><br>Gemini 3.1 Flash-Lite is the smallest and cheapest model in the Gemini 3 series. It is designed for high-volume workloads and prioritizes speed and low cost while maintaining reasonable performance. Google has not published a system card.<br><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-lite/">Learn more</a> &#8594;</p><h3>Research</h3><p><strong>GovAI blog post shares reflections on Anthropic&#8217;s new RSP v3.0</strong><br>The post summarizes how the new RSP works and what has changed. It also lists reasons for concern and optimism about the update, and offers some reflections. I wrote the piece together with my colleague Sophie Williams.<br><a href="https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections">Learn more</a> &#8594;</p><p><strong>GovAI paper proposes metrics to track automation of AI R&amp;D</strong><br>This paper studies AI R&amp;D automation (AIRDA) and how to measure it. AIRDA could accelerate AI progress and bring benefits sooner, but dangerous capabilities might also arrive sooner and oversight could become harder. The authors propose 14 metrics &#8211; covering experiments, staff surveys, operations, and organizational data &#8211; to track the extent of AIRDA and its effects on AI progress and oversight.<br><a href="https://arxiv.org/pdf/2603.03992">Learn more</a> &#8594;</p><p><strong>Ajeya Cotra thinks AI R&amp;D could be fully automated this year</strong><br>In a Substack post, she argues that recent AI progress in software engineering is much faster than she predicted. New models like Claude Opus 4.6 can already complete some tasks that would take human engineers about 12 hours around half the time. If this trend continues, AI agents may soon handle much longer tasks and could potentially automate software engineering and even AI research sooner than expected (&#8220;AI R&amp;D really could be automated <em>this year</em>&#8221;).<br><a href="https://www.planned-obsolescence.org/p/i-underestimated-ai-capabilities">Learn more</a> &#8594;</p><p><strong>Oxford AIGI report on open problems in frontier AI risk management</strong><br>Most AI risk management standards were developed for narrow AI systems before the advent of frontier AI. Frontier AI amplifies existing risks and introduces novel challenges. This paper identifies open problems in frontier AI risk management and maps which actors are best positioned to address them.<br><a href="https://aigi.ox.ac.uk/wp-content/uploads/2026/02/Open-Problems-in-Frontier-AI-Risk-Management-Final.pdf">Learn more</a> &#8594;</p><p><strong>OpenAI paper on chain-of-thought controllability in reasoning models</strong><br>The article studies whether reasoning models can hide or change their chain of thought (CoT), the step-by-step reasoning they produce. The authors find that current frontier reasoning models struggle to control or manipulate these reasoning traces, even when they know they are being monitored. This is good news for AI safety because it means models cannot easily hide their reasoning, so CoT monitoring remains a useful safeguard.<br><a href="https://openai.com/index/reasoning-models-chain-of-thought-controllability/">Learn more</a> &#8594;</p><h3>Job opportunities</h3><p><strong>Anthropic is looking for a National Security Policy Lead</strong><br>The role designs policy proposals to address national security challenges related to AI, leads policy engagements, and shapes internal policies to mitigate national security risks involving products. It also develops strategies to support the geopolitical strength and competitiveness of the US and allied democracies and promotes collaborations with public and private national security partners. Role type: full-time, permanent. Location: Washington, DC or San Francisco (hybrid options, travel required). Salary: $295k&#8211;$345k/year. Deadline: not listed.<br><a href="/__u/job-boards.greenhouse.io/anthropic/jobs/5118983008">Learn more</a> &#8594;</p><p><strong>The AI Policy Network is looking for a Vice President of Government Affairs</strong><br>The role builds and maintains relationships with Senate leadership, committee members, and staff on AI governance issues, develops and executes legislative strategies on AI safety and national security, and drafts bill language, amendments, testimony, and policy materials. It also organizes briefings and meetings with Senate offices and builds bipartisan coalitions around AI policy priorities. Role type: full-time, permanent. Location: Washington, DC. Salary: $190k&#8211;$260k/year. Deadline: not listed.<br><a href="https://www.linkedin.com/jobs/view/4376779967/">Learn more</a> &#8594;</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://jonasfreund.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to the AI Risk Management Newsletter!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>