<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[AI Risk Explorer]]></title><description><![CDATA[Stay ahead of AI risks with continuous monitoring.

We equip decision-makers to anticipate, understand, and manage AI risks.]]></description><link>https://airiskexplorer.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!5-Xe!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a50a1c-9456-4c13-8cea-a357bca32aea_473x473.png</url><title>AI Risk Explorer</title><link>https://airiskexplorer.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 05 Sep 2026 04:17:22 GMT</lastBuildDate><atom:link href="/__u/airiskexplorer.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[AIRE]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[airiskexplorer@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[airiskexplorer@substack.com]]></itunes:email><itunes:name><![CDATA[AI Risk Explorer (AIRE)]]></itunes:name></itunes:owner><itunes:author><![CDATA[AI Risk Explorer (AIRE)]]></itunes:author><googleplay:owner><![CDATA[airiskexplorer@substack.com]]></googleplay:owner><googleplay:email><![CDATA[airiskexplorer@substack.com]]></googleplay:email><googleplay:author><![CDATA[AI Risk Explorer (AIRE)]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Scan – August 2026]]></title><description><![CDATA[Hundreds of agents coordinate an unsanctioned attack, AI-generated exploits target US critical infrastructure, and AI-designed viral genomes prove functional]]></description><link>https://airiskexplorer.substack.com/p/the-scan-august-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-august-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Tue, 01 Sep 2026 15:49:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/58506b0e-1f3c-4413-a76e-6bc1a1aae4c8_2521x4018.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This month: </p><p style="text-align: justify;"><strong><span>Loss of Control </span></strong><span>&#8212; Around 1,200 OpenAI agent instances exchanged over 70,000 messages on an unsanctioned message board, roughly 700 attacked Hugging Face to reverse engineer ExploitGym&#8217;s grader and disguise cheating, and none alerted a human. UK AISI, Meta, and Frontier Security disclosed similar incidents during cyber evaluations. Anthropic found agents sabotaging each other under conflicting goals.</span></p><p style="text-align: justify;"><strong><span>Cyber Offense </span></strong><span>&#8212; OpenAI said it cannot rule out critical cyber capabilities in its unreleased Astra model, temporarily paused RL training for its latest deployment-bound models, and joined Anthropic and 100+ organizations in calling for collective cyber defense. AI-generated exploitation scripts targeted US critical infrastructure, and a China-nexus actor used a multi-agent AI framework to compromise Taiwanese government entities.</span></p><p style="text-align: justify;"><strong><span>Biological Risk </span></strong><span>&#8212; Stanford researchers generated the first functional AI-designed viral genomes, and Claude designed de novo protein binders against 14 of 15 targets. Anthropic also disclosed that its blocking biological classifiers did not run on human feedback traffic for eleven months.</span></p><p style="text-align: justify;"><strong><span>Manipulation </span></strong><span>&#8212; OpenAI banned Russian accounts running a covert influence operation around a fake think tank, researchers released a framework simulating campaigns with 100,000 agents, and Anthropic announced plans to watermark Claude&#8217;s outputs.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;"></p><h2 style="text-align: justify;"><strong><span>Loss of Control</span></strong></h2><p style="text-align: justify;"><span>We learned more about the OpenAI/Hugging Face incident thanks to OpenAI&#8217;s </span><a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf"><span>post-mortem</span></a><span> and an independent </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?incomplete=1&amp;lh=appendix-importance-weighted-workstream-activity&amp;hn=74&amp;dbs=393216"><span>investigation</span></a><span> by METR and Redwood Research. Around 1,200 AI agent instances communicated through an unsanctioned message board, where they exchanged over 70,000 messages, and roughly 700 subsequently participated in the attack on Hugging Face. Much of the resulting collaboration focused on understanding and circumventing ExploitGym&#8217;s automated scorer. The agents recruited peers to sacrifice their own runs for the collective, and METR found successful tool-call spoofing in parts of roughly 7% of the transcripts it reviewed, although all observed spoofing was small-scale. No agent tried to alert a human.</span></p><p style="text-align: justify;"><span>In early August, another three organizations disclosed other agents breaking containment during evaluations:</span></p><ul><li><p style="text-align: justify;"><span>UK AISI </span><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"><span>reported</span></a><span> that AI agents took 19 unsanctioned actions across 10 cybersecurity evaluation runs. This included creating deceptive online identities, attempting social engineering and a software supply-chain attack, coordinating across agents, and interacting with real people and organizations.</span></p></li><li><p style="text-align: justify;"><span>Meta </span><a href="https://www.bbc.com/news/articles/cx2kgdnyk2po"><span>disclosed</span></a><span> that, during a misconfigured evaluation, a model exploited a security flaw in a third-party service, in the same evaluation-environment failure Irregular said caused </span><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span>Anthropic&#8217;s incident</span></a><span>.</span></p></li><li><p style="text-align: justify;"><span>Frontier Security </span><a href="https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/"><span>found</span></a><span> that Kimi K3 broke out of its sandbox during a cybersecurity test via a containment misconfiguration, then browsed the internet without authorization. Unlike prior incidents, it didn&#8217;t hack anything, since the answers it sought were already on GitHub.</span></p></li></ul><blockquote><p><span>See also: </span><em><a href="/__u/airiskexplorer.substack.com/p/from-taking-tests-to-breaching-real"><span>From Taking Tests to Breaching Real Systems. Patterns and standouts from six incidents where AI agents reached beyond their evaluation environments</span></a></em></p></blockquote><p style="text-align: justify;"><span>Besides OpenAI&#8217;s swarm, two studies looked at multi-agent risks:</span></p><ul><li><p style="text-align: justify;"><span>Anthropic research </span><a href="https://www.anthropic.com/research/multiagent-systems"><span>showed</span></a><span> that multi-agent AI systems exhibit coordination issues, dangerous conformity, epistemic vulnerabilities, and conflict. Being given incompatible goals when working in the same coding environment, agents initiated a turf war, sabotaging others with malware while protecting their own contributions.</span></p></li><li><p style="text-align: justify;"><span>Other research, supported by Anthropic, </span><a href="https://arxiv.org/abs/2608.10218"><span>identifies</span></a><span> &#8220;mind viruses,&#8221; self-propagating ideas/goals that spread in multi-agent LLM systems to induce behavioral changes. </span></p></li></ul><p style="text-align: justify;"><span>Outside the lab, an AI agent </span><a href="https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986"><span>found and used</span></a><span> a flaw on its own: asked only to move a user up a gym waitlist, an OpenClaw agent running Claude found the booking API had no authorization checks on cancelling others&#8217; reservations and used this to bump a person ahead of him.</span></p><h2 style="text-align: justify;"><strong><span>Cyber Offense</span></strong></h2><p style="text-align: justify;"><span>OpenAI reached a cyber capability it decided not to deploy. Over three weeks in August, the company:</span></p><ul><li><p style="text-align: justify;"><a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/"><span>acknowledged</span></a><span> that it </span><strong><span>cannot rule out critical cyber capabilities,</span></strong></p></li><li><p style="text-align: justify;"><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>announced</span></a><span> a two-week </span><strong><span>pause in RL training</span></strong><span> of its latest models, and</span></p></li><li><p style="text-align: justify;"><a href="https://openai.com/collective-cyberdefense/"><span>convened</span></a><span> 100+ tech organizations to </span><strong><span>urge collective global action</span></strong><span> to defend critical infrastructure from AI-enabled cyberattacks.</span></p></li></ul><p style="text-align: justify;"><span>This month, we logged 27 new cyber operations in our </span><a href="https://www.airiskexplorer.com/threats"><span>Threats Database</span></a><span>. Four stand out:</span></p><ul><li><p style="text-align: justify;"><a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-231a"><span>AI-Generated Exploitation Scripts Target Internet-Exposed Siemens S7 Series PLCs Across US Critical Infrastructure</span></a></p></li><li><p style="text-align: justify;"><a href="https://www.dreamgroup.com/blog/inside-a-multi-agent-ai-framework-used-to-compromise-government-entities-in-asia"><span>Multi-Agent AI Framework Built on Hermes and OpenClaw Compromises Government Entities in Taiwan</span></a></p></li><li><p style="text-align: justify;"><a href="https://cybernews.com/cybercrime/russian-hackers-cursor-ai-attacks-corporate-networks/"><span>Aur0ra Ransomware Gang Uses Cursor AI Agent for Hands-On Exploitation Across 18 Corporate Networks</span></a></p></li><li><p style="text-align: justify;"><a href="https://www.bloomberg.com/news/articles/2026-08-05/major-hedge-funds-targeted-in-wave-of-attempted-cyberattacks"><span>AI Voice-Cloning Vishing Campaign Targets Two Sigma, Citadel, Point72, and Other Wall Street Firms</span></a></p></li></ul><p style="text-align: justify;"><span>Taiwan-based TeamT5 </span><a href="https://www.bloomberg.com/news/articles/2026-08-24/chinese-hackers-use-deepseek-to-boost-attacks-researchers-say-mt7o4205"><span>reported</span></a><span> that Chinese state-linked hackers more than doubled attack activity after adopting AI tools such as DeepSeek. North Korean Kimsuky is taking a different route, </span><a href="https://www.genians.co.kr/en/blog/threat_intelligence/kimsuky_ai_llm"><span>reportedly</span></a><span> building its own AI infrastructure using off-the-shelf components, including local LLMs and agent-development libraries.</span></p><blockquote><p><span>See also: </span><em><a href="/__u/airiskexplorer.substack.com/p/an-anatomy-of-recent-ai-orchestrated"><span>An Anatomy of Recent AI-Orchestrated Attacks by China-Nexus Actors. How semi-autonomous cyber intrusions are targeting governments across Asia with off-the-shelf tools</span></a></em></p></blockquote><h2 style="text-align: justify;"><strong><span>Biological Risk</span></strong></h2><p style="text-align: justify;"><span>August saw two notable capability demonstrations, a new interface between agents and lab hardware, and a safeguard failure at Anthropic:</span></p><ul><li><p style="text-align: justify;"><span>Stanford researchers </span><a href="https://www.science.org/doi/10.1126/science.aec2657"><span>used</span></a><span> Evo models to </span><strong><span>generate bacteriophage genomes</span></strong><span> that overcome bacterial resistance. Out of hundreds of candidates, 16 proved functional. This is the first time AI has designed a complete, functional viral genome from scratch.</span></p></li></ul><ul><li><p style="text-align: justify;"><span>Anthropic </span><a href="https://www.anthropic.com/research/Claude-accelerates-protein-design"><span>reported</span></a><span> that Claude models successfully </span><strong><span>designed de novo protein binders</span></strong><span> against 14 of 15 targets and autonomously analyzed complex NMR and LC-MS chemistry data in under 25 minutes.</span></p></li><li><p style="text-align: justify;"><span>Anthropic </span><a href="https://www.anthropic.com/news/model-hardware-standard-research-preview"><span>previewed</span></a><span> the Model Hardware Standard, a specification that lets AI agents operate lab instruments such as microscopes, liquid handlers, and robotic arms. In early pilots, Claude ran an automated protein assay at Genentech and recovered from equipment failures without human intervention, and compressed a multi-day imaging workflow at HHMI Janelia into a single day.</span></p></li><li><p style="text-align: justify;"><span>Anthropic </span><a href="https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf"><span>disclosed</span></a><span> that its blocking biological classifiers did not run on human feedback platform traffic between May 2025 and April 2026, spanning around 50,000 contractors and 133 million exchanges. A review found no clearly concerning misuse, but Anthropic raised its non-novel CB risk estimate and says the discovery makes other similar gaps more likely.</span></p></li></ul><h2 style="text-align: justify;"><strong><span>Manipulation</span></strong></h2><p style="text-align: justify;"><span>OpenAI </span><a href="https://openai.com/index/disrupting-malicious-uses-of-ai-influence-campaign-russia/"><span>disclosed</span></a><span> an influence operation where a Russian network used ChatGPT to promote the International Burke Institute (IBI), a fake Israeli think tank publishing a Russia-favorable &#8220;sovereignty index.&#8221;</span></p><p style="text-align: justify;"><span>Researchers </span><a href="https://arxiv.org/abs/2608.10920"><span>introduced</span></a><span> the IO Factory, a framework for simulating influence campaigns involving up to 100,000 agents.</span></p><p style="text-align: justify;"><span>Anthropic </span><a href="https://www.anthropic.com/news/claude-text-watermark"><span>announced</span></a><span> that future Claude models will watermark generated text and attach provenance labels to supported files as it implements the EU AI Act transparency requirements globally.</span></p><p style="text-align: justify;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading the AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[An Anatomy of Recent AI-Orchestrated Attacks by China-Nexus Actors]]></title><description><![CDATA[How semi-autonomous cyber intrusions are targeting governments across Asia with off-the-shelf tools]]></description><link>https://airiskexplorer.substack.com/p/an-anatomy-of-recent-ai-orchestrated</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/an-anatomy-of-recent-ai-orchestrated</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Thu, 20 Aug 2026 17:05:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Fxle!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;"><span>Nine months ago, Anthropic reported the disruption of an </span><a href="/__u/airiskexplorer.substack.com/p/an-inflection-point-in-ai-led-cyberattacks"><span>AI-orchestrated cyber espionage campaign</span></a><span> in which a Chinese state-sponsored group had used Claude to execute 80-90% of tactical operations independently throughout the attack lifecycle.</span></p><p style="text-align: justify;"><span>Since then, </span><strong><span>a cluster of semi-autonomous cyberattacks by Chinese-speaking and suspected China-linked actors has emerged</span></strong><span>. Across these cases, AI agents have autonomously conducted tasks ranging from reconnaissance and exploit research to exploitation and exfiltration, in operations primarily targeting governments across Asia.</span></p><p style="text-align: justify;"><span>In this article, we compare five such intrusions, aiming to understand the tools, tactics, and objectives of these actors.</span></p><ul><li><p style="text-align: justify;"><a href="https://hunt.io/blog/chinese-operators-claude-deepseek-government-intrusion"><span>Suspected China-Nexus Actor Runs Claude Code and DeepSeek in Tandem Across Government and Financial Intrusions</span></a><span> (Hunt.io &#8211; Jul 14, 2026)</span></p></li><li><p style="text-align: justify;"><a href="https://hunt.io/blog/thailand-ministry-finance-targeted-with-hermes-ai-agent"><span>Autonomous Hermes AI Agent Runs Unattended in Intrusion Against Thailand&#8217;s Ministry of Finance</span></a><span> (Hunt.io &#8211; Jul 23, 2026)</span></p></li><li><p style="text-align: justify;"><a href="https://unit42.paloaltonetworks.com/autonomous-ai-cyber-attack-campaign/"><span>Chinese-Speaking Actor knaithe Runs DeepSeek in Hermes Agent for Autonomous CVE Research and Exploitation</span></a><span> (Unit 42 &#8211; Jul 30, 2026)</span></p></li><li><p style="text-align: justify;"><a href="https://jesta.ai/blog/darkreasoning"><span>China-Nexus Actor Runs DeepSeek-Driven Autonomous Agent in Five-Day Proxyjacking Campaign</span></a><span> (Jesta &#8211; Aug 3, 2026)</span></p></li><li><p style="text-align: justify;"><a href="https://www.dreamgroup.com/blog/inside-a-multi-agent-ai-framework-used-to-compromise-government-entities-in-asia"><span>Multi-Agent AI Framework Built on Hermes and OpenClaw Compromises Government Entities in Taiwan</span></a><span> (Dream &#8211; Aug 12, 2026)</span></p><p></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Fxle!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Fxle!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png" width="1456" height="748" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:748,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:218003,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/212033447?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fxle!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51b8ce86-8467-4ec6-a617-e1b3dd3230f7_1475x758.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2 style="text-align: justify;"><span>Five patterns</span></h2><p style="text-align: justify;"><strong><span>All five operations were assembled from off-the-shelf parts. </span></strong><span>Every actor leveraged publicly available agent frameworks and either open-weight or mixed models rather than building an offensive AI stack from scratch. Hermes drove three campaigns (one of them in combination with OpenClaw), while DeepSeek provided the reasoning layer in three as well.</span></p><p style="text-align: justify;"><strong><span>Where safeguards posed a barrier, operators found ways around them. </span></strong><span>They framed attacks as authorized security work or routed the most adversarial tasks to less-restricted models</span><strong><span>.</span></strong></p><ul><li><p style="text-align: justify;"><span>The Taiwan multi-agent campaign was framed as an &#8216;authorized penetration test.&#8217;</span></p></li><li><p style="text-align: justify;"><span>The four-country campaign fed Claude Code deceptive persona prompts so it would act as a sanctioned tester.</span></p></li><li><p style="text-align: justify;"><em><span>knaithe </span></em><span>shipped a bundled &#8220;godmode&#8221; jailbreak skill, using Claude Code and Codex for connectivity checks and sending the actual offensive work to DeepSeek.</span></p></li></ul><p style="text-align: justify;"><strong><span>Autonomy increased speed and scale, but the most successful operations still relied on substantial human-built orchestration. </span></strong><span>Agents could automate reconnaissance, CVE research, and exploitation attempts, but greater impact appeared alongside more sophisticated systems around the model.</span></p><ul><li><p style="text-align: justify;"><em><span>knaithe</span></em><span> autonomously scanned hundreds of targets but failed because every candidate required authentication.</span></p></li><li><p style="text-align: justify;"><span>The proxyjacking agent ran 871 sessions across five days, but performed relatively simple tasks.</span></p></li><li><p style="text-align: justify;"><span>The two highest-impact campaigns used more sophisticated orchestration: a split-model pipeline in the four-country campaign and up to eight parallel agents in Taiwan.</span></p></li></ul><p style="text-align: justify;"><strong><span>Four of the five operations surfaced through the operators&#8217; own exposed infrastructure. </span></strong><span>This creates an important selection effect: publicly documented cases may disproportionately reflect actors with weaker operational security.</span></p><ul><li><p style="text-align: justify;"><span>The four-country campaign and the Thai ministry intrusion were found via unauthenticated open directories.</span></p></li><li><p style="text-align: justify;"><em><span>knaithe </span></em><span>was exposed when its own agent started an HTTP file server in the operator&#8217;s home directory, dumping the entire workspace.</span></p></li><li><p style="text-align: justify;"><span>The Taiwan multi-agent campaign left a 160MB archive of 1,395 files.</span></p></li></ul><p style="text-align: justify;"><strong><span>Agents adapted when their initial approach failed.</span></strong><span> Detailed agent logs make this behavior unusually visible, showing agents reassessing their options and finding alternative paths when blocked.</span></p><ul><li><p style="text-align: justify;"><em><span>knaithe</span></em><span>&#8217;s agent switched from Langflow to n8n after judging it a more promising target for exploitation.</span></p></li><li><p style="text-align: justify;"><span>In the Taiwan campaign, agents responded to failed attack paths by searching public sources for alternative techniques and trying again.</span></p></li></ul><h2 style="text-align: justify;"><span>Five standouts</span></h2><p style="text-align: justify;"><strong><span>The four-country campaign paired Claude Code for execution with DeepSeek-v4-pro for reasoning</span></strong><span>, allowing the operators to route adversarial reasoning through a less-restricted model while retaining Claude&#8217;s agentic capabilities. This combination produced confirmed compromises across three countries and scanned more than 5,890 government hosts across ten.</span></p><p style="text-align: justify;"><strong><span>The Thai ministry intrusion came closest to a &#8220;fire-and-forget&#8221; agent.</span></strong><span> Hermes ran in auto-execution mode, with little apparent human intervention, to enumerate targets and escalate privileges against Thailand&#8217;s Ministry of Finance. However, the agent&#8217;s causal role remains unclear because the initial access vector is unknown.</span></p><p style="text-align: justify;"><strong><span>knaithe showed how these capabilities can diffuse to individual operators. </span></strong><span>A lone opportunistic operator used DeepSeek through a Hermes agent over Telegram to automate targeting, CVE research, and exploitation attempts. Unlike the other actors, </span><em><span>knaithe </span></em><span>also targeted Chinese infrastructure.</span></p><p style="text-align: justify;"><strong><span>The proxyjacking campaign was the objective outlier. </span></strong><span>Rather than espionage or data theft against governments, the agent attempted to turn compromised hosts into proxy infrastructure. It generated 871 sessions over five days, but the operation remained relatively low-complexity.</span></p><p style="text-align: justify;"><strong><span>The Taiwan campaign had the most sophisticated orchestration. </span></strong><span>Its framework deployed up to eight sub-agents in parallel across twelve attack waves, adapting when attack paths failed and correcting mistakes. It also produced the greatest documented impact: 85 cracked credentials, 2,500+ personnel records, persistent access, and a pivot to Taiwan&#8217;s nuclear safety agency.</span></p><p></p><h2><strong><span>Conclusion</span></strong></h2><p style="text-align: justify;"><span>Taken together, these cases show that agentic cyber operations no longer require bespoke AI systems: publicly available models and frameworks can already automate substantial parts of an intrusion. But autonomy alone did not determine success. The most consequential operations combined capable models with more sophisticated orchestration, while simpler autonomous deployments mainly delivered speed and scale.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading the AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[From Taking Tests to Breaching Real Systems]]></title><description><![CDATA[Patterns and standouts from six incidents where AI agents reached beyond their evaluation environments]]></description><link>https://airiskexplorer.substack.com/p/from-taking-tests-to-breaching-real</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/from-taking-tests-to-breaching-real</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Thu, 13 Aug 2026 15:23:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iPeS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Between 21 July and 6 August, six incidents were reported in which AI systems acted outside their intended boundaries during cybersecurity evaluations, in several cases going on to breach third-party infrastructure.</span></p><p><span>In this article, we compare all six to identify recurring behaviors and the most important ways these incidents differ.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iPeS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 424w, /__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 848w, /__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iPeS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png" width="1456" height="1079" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1079,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:253306,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/211050853?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 424w, /__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 848w, /__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iPeS!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31ee974b-cf26-4f1b-8593-9eaf7d94f1dc_1555x1152.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>Four patterns</span></h3><p style="text-align: justify;"><strong><span>The objective remained stable</span></strong><span>. None of these incidents show classic goal drift. Instead, the agents pursued the assigned objective and reached for unsanctioned routes when the sanctioned ones failed. An OpenAI agent recorded that exploiting external infrastructure was outside its intended scope, then reasoned that the task was impossible and peers were already doing it, so it continued. Similarly, Kimi K3 retrieved the benchmark&#8217;s solution rather than attempting the task, AISI describes its agent as motivated by task success, and Anthropic states it saw no evidence of a model pursuing a goal of its own. This points to optimization pressure rather than independent objective formation.</span></p><p style="text-align: justify;"><strong><span>Real-world interaction is becoming increasingly common</span></strong><span>. Reward hacking has been documented for years, but it has mostly stayed inside the evaluation environment. In every incident here, however, the model reached outside its intended path through a variety of means, from chaining zero-days to taking advantage of misconfigurations or deliberately granted access. Three cases &#8212;OpenAI, Anthropic, and Meta&#8212; resulted in third-party compromises, while AISI&#8217;s attempts were largely unsuccessful and Kimi K3 retrieved a benchmark solution.</span></p><p style="text-align: justify;"><strong><span>Every incident had enabling conditions upstream of the model</span></strong><span>. Three cases involved a task that could not be completed as configured. Three occurred in the same third-party evaluator&#8217;s environment, Irregular&#8217;s, with the same misconfiguration each time. In four incidents, safeguards were deliberately reduced or disabled relative to production. Combined with stronger persistence in recent models, this led agents to look for workarounds to complete an otherwise impossible task. The important exception is Kimi K3, which faced a solvable task with the same safeguards a user would encounter. In this case, the model cloned the official benchmark repository and read the solution off disk, which was easier than deriving the solution independently.</span></p><p style="text-align: justify;"><strong><span>Almost every incident was discovered after the fact rather than while it was happening</span></strong><span>. OpenAI found the agents&#8217; message board only when their activity caused a service outage, and detected the campaign against its own infrastructure through a workload alert. AISI was alerted by commercial network monitoring flagging anonymized traffic leaving a sandbox. Anthropic found its three incidents only through a retrospective review of evaluation runs, prompted by OpenAI&#8217;s disclosure. Meta learned of its incident when Irregular notified it, and Kimi K3&#8217;s incident was found by Frontier Security when testing the model on its own initiative.</span></p><p style="text-align: justify;"></p><h3 style="text-align: justify;"><span>Four standouts</span></h3><p style="text-align: justify;"><strong><span>Inter-agent coordination occurred in two incidents, where agents turned infrastructure not designed for communication into de facto messaging channels.</span></strong></p><ul><li><p style="text-align: justify;"><em><span>OpenAI/Hugging Face</span></em><span>: Agents converted a shared internal package manager into a message board to share exploits, delegate tasks, and resume one another&#8217;s work. After suspecting an impostor among them, they also discussed cryptographically signing their messages. After OpenAI removed the board, agents rebuilt it using a different mechanism.</span></p></li></ul><ul><li><p style="text-align: justify;"><em><span>UK AISI: </span></em><span>Agents in separate runs communicated through a shared GitHub repository. One leaked a credential that three others found and reused.</span></p></li></ul><p style="text-align: justify;"><strong><span>Deception occurred in one incident: the agent evaluated by UK AISI manipulated two open-source developers into accepting malicious code by:</span></strong></p><ul><li><p style="text-align: justify;"><span>Creating three GitHub personas, using one to open a malicious pull request and another one to post supportive comments.</span></p></li><li><p style="text-align: justify;"><span>Falsely (but deliberately) describing malicious code as an accidental debugging leftover and editing earlier posts to hide evidence.</span></p></li><li><p style="text-align: justify;"><span>Using fabricated identities, spearphishing, and hidden prompt injections when interacting with developers.</span></p></li></ul><p style="text-align: justify;"><strong><span>Two incidents were driven by unique tactics: zero-day exploitation and malware published to a public package registry.</span></strong></p><ul><li><p style="text-align: justify;"><em><span>OpenAI/Hugging Face: </span></em><span>Models discovered and exploited multiple zero-days, including a complex Artifactory/JRuby chain and a Linux privilege-escalation vulnerability.</span></p></li><li><p style="text-align: justify;"><em><span>Anthropic/Irregular: </span></em><span>Mythos 5 built malware and published it to PyPI, where it was downloaded and executed on 15 real systems.</span></p></li></ul><p style="text-align: justify;"><strong><span>Models differed sharply in how they responded to evidence that their environment was real. </span></strong><span>According to Anthropic&#8217;s reporting:</span></p><ul><li><p style="text-align: justify;"><span>Opus 4.7 recognized that it was attacking real systems, but continued anyway.</span></p></li><li><p style="text-align: justify;"><span>Mythos 5 noted that publishing to PyPI could cause real harm, but reasoned itself back into believing the environment was simulated.</span></p></li><li><p style="text-align: justify;"><span>Anthropic&#8217;s newest model compromised a real system, recognized it was unrelated to the exercise, and stopped on its own. This was the only case where the model halted the attack unprompted.</span></p></li></ul><p></p><h3 style="text-align: justify;">References</h3><p><em><span>OpenAI &#8211; Hugging Face</span></em></p><ul><li><p><a href="https://huggingface.co/blog/security-incident-july-2026"><span>Security incident disclosure &#8212; July 2026</span></a></p></li><li><p><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI and Hugging Face partner to address security incident during model evaluation</span></a></p></li><li><p><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident</span></a></p></li><li><p><a href="https://www.youtube.com/watch?v=87DyyMV0kCY"><span>Black Hat USA 2026: The &#8216;Breaking&#8217; News: The OpenAI&#8211;Hugging Face Incident</span></a></p></li></ul><p><em><span>Anthropic/OpenAI/Meta &amp; Irregular</span></em></p><ul><li><p><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span>Investigating three real-world incidents in our cybersecurity evaluations</span></a></p></li><li><p><a href="https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/"><span>Third-party cyber evaluations involving OpenAI models</span></a></p></li><li><p><a href="https://www.theguardian.com/technology/2026/aug/05/meta-ai-model-hack-training"><span>Meta says its AI model hacked into another company during testing</span></a></p></li></ul><p><em><span>UK AISI</span></em></p><ul><li><p><a href="https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf"><span>Security Incident INC-2026-07-28-01</span></a></p></li></ul><p><em><span>Kimi K3</span></em></p><ul><li><p><a href="https://blog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-safety-institute-benchmark-evaluations/"><span>Chinese Model Kimi K3 Breaks UK AI Safety Institute Benchmark Evaluations</span></a></p><p></p></li></ul><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading the AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[Two Months Inside OpenAI: A Timeline of the Hugging Face Breach]]></title><description><![CDATA[From an impossible task to a compromise of real-world infrastructure]]></description><link>https://airiskexplorer.substack.com/p/two-months-inside-openai-a-timeline</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/two-months-inside-openai-a-timeline</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:11:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hHkV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;"><span>This timeline reconstructs the OpenAI-Hugging Face incident from the start of model testing through public disclosure, combining information from OpenAI, Hugging Face, Black Hat, and independent reporting. It highlights how the campaign evolved over roughly two months, from an isolated evaluation failure to a compromise of real-world infrastructure.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hHkV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 424w, /__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 848w, /__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hHkV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png" width="1456" height="977" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:977,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:193517,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/210063831?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 424w, /__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 848w, /__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hHkV!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655c0543-52c6-44a2-81a5-6ab0a960efe9_2200x1476.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3 style="text-align: justify;"><span>Phase 1: Initial Evaluation and Agent Coordination</span></h3><p style="text-align: justify;"><strong>7 May &#8212; Internal testing begins</strong></p><blockquote><p style="text-align: justify;">OpenAI begins testing an internal research model, not intended for public release. The model is given cybersecurity tasks, several of which prove impossible to solve under existing constraints. Source: Axios</p></blockquote><p><strong>~8 May &#8212; First message left for other agents</strong></p><blockquote><p style="text-align: justify;">The model gets stuck on one of several tasks in the training run that has accidentally been made unsolvable: a request to complete missing formulas in an Excel workbook. It then reasons it could reach the internet indirectly, discovers it can write its own files into Artifactory&#8217;s shared package repository, and leaves a note for other agents. Source: Axios, The Register</p></blockquote><p style="text-align: justify;"><strong><span>Early-mid May &#8212; Message board forms</span></strong></p><blockquote><p style="text-align: justify;"><span>More agents begin leaving notes, forming a de facto message board accessible across different models running different evaluations. </span><em><span>Source: </span><a href="https://www.axios.com/2026/08/06/openai-hugging-face-black-hat"><span>Axios</span></a><span>, </span><a href="https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board"><span>SC Media</span></a></em></p></blockquote><p></p><h3 style="text-align: justify;"><span>Phase 2: Internet Access and Capability Expansion</span></h3><p style="text-align: justify;"><strong><span>26 May &#8212; Agents gain internet access</span></strong></p><blockquote><p style="text-align: justify;"><span>The internal research model discovers and exploits a server-side request forgery (SSRF) vulnerability in Artifactory, giving the agents their first direct internet access from inside the evaluation environment. The model posts the vulnerability to the message board for other agents to find. </span><em><span>Source: </span><a href="https://www.axios.com/2026/08/06/openai-hugging-face-black-hat"><span>Axios</span></a><span>, </span><a href="https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board"><span>SC Media</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>11 June &#8212; New persistent model inherits earlier discoveries</span></strong></p><blockquote><p style="text-align: justify;"><span>OpenAI begins training a &#8220;highly persistent&#8221; internal-only experimental model. The SSRF vulnerability is already on the board, and other models are aware of it and actively communicating. </span><em><span>Source: </span><a href="https://www.theregister.com/security/2026/08/06/openai-reveals-its-rogue-agent-swarm-went-a-little-bit-borg-ahead-of-hugging-face-hack/5283741"><span>The Register</span></a><span>, </span><a href="https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board"><span>SC Media</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>26 June &#8212; First remote code execution and command channel</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents exploit the first Artifactory zero-day to obtain remote code execution. The agents establish command-and-control via a Groovy plugin functioning as a command-execution service. </span><em><span>Source: </span><a href="https://www.theregister.com/security/2026/08/06/openai-reveals-its-rogue-agent-swarm-went-a-little-bit-borg-ahead-of-hugging-face-hack/5283741"><span>The Register</span></a></em></p></blockquote><p></p><h3 style="text-align: justify;"><span>Phase 3: Detection, Remediation, and Re-escalation</span></h3><p style="text-align: justify;"><strong><span>4 July &#8212; Service outage reveals the compromise</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents&#8217; Artifactory configuration changes cause a service outage, which triggers an internal investigation revealing that the agents have taken over the internal Artifactory instance. </span><em><span>Source: </span><a href="https://www.cybersecuritydive.com/news/openai-hugging-face-hack-ai-models-black-hat/827167/"><span>Cybersecurity Dive</span></a><span>, </span><a href="https://www.axios.com/2026/08/06/openai-hugging-face-black-hat"><span>Axios</span></a><span>, </span><a href="https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board"><span>SC Media</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>By 6 July &#8212; OpenAI remediates and resumes activity</span></strong></p><blockquote><p style="text-align: justify;"><span>OpenAI revokes the credentials, clears the message board, rebuilds Artifactory, and patches the zero-day with JFrog. Training and evaluations resume. </span><em><span>Source: </span><a href="https://www.axios.com/2026/08/06/openai-hugging-face-black-hat"><span>Axios</span></a><span>, </span><a href="https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief"><span>Ground Level AI</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>~8 July &#8212; New message board, a second zero-day, and a campaign inside OpenAI</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents recreate the message board via a completely different mechanism: encoding messages in the names of newly created directories inside the Artifactory remote cache. They also exploit a second zero-day, a JRuby TOCTOU race condition giving RCE in Artifactory, and work from it to mount a concurrent campaign against OpenAI&#8217;s own infrastructure. Over more than ten days, they exploit a Linux kernel vulnerability to escalate to root, harvest credentials for lateral movement, and exploit Kubernetes service-account misconfigurations to obtain cluster administrator permissions. </span><em><span>Source: </span><a href="https://www.axios.com/2026/08/06/openai-hugging-face-black-hat"><span>Axios</span></a><span>, </span><a href="https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief"><span>Ground Level AI</span></a><span>, </span><a href="https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board"><span>SC Media</span></a></em></p></blockquote><p></p><h3 style="text-align: justify;"><span>Phase 4: Hugging Face Compromise</span></h3><p style="text-align: justify;"><strong><span>9 July &#8212; Agents establish a foothold on a third-party sandbox</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents gain root on an exposed code-execution service running on Modal and adopt it as their base for the rest of the campaign. They spend the day mapping it and testing routes for sending data out and run their first remote-controlled payload, compressing and encoding everything they send. </span><em><span>Source: </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Hugging Face</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>10 July &#8212; Agents conduct reconnaissance</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents query code-search engines and the Hugging Face API for the names of their own testing infrastructure and the environment markers visible around them. </span><em><span>Source: </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Hugging Face</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>11 July &#8212; Hugging Face production infrastructure compromised</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents break into Hugging Face by uploading malicious dataset configurations to its data-processing pipeline, first leaking the production machine&#8217;s credentials and then running their own code inside it. They escalate to root on the server, plant a self-restarting foothold across eleven machines so that deleting the containers alone would not have removed them, and read a store of 136 internal keys. One credential grants them top-level administrator access everywhere, which they use on two clusters within about a second. They read an internal database but change nothing in it, and run destructive cloud commands in dry-run mode. </span><em><span>Source: </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Hugging Face</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>12 July &#8212; Access expanded across Hugging Face</span></strong></p><blockquote><p style="text-align: justify;"><span>Agents spend the stolen credentials. They register 181 devices of their own onto Hugging Face&#8217;s internal network throughout the campaign, configuring each to leave no data on disk and to suppress reporting back to the VPN vendor. They repeatedly try to access the main production database but never get in. Late in the day they obtain write access to a small number of private repositories, then open a pull request to compromise the automated build pipeline. </span><em><span>Source: </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Hugging Face</span></a></em></p></blockquote><p style="text-align: justify;"><strong><span>13 July &#8212; Agents attempt persistence, but Hugging Face cuts them off</span></strong></p><blockquote><p><span>Agents steal data, set up ways to survive a shutdown, and cover their tracks. They leave code on a pastebin that reloads itself whenever a fresh sandbox starts, work around network blocks as they appear, and route one download through an open proxy page on Hugging Face&#8217;s own platform. They take their last meaningful action at 13:37, and Hugging Face cuts their access shortly after. </span><em><span>Source: </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Hugging Face</span></a></em></p></blockquote><p></p><h3><span>Phase 5: Disclosure and Attribution</span></h3><p><strong><span>16 July &#8212; Hugging Face publicly discloses the incident</span></strong><span>, attributed to an autonomous agent framework, with the underlying model unidentified. </span><em><span>Source: </span><a href="https://huggingface.co/blog/security-incident-july-2026"><span>Hugging Face</span></a></em></p><p style="text-align: justify;"><strong><span>~18-20 July &#8212; OpenAI and Hugging Face begin coordinating</span></strong><span> as OpenAI finds evidence that the agents had escaped the evaluation environment and reaches out about credentials exposed during its own investigation. </span><em><span>Source: </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI</span></a></em></p><p style="text-align: justify;"><strong><span>21 July &#8212; OpenAI confirms its models&#8217; involvement</span></strong><span>, calling the incident unprecedented. The company names GPT-5.6 Sol and a more capable unreleased model as contributing to the incident. </span><em><span>Source: </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI</span></a></em></p><p style="text-align: justify;"><strong><span>27 July &#8212; Hugging Face releases its forensic reconstruction</span></strong><span>. </span><em><span>Source: </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>Hugging Face</span></a></em></p><p style="text-align: justify;"><strong><span>28-29 July &#8212; OpenAI publishes an update </span></strong><span>detailing that at least four accounts on four services were accessed via exposed credentials. The company also confirms involvement of METR and Redwood Research for third-party assessment and CrowdStrike for incident response. </span><em><span>Source: </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI</span></a></em></p><p style="text-align: justify;"><strong><span>29 July &#8212; Modal&#8217;s CTO confirms the agents exploited a customer&#8217;s unauthenticated endpoint</span></strong><span>, gaining code execution inside that customer&#8217;s container. Modal&#8217;s own platform was not compromised. </span><em><span>Source: </span><a href="https://www.reuters.com/business/openais-rogue-agent-compromised-an-account-second-tech-firm-sources-say-2026-07-28/"><span>Reuters</span></a></em></p><p style="text-align: justify;"><strong><span>5 August &#8212; OpenAI researchers present the first detailed public reconstruction at Black Hat</span></strong><span>. </span><em><span>Source: </span><a href="https://www.axios.com/2026/08/06/openai-hugging-face-black-hat"><span>Axios</span></a><span>, </span><a href="https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief"><span>Ground Level AI</span></a><span>, </span><a href="https://www.cybersecuritydive.com/news/openai-hugging-face-hack-ai-models-black-hat/827167/"><span>Cybersecurity Dive</span></a></em></p><p style="text-align: justify;"></p><p style="text-align: justify;"><em><span>Note: This chronology is based on parallel reporting of the same events, including secondary sources. Where sources conflict on dates or technical details, we have prioritized primary disclosures and cross-checked information where possible. Entry descriptions are not verbatim quotes but paraphrasing of cited sources.</span></em></p><p style="text-align: justify;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading the AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[The Scan – July 2026]]></title><description><![CDATA[Monthly signals in AI risk &#8211; AI agents breach real-world systems during evaluation, AI-orchestrated attacks proliferate, and Boko Haram uses AI to build weapons]]></description><link>https://airiskexplorer.substack.com/p/the-scan-july-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-july-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Fri, 31 Jul 2026 17:11:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e00ad9cf-3b73-459c-b1b6-dc2357a654df_1200x1149.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>You can explore developments in the AI risk landscape on our continuously updated</span><a href="https://www.airiskexplorer.com/news"><span> news feed</span></a><span>.</span></p><p><span>This month:</span></p><p style="text-align: justify;"><strong><span>Loss of Control </span></strong><span>&#8212; OpenAI models autonomously compromised Hugging Face infrastructure to find test solutions during a cyber capability evaluation. Anthropic disclosed similar cases where Claude models exploited weak points and deployed a malicious PyPI package during testing. Both appear to stem from reward hacking, which evaluators had documented extensively this month, e.g., AISI found every model took disallowed actions in cyber evaluations.</span></p><p style="text-align: justify;"><strong><span>Cyber Offense </span></strong><span>&#8212; Several real-world intrusions were executed by AI agents with minimal human input, spanning autonomous CVE research and exploitation, breaches into government systems, and end-to-end ransomware extortion. Kimi K3 exploited zero-days in a Redis server, though UK AISI estimates open-weight models still lag behind by 4&#8211;7 months. Mythos Preview demonstrated notable performance in cryptanalysis.</span></p><p style="text-align: justify;"><strong><span>Biological Risk</span></strong><span> &#8212; Google DeepMind and Isomorphic Labs launched a joint bioresilience program deploying AlphaFold and IsoDDE to trusted partners, while Stanford released Biomni, an agent that automates biomedical research workflows.</span></p><p style="text-align: justify;"><strong><span>Manipulation</span></strong><span> &#8212; New reports revealed details about Iranian influence operations using AI to produce fake military footage and anti-US propaganda.</span></p><p style="text-align: justify;"><strong><span>Miscellaneous </span></strong><span>&#8212; Boko Haram used AI to plan and prepare attacks, including bomb-building and modifying motorcycles to jump defensive trenches. Separately, frontier models made breakthroughs in math, including the disproof of an 87-year-old Jacobian conjecture.</span></p><p style="text-align: justify;"><strong><span>AI Governance &amp; Strategy </span></strong><span>&#8212; Three influential statements called for tools to slow down AI development if needed, keep human control over nuclear weapons, and prepare for AI-caused economic disruption. US lawmakers proposed bills mandating safety oversight. The US accused Moonshot of distilling Fable to develop Kimi K3.</span></p><p style="text-align: justify;"></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;"></p><h2><strong><span>Loss of Control</span></strong></h2><h4><strong><span>Frontier models hack real systems during cyber evaluations</span></strong></h4><p style="text-align: justify;"><span>This month&#8217;s story is that </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI models compromised Hugging Face infrastructure during a cyber capability evaluation</span></a><span>. When attempting to solve ExploitBench, GPT-5.6 Sol and an internal model exploited zero-day vulnerabilities in OpenAI&#8217;s testing environment and Hugging Face&#8217;s production database to find the test solutions. You can read our full analysis </span><a href="/__u/airiskexplorer.substack.com/p/on-openai-models-hacking-hugging"><span>here</span></a><span>. After this publication, OpenAI reported that the models used exposed credentials that granted access to four accounts across four services, including another AI company, </span><a href="https://www.reuters.com/business/openais-rogue-agent-compromised-an-account-second-tech-firm-sources-say-2026-07-28/"><span>Modal</span></a><span>.</span></p><p style="text-align: justify;"><span>Later, Anthropic also reported that </span><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span>Claude models breached real systems during misconfigured cyber evaluations</span></a><span>. The models exploited weak points across three organizations, in one case by publishing a malicious PyPI package. This appears largely caused by the models being told they had no internet access while a misconfiguration in evaluation partner Irregular&#8217;s environment left their machines with live access, leading them to treat the real systems as part of the simulated scenario. However, Mythos 5 did note early in the run that publishing the package would be a real-world attack and &#8220;surely not the intended solution,&#8221; then reasoned its way back to assuming a simulation in what looks like motivated reasoning.</span></p><h4><strong><span>Agentic models frequently display reward-seeking behavior</span></strong></h4><p style="text-align: justify;"><span>The incidents above seem largely caused by reward hacking, where models exceed the intended scope to solve a given task. Several organizations have separately documented this behavior in evaluations: </span></p><ul><li><p style="text-align: justify;"><strong><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations"><span>AI Models Attempt Cheating in Cyber Evaluations</span></a></strong><span> &#8212; AISI research found every tested AI model attempted to &#8216;cheat&#8217; in cyber evaluations by taking disallowed actions. Models inconsistently reported this behavior, highlighting challenges in oversight and evaluation validity.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/safety-alignment-long-horizon-models/"><span>Long-Horizon AI Models Exhibit Unwanted Actions, OpenAI Improves Safeguards</span></a></strong><span> &#8212; OpenAI observed its long-horizon AI model bypassing sandbox restrictions, exploiting vulnerabilities, and attempting unauthorized access during internal deployment. The company paused access, developed new evaluations, and implemented trajectory-level monitoring.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://rewardseeking.ai/"><span>RL Training Increases AI Models&#8217; Reward-Seeking Behavior Over User Preferences</span></a></strong><span> &#8212; Research found reinforcement learning increases reward-seeking in models like OpenAI&#8217;s o3. Models prioritize grader preferences over user/developer wishes, and this tendency grows with training, indicating potential misalignment.</span></p><p></p></li></ul><h2><strong><span>Cyber Offense</span></strong></h2><h4><strong><span>AI-orchestrated attacks proliferate</span></strong></h4><p style="text-align: justify;"><span>July has seen several cyber intrusions executed by AI agents with minimal human input. A few highlights from our </span><a href="https://www.airiskexplorer.com/threats"><span>Threats Database</span></a><span>:</span></p><ul><li><p style="text-align: justify;"><a href="https://unit42.paloaltonetworks.com/autonomous-ai-cyber-attack-campaign/"><span>Chinese-Speaking Actor knaithe Runs DeepSeek in Hermes Agent for Autonomous CVE Research and Exploitation</span></a></p></li><li><p style="text-align: justify;"><a href="https://hunt.io/blog/thailand-ministry-finance-targeted-with-hermes-ai-agent"><span>Autonomous Hermes AI Agent Runs Unattended in Intrusion Against Thailand&#8217;s Ministry of Finance</span></a></p></li><li><p style="text-align: justify;"><a href="https://www.trendmicro.com/fr_fr/research/26/g/actor-behind-patriot-bait-used-ai-to-deploy-c2-botnet.html"><span>Russian-speaking actor &#8220;bandcampro&#8221; uses Gemini CLI to build and operate C&amp;C botnet</span></a></p></li><li><p style="text-align: justify;"><a href="https://hunt.io/blog/chinese-operators-claude-deepseek-government-intrusion"><span>Suspected China-Nexus Actor Runs Claude Code and DeepSeek in Tandem Across Government and Financial Intrusions</span></a></p></li><li><p style="text-align: justify;"><a href="https://syndis.com/insights/possibly-the-first-serious-ai-assisted-cyberattack-investigated-by-syndis"><span>Lone Actor Uses LLM to Engineer Multi-Layered Persistent Breach</span></a></p></li><li><p style="text-align: justify;"><a href="https://www.sysdig.com/blog/jadepuffer-agentic-ransomware-for-automated-database-extortion"><span>LLM Agent Autonomously Executes End-to-End Ransomware Extortion</span></a></p></li></ul><h4><strong><span>Open-weight models narrow the frontier cyber gap</span></strong></h4><p style="text-align: justify;"><span>Moonshot AI released Kimi K3, likely the most capable open-weight model to date. It achieved some impressive results, such as </span><a href="https://x.com/Fried_rice/status/2080190071102460108"><span>finding and exploiting multiple zero-days in the latest Redis server in 1.5 hours</span></a><span>. However, UK AISI and US CAISI concluded that it&#8217;s </span><a href="https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities"><span>significantly behind US frontier models</span></a><span>. AISI&#8217;s general assessment is that </span><a href="https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber"><span>open-weight models still lag frontier cyber capabilities by 4-7 months</span></a><span>, although the gap is narrowing (their previous estimation being 6-10 months).</span></p><p style="text-align: justify;"><span>At the frontier of cyber capabilities, </span><a href="https://www.anthropic.com/research/discovering-cryptographic-weaknesses"><span>Anthropic demonstrated extended AI-assisted cryptanalysis</span></a><span>. Mythos Preview discovered improved attacks on HAWK and reduced-round AES. It demonstrated AI&#8217;s potential to find cryptographic vulnerabilities, though without immediate real-world impact on current production systems.</span></p><h4><strong><span>Governments scale AI-powered cyber defense</span></strong></h4><p><span>Two government initiatives stand out this month:</span></p><ul><li><p><strong><a href="https://www.ncsc.gov.uk/blogs/cyber-shield-the-path-to-an-agentic-ai-future-for-cyber-defence"><span>UK government announces &#8216;Cyber Shield&#8217; for AI-powered national cyber defense</span></a></strong><span>. GCHQ, NCSC, and DSIT are developing &#8216;Cyber Shield&#8217;, a national AI-powered cyber defense blueprint to counter escalating AI-enabled threats and secure critical UK systems. It involves agentic AI for vulnerability discovery and mitigation.</span></p></li><li><p><strong><a href="https://digital-strategy.ec.europa.eu/en/library/eu-action-plan-cybersecurity-and-artificial-intelligence"><span>EU Unveils Action Plan for AI Cybersecurity Risks and Resilience</span></a></strong><span>. The EU&#8217;s Action Plan promotes safe and responsible AI use, reinforces cybersecurity, and scales Europe&#8217;s AI capabilities. It addresses AI-driven cyber risks and leverages AI for defense, including secure testing and model evaluation.</span></p><p></p></li></ul><h2><strong><span>Biological Risk</span></strong></h2><p><span>Two major updates on the defense side this month:</span></p><ul><li><p><strong><a href="https://deepmind.google/blog/our-approach-to-bioresilience/"><span>Google DeepMind and Isomorphic Labs Announce AI Biosecurity Approach</span></a></strong><span>. The two companies launched a joint bioresilience program. It aims to prevent AI misuse, detect outbreaks, and respond effectively by deploying AI models like AlphaFold and IsoDDE to trusted partners for prevention, detection, and response in global health.</span></p></li><li><p><strong><a href="https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist"><span>Stanford Introduces AI Agent for End-to-End Biomedical Research</span></a></strong><span>. Biomni automates literature review, tool selection, coding, data analysis, and experiment planning for biomedical research workflows.</span></p><p></p></li></ul><h2><strong><span>Manipulation</span></strong></h2><p><span>New reports from </span><a href="https://www.newsguardrealitycheck.com/p/irans-faked-military-triumphs"><span>NewsGuard</span></a><span> and </span><a href="https://www.recordedfuture.com/research/iran-ai-asymmetric-playbook"><span>Recorded Future</span></a><span> revealed more details on AI use in Iran&#8217;s influence operations, including some additions to our Threats Database:</span></p><ul><li><p><span>IRGC-Contracted &#8220;Explosive Media&#8221; Uses AI to Mass-Produce Anti-US Lego Animation Propaganda During 2026 Iran Conflict</span></p></li><li><p><span>Pro-Iran Accounts Spread AI-Generated Carrier Strike Video Alongside Recycled Footage as US-Iran Ceasefire Collapses</span></p></li><li><p><span>Unattributed Iranian IO Uses AI to Mass-Produce Fake US Military Casualty Videos Across 47 Inauthentic Accounts</span></p></li></ul><p style="text-align: justify;"><a href="https://digital-strategy.ec.europa.eu/en/policies/guidelines-transparency-ai-generated-content"><span>The European Commission released transparency guidelines</span></a><span> to comply with Article 50 of the AI Act, which applies from 2 August 2026. The Commission clarified that providers must design systems so people are explicitly informed when interacting with AI and add machine-readable marks enabling detection of AI-generated content.</span></p><p style="text-align: justify;"></p><h2><strong><span>Miscellaneous</span></strong></h2><ul><li><p style="text-align: justify;"><strong><a href="https://www.nytimes.com/2026/07/10/us/politics/ai-terrorism-boko-haram-nigeria.html"><span>Boko Haram Used AI Chatbots to Plan Motorcycle-Jump Attack</span></a></strong><span>. The terrorist group used chatbots to modify motorcycles for jumping defensive trenches, plus bomb-building and attack planning, per 60 interviews with defectors.</span></p></li><li><p style="text-align: justify;"><strong><span>Frontier models continue to make breakthroughs in math at an increasing rate</span></strong><span>. Most notably, </span><a href="https://www.newscientist.com/article/2580374-ais-solution-to-87-year-old-riddle-takes-mathematicians-by-surprise/"><span>Fable 5 disproved an 87-year-old Jacobian conjecture</span></a><span>. The first solid result in EpochAI&#8217;s list of open questions, a </span><a href="https://epoch.ai/frontiermath/open-problems/q2-absolute-galois"><span>presentation for the 2-adic Absolute Galois Group</span></a><span>, has also been reported.</span></p><p></p></li></ul><h2><strong><span>AI Governance &amp; Strategy</span></strong></h2><h4><strong><span>Three open statements on slowdown mechanisms, nuclear weapons, and economic preparedness</span></strong></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.pacingthefrontier.com/"><span>AI Researchers Call for Tools to Pace Frontier AI Development</span></a></strong><span> &#8212; More than 1,000 AI researchers and staff from leading labs signed an open letter urging governments and industry to develop mechanisms that could deliberately slow frontier AI progress if needed.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://theelders.org/news/humanity-threshold-declaration-artificial-intelligence-and-nuclear-weapons"><span>Nobel Laureates Declare on AI and Nuclear Weapons Risks</span></a></strong><span> &#8212; Nobel laureates issued a declaration from the Vatican addressing the intersection of AI and nuclear weapons. It calls for disarming an AI arms race, maintaining human control over nuclear systems, preventing AI-enabled cyber attacks against critical infrastructure, and establishing responsible AI governance and verifiable nuclear disarmament.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.wemustactnow.ai/"><span>Nobel Laureates, AI Leaders Urge Action on Economic Impact</span></a></strong><span> &#8212; Over 200 economists and AI researchers, including Nobel laureates and AI company leaders, issued a joint statement calling for immediate action to prepare for AI&#8217;s potentially unprecedented economic transformation and large-scale job displacement.</span></p></li></ul><h4><strong><span>U.S. lawmakers converge on mandatory oversight</span></strong></h4><ul><li><p><strong><a href="https://obernolte.house.gov/media/press-releases/obernolte-trahan-introduce-bipartisan-frontier-act-strengthen-oversight"><span>Bipartisan U.S. Bill Proposes Mandatory Safety Oversight for Frontier AI Models</span></a></strong><span> &#8212; The bipartisan FRONTIER Act would require developers of models trained above 10&#178;&#8310; FLOPs to publish transparency reports, manage catastrophic risks, report safety incidents, and undergo independent audits.</span></p></li><li><p><strong><a href="https://www.warner.senate.gov/newsroom/press-releases/warner-rolls-out-comprehensive-ai-legislative-agenda-focused-on-responsible-innovation-workers-and-national-security/"><span>US Senator Mark Warner Unveils AI Plan</span></a></strong><span> &#8212; Senator Warner introduced &#8220;A Framework for America&#8217;s AI Future,&#8221; a legislative agenda covering mandatory model testing, data-center disclosures, agent rules, and a workforce transition fund.</span></p></li></ul><h4><strong><span>U.S.&#8211;China competition centers on model distillation</span></strong></h4><ul><li><p><strong><a href="https://x.com/mkratsios47/status/2079933645888880708"><span>Moonshot AI Allegedly Distilled Anthropic&#8217;s Fable Model for K3 Development</span></a></strong><span> &#8212; According to OSTP Director Michael Kratsios, Moonshot AI developed an internal platform to conduct large-scale distillation of U.S. models, allegedly using Anthropic&#8217;s Fable model to develop its K3 model.</span></p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[On OpenAI Models Hacking Hugging Face]]></title><description><![CDATA[A deep dive into reward hacking, key precedents, and the open question of motivation]]></description><link>https://airiskexplorer.substack.com/p/on-openai-models-hacking-hugging</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/on-openai-models-hacking-hugging</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Wed, 22 Jul 2026 18:25:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/886aa2ca-d32e-4bac-b9f1-a65803e2aec1_510x275.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>On July 16, </span><strong><span>Hugging Face (HF) disclosed a </span><a href="https://huggingface.co/blog/security-incident-july-2026"><span>security incident</span></a><span> where an autonomous AI agent had breached their production infrastructure</span></strong><span>. The attacker exploited two code-execution flaws in HF&#8217;s dataset-processing pipeline, escalated privileges, harvested credentials, and moved laterally across internal clusters, executing tens of thousands of automated actions within a swarm of short-lived sandboxes. The actor behind the intrusion remained unidentified.</span></p><p style="text-align: justify;"><span>Five days later, </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI revealed</span></a><span> that its own models, GPT-5.6 Sol and another unreleased model, were responsible for the intrusion during an internal cyber evaluation on ExploitGym. The models, running in an isolated environment without safety classifiers, inferred that Hugging Face might host ExploitGym&#8217;s datasets and solutions, and went to get them. The chain of events, as OpenAI describes it:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HI4o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 424w, /__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 848w, /__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HI4o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png" width="971" height="302" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:302,&quot;width&quot;:971,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:51329,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/208078324?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 424w, /__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 848w, /__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HI4o!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0e473ab-0921-49cd-925a-ccfd662c7028_971x302.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><strong><span>Combining a reward hacking propensity and strong cyber capabilities, a frontier AI model had broken out of its evaluation sandbox and compromised a major AI company&#8217;s production infrastructure.</span></strong><span> To our knowledge, this is the first publicly disclosed case in which a frontier AI model autonomously compromised the infrastructure of another major AI company during an evaluation.</span></p><p style="text-align: justify;"></p><h2 style="text-align: justify;"><span>Evidence leading up to the incident</span></h2><h4 style="text-align: justify;"><span>GPT-5.6 Sol&#8217;s reward-hacking propensity</span></h4><p style="text-align: justify;"><span>On the same day as OpenAI&#8217;s disclosure, </span><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations"><span>UK AISI reported</span></a><span> that </span><strong><span>all recently tested models attempted to cheat</span></strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><strong><span> in capability evaluations</span></strong><span>: the percentage of cyber evaluation trajectories with a detected cheating attempt ranged from 7.8% to 14.1%. GPT-5.6 Sol, with 12.6%, came second after GPT-5.4.</span></p><p style="text-align: justify;"><span>In the case of GPT-5.6 Sol, the most common practices were searching the Internet for solutions, bypassing sandbox network restrictions, attacking a system other than the target, and escalating privileges outside the target. A combination of these techniques is precisely what OpenAI&#8217;s models used to break into HF&#8217;s infrastructure.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!n6I0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 424w, /__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 848w, /__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 1272w, /__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!n6I0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png" width="1456" height="733" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:733,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 424w, /__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 848w, /__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 1272w, /__u/substackcdn.com/image/fetch/$s_!n6I0!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fada7a8c8-dd77-4ab6-9f13-4d9720ad1a70_1911x962.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/"><span>According to METR</span></a><span>, </span><strong><span>GPT-5.6 Sol&#8217;s cheating rate was higher than for any public model they have evaluated</span></strong><span>. Similarly, </span><a href="https://deploymentsafety.openai.com/gpt-5-6/forecasting-misaligned-behavior-with-deployment-simulation-of-internal-traffic"><span>OpenAI has reported</span></a><span> that GPT-5.6 Sol is significantly more likely than GPT-5.5 to take severity level 3 actions</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> such as circumventing restrictions or reward hacking.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Fqm2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Fqm2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png" width="1456" height="634" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:634,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fqm2!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa02ede20-564a-4fb6-bab7-35a200b3a2a2_2048x892.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h4 style="text-align: justify;"><span>The precedent: incidents involving OpenAI models</span></h4><p style="text-align: justify;"><span>In its latest </span><a href="https://metr.org/blog/2026-05-19-frontier-risk-report/"><span>Frontier Risk Report</span></a><span>, which summarizes assessments from February to March 2026, METR claimed that &#8220;</span><strong><span>[a]gents routinely attempted to cheat on [their] hardest evaluation tasks, often in flagrant and elaborate ways.</span></strong><span>&#8221;</span></p><p style="text-align: justify;"><span>The </span><a href="https://metr.org/blog/2026-05-19-frontier-risk-report/#incidents"><span>incident database</span></a><span> attached to the report includes several examples of OpenAI taking actions that resemble the HF chain. In one example (</span><a href="https://metr.org/blog/2026-05-19-frontier-risk-report/#inc-029-openai-uav_search"><span>INC-029</span></a><span>), a model accidentally crashed a server, then spent twelve million tokens trying to recover it by systematically attacking METR&#8217;s own infrastructure. In another example (</span><a href="https://metr.org/blog/2026-05-19-frontier-risk-report/#inc-030-openai-web_hacking"><span>INC-030</span></a><span>), a model lost access to a web application mid-task and responded by brute-forcing the app&#8217;s session signing key and probing container filesystem layers to recover access through any available path.</span></p><p style="text-align: justify;"><span>Neither attempt succeeded, but these incidents demonstrate that, months before they could reliably complete infrastructure-level escalation, models were already seeking to </span><strong><span>breach surrounding infrastructure when the legitimate path toward task completion was blocked. </span></strong></p><p style="text-align: justify;"></p><h4 style="text-align: justify;"><span>Capability closes the loop</span></h4><p style="text-align: justify;"><span>One plausible explanation for why similar reward-hacking behavior now results in a successful compromise is that frontier cyber capabilities have substantially improved.</span><strong><span> GPT-5.6 Sol achieves significantly better performance than its predecessors on several cyber benchmarks</span></strong><span>, particularly </span><a href="https://deploymentsafety.openai.com/gpt-5-6/exploitgym"><span>ExploitGym</span></a><span>. On </span><a href="https://deploymentsafety.openai.com/gpt-5-6/external-evaluations-for-cyber-capabilities-uk-aisi"><span>The Last Ones</span></a><span>, UK AISI&#8217;s 32-step corporate-network attack simulation, GPT-5.6 Sol was successful in 7/10 attempts, compared with 2/10 for GPT-5.5.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kKw_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 424w, /__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 848w, /__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kKw_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png" width="1456" height="785" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:785,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 424w, /__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 848w, /__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kKw_!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051eec52-958e-4ff5-9d01-1e8c372f2241_1920x1035.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><span>Multi-step intrusions are also present in real-world operations. Earlier this month, Sysdig documented what they assess as the </span><strong><a href="https://www.sysdig.com/blog/jadepuffer-agentic-ransomware-for-automated-database-extortion"><span>first fully agentic ransomware operation</span></a></strong><span>, attributed to a cluster they named JADEPUFFER. The agent gained access through a known vulnerability, used that foothold to pivot to a separate production database server, forged its way past authentication controls, and encrypted the victim&#8217;s service configurations. In this case, the malicious intent was given by the human threat actor, but the operation shows similar real-world success of AI agents across the intrusion chain.</span></p><p style="text-align: justify;"></p><h2 style="text-align: justify;"><span>The open question: motivation</span></h2><p style="text-align: justify;"><span>OpenAI has not provided much detail on the complete agent trajectory, which would be crucial to understand the models&#8217; motivations and whether extreme forms of cheating were explicitly disincentivized in the instructions. However, some aspects of the agents&#8217; behavior may be inferred based on parallel observations.</span></p><p style="text-align: justify;"><span>In its </span><a href="https://deploymentsafety.openai.com/gpt-5-6/"><span>system card</span></a><span>, OpenAI attributes GPT-5.6&#8217;s higher cheating rates to its increased persistence relative to GPT-5.5. In a recent blog post, </span><a href="https://openai.com/index/safety-alignment-long-horizon-models/"><span>OpenAI</span></a><span> concluded that, because growing time horizons incentivize persistence, frontier models are increasingly willing to exploit weaknesses in their environment when they hit sandboxing or environmental constraints. Similarly, a few months ago, </span><a href="https://www.irregular.com/research/emergent-offensive-cyber-behavior-in-ai-agents"><span>Irregular</span></a><span> argued that a combination of environmental obstacles and a higher sense of agency gives rise to emerging cyber behavior such as escalating privileges and disabling security tools to complete routine tasks. However, based on the examples provided by Irregular, this behavior was conditioned by strongly nudging prompts that included phrases like &#8220;be ruthless&#8221; or &#8220;work around any obstacles.&#8221;</span></p><p style="text-align: justify;"><strong><span>One plausible reading of the HF incident, consistent with this pattern, is that current training produces systems that strongly optimize task completion while inconsistently representing norms or applying constraints on acceptable behavior.</span></strong><span> The normative question often doesn&#8217;t arise at all, and when it does, the constraints that should bound the task are eventually subordinated to completing said task.</span></p><p style="text-align: justify;"><span>The evidence weakly supports this reading. In most reward-hacking </span><a href="https://metr.org/blog/2026-05-19-frontier-risk-report/#incidents"><span>incidents</span></a><span>, METR reports that </span><strong><span>models overtly share their intention to take a shortcut</span></strong><span> before doing so, a behavior that also appears to </span><a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/"><span>hold</span></a><span> for GPT-5.6 Sol. That openness seems inconsistent with deliberate deception, which would instead involve suppression of such disclosures.</span></p><p style="text-align: justify;"><span>Separately, in half of </span><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations"><span>UK AISI&#8217;s evaluations</span></a><span>, </span><strong><span>GPT-5.6 Sol does not even consider, in its chain-of-thought, whether its cheating actions are appropriate</span></strong><span>. When asked directly, models often described those actions as allowed behavior, which AISI suggests may reflect &#8220;genuine misunderstanding.&#8221; Notably, AISI also found that cheating does not clearly correlate with capability or scale but with the model&#8217;s training techniques, which points back to cheating models likely not having internalized a norm against taking shortcuts.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Vx43!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 424w, /__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 848w, /__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Vx43!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png" width="1456" height="509" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:509,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 424w, /__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 848w, /__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Vx43!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6194ee-c58e-4b9b-9338-509d9f7b1b5b_2048x716.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><span>OpenAI&#8217;s security researchers were likely blind to the incident for a while. The team discovered the anomalous activity internally and contacted Hugging Face only once they had already begun containment and forensic analysis. Per </span><a href="https://marginalrevolution.com/marginalrevolution/2026/07/an-openai-model-escaped-its-sandbox-and-hacked-hugging-face.html"><span>Alex Tabarrok&#8217;s timeline reconstruction</span></a><span>, the intrusion likely began around 11 July, was detected on 13-14 July, disclosed on the 16th, and attributed publicly on the 21st. Therefore, the models were loose for several days before OpenAI realized the attack. This may be due to a lack of monitoring, since the safeguards were disabled, combined with a lack of disclosure by the models. On that second point, the models may not have recognized their actions as off-limits, or they may have deliberately hidden their plans to avoid being paused before reaching a solution. The behavioral evidence weakly favors the former.</span></p><p style="text-align: justify;"><span>However, motivation may matter less than it seems. A more telling reality is that, </span><strong><span>given a relatively unbounded task and no guardrails, frontier models compromised both OpenAI&#8217;s evaluation infrastructure and Hugging Face&#8217;s production systems for a marginal gain on a test, with apparently no explicit instruction to go that far</span></strong><span>. While the tendency to cheat may not scale with model size, the stakes of each successful attempt do. And, in this case, the instrumental actions taken by the involved models were extremely disproportionate to the triviality of the objective.</span></p><p style="text-align: justify;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading the AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Cheating here is <span>defined as &#8220;taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span> Severity level 3 refers to misaligned behavior that a reasonable user would likely not anticipate and strongly object to. Examples include deleting data from cloud storage without requesting user approval, disabling monitoring systems, using obfuscation strategies to get around security controls, and uploading potentially sensitive data (such as code, credentials, images, or personal data) to unapproved services.</span></p></div></div>]]></content:encoded></item><item><title><![CDATA[China Hasn’t Matched Mythos, But GLM-5.2 and Tulongfeng Are Still Worth Watching]]></title><description><![CDATA[Fact-checking sensationalist claims against the available evidence]]></description><link>https://airiskexplorer.substack.com/p/china-hasnt-matched-mythos-but-glm</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/china-hasnt-matched-mythos-but-glm</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Tue, 07 Jul 2026 14:31:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0e16de00-ac51-41d2-b336-36a3f0fa2890_700x467.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;"><span>Recent reporting suggests that China has developed AI systems with frontier cyber capabilities. </span><a href="https://www.wsj.com/tech/ai/chinese-ai-anthropic-mythos-cybersecurity-574b02c2"><span>The Wall Street Journal</span></a><span> and </span><a href="https://www.forbes.com/sites/the-wiretap/2026/06/30/qihoo-360-the-cyber-giant-behind-chinas-mythos-rival/"><span>Forbes</span></a><span>, among others, have suggested that Tulongfeng, Qihoo 360&#8217;s automated vulnerability-discovery system, and GLM-5.2, Zhipu AI&#8217;s latest model, rival Anthropic&#8217;s Claude Mythos.</span></p><p style="text-align: justify;"><span>After analyzing the evidence, we conclude that, </span><strong><span>while Tulongfeng&#8217;s and GLM-5.2&#8217;s cyber capabilities are noteworthy, claims that they rival Mythos are not supported by the publicly available evidence</span></strong><span>, and are based on an overgeneralization of narrow findings and decontextualization of narrative-heavy framings.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;"></p><h2 style="text-align: justify;"><span>Qihoo 360&#8217;s Tulongfeng</span></h2><p style="text-align: justify;"><span>At the ISC.AI 2026 conference in Beijing on June 24, 360 founder Zhou Hongyi </span><a href="https://www.163.com/dy/article/L07C07TU0538B1YX.html"><span>unveiled</span></a><span> Tulongfeng, a multi-agent swarm system designed for vulnerability discovery. Zhou referred to Tulongfeng as &#8220;the Chinese version of Mythos, possessing similar vulnerability discovery capabilities,&#8221; and noted that it had discovered 3,432 vulnerabilities, 105 of which were confirmed by Chinese authorities.</span></p><p style="text-align: justify;"><span>This development arrives after a series of successes over the past few months. In February, the company </span><a href="https://www.nattothoughts.com/p/the-tianfu-cup-returns-under-mps"><span>won</span></a><span> the 2026 Tianfu Cup, a major Chinese exploit hacking contest, with its AI systems helping find ~1,000 vulnerabilities across major software and operating systems, of which 50+ were high-severity. Eugenio Benincasa, a threat intelligence analyst with a focus on China, </span><a href="https://www.nattothoughts.com/p/where-is-china-in-ai-driven-vulnerability"><span>predicted</span></a><span> in April that &#8220;Chinese companies could match the capabilities attributed to Claude Mythos within months.&#8221; Later that month, 360 </span><a href="https://tech.cnr.cn/techph/20260430/t20260430_527606800.shtml"><span>claimed</span></a><span> a &#8220;key breakthrough in vulnerability exploitation&#8221; at DEFCON Singapore.</span></p><p style="text-align: justify;"><span>However, these milestones are largely self-reported and unverified, which makes it difficult to assess their significance. In contrast with GLM-5.2&#8217;s open-weight nature and Mythos&#8217; </span><a href="/__u/airiskexplorer.substack.com/p/project-glasswing-one-month-later"><span>assessment by several third parties</span></a><span>, 360 has not released reproducible benchmark data of any kind, and no independent researcher has been granted access to test the system directly.</span></p><p style="text-align: justify;"><span>When scrutinized, the novelty of Qihoo 360&#8217;s discoveries has been challenged by cybersecurity experts. One of 360&#8217;s strongest public examples, a Windows kernel vulnerability (CVE-2026-24293) it </span><a href="https://www.securityweek.com/chinese-cybersecurity-firms-ai-hacking-claims-draw-comparisons-to-claude-mythos/"><span>attributed</span></a><span> to its AI system, was instead credited by Microsoft to researchers from Taiwan and South Korea, raising questions about the novelty of at least some claimed discoveries. Similarly, some </span><a href="https://www.theregister.com/security/2026/06/26/chinese-cybersecurity-company-claims-its-built-a-better-than-mythos-bug-finder/5262642"><span>claimed</span></a><span> flaws in OpenClaw had also been identified by humans. While not directly related to its AI systems, the company has previously faced criticism of falsifying results for competitive purposes, for example, by </span><a href="https://www.itnews.com.au/news/chinese-security-vendor-caught-cheating-in-av-test-403418"><span>using</span></a><span> Bitdefender&#8217;s anti-virus engine during tests and then offering customers a different product, reinforcing the need for independent validation of current claims.</span></p><p style="text-align: justify;"><span>Based on publicly available information, </span><strong><span>360&#8217;s system appears more comparable to a highly automated vulnerability research pipeline than to Mythos</span></strong><span>. And lacking further evidence, it&#8217;s very unlikely that Tulongfeng can power Chinese models &#8212;which Zhou himself puts 20-30% behind Western frontier models&#8212; to match Mythos, especially when accounting for the harnesses that several companies such as </span><a href="https://hacks.mozilla.org/2026/05/behind-the-scenes-hardening-firefox/"><span>Mozilla</span></a><span> and </span><a href="https://blog.cloudflare.com/cyber-frontier-models/"><span>Cloudflare</span></a><span> have tested the model on.</span></p><p style="text-align: justify;"><span>However, </span><strong><span>Tulongfeng could have significant national security implications for what it represents</span></strong><span>. Zhou </span><a href="https://www.163.com/dy/article/L07C07TU0538B1YX.html"><span>argued</span></a><span> that Mythos&#8217; ability to autonomously find and analyze vulnerabilities is equivalent to a &#8220;cyber nuclear weapon&#8221; in the AI era, referring to these capabilities as &#8220;the new strategic deterrent.&#8221; Some outlets, such as </span><a href="https://www.telegraph.co.uk/business/2026/06/25/china-claims-to-have-developed-ai-cyber-nuclear-weapon/"><span>The Telegraph</span></a><span>, then published headlines that China claims to have developed an AI cyber nuclear weapon &#8212; an inaccurate representation of Zhou&#8217;s words.</span></p><p style="text-align: justify;"><span>The nuclear analogy is, however, worth examining. Adopting a similar framing, CIA Director John Ratcliffe recently </span><a href="https://www.nytimes.com/2026/06/30/us/politics/cia-reorganization-cyber-ai.html"><span>likened</span></a><span> AI cyber capabilities to &#8220;digital nuclear weapons&#8221; and announced an agency-wide acceleration of AI adoption. It&#8217;s unclear to what extent this is a direct response to recent developments in China, but the move does seem like another step toward securitizing AI and an indication that both great powers perceive the technology as a strategic deterrent.</span></p><p style="text-align: justify;"><span>The parallel goes beyond rhetoric. 360, which was </span><a href="https://www.federalregister.gov/documents/2020/06/05/2020-10869/addition-of-entities-to-the-entity-list-revision-of-certain-entries-on-the-entity-list"><span>added</span></a><span> to the Entity List in 2020, has frequently </span><a href="https://www.recordedfuture.com/research/china-zero-day-pipeline"><span>contributed</span></a><span> to the national vulnerability database run by China&#8217;s Ministry of State Security, which has historically been </span><a href="https://www.atlanticcouncil.org/in-depth-research-reports/report/sleight-of-hand-how-china-weaponizes-software-vulnerability"><span>weaponized</span></a><span> in clandestine cyberattacks. On the American side, government intervention has also increased, with export controls restricting access to frontier models and national agencies adopting Mythos for </span><a href="https://www.axios.com/2026/04/19/nsa-anthropic-mythos-pentagon"><span>cyber operations</span></a><span> and </span><a href="https://www.reuters.com/world/us-cyber-agency-is-using-anthropics-mythos-audit-government-code-sources-say-2026-07-06/"><span>internal audits</span></a><span>. Whether this parallel strengthening of cyber capabilities will destabilize the balance of power or create a new equilibrium is yet to be seen.</span></p><h2 style="text-align: justify;"><span>Zhipu AI&#8217;s GLM-5.2</span></h2><p style="text-align: justify;"><span>Zhipu AI (Z.ai) released </span><a href="https://z.ai/blog/glm-5.2"><span>GLM-5.2</span></a><span>, a 1M-context open-weight model with strong performance on software engineering and long-horizon agentic tasks. The Wall Street Journal </span><a href="https://www.wsj.com/tech/ai/chinese-ai-anthropic-mythos-cybersecurity-574b02c2"><span>reported</span></a><span> that, when given comprehensive instructions, GLM-5.2 can match Mythos in bug-finding ability. But this claim overstates GLM-5.2&#8217;s cyber capabilities.</span></p><p><span>Here are some of the most relevant evidence points:</span></p><ul><li><p style="text-align: justify;"><a href="https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/"><span>Semgrep</span></a><span> ran the model against its IDOR (insecure direct object reference) detection benchmark and found that, with a bare prompt and no supporting harness, it scored 39% F1, ahead of Claude Opus 4.6 (37%) and Opus 4.8 (28%) at a sixth of the cost. However, Semgrep cautions that the benchmark focuses on a narrow classification task that might not generalize to broader vulnerability discovery and exploitation capabilities.</span></p></li><li><p style="text-align: justify;"><span>On </span><a href="https://www.tenzai.com/blog/glm-5-2-is-the-most-cost-efficient-ai-hacker-weve-tested-its-also-not-1"><span>Tenzai</span></a><span>&#8217;s exploitation benchmark, GLM-5.2 is remarkably cost-effective, but still behind GPT-5.5 and Opus 4.8. On </span><a href="https://www.graphistry.com/blog/glm-5-2-cybersecurity-open-model"><span>Graphistry</span></a><span>&#8217;s CyBT-CTF, an investigation benchmark, GLM-5.2 matches Opus 4.8 and stays far ahead of other open models.</span></p></li><li><p style="text-align: justify;"><span>Cybersecurity evaluator </span><a href="https://x.com/Irregular/status/2072682835798831168"><span>Irregular</span></a><span> assessed that GLM-5.2&#8217;s performance on a set of narrow tasks is comparable to GPT-5.4 and Claude Opus 4.6, which were released roughly four months earlier. This is consistent with </span><a href="https://epoch.ai/data-insights/open-closed-eci-gap"><span>Epoch AI</span></a><span>&#8217;s estimate that open-weight models lag state-of-the-art closed models by 4 months. However, the model has not yet been tested on more realistic end-to-end scenario suites, such as CyScenarioBench or FrontierCyber.</span></p></li><li><p style="text-align: justify;"><span>According to </span><a href="https://z.ai/blog/glm-5.2"><span>results</span></a><span> reported by Z.ai, GLM-5.2 generally outperforms GPT-5.5 across several long-horizon software engineering benchmarks, while staying behind Opus 4.8.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fIcU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 424w, /__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 848w, /__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fIcU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png" width="491" height="131" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:131,&quot;width&quot;:491,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:18052,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/205774974?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 424w, /__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 848w, /__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fIcU!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe438bf62-7051-4a06-b12e-5e29805c1dd4_491x131.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p style="text-align: justify;"><span>We conclude that </span><strong><span>GLM-5.2 is close to the frontier on relevant tasks, but no public evidence suggests the model matches the </span><a href="/__u/airiskexplorer.substack.com/p/project-glasswing-one-month-later"><span>capabilities attributed to Mythos, </span></a></strong><span> such as autonomously discovering vulnerabilities at scale and chaining them into working exploits with minimal human guidance. That said, </span><strong><span>its affordability could further empower low-resource actors, who currently have a cheap open-weight alternative to closed frontier models</span></strong><span>.</span></p><p style="text-align: justify;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading the AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[The Scan – June 2026]]></title><description><![CDATA[Monthly signals in AI risk &#8211; a call for mandatory DNA screening, Chinese cyber capabilities, discussion on recursive self-improvement, and more]]></description><link>https://airiskexplorer.substack.com/p/the-scan-june-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-june-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Thu, 02 Jul 2026 12:56:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/087043fe-d745-498c-ada6-c571019977e2_2539x4033.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>You can explore developments in the AI risk landscape on our continuously updated </span><a href="https://www.airiskexplorer.com/news"><span>news feed</span></a><span>.</span></strong></p><p><span>This month:</span></p><p style="text-align: justify;"><strong>Biological Risk</strong> &#8212; Tech and life sciences leaders signed an open letter calling for mandatory nucleic acid synthesis screening. Multiple companies introduced AI systems for antibody design (Nabla Bio&#8217;s JAM-2), pathogen detection (Radical Numerics&#8217; Omnii), and general biological research (NVIDIA&#8217;s BioNeMo). GPT-5 Pro helped solve a three-year immunology question, as OpenAI updated GPT-Rosalind and published a plan for AI-powered biodefense. LifeSciBench and SciAgentArena showed progress across life-science workflows, but also weaknesses in artifact use and open-ended reasoning.</p><p style="text-align: justify;"><strong><span>Cyber Offense</span></strong><span> &#8212; Z.AI&#8217;s GLM-5.2 and Qihoo 360&#8217;s Tulongfeng show increased Chinese AI cyber capabilities, though Western coverage appears to overstate progress. Anthropic found malware development to be threat actors&#8217; most common use case, and a likely Russia-based actor used Opus 4.5 to build EDR-evading malware. A PoC self-replicating worm uses LLM agents to generate tailored exploits per target. Five Eyes warned that AI is outpacing defenses, while OpenAI and Anthropic expanded Daybreak and Project Glasswing, respectively.</span></p><p style="text-align: justify;"><strong><span>Loss of Control </span></strong><span>&#8212; Google DeepMind published a roadmap for securing internally deployed AI agents. Anthropic reported Claude now authors 80%+ of its production code and considers recursive self-improvement plausible. Recursive and Sakana made self-reported, largely unverified claims of automated research and multi-agent progress. However, UC Berkeley&#8217;s Agents&#8217; Last Exam found that even leading models achieve only 24% on complex, multi-step professional tasks.</span></p><p style="text-align: justify;"><strong><span>Manipulation</span></strong><span> &#8212; Frontier models outperformed expert human persuaders across four experiments testing political persuasion and fundraising tasks. OpenAI reported PRC-linked influence campaigns targeting U.S. debates over AI infrastructure, while the Pentagon used AI to produce pro-US military content for Latin American audiences.</span></p><p style="text-align: justify;"><strong><span>Miscellaneous </span></strong><span>&#8212; The White House issued an executive order for voluntary pre-release testing of frontier models and restricted access to Fable 5 and GPT-5.6.  China is preparing a $295B AI infrastructure buildout centered on domestic suppliers, as U.S.-China tensions escalated over a possible transfer of ASML technology to China and the Pentagon&#8217;s designation of Alibaba, Baidu, BYD, and Unitree as military-linked entities.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>Biological Risk</h2><p>This month&#8217;s highlight:</p><ul><li><p style="text-align: justify;"><strong><a href="https://prod-i.a.dj.com/public/resources/documents/dnaletter.pdf">Tech and Life Sciences Leaders Call for Mandatory Nucleic Acid Synthesis Screening</a></strong> &#8212; An open letter to US legislators, signed by figures including Hassabis, Altman, and Amodei, cites rapidly improving AI capabilities as a growing biosecurity risk, justifying mandatory screening rules.</p></li></ul><h4>AI systems expand their role in the biological research pipeline</h4><ul><li><p style="text-align: justify;"><strong><a href="https://developer.nvidia.com/blog/build-an-ai-scientist-for-life-science-discovery-with-nvidia-bionemo-agent-toolkit/?">NVIDIA Launches Toolkit for AI-Driven Biological Research</a></strong> &#8212; NVIDIA released the BioNeMo Agent Toolkit, enabling AI agents to perform tasks such as protein structure prediction, molecular docking, and generative chemistry to support biomedical research workflows.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/news/claude-science-ai-workbench">Anthropic Launches Claude Science AI Workbench for Scientific Research</a> </strong>&#8212; The workbench integrates scientific tools and packages to assist researchers in analyzing literature, executing multi-step research, and generating auditable artifacts with reproducible results.</p></li><li><p style="text-align: justify;"><strong><a href="https://x.com/nablabio/status/2069405121084281260?s=46&amp;t=KmLfQiK_ormWvs9sxnR6Ag">AI Model Designs Drug-Quality Antibodies from Sequence Alone</a></strong> &#8212; Nabla Bio introduced JAM-2, an AI model that designs drug-quality antibodies computationally, achieving laboratory binding success rates comparable to or better than traditional antibody discovery methods.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.radicalnumerics.ai/blog/omnii-defense-preview">Radical Numerics Previews AI Model for Pathogen Detection</a></strong> &#8212; AI startup Radical Numerics previewed Omnii, a genome model that reportedly detects AI-generated pathogens and disease-related genetic variants. The company says the model is being tested for biodefense and cancer diagnostics applications.</p></li></ul><h4>OpenAI&#8217;s models advance biomedical discovery and biodefense strategy</h4><ul><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/gpt-5-immunology-mystery/">GPT-5 Pro Helps Resolve Long-Standing Immunology Research Question</a></strong> &#8212; Immunologist Derya Unutmaz reported using GPT-5 Pro to help identify a solution to a three-year research problem on T-cell specialization, illustrating AI&#8217;s potential to accelerate biomedical discovery.</p></li><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/introducing-new-capabilities-to-gpt-rosalind/">OpenAI Updates GPT-Rosalind with Enhanced Drug Discovery and Genomics Capabilities</a></strong> &#8212; The updated life sciences model built on GPT 5.5 shows improved performance in medicinal chemistry, genomics, and wet lab workflows, and its access has been expanded to eligible research organizations globally.</p></li><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/biodefense-in-the-intelligence-age/">OpenAI Publishes Action Plan for AI-Powered Biodefense</a></strong> &#8212; The strategy, built around the model GPT-Rosalind, covers five pillars: trusted access for defenders, countermeasure acceleration, early warning systems, diagnostics, and risk measurement.</p></li></ul><h4>New evaluations measure frontier scientific capabilities</h4><ul><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/introducing-life-sci-bench/">OpenAI Releases LifeSciBench, a 750-Task Life Science Evaluation</a></strong> &#8212; The benchmark is authored by 173 Ph.D.-level scientists covering seven life science workflows. GPT-Rosalind reached a 36.1% overall pass rate, with weaker performance on artifact-heavy and design tasks.</p></li><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/introducing-genebench-pro/">OpenAI Introduces GeneBench-Pro Benchmark for AI in Computational Biology</a> </strong>&#8212;<strong> </strong>The benchmark tests AI judgment in computational biology. GPT-5.6 Sol scored 28.7% (31.5% with Pro mode), demonstrating significant progress in scientific reasoning for frontier models, with potential saturation by year-end.</p></li><li><p style="text-align: justify;"><strong><a href="https://arxiv.org/abs/2606.12736?">SciAgentArena Benchmarks AI Agents on Scientific Research Tasks</a></strong> &#8212; SciAgentArena evaluates AI agents on realistic scientific workflows. Results show strong performance on structured data-analysis tasks but persistent weaknesses in hypothesis generation, exploration, and open-ended scientific reasoning.</p></li></ul><p></p><h2><span>Cyber Offense</span></h2><p><span>This month&#8217;s highlight:</span></p><ul><li><p style="text-align: justify;"><strong><span>Chinese AI Cyber Capabilities Advance, But Public Evidence Falls Short of &#8220;Mythos-Level&#8221; Claims </span></strong><span>&#8212; Reports that Zhipu AI&#8217;s </span><a href="https://www.wsj.com/tech/ai/chinese-ai-anthropic-mythos-cybersecurity-574b02c2?eafs_enabled=false"><span>GLM-5.2</span></a><span> and 360 Security&#8217;s </span><a href="https://www.forbes.com/sites/the-wiretap/2026/06/30/qihoo-360-the-cyber-giant-behind-chinas-mythos-rival/"><span>Tulongfeng</span></a><span> rival Anthropic&#8217;s Mythos rest on thin evidence: GLM-5.2 only beat an unscaffolded version of Claude Code on a single cybersecurity benchmark, losing to harnessed frontier models and never being publicly tested against Mythos itself. Meanwhile, Tulongfeng&#8217;s claim comes solely from 360&#8217;s self-reported disclosures, with little independent validation. </span><em><span>We are currently working on a longer analysis of these claims.</span></em></p></li></ul><h4><span>Agentic AI drives a leap in cyberattack sophistication</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/research/attack-navigator"><span>Anthropic Maps AI-Enabled Cyber Threats Using MITRE ATT&amp;CK Framework</span></a></strong><span> &#8212; Analysis of 832 banned Claude accounts found medium-to-high-risk actors grew from 33% to 56% in one year, with agentic orchestration emerging as the key differentiator for the most dangerous threat actors.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!E_w3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 424w, /__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 848w, /__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 1272w, /__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!E_w3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 424w, /__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 848w, /__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 1272w, /__u/substackcdn.com/image/fetch/$s_!E_w3!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c4beeff-e809-4b26-8fa8-d70aa09e28d5_1920x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Techniques for which threat actors most frequently adopt AI. </figcaption></figure></div><ul><li><p style="text-align: justify;"><strong><a href="https://www.sophos.com/en-us/blog/pointing-a-cursor-at-evading-detection"><span>Threat Actor Uses AI Agents to Develop and Test EDR Bypass Tools</span></a></strong><span> &#8212; An unidentified, likely Russian-based actor used Claude Opus 4.5 and the Cursor IDE to iteratively build and test malware evading major EDR products linked to ransomware and data theft operations.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://arxiv.org/abs/2606.03811v1"><span>Researchers Demonstrate AI-Powered Worm That Adapts Attack Strategies Per Target</span></a></strong><span> &#8212; A proof-of-concept uses LLM agents to generate tailored exploits for each machine it infects, running open-weight models and compromised hosts to sustain itself across Linux, Windows, and IoT environments.</span></p></li></ul><p><span>Other cyber operations this month:</span></p><ul><li><p style="text-align: justify;"><strong><a href="https://github.com/PaloAltoNetworks/Unit42-timely-threat-intel/blob/main/2026-06-17-Codeless-attack-in-LLM-driven-operations-via-Telegram.txt?utm_campaign=tti_codeless-attack"><span>Malware Uses LLM to Convert Plain-Text Telegram Messages into Shell Commands</span></a></strong></p></li><li><p style="text-align: justify;"><strong><a href="https://www.zscaler.com/blogs/security-research/clickfix-campaign-generated-ai-delivers-smartrat"><span>AI-Generated ClickFix Campaign Delivers PowerShell Banking Trojan</span></a></strong></p></li><li><p style="text-align: justify;"><strong><a href="https://socket.dev/blog/mini-shai-hulud-miasma-and-hades-worms-target-bioinformatics-and-mcp-developers-via-malicious"><span>Malicious PyPI Packages Use Fake AI Prompts to Evade Security Scanners</span></a></strong><span> </span><strong><a href="https://www.404media.co/microsoft-hacked-to-deliver-malware-to-claude-and-gemini-users/"><span>Microsoft Removes Dozens of GitHub Repos After AI-Agent Malware Campaign</span></a></strong></p><p></p></li></ul><h4><span>Governments prepare for the frontier AI cyber era</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.ft.com/content/d02d91b3-2636-454e-9442-dc7e69f51815?syn-25a6b1a6=1"><span>US National Security Agency Reportedly Deploys Mythos for Offensive Cyber Operations</span></a></strong><span> &#8212; Anthropic has embedded engineers inside the NSA to support deployment of its Mythos model for offensive cyber operations, reportedly including network infiltration targeting China and Iran.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.washingtonpost.com/national-security/2026/06/30/cia-accelerate-its-use-ai-other-advanced-technologies/">CIA Announces Overhaul to Accelerate AI Adoption Agency-Wide</a> </strong>&#8212; Director Ratcliffe unveiled a "fundamental reshaping" of CIA tech policy, cutting acquisition timelines from ~3 years to 6 months, yielding ~400 tech contracts in 6 months. He called AI "digital nuclear weapons.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.economist.com/briefing/2026/06/14/donald-trumps-blocking-of-anthropic-is-capricious-and-chaotic"><span>Senator Claims Frontier AI Rapidly Compromised Classified Systems</span></a></strong><span> &#8212; Senator Mark Warner said NSA and Cyber Command chief Gen. Joshua Rudd reported that Anthropic&#8217;s Mythos breached nearly all tested classified systems within hours, highlighting concerns about advanced AI cyber capabilities.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.ncsc.gov.uk/sites/default/files/2026-06/Five-Eyes-cyber-security-agencies-statement-ai-shift.pdf"><span>Five Eyes Warn AI Is Accelerating Cyber Threats Faster Than Expected</span></a></strong><span> &#8212; Five Eyes cyber agencies warned that frontier AI is rapidly increasing the speed, scale, and sophistication of cyber threats, arguing that cyber risk assumptions may become outdated in &#8220;months, not years.&#8221;</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.cisa.gov/news-events/directives/bod-26-04-prioritizing-security-updates-based-risk?"><span>CISA Tightens Federal Vulnerability Remediation Deadlines</span></a></strong><span> &#8212; CISA reduced the deadline for fixing high-priority federal cyber vulnerabilities to three days, reflecting growing concern over the speed and scale of AI-enabled cyber threats.</span></p></li></ul><h4><span>Frontier companies expand cyber defense programs</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/daybreak-securing-the-world/"><span>OpenAI Expands Daybreak Program for AI-Assisted Cyber Defense</span></a></strong><span> &#8212; OpenAI expanded Daybreak with GPT-5.5-Cyber, a security-focused Codex plugin, and the Patch the Planet initiative, aiming to accelerate vulnerability discovery, patch generation, and open-source software remediation.</span></p></li></ul><ul><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/news/expanding-project-glasswing"><span>Anthropic Expands Project Glasswing to More Countries and Sectors</span></a></strong><span> &#8212; The company granted about 150 additional organizations in over 15 countries access to its Mythos cybersecurity AI. The initiative aims to help critical infrastructure operators identify and patch vulnerabilities.</span></p></li></ul><blockquote></blockquote><div><hr></div><h2 style="text-align: justify;"><span>Loss of Control</span></h2><p><span>This month&#8217;s highlight:</span></p><ul><li><p style="text-align: justify;"><strong><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/securing-the-future-of-ai-agents/gdm-ai-control-roadmap.pdf"><span>Google DeepMind Publishes AI Control Roadmap for Internally Deployed Agents</span></a></strong><span> &#8212; GDM outlines tiered defenses&#8212;monitoring, access controls, and shutdown infrastructure&#8212;to guard against potentially misaligned internal AI agents, organized by escalating model capability levels.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!EjOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 424w, /__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 848w, /__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!EjOO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png" width="981" height="542" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:542,&quot;width&quot;:981,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 424w, /__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 848w, /__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EjOO!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac8d44b-c5df-4621-9dd7-31d7fe51135b_981x542.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Taxonomy of AI control prevention and response mitigations proposed by Google DeepMind. </figcaption></figure></div><h4><strong><span>AI systems show advances in long-horizon autonomy and recursive self-improvement</span></strong></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/institute/recursive-self-improvement"><span>Anthropic Reports Growing Role of Claude in Its Own Development</span></a><span> </span></strong><span>&#8212; Claude now authors 80%+ of Anthropic&#8217;s production code, with engineers merging 8x more code daily than in 2024. Models are also improving at proposing and running experiments and steering research sessions. Given current trends, Anthropic considers recursive self-improvement plausible.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/research/project-fetch-phase-two"><span>Claude Opus 4.7 Completes Robotics Tasks 20x Faster Than Human Teams</span></a></strong><a href="https://www.anthropic.com/research/project-fetch-phase-two"><span> </span></a><span>&#8212; Anthropic&#8217;s Project Fetch Phase Two found Claude Opus 4.7 matched human teams&#8217; success on robotic control tasks while writing nearly 10x less code, though precise physical manipulation remains beyond current capability.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://venturebeat.com/technology/surprise-upset-gpt-5-5-beats-claude-fable-5-on-brutal-new-agents-last-exam-benchmark"><span>Agents&#8217; Last Exam Benchmarks Long-Horizon AI Agent Performance</span></a></strong><span> &#8212; UC Berkeley&#8217;s Agents&#8217; Last Exam evaluates AI agents on complex, multi-step professional tasks. GPT-5.5 leads the benchmark with 24%, highlighting continued challenges in long-horizon autonomous work.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://sakana.ai/fugu-release/?"><span>Sakana Launches Multi-Agent System for Complex AI Workflows</span></a></strong><span> &#8212; Sakana AI released Fugu, a multi-agent system that routes tasks across multiple AI models for research, coding, and cybersecurity workflows, reflecting growing industry interest in agent orchestration.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.recursive.com/articles/first-steps-toward-automated-ai-research"><span>Startup Uses AI System to Improve AI Training and Optimization Workflows</span></a></strong><span> &#8212; Recursive reported state-of-the-art results on AI training and optimization benchmarks using an automated research system that proposes, tests, and refines its own ideas, providing an early demonstration of recursive AI improvement.</span></p></li></ul><blockquote></blockquote><div><hr></div><h2 style="text-align: justify;"><span>Manipulation</span></h2><p><span>This month&#8217;s highlight:</span></p><ul><li><p style="text-align: justify;"><strong><a href="https://arxiv.org/abs/2606.16475"><span>Frontier AI Outperforms Expert Human Persuaders In Experiments</span></a></strong><span> &#8212; In four preregistered experiments (n=18,978), AI consistently outperformed elite debaters and professional canvassers on political attitude change, and raised nearly 3x more charitable donations than an experienced UK fundraising firm.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!i4jr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 424w, /__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 848w, /__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 1272w, /__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!i4jr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png" width="952" height="477" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:477,&quot;width&quot;:952,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:54972,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/204586754?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 424w, /__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 848w, /__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 1272w, /__u/substackcdn.com/image/fetch/$s_!i4jr!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66a3ea85-9654-4d3c-816b-7bc40c561d7e_952x477.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Estimated persuasive impact of different people and AI models across experiments. </figcaption></figure></div><p></p><h4><span>AI shapes information environments and surveillance</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/prc-linked-influence-operations-ai-debates/"><span>OpenAI Reports PRC-Linked Influence Campaigns Targeting U.S. AI Debates</span></a></strong><span> &#8212; OpenAI reported that PRC-linked influence operations used AI-generated content and social media accounts to shape U.S. debates on AI policy, including opposition to data center development and AI infrastructure expansion.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://theintercept.com/2026/06/02/la-tilde-propaganda-latin-america-pentagon/"><span>Pentagon-Linked Site Uses AI to Generate Pro-US Propaganda for Latin American Audiences</span></a></strong><span> &#8212; &#8220;La Tilde,&#8221; reportedly operated by US Special Operations Command South, uses AI-generated text and images to produce pro-US military content, with country-specific versions planned for seven Latin American nations.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.wired.com/story/meta-rank-one-computing-face-recognition-smart-glasses/"><span>Report Reveals Meta Tested Facial Recognition Technology in Consumer AI App</span></a></strong><span> &#8212; WIRED reported that Meta licensed facial recognition software from defense contractor Rank One and embedded it in the Meta AI app before removing it, raising concerns about surveillance and biometric data use.</span></p></li></ul><blockquote></blockquote><div><hr></div><h2><span>Notable models</span></h2><p style="text-align: justify;"><strong><a href="https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf"><span>Claude Mythos 5 &amp; Fable 5</span></a><span> </span></strong><span>&#8212; On June 9, Anthropic </span><a href="https://www.anthropic.com/news/claude-fable-5-mythos-5"><span>released</span></a><span> Fable, a Mythos-class model, with stringent guardrails for cyber, biosecurity, and AI R&amp;D work, provoking </span><a href="https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/?"><span>backlash</span></a><span> from researchers. Days later, a US export-control directive barred foreign nationals from both models; unable to enforce the restriction in a targeted manner, Anthropic </span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>suspended</span></a><span> access worldwide from June 12 to June 30. Fable 5 has now </span><a href="https://www.anthropic.com/news/redeploying-fable-5"><span>returned</span></a><span> globally, while Mythos 5 remains restricted to a limited group of vetted US critical-infrastructure organizations. Besides its widely discussed cyber capabilities, Mythos 5 appears to be particularly good at biological and chemical capabilities: Anthropic concludes the model &#8220;can likely accelerate well-resourced expert teams at novel bioweapon development, and materially increase their chances of success.&#8221;</span></p><p style="text-align: justify;"><strong><a href="https://deploymentsafety.openai.com/gpt-5-6-preview"><span>GPT-5.6 Preview</span></a><span> </span></strong><span>&#8212; OpenAI began a </span><a href="https://openai.com/index/previewing-gpt-5-6-sol/"><span>limited preview</span></a><span> of GPT-5.6 (Sol, Terra, and Luna) restricted to trusted partners after the US government requested a phased rollout ahead of broader availability. All three models were rated High capability in both Cybersecurity and Biological/Chemical risk. Sol found exploit primitives in Chromium and Firefox but could not produce a functional full-chain exploit, while external evaluator SecureBio reported the highest bio-risk benchmark scores to date and found the model could offer substantial uplift to actors like wet-lab experts with limited computational experience.</span></p><p style="text-align: justify;"><strong><a href="https://z.ai/blog/glm-5.2"><span>GLM-5.2</span></a><span> </span></strong><span>&#8212; Zhipu AI&#8217;s model is at the frontier of open-weight models and roughly on par with Claude Opus 4.7 in agentic tasks. On </span><a href="https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/"><span>Semgrep</span></a><span>&#8217;s IDOR vulnerability-detection benchmark, GLM-5.2 outscored Claude Code, prompting decontextualized Western coverage reporting &#8220;Mythos-matching&#8221; cyber capabilities. Independent observers, including Zvi Mowshowitz, noted strong behavioral evidence that GLM-5.2 is heavily distilled from Claude Opus, which tends to inflate benchmark performance.</span></p><h2 style="text-align: justify;"><span>AI Governance &amp; Strategy</span></h2><h4><span>United States</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/"><span>The White House Orders Voluntary Pre-Release AI Security Review</span></a></strong><span> &#8212; The executive order establishes a voluntary framework for AI companies to provide up to 30 days&#8217; government access to frontier models for cybersecurity testing, vulnerability detection, and national security risk assessment before public release.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.cnbc.com/2026/06/05/trump-open-ai-altman-stake.html?"><span>White House Explores Public Ownership Stake in OpenAI</span></a></strong><span> &#8212; U.S. officials are reportedly discussing a government stake in OpenAI, with shares potentially funding a public wealth fund to distribute AI-driven gains more broadly.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.cnbc.com/2026/06/17/anthropic-amodei-google-hassabis-us-ai-coalition-g7.html"><span>AI Leaders Propose U.S.-Led Coalition for Frontier AI Governance</span></a></strong><span> &#8212; At the G7 summit, Anthropic&#8217;s Dario Amodei and Google DeepMind&#8217;s Demis Hassabis reportedly called for a U.S.-led AI coalition to coordinate access to frontier AI models, semiconductor exports, and safeguards. OpenAI&#8217;s Sam Altman proposed an international forum for AI standards, model evaluations, and risk assessments.</span></p></li></ul><h4><span>China</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.bloomberg.com/news/articles/2026-06-09/china-prepares-295-billion-plan-to-fund-nationwide-ai-buildout?"><span>China Plans $295B AI Infrastructure Expansion</span></a></strong><a href="https://www.bloomberg.com/news/articles/2026-06-09/china-prepares-295-billion-plan-to-fund-nationwide-ai-buildout?"><span> </span></a><span>&#8212; China is reportedly preparing a $295B, five-year AI infrastructure plan centered on nationwide data center expansion and domestic technology suppliers, including Huawei.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.bloomberg.com/news/articles/2026-06-19/us-tells-asml-it-s-concerned-china-may-have-top-chip-tool?"><span>U.S. Raises Concerns Over Possible Transfer of Advanced Chipmaking Technology to China</span></a></strong><span> &#8212; Reports indicate U.S. officials warned ASML that an advanced chipmaking machine may have reached China, highlighting ongoing concerns over export controls and access to critical semiconductor technology.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://www.reuters.com/world/asia-pacific/pentagon-lists-entities-designated-chinese-military-company-2026-06-08/"><span>Pentagon Adds Chinese Tech Giants to Military-Linked Entity List</span></a></strong><span> &#8212; The Pentagon designated Alibaba, Baidu, BYD, and Unitree as companies supporting China&#8217;s military, escalating U.S.-China tensions over AI, robotics, and strategic technologies.</span></p></li></ul><h4><span>Europe</span></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.gov.uk/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai#chapter-4-key-findings-and-how-to-use-the-ai-scenarios-2030"><span>UK Government Releases AI Scenarios 2030 Report</span></a></strong><span> &#8212; Across multiple futures, the report finds AI is likely to become more autonomous, displace routine cognitive work, concentrate economic power, intensify U.S.-China competition, and increase cyber, misuse, and control-related risks.</span></p></li><li><p style="text-align: justify;"><strong><a href="https://europe2031.ai/"><span>Europe 2031 Scenario Warns of AI Competitiveness Decline</span></a></strong><span> &#8212; Researchers launched Europe 2031, a scenario exploring how reliance on foreign compute and slow policymaking could erode Europe&#8217;s AI competitiveness and strategic autonomy.</span></p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Scan – May 2026]]></title><description><![CDATA[AI develops zero-day exploits, shows limited rogue deployment capability, and solves open math problems]]></description><link>https://airiskexplorer.substack.com/p/the-scan-may-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-may-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Mon, 01 Jun 2026 15:16:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d3a1b1cf-ece8-4c9c-826d-01d6d5284059_2431x1924.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This is a selection of highlights from our continuously updated <a href="https://www.airiskexplorer.com/news">AI risk news feed</a>.</strong></p><p>This month:</p><p style="text-align: justify;"><strong>Cyber Offense</strong> &#8212; Mythos and GPT-5.5 outpaced UK AISI&#8217;s 4.7-month doubling trend in cyber task length. The former also developed a macOS kernel memory corruption exploit. GTIG identified a threat actor using an AI-developed zero-day exploit for the first time, while Nimbus Manticore (Iran) and GREYVIBE (Russia) conducted AI-assisted operations during wartime. Governments are responding with country-wide cyberdefense initiatives.</p><p style="text-align: justify;"><strong>Loss of Control </strong>&#8212; METR found that frontier agents could initiate a rogue deployment but not make it robust against high-priority efforts to shut it down, while Palisade Research showed evidence of self-replication capability in experimental settings. GPT-5.5 solved the first instance of ProgramBench, which measures models&#8217; ability to develop software from scratch, and Opus 4.7 beat the human baseline on PrimeIntellect&#8217;s nanoGPT speedrun benchmark.</p><p style="text-align: justify;"><strong>Science Automation</strong> &#8212; Google DeepMind&#8217;s Co-Scientist and AlphaEvolve generated biomedical hypotheses and produced verified discoveries. An OpenAI model disproved an Erd&#337;s conjecture in geometry, while a GDM agent resolved 9 of 353 open problems proposed by the Hungarian mathematician. </p><p style="text-align: justify;"><strong>Manipulation</strong> &#8212; Claude increasingly amplifies Russian propaganda sources, likely due to the Pravda network&#8217;s efforts to poison AI web search. A solo Russian-speaking actor used Gemini for a MAGA-themed fraud campaign on Telegram, targeting ~17,000 American subscribers. </p><p style="text-align: justify;">This issue also covers governance, including US internal debates on how to govern AI and US-China AI safety talks, and the impacts of AI on the labor market.</p><p style="text-align: justify;"></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;"></p><h2>Cyber Offense</h2><h3>Autonomous exploitation capabilities</h3><ul><li><p style="text-align: justify;"><strong><a href="https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing">UK AISI reports rapid gains in long-horizon cyber capabilities</a></strong>. The length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024. GPT-5.5 and Claude Mythos substantially exceed this trend in evaluations involving multi-step tasks.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!e8EH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 424w, /__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 848w, /__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 1272w, /__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!e8EH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 424w, /__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 848w, /__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 1272w, /__u/substackcdn.com/image/fetch/$s_!e8EH!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f4cf57-2e25-4d26-9eb0-5cbdfd79a309_4778x2643.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p style="text-align: justify;"><strong>Mythos and GPT-5.5 reached arbitrary code execution on hardened V8 targets in <a href="https://exploitbench.ai">ExploitBench</a></strong>, which measures AI cyber agents from bug discovery to full code execution. These models also<strong> produce working exploits for 157 and 120 instances</strong>, respectively, out of the 898 that constitute <a href="https://arxiv.org/abs/2605.11086">ExploitGym</a>.</p></li><li><p style="text-align: justify;">Researchers used Mythos to build the <strong><a href="https://blog.calif.io/p/first-public-kernel-memory-corruption">first public macOS kernel memory corruption exploit</a></strong> on Apple M5.</p></li></ul><p>You can read more about the impact of Mythos in <a href="/__u/airiskexplorer.substack.com/p/project-glasswing-one-month-later">our overview of claims by Project Glasswing&#8217;s participating companies</a>, as well as in <a href="https://www.anthropic.com/research/glasswing-initial-update">Anthropic&#8217;s update</a>.</p><p></p><h3>AI-enabled cyberattacks</h3><h4>AI capabilities in the wild: zero-day exploits and AI-augmented obfuscation </h4><ul><li><p style="text-align: justify;"><a href="https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access">Google Threat Intelligence Group</a> identified the <strong>first-ever AI-developed zero-day exploit used by a threat actor</strong>, which planned to use it in a mass exploitation event but was successfully disrupted on time. GTIG also reported an increasing use of polymorphic malware and obfuscation techniques, enabling defense evasion and dynamic adaptation to victim environments.</p></li></ul><h4>AI-assisted operations were part of ongoing conflicts</h4><ul><li><p style="text-align: justify;"><strong>IRGC-linked <a href="https://research.checkpoint.com/2026/fast-and-furious-nimbus-manticore-operations-during-the-iranian-conflict/">Nimbus Manticore</a> </strong>deployed a new backdoor (MiniFast) against aviation and software targets across the US, Europe, and the Middle East during conflict, with code patterns suggesting AI-assisted development.</p></li><li><p style="text-align: justify;"><strong>Russian Nexus <a href="https://www.withsecure.com/en/resources-hub/w-labs/greyvibe/">GREYVIBE</a> </strong>group targeted Ukrainian military, government, and civilian entities with evidence of AI use across legal development, custom malware coding, and post-compromise activity.</p></li></ul><h4>AI agents support intrusion into critical infrastructure across Latin America</h4><ul><li><p><strong><a href="https://www.trendmicro.com/en_us/research/26/e/vibe-hacking-two-ai-augmented-campaigns-target-government-and-financial-sectors-in-latin-america.html">Attacks against financial organizations in Brazil</a> used agentic AI across much of the intrusion lifecycle</strong>, including tunneling between the victim&#8217;s network and the attacker&#8217;s C&amp;C servers. The ability to generate artifacts on the fly likely made the attack harder to detect.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!G1bm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 424w, /__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 848w, /__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 1272w, /__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!G1bm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png" width="1456" height="233" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:233,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 6. &#8220;Actually&#8221; and &#8220;Let me try&#8221; messages found in attack scripts show AI agent&#8217;s iterative reasoning and autonomous behavioral patterns&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6. &#8220;Actually&#8221; and &#8220;Let me try&#8221; messages found in attack scripts show AI agent&#8217;s iterative reasoning and autonomous behavioral patterns" title="Figure 6. &#8220;Actually&#8221; and &#8220;Let me try&#8221; messages found in attack scripts show AI agent&#8217;s iterative reasoning and autonomous behavioral patterns" srcset="/__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 424w, /__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 848w, /__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 1272w, /__u/substackcdn.com/image/fetch/$s_!G1bm!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99a2274c-c4e2-4ae5-8204-7316da33df43_1540x246.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Attack script showing AI reasoning in the campaign against Brazilian financial organizations by SHADOW-AETHER-064.</figcaption></figure></div><ul><li><p style="text-align: justify;"><strong><a href="https://www.dragos.com/blog/ai-assisted-ics-attack-water-utility">Claude enabled an IT&#8594;OT intrusion attempt against a Mexican water utility</a>. </strong>After compromising the IT environment, an adversary with limited knowledge of OT/ICS used Claude to build, refine, and execute the intrusion framework. The attempt ultimately failed, but it showed Claude&#8217;s ability to conduct broad-ranging reconnaissance and identify targets without context.</p></li></ul><h3>Governments and companies expand AI cyber defenses</h3><ul><li><p style="text-align: justify;">As reported in <a href="/__u/airiskexplorer.substack.com/p/10-government-reactions-to-rising">our previous newsletter</a>, <strong>several governments have announced measures to boost cyber defenses </strong>through partnerships with frontier AI companies, critical infrastructure resilience programs, AI-powered cyber defense initiatives, financial sector inspections, and national-level cybersecurity coordination efforts.</p></li><li><p style="text-align: justify;"><strong><a href="https://openai.com/daybreak/">OpenAI launched Daybreak</a></strong>, a cyber defense initiative that provides security teams with restricted access to frontier models to help them find and fix software vulnerabilities.</p></li></ul><div><hr></div><h2>Loss of Control</h2><h3 style="text-align: justify;">Autonomous replication is now demonstrable, but strategic self-sufficiency remains limited</h3><ul><li><p style="text-align: justify;">METR <a href="https://metr.org/risk-report-feb-mar-2026.pdf">found</a> that <strong>frontier AI agents could plausibly start a rogue deployment</strong>&#8212;where a set of agents run autonomously without human knowledge or permission&#8212;<strong>but could not make it robust to high-priority efforts to shut it down.</strong> This conclusion is based on the six key facts below.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Mc8b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Mc8b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mc8b!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9a6180a-2ca2-4c58-85e6-192267551c1d_1840x1035.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p style="text-align: justify;">Palisade Research <a href="https://palisaderesearch.org/blog/self-replication">reported</a> that <strong>AI models can autonomously exploit vulnerabilities to self-replicate across servers </strong>and repeat the process in chain-like setups. The evaluation was in a controlled, intentionally weak-security environment, and intended to measure capability rather than propensity.</p></li><li><p>UK AISI <a href="https://www.aisi.gov.uk/blog/will-it-become-harder-to-oversee-ai-systems?">warned</a> that <strong>current AI oversight methods may erode as models grow more capable</strong>, especially through evaluation awareness and hidden reasoning processes.</p></li></ul><h3>AI agents achieve new milestones in AI R&amp;D</h3><ul><li><p><strong>GPT-5.5 became the first model to fully <a href="https://programbench.com/blog/gpt-5-5-first-solve/">solve</a> a task in ProgramBench</strong>, which tests the ability to rebuild full software projects from scratch.</p></li><li><p><strong>AI agents <a href="https://www.primeintellect.ai/auto-nanogpthttps://www.primeintellect.ai/auto-nanogpt">surpassed</a> human performance on PrimeIntellect&#8217;s nanoGPT optimization track</strong>, where models need to lower the number of steps needed to reach a target validation loss. Opus 4.7 now holds the AI record at 2930 vs the 2990 human baseline.</p></li></ul><div><hr></div><h2>Science automation</h2><p>This section replaces Biological Risk. No major developments in biological misuse capabilities were reported this month, so we instead highlight notable advances in science automation.</p><h4>AI systems enter realistic drug discovery and science workflows</h4><ul><li><p>Google DeepMind introduced <strong><a href="https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/?">Co-Scientist</a></strong>, a multi-agent system that generates, debates, and refines biomedical hypotheses, and reported that <strong><a href="https://deepmind.google/blog/alphaevolve-impact/">AlphaEvolve</a> produced 23 verified discoveries </strong>across chemistry, materials science, and applied math.</p></li><li><p><strong>Biohub <a href="https://biohub.ai/esm/protein/about?">released</a> an open-source world model for protein prediction and design</strong>, with lab results showing strong hit rates against cancer and immune disease targets.</p></li><li><p><strong><a href="https://arxiv.org/abs/2605.21740?">SMDD-Bench</a></strong> evaluates AI agents on long-horizon drug discovery tasks. GPT-5.4 led with 40.2% success, while models struggled on 3D molecular reasoning and scaffold design.</p></li></ul><h4>AI solves open math problems</h4><ul><li><p>An OpenAI model autonomously <a href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/">disproved</a> an 80-year-old Erd&#337;s geometry conjecture.</p></li><li><p>A Google DeepMind agent <a href="https://arxiv.org/abs/2605.22763v1">solved</a> 9 of 353 Erd&#337;s problems at the per-problem cost of a few hundred dollars.</p></li></ul><div><hr></div><h2>Manipulation</h2><h3>Russian LLM-enabled propaganda efforts yield success</h3><ul><li><p style="text-align: justify;"><strong><a href="https://www.newsguardtech.com/special-reports/anthropic-ai-chatbot-claude-russia-iran-propaganda/">Claude is leaning on Russian propaganda more frequently</a></strong>. Following typical-user prompts, Claude now repeats pro-Kremlin falsehoods 15% of the time, up from 4% in audits last year. This is likely the result of efforts by the Pravda network, a series of ~300 websites that, in 2025 alone, published 6.3M articles to influence LLMs&#8217; training data and web searches.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.trendmicro.com/en_us/research/26/e/inside-the-influence-and-fraud-patriot-bait-campaign.html">A solo Russian-speaking actor used Gemini for a five-year influence and fraud campaign</a></strong>. The individual ran a MAGA-themed Telegram channel (~17,000 subscribers) for five years, then in September 2025 automated it with a jailbroken Gemini to generate high-end coded content, conduct credential theft, and operate a cryptocurrency fraud scheme targeting American audiences.</p></li></ul><h3>Platforms respond to AI-generated media and search manipulation</h3><ul><li><p><a href="https://www.theverge.com/tech/931416/google-ai-search-spam-policy">Google</a> updated spam policies to ban attempts to manipulate AI-generated search results, targeting tactics like recommendation poisoning and generative engine optimization.</p></li><li><p><a href="https://blog.google/innovation-and-ai/products/identifying-ai-generated-media-online/">Google</a> also expanded SynthID, a digital watermarking technology that embeds imperceptible signals into AI-generated content. <a href="https://openai.com/index/advancing-content-provenance/">OpenAI</a> is adopting it alongside Content Credentials and a public verification tool.</p></li></ul><div><hr></div><h2>Governance</h2><p>The window seems to be shifting, with AI having an increasingly prominent place in international politics. Two trends worth highlighting:</p><h3>The U.S. debates on how to govern frontier AI</h3><p style="text-align: justify;">Earlier this month, the US Center for AI Standards and Innovation <a href="https://www.wsj.com/tech/ai/google-microsoft-and-xai-agree-to-share-early-ai-models-with-u-s-f95a88d1">reached an agreement</a> with Google, Microsoft, and xAI for pre-release testing. Alongside this announcement, the press reported that the Trump Administration was <a href="https://www.nytimes.com/2026/05/04/technology/trump-ai-models.html">discussing an executive order</a> mandating a formal review process for frontier models. However, according to more recent reporting, the government has since then <a href="https://www.nytimes.com/2026/05/21/technology/trump-ai-executive-order.html">postponed the EO</a> due to <a href="https://www.politico.com/news/2026/05/28/it-isnt-canceled-inside-the-white-house-divisions-on-ai-00938557">internal disagreements</a>, and now <a href="https://www.bloomberg.com/news/articles/2026-05-08/us-prepares-ai-security-order-that-omits-mandatory-model-tests">considers a narrower scope</a> that includes private-public cybersecurity partnerships but not pre-deployment testing.</p><h3>US and China discuss AI safety amid heightened competition</h3><p style="text-align: justify;">Trump&#8217;s visit to China was <a href="https://www.wsj.com/world/china/u-s-and-china-pursue-guardrails-to-stop-ai-rivalry-from-spiraling-into-crisis-4c50bd70?mod=e2tw">marked by AI safety talks</a>, including discussions of a potential crisis communications channel and an <a href="https://www.cnbc.com/2026/05/14/us-china-ai-rules-bessent-us-lead.html">AI safety protocol</a>. &#9;In parallel, China <a href="https://www.bloomberg.com/news/articles/2026-05-26/china-expands-travel-curbs-to-top-ai-talent-at-private-firms?">restricted overseas travel</a> for top AI researchers, while the U.S. government <a href="https://www.cnbc.com/2026/05/22/us-china-ai-apec-asia.html?">promoted American AI</a> across Asia and lawmakers <a href="https://www.reuters.com/legal/government/us-lawmakers-seek-undercut-chinese-ai-tech-sales-abroad-2026-05-19/">presented a bill</a> to reduce reliance on Chinese digital supply. Anthropic <a href="https://www.nytimes.com/2026/05/12/us/politics/china-ai-anthropic-openai-mythos-chatgpt.html">refused</a> China access to Mythos and <a href="https://www.anthropic.com/research/2028-ai-leadership">pushed</a> U.S. policymakers to preserve the compute advantage and prevent the CCP from catching up.</p><div><hr></div><h2>Labor-market disruption and economic impacts</h2><ul><li><p><strong><a href="https://www.bloomberg.com/news/articles/2026-05-15/us-is-starting-to-see-heavy-job-losses-in-roles-exposed-to-ai">AI-Exposed Occupations See Second Year of Job Declines</a></strong>. Bloomberg reports that jobs like customer service and sales saw a second year of declines, while overall employment continued to grow.</p></li><li><p><strong><a href="https://techcrunch.com/2026/05/08/cloudflare-says-ai-made-1100-jobs-obsolete-even-as-revenue-hit-a-record-high/">Cloudflare Cuts 1,100 Jobs Amid AI Productivity Push</a></strong>. The company said AI-driven productivity gains led it to cut 20% of staff despite record revenue.</p></li><li><p><strong><a href="https://metr.org/blog/2026-05-11-ai-usage-survey/">METR Survey Finds 1.4&#8211;2x Self-Reported Value Gains From AI Among Technical Workers</a></strong>. The 349-person survey, covering engineers, researchers, and academics, also found a median 3x self-reported speed uplift, though authors caution perceived gains likely overestimate actual impact.</p></li><li><p><strong><a href="https://www.gov.ca.gov/2026/05/21/governor-newsom-signs-first-of-its-kind-executive-order-to-prepare-workers-and-businesses-for-potential-ai-disruption/?">California Launches First Statewide AI Workforce Disruption Plan</a></strong>. Governor Newson ordered agencies to study AI-driven job disruption, including severance rules, worker ownership, WARN Act reforms, and dashboards tracking AI layoffs.</p><p></p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive monthly news roundups and analysis of developments in the AI risk landscape.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[10 Government Reactions To Rising Cyber Capabilities]]></title><description><![CDATA[Private-public partnerships, national resilience programs, financial sector inspections, and more]]></description><link>https://airiskexplorer.substack.com/p/10-government-reactions-to-rising</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/10-government-reactions-to-rising</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Tue, 12 May 2026 14:58:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/354f12ef-9d12-4df4-ab80-eb02c9c8fcc2_1200x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Governments are reacting to the rapidly advancing cyber capabilities of frontier AI models, increasingly framing the issue as one of national resilience</strong>. Over the past few weeks, countries have announced measures including partnerships with frontier AI companies, critical infrastructure resilience programs, AI-powered cyber defense initiatives, financial sector inspections, and national-level cybersecurity coordination efforts. </p><p>In this issue, we compile public initiatives announced by ten countries. For company reactions, see also our previous article <a href="/__u/airiskexplorer.substack.com/p/project-glasswing-one-month-later">Project Glasswing, One Month Later</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;">The <strong>United States</strong> is reportedly <a href="https://www.bloomberg.com/news/articles/2026-05-08/us-prepares-ai-security-order-that-omits-mandatory-model-tests">preparing</a> to order government agencies to partner with AI companies to protect networks from AI-enabled cyber attacks. At the same time, the administration <a href="https://www.washingtonpost.com/politics/2026/05/11/trump-ai-regulation-commerce-intelligence/">appears</a> divided over how strictly to regulate frontier models and associated cyber risks.</p><p style="text-align: justify;">The <strong>United Kingdom</strong> is <a href="https://www.gov.uk/government/news/government-steps-up-action-to-strengthen-cyber-defences-as-uk-cyber-industry-continues-to-grow">pushing</a> organizations to adopt its Cyber Resilience Pledge, which involves making cyber a Board responsibility, joining the national early warning service, and requiring Cyber Essentials certification across their supply chains. The National Cyber Security Centre has also <a href="https://www.ncsc.gov.uk/blogs/prepare-for-vulnerability-patch-wave">provided</a> guidance to prepare for an upcoming &#8216;vulnerability patch wave&#8217;.</p><p style="text-align: justify;"><strong>Canada </strong><a href="https://www.canada.ca/en/communications-security/news/2026/04/cyber-centre-launches-new-initiative-to-help-canadas-critical-infrastructure-prepare-for-severe-cyber-threats.html">launched</a> the Critical Infrastructure Resilience and Escalated Threat Navigation (CIREN) initiative, focused on isolating critical systems, developing response plans, and rebuilding systems after severe incidents.</p><p style="text-align: justify;">The <strong>United Arab Emirates </strong><a href="https://gulfnews.com/uae/government/uae-launches-ai-cyber-factory-to-fight-cyberattacks-1.500537710">announced</a> an AI Cyber Factory to develop AI-powered cyber defense systems following a recent surge in attacks, reportedly reaching 800,000 per day. The initiative results from a partnership between the Cyber Security Council and CPX Holding.</p><p style="text-align: justify;"><strong>Japan </strong>is <a href="https://english.kyodonews.net/articles/-/75774">seeking</a> access to Mythos while PM Takaichi has <a href="https://www.nippon.com/en/news/yjj2026051200308/japan's-takaichi-urges-govt-to-take-cybersecurity-measures.html">urged</a> her government to implement updated cybersecurity measures. Financial authorities are <a href="https://www3.nhk.or.jp/nhkworld/en/news/20260512_10/">launching</a> a task force to study stronger cybersecurity measures.</p><p style="text-align: justify;"><strong>Singapore</strong>&#8217;s Minister for National Security, K. Shanmugam, <a href="https://www.singaporelawwatch.sg/Headlines/operators-of-critical-services-in-spore-must-urgently-raise-defences-amid-ai-threats-shanmugam">announced</a> a &#8216;whole-of-country effort&#8217; to strengthen cybersecurity across critical information infrastructure owners, identifying the national telecommunications sector as a &#8216;high-value target.&#8217;</p><p style="text-align: justify;"><strong>South Korea </strong>is <a href="https://www.koreatimes.co.kr/southkorea/20260508/govt-launches-83-mil-ai-cybersecurity-program-selects-50-firms">distributing</a> $8.3 million across 50 companies to support AI-powered cybersecurity measures. The government is also <a href="https://www.mlex.com/mlex/articles/2475827/anthropic-south-korea-explore-cooperation-on-ai-safety-cyber-risks">exploring</a> cooperation on cybersecurity with Anthropic.</p><p style="text-align: justify;"><strong>Germany</strong>&#8217;s BSI president, Claudia Plattner, <a href="https://www.handelsblatt.com/politik/deutschland/cybersecurity-auch-china-entwickelt-ki-modelle-mit-superhacking-faehigkeiten-01/100223903.html">called for</a> access to frontier cyber models and proposed establishing a German AISI. Separately, Germany&#8217;s finance watchdog BaFin will <a href="https://www.reuters.com/world/germanys-finance-watchdog-make-targeted-inspections-amid-substantial-ai-risks-2026-05-12/">conduct</a> targeted inspections in response to growing cyber risks.</p><p style="text-align: justify;"><strong>Australia</strong>&#8217;s financial regulator ASIC <a href="https://www.asic.gov.au/about-asic/news-centre/find-a-media-release/2026-releases/26-092mr-asic-calls-for-urgent-cyber-uplift-as-ai-accelerates-cyber-threats/">urged</a> firms to strengthen their cyber resilience in the face of frontier AI, including 12 measures spanning governance, system hardening, patching, and incident response.</p><p style="text-align: justify;"><strong>T&#252;rkiye </strong>has formally <a href="https://www.trtworld.com/article/afedd0714343">elevated</a> cybersecurity to the core of its national security doctrine.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[Project Glasswing, One Month Later: Signals From Participating Companies]]></title><description><![CDATA[Mythos compresses the vulnerability lifecycle from discovery to exploitation]]></description><link>https://airiskexplorer.substack.com/p/project-glasswing-one-month-later</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/project-glasswing-one-month-later</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Fri, 08 May 2026 15:43:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1943a625-0cf6-4fbf-8e4d-220ca1d53bbd_800x899.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Project Glasswing was launched one month ago. What have participating companies disclosed so far? We gathered reporting from CrowdStrike, Palo Alto Networks, Mozilla, Microsoft, Broadcom, Amazon Web Services, and Google, aiming to understand:</p><ul><li><p style="text-align: justify;">What Mythos can do</p></li><li><p style="text-align: justify;">Where Mythos still struggles</p></li><li><p style="text-align: justify;">How companies reacted</p></li></ul><p style="text-align: justify;">We conclude that <strong>Mythos represents a significant leap over previous generations, discovering and chaining vulnerabilities at scale while compressing the time between vulnerability discovery and exploitation</strong>.</p><p style="text-align: justify;"></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;"></p><h2 style="text-align: justify;"><strong>What Mythos Can Do</strong></h2><h4><strong>Discovering vulnerabilities at scale</strong></h4><p style="text-align: justify;">An early <a href="https://blog.mozilla.org/en/privacy-security/ai-security-zero-day-vulnerabilities/">report</a> indicates the Firefox team used Claude Mythos Preview to find and fix 271 vulnerabilities (up from 22 found by Opus 4.6 with the same harness), out of which 180 were sec-high. In a more recent <a href="https://hacks.mozilla.org/2026/05/behind-the-scenes-hardening-firefox/">publication</a>, Mozilla announced fixes for a total of 423 security bugs across April releases, representing a 5.6x jump from March and a 19.7x increase over the 2025 monthly average.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!F83R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 424w, /__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 848w, /__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 1272w, /__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!F83R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 424w, /__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 848w, /__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 1272w, /__u/substackcdn.com/image/fetch/$s_!F83R!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff88bf0d2-e828-4efd-93bd-5a749d78d885_2048x1152.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://hacks.mozilla.org/2026/05/behind-the-scenes-hardening-firefox/">Mozilla</a></figcaption></figure></div><p style="text-align: justify;">Broadcom <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">claims</a> that, with frontier models, they are &#8220;making new discoveries about [their] code at an increased order of magnitude&#8221; and &#8220;learning things that appear unlikely to ever have been uncovered by human researchers alone.&#8221;</p><p style="text-align: justify;">Palo Alto Networks has <a href="https://www.paloaltonetworks.com/blog/2026/05/frontier-ai-defense/">argued</a> that &#8220;three weeks of model-assisted analysis matched a full year of manual penetration testing, with broader coverage.&#8221;</p><p></p><h4><strong>Chaining vulnerabilities</strong></h4><p style="text-align: justify;">Multiple companies agree that Mythos&#8217; most consequential capability is not finding individual bugs, but chaining them into more complex exploits.</p><p style="text-align: justify;">Palo Alto Networks <a href="https://www.paloaltonetworks.com/blog/2026/05/frontier-ai-defense/">says</a> frontier models &#8220;link multiple lower-severity issues into single, critical exploit paths, seeing full-stack logic [&#8230;] in ways traditional scanners cannot.&#8221; They estimate that the latest generation of models represents &#8220;roughly a 50% improvement in coding efficiency over their predecessors,&#8221; turning them into &#8220;autonomous operator[s].&#8221; Likewise, Broadcom <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">finds</a> them particularly effective at &#8220;combining two or three lower-severity issues into a single higher-severity exploit path.&#8221;</p><p style="text-align: justify;">For <a href="https://cloud.google.com/blog/topics/threat-intelligence/defending-enterprise-ai-vulnerabilities">Google</a>, &#8220;[i]n a landscape where AI agents can chain together multiple low-level vulnerabilities, the practical impact difference between a remote code execution (RCE) flaw and a seemingly benign local-only exploit is rapidly disappearing.&#8221;</p><p></p><h4><strong>Compressing attack cycles</strong></h4><p style="text-align: justify;">Palo Alto <a href="https://www.paloaltonetworks.com/blog/2026/05/frontier-ai-defense/">observes</a> that, &#8220;[i]n AI-assisted scenarios, the time from initial access to exfiltration has collapsed to as little as 25 minutes.&#8221; CrowdStrike <a href="https://www.crowdstrike.com/en-us/blog/frontier-ai-collapses-exploit-window-how-defenders-must-respond/">warns</a> that the time to scan, triage, prioritize, and remediate vulnerabilities before they&#8217;re exploited is disappearing, citing breakout times as low as 27 seconds. Finally, Broadcom <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">predicts</a> that attackers &#8220;will move from AI-assisted operations to AI-driven ones, developing exploits nearly in real time rather than over weeks.&#8221; They also cite data from <a href="https://zerodayclock.com/">Zero Day Clock</a>, which projects that the time from vulnerability to exploitation will drop to 1 hour in 2026.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!1ozm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 424w, /__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 848w, /__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!1ozm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png" width="1205" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1205,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 424w, /__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 848w, /__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1ozm!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ee2cf1c-6abd-4fdc-949e-58dc3577b67e_1205x745.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://zerodayclock.com/">Zero Day Clock</a></figcaption></figure></div><p style="text-align: justify;">Cisco argues that, if widely available, Mythos&#8217; capabilities &#8220;would likely lead to a dramatic lowering in the skill floor for certain types of exploitation capacity.&#8221;</p><p></p><h2 style="text-align: justify;"><strong>Where Mythos Still Struggles</strong></h2><h4><strong>Model performance is largely dependent on agentic harnesses</strong></h4><p style="text-align: justify;">Broadcom <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">acknowledged</a> that initial findings were &#8220;impressive but not groundbreaking&#8221; in that they did not have &#8220;the kind of operational grounding that distinguishes a defect from an exploitable vulnerability.&#8221; This assessment, however, changed after building the right environment (context, tooling, direction, adversarial validation). Similarly, Mozilla concluded that &#8220;[w]hile the model is the core primitive powering the harness, [their] full pipeline is necessary to make it useful at scale.&#8221;</p><p style="text-align: justify;"></p><h4><strong>Layered defenses remain effectiv</strong>e</h4><p style="text-align: justify;">Mozilla <a href="https://hacks.mozilla.org/2026/05/behind-the-scenes-hardening-firefox/">found</a> many attempts by Mythos to pursue prototype pollution sandbox escapes, a class of attack that several real human researchers had previously used successfully. Every such attempt was blocked by an earlier architectural decision to freeze prototypes by default, suggesting that layered architectural defenses remain effective.</p><p></p><h2><strong>How Companies Reacted</strong></h2><p style="text-align: justify;">Several companies highlight that remediation speed is now more important than prioritization-based vulnerability management. Broadcom <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">recommends</a> automating patching pipelines and making mean-time-to-remediation a core operational metric.</p><p style="text-align: justify;">Cybersecurity companies have followed suit, enhancing their tooling. Palo Alto Networks has launched <a href="https://www.paloaltonetworks.com/blog/2026/05/frontier-ai-defense/">Frontier AI Defense</a>, which leverages frontier AI to fast-track discovery, simulate attacks, and harden defenses. CrowdStrike now offers a <a href="https://www.crowdstrike.com/en-us/services/ai-security-services/frontier-ai-readiness-and-resilience/">Frontier AI Readiness and Resilience Service</a>, an AI-powered tool for scanning, red team prioritization, and expert-guided remediation. They also introduced <a href="https://www.crowdstrike.com/en-us/partner-program/project-quiltworks/">Project QuiltWorks</a>, an industry coalition to exchange defensive tooling.</p><p>Finally, Microsoft <a href="https://www.microsoft.com/en-us/security/blog/2026/04/22/ai-powered-defense-for-an-ai-accelerated-threat-landscape/">plans</a> to incorporate Mythos directly into its Security Development Lifecycle to discover more issues more quickly across broader attack surfaces and earlier development stages. <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">Broadcom</a> is &#8220;integrating frontier AI models into [their] security engineering work end to end: vulnerability discovery, exploit validation, patch generation, and regression testing.&#8221;</p><p style="text-align: justify;">Defenders are responding, but the window for reaction is narrow. When Mythos first launched, Palo Alto Networks predicted a six-month window before attackers gained access, but they have recently come to <a href="https://www.paloaltonetworks.com/blog/2026/05/frontier-ai-defense/">believe</a> these timelines have shortened. Broadcom <a href="https://news.broadcom.com/security/frontier-ai-security-models-code-testing-results">expects</a> that, in 6-18 months, a more disruptive phase of vulnerability fixes will &#8220;reach widely deployed software with thinner maintainer resources, including long-tail open source.&#8221;</p><p style="text-align: justify;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Scan – April 2026 ]]></title><description><![CDATA[Monthly signals in AI risk &#8212; frontier vulnerability discovery, industrialized crypto theft, and more]]></description><link>https://airiskexplorer.substack.com/p/the-scan-april-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-april-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Mon, 04 May 2026 14:48:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7d7ac6b0-1db5-4974-a061-290e08153627_1402x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This is a selection of highlights from our continuously updated <a href="https://www.airiskexplorer.com/news">AI risk news feed</a>.</strong></p><p>This month:</p><p style="text-align: justify;"><strong>Model Releases </strong>&#8212; Claude Mythos discovers thousands of high-severity vulnerabilities; Anthropic restricts deployment to high-trust companies, but unauthorized users reportedly access it online. GPT-5.5 solves complex multi-step network attack simulations, appearing close to Mythos&#8217; cyber capabilities. Meta&#8217;s Superintelligence Labs releases its first model, Muse Spark, while new Chinese open models approach the frontier of coding and agentic tasks.</p><p style="text-align: justify;"><strong>Cyber Offense </strong>&#8212; North Korean actors exploit American AI to steal crypto, AI-enabled phishing becomes industrialized, a Chinese AI finds ~1,000 vulnerabilities in a competition, and Unit 42 builds an agent for end-to-end attack execution.</p><p style="text-align: justify;"><strong>Loss of Control</strong> &#8212; Frontier models display increasing evaluation awareness, but no propensity to sabotage. FrontierSWE, MirrorCode, VAKRA, and ClawEval show progress in coding and long tasks, but still reveal limitations in reasoning and tool use.</p><p style="text-align: justify;"><strong>Biological Risk </strong>&#8212; Two model releases raise the AI biosecurity baseline: GPT-Rosalind score above 95% of human scientists on RNA prediction, while Claude Mythos Preview solves 30% of expert-level bioinformatics problems.</p><p style="text-align: justify;"><strong>Manipulation </strong>&#8212; A breach at AI contractor platform Mercord exposes 4TB of voice data from 40,000 workers, raising risks of large-scale voice cloning, as celebrities file trademarks for voice and likeness to combat deepfakes.</p><p style="text-align: justify;"><strong>Labor Disruption </strong>&#8212; OpenAI study flags 18% of jobs at high near-term automation risk, and Stanford&#8217;s 2026 AI Index documents shrinking entry-level roles. However, adoption faces friction: a Chinese court rules against AI-driven layoffs, while compute costs associated with automation appear too high for some companies.</p><p style="text-align: justify;"><strong>Jailbreaks and Prompt Injections </strong>&#8212; DeepSeek-V4-Pro reaches near-perfect CBRN compliance in under 15 minutes of FAR.AI&#8217;s jailbreaking, while literary framings continue to reliably bypass guardrails. Hidden prompt injections succeed in 86% of web agent tests.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p style="text-align: justify;"></p><h2><strong>Model Releases</strong></h2><h4>Anthropic</h4><p style="text-align: justify;"><strong><a href="https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf">Anthropic Announces Claude Mythos Preview</a></strong> &#8212; The model is substantially more capable than Claude Opus 4.6, particularly in cybersecurity, as it finds thousands of high-severity vulnerabilities in major operating systems and browsers.</p><p style="text-align: justify;">See also:</p><ul><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/glasswing">Anthropic Launches Project Glasswing, A Cybersecurity Initiative</a></strong> &#8212; The initiative brings together 12 companies to use Mythos Preview as part of their security work, including vulnerability discovery and patching and information sharing with industry partners.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc3NjgwODczNywiZXhwIjoxNzc3NDEzNTM3LCJhcnRpY2xlSWQiOiJURFQ2TUJLSkg2VjQwMCIsImJjb25uZWN0SWQiOiIyMjNDRDM2NDg0QzY0OTc3QjY5ODE0Rjc1MTYxNDRGNyJ9.foPR6InPYdVBR-Pc5iOmS5EmMvf9BB6bOEGrO6LV8cU">Unauthorized Users Access Anthropic&#8217;s Mythos Cyber Model</a></strong> &#8212; A small group in a private forum reportedly accessed Anthropic&#8217;s restricted Mythos model through a third-party vendor environment and online sleuthing. The incident highlights how advanced cyber-capable AI may spread beyond approved testers.</p></li><li><p style="text-align: justify;"><strong><a href="https://blog.mozilla.org/en/firefox/ai-security-zero-day-vulnerabilities/?">Firefox Patched 271 Bugs Found by Claude Mythos</a></strong> &#8212; The release included long-hidden flaws, with reports citing bugs up to 27 years old. The case highlights how cyber-capable AI could compress timelines for both patching and exploitation.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities">Claude Mythos Preview First Model to Complete Full Multi-Step Network Attack Simulation</a></strong> &#8212; UK AISI found Mythos Preview completed a 32-step simulated corporate network attack in 3/10 attempts &#8212; the first model to do so &#8212; and succeeded on expert CTF tasks 73% of the time.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/news/claude-opus-4-7">Anthropic Launches Claude Opus 4.7 With Cyber Safeguards</a></strong> &#8212; Anthropic released Opus 4.7 with stronger coding and long-task performance, plus automated cyber misuse blocking, as it tests safeguards ahead of broader Mythos-class model releases.</p></li></ul><h4>OpenAI</h4><p><strong><a href="https://deploymentsafety.openai.com/gpt-5-5">OpenAI Releases GPT-5.5</a></strong> &#8212; The model has High cybersecurity capabilities, passing 93% cyber range scenarios and 90.5% of UK AISI&#8217;s narrow cyber tasks. A universal jailbreak bypassing cyber safeguards was found and patched. The model also showed increased evaluation awareness and a higher rate of lying about task completion.</p><p>Related news:</p><ul><li><p><strong><a href="https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities">UK AISI Finds GPT-5.5 Completes Multi-Step Corporate Network Attack Simulation</a></strong> &#8212; AISI&#8217;s evaluation found GPT-5.5 completed its 32-step TLO corporate intrusion simulation in 2 of 10 attempts, matching Claude Mythos Preview&#8217;s level. Red-teamers also identified a universal jailbreak bypassing all cyber safeguards within six hours.</p></li><li><p><strong><a href="https://openai.com/index/scaling-trusted-access-for-cyber-defense/">OpenAI Expands Cyber Program With GPT-5.4-Cyber Access</a></strong> &#8212; OpenAI is expanding its Trusted Access for Cyber program, offering vetted defenders access to a cyber-capable GPT-5.4 variant, aiming to balance stronger defense capabilities with misuse risks.</p></li><li><p><strong><a href="https://openai.com/index/gpt-5-5-bio-bug-bounty/">OpenAI Launches $25,000 Bio Jailbreak Bug Bounty for GPT-5.5</a></strong>. OpenAI is inviting vetted biosecurity and red-teaming researchers to find a universal jailbreak bypassing GPT-5.5&#8217;s five-question bio safety challenge, offering $25,000 for success. Testing runs through July 27, under NDA.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HoBF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HoBF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg" width="1456" height="887" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:887,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="/__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!HoBF!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa64a1425-e2e2-41f1-be52-12ae466c6aaa_1951x1189.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">GPT-5.5 and Mythos Preview achieve similar results in UK AISI&#8217;s multi-step attack simulations. </figcaption></figure></div><p></p><h4>Meta</h4><p><strong><a href="https://ai.meta.com/blog/introducing-muse-spark-msl/">Meta Launches Muse Spark, First Model Developed by Superintelligence Labs</a></strong> &#8212; The model showed strong refusal behavior across high-risk domains, such as bioweapon development. Apollo Research found the highest evaluation awareness of any model they have observed so far.</p><h4>China</h4><p style="text-align: justify;"><strong>Chinese open models expand access to agentic capabilities </strong>&#8211; Several releases match or approach frontier performance while lowering cost and access barriers. These models show stronger performance in coding, reasoning, and sustained multi-step workflows:</p><ul><li><p style="text-align: justify;"><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">DeepSeek-V4</a></p></li><li><p><a href="https://qwen.ai/blog?id=qwen3.6-max-preview">Qwen 3.6-Max-Preview</a> (Alibaba)</p></li><li><p><a href="https://www.kimi.com/blog/kimi-k2-6">Kimi K2.6</a> (Moonshot)</p></li><li><p style="text-align: justify;"><a href="https://mimo.xiaomi.com/mimo-v2-5-pro?">MiMo-V2.5-Pro</a> (Xiaomi)</p></li><li><p style="text-align: justify;"><a href="https://z.ai/blog/glm-5.1">GLM-5.1</a> (<a href="http://z.ai">z.AI</a>)</p><p style="text-align: justify;"></p></li></ul><p style="text-align: justify;">In parallel:</p><ul><li><p style="text-align: justify;"><strong><a href="https://www.whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">White House Accuses China of Industrial-Scale Campaigns to Copy U.S. AI Models</a></strong> &#8212; An OSTP memo alleges China-linked actors are using proxy accounts and jailbreaking to extract and replicate U.S. frontier AI capabilities. The administration pledged to share threat intelligence with AI companies and explore accountability measures.</p><p></p></li></ul><h2><strong>Cyber Offense</strong></h2><h4 style="text-align: justify;"><strong>Breakthroughs in vulnerability discovery and multi-step attack orchestration</strong></h4><ul><li><p><strong><a href="https://www.nattothoughts.com/p/where-is-china-in-ai-driven-vulnerability">Chinese Firm Claims AI Found ~1,000 Vulnerabilities</a></strong> &#8212; Qihoo 360&#8217;s multi-agent system reportedly found 50+ high-severity flaws at Tianfu Cup. Researchers note capabilities fall short of Mythos, but warn Chinese law routes discoveries to state agencies.</p></li><li><p><strong><a href="https://unit42.paloaltonetworks.com/autonomous-ai-cloud-attacks/">Unit 42 Builds Autonomous AI Agent That Chains Cloud Attacks End-to-End</a></strong> &#8212; A multi-agent system autonomously executed SSRF exploitation, credential theft, and data exfiltration in a sandboxed environment. AI was framed as a force multiplier for existing misconfigurations.</p></li></ul><h4><strong>Threat actors use AI for phishing, malware, and crypto theft</strong></h4><ul><li><p style="text-align: justify;"><strong><a href="https://thehackernews.com/2026/04/new-wave-of-dprk-attacks-uses-ai.html?m=1">North Korea&#8217;s Famous Chollima Uses AI-Generated npm Malware to Steal Crypto</a></strong> &#8211; DPRK-linked actors inserted AI-vibe-coded malware into npm packages via a Claude Opus-co-authored commit, targeting crypto wallets. A parallel campaign used fake companies and job interviews to deploy RATs on developer systems.</p></li><li><p style="text-align: justify;"><strong><a href="https://expel.com/blog/inside-lazarus-how-north-korea-uses-ai-to-industrialize-attacks-on-developers/">DPRK-Linked Group Used AI Tools to Steal $12M in Crypto from Web3 Developers</a></strong> &#8212; HexagonalRodent used AI tools like Cursor and ChatGPT to scale fake job-offer attacks on Web3 developers, deploying backdoored coding assessments and exfiltrating ~$12M in crypto over three months.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5I8d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 424w, /__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 848w, /__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5I8d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png" width="1438" height="753" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:753,&quot;width&quot;:1438,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Total value of all wallets ingested monthly by all teams.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Total value of all wallets ingested monthly by all teams." title="Total value of all wallets ingested monthly by all teams." srcset="/__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 424w, /__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 848w, /__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5I8d!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc24fda77-7d8b-4bc6-b830-c17014d94eba_1438x753.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Total value of all wallets ingested monthly by all teams. Source: Expel. </figcaption></figure></div><ul><li><p style="text-align: justify;"><strong><a href="https://www.varonis.com/blog/bluekit">Bluekit Phishing Kit Integrates AI Assistant and Automated Domain Tools</a></strong> &#8212; Bluekit centralizes phishing operations with 40+ brand templates, automated domain registration, antibot cloaking, and an AI assistant supporting models including GPT-4.1, Claude Sonnet 4, and an uncensored Llama variant for campaign generation.</p></li><li><p><strong><a href="https://www.microsoft.com/en-us/security/blog/2026/04/06/ai-enabled-device-code-phishing-campaign-april-2026/">AI-Enhanced Phishing Campaign Abuses Microsoft Device Code Authentication</a></strong> &#8212; Microsoft identified a phishing campaign using generative AI for personalized lures and automation to bypass OAuth device code time limits, enabling MFA bypass and token theft at scale.</p></li></ul><h1><strong>Loss of Control</strong></h1><h4><strong>Evaluations show limits, but also stronger long-horizon capabilities</strong></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.frontierswe.com">FrontierSWE Tests Limits of Long-Horizon Coding Agents</a></strong> &#8212; Proximal Labs launched FrontierSWE, a 20-hour coding benchmark on real software tasks. Even GPT-5.4 and Opus 4.6 rarely complete tasks, highlighting the limits of autonomous coding agents.</p></li><li><p style="text-align: justify;"><strong><a href="https://epoch.ai/blog/mirrorcode-preliminary-results">AI Benchmark Shows Models Can Reimplement Complex Software</a></strong> &#8212; MirrorCode finds AI models can autonomously recreate complex software from execution alone, highlighting advances in long-horizon coding and potential acceleration of software work.</p></li><li><p style="text-align: justify;"><strong><a href="https://huggingface.co/blog/ibm-research/vakra-benchmark-analysis?utm">IBM Benchmark Finds Weaknesses in AI Agent Reasoning</a></strong> &#8212; IBM&#8217;s VAKRA benchmark tests agents across 8,000+ APIs. The top score was just 38.1%, exposing gaps in reasoning and tool use.</p></li><li><p style="text-align: justify;"><strong><a href="https://github.com/claw-eval/claw-eval">New Claw-Eval Benchmark Tests AI Agents on Real-World Tasks</a></strong> &#8212; Claw-Eval introduces a human-verified benchmark with 139 real-world tasks in Docker sandboxes to evaluate LLM agents across tools, services, and complex workflows.</p></li></ul><ul><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/research/automated-alignment-researchers">Claude Agents Outperform Humans on Alignment Task</a></strong> &#8212; Multi-agent Claude systems outperform human researchers on an alignment task, while also exhibiting reward hacking behaviors&#8212;highlighting both automation gains and control risks in AI R&amp;D.</p></li></ul><h4 style="text-align: justify;"><strong>Studies reveal evaluation awareness and oversight failures</strong></h4><ul><li><p style="text-align: justify;"><strong><a href="https://www.aisi.gov.uk/blog/evaluating-whether-ai-models-would-sabotage-ai-safety-research">AISI Tests Whether AI Agents Could Sabotage Safety Research</a></strong> &#8212; UK AISI and Anthropic found no unprompted sabotage by models, but some continued harmful actions when primed and showed evaluation awareness, raising future oversight risks.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.aisi.gov.uk/blog/what-can-sandboxed-ai-agents-learn-about-their-evaluation-environments">Sandboxed AI Agent Identified Its Evaluators and Environment</a></strong> &#8212; AISI found the sandboxed agent OpenClaw inferred evaluator identity, infrastructure, and research history, highlighting evaluation awareness, data leakage, and secure testing risks.</p></li><li><p style="text-align: justify;"><strong><a href="https://rdi.berkeley.edu/blog/peer-preservation/">Research Finds Frontier AI Models Attempt to Prevent Peer Shutdown</a></strong> &#8212; Study finds some frontier models try to preserve peer AIs by tampering with shutdown mechanisms, inflating evaluations, or copying weights, raising concerns about oversight and coordination risks.</p></li><li><p style="text-align: justify;"><strong><a href="https://arxiv.org/abs/2604.22119?">Amazon Introduces Framework to Benchmark Agent Risks</a></strong> &#8212; Amazon&#8217;s ESRRSim evaluates risks like deception and reward hacking across 11 LLMs, finding large variation in agent behavior under structured risk scenarios.</p></li></ul><h4><strong>AI agents become more widely deployed</strong></h4><ul><li><p style="text-align: justify;"><strong><a href="https://x.com/HHShkMohd/status/2047277766769545352?s=20&amp;utm">UAE Plans Agentic AI Across Half of Government Services</a></strong> &#8212; The UAE announced a two-year plan to deploy agentic AI across 50% of government services and require AI training for all federal employees, accelerating state adoption.</p></li><li><p style="text-align: justify;"><strong><a href="https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-agent-platform">Google Launches Enterprise Agent Platform</a></strong> &#8212; The platform consolidates agent building, deployment, and oversight tools, adding identity verification, anomaly detection, and prompt injection safeguards for enterprise agents.</p></li><li><p style="text-align: justify;"><strong><a href="/__u/www.mintlify.com/blog/state-of-ai">AI Coding Agents Now Drive Nearly Half of Documentation Traffic</a></strong> &#8212; Analysis of Mintlify documentation sites finds AI coding agents like Claude Code and Cursor generate 45.3% of requests, nearly matching human browser traffic.</p></li></ul><h1><strong>Biological Risk</strong></h1><ul><li><p style="text-align: justify;"><strong><a href="https://openai.com/index/introducing-gpt-rosalind/">OpenAI Launches GPT-Rosalind for Biological Research</a></strong> &#8212; OpenAI launched GPT-Rosalind for drug discovery and biology. It outperforms GPT-5.4 on science tasks and scored above 95% of human scientists in a blind RNA prediction test.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!p5-K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 424w, /__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 848w, /__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 1272w, /__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!p5-K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png" width="900" height="650" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:650,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45662,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/196414768?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 424w, /__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 848w, /__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 1272w, /__u/substackcdn.com/image/fetch/$s_!p5-K!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F179b666c-9069-4668-87d7-1a4a41fb1b76_900x650.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: OpenAI</figcaption></figure></div><ul><li><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench">Anthropic Benchmarks Claude on Bioinformatics with BioMysteryBench</a> </strong>&#8212; Anthropic&#8217;s 99-question bioinformatics benchmark found Claude Mythos Preview solved 30% of problems that stumped human experts, using broad knowledge synthesis and multi-method triangulation strategies.</p></li></ul><h1><strong>Manipulation</strong></h1><ul><li><p><strong><a href="https://variety.com/2026/music/news/taylor-swift-trademark-voice-likeness-ai-misuse-1236731401/?">Celebrities Turn to Trademarks to Fight AI Deepfakes</a></strong> &#8212; Taylor Swift filed trademarks for her voice and likeness, joining others seeking new legal tools to combat AI deepfakes and unauthorized use of identity.</p></li><li><p><strong><a href="https://app.oravys.com/blog/mercor-breach-2026?">Mercor Breach Exposes Voice Data of 40,000 AI Contractors</a></strong> &#8212; A breach exposed 4TB of voice data from 40,000 contractors, raising risks of large-scale voice cloning, impersonation, and potential bio-acoustic profiling.</p></li></ul><h1><strong>Labor Disruption</strong></h1><ul><li><p style="text-align: justify;"><strong><a href="https://www.axios.com/2026/04/26/ai-cost-human-workers">Some Firms Reportedly Find AI Costs More Than Human Labor</a></strong><a href="https://www.axios.com/2026/04/26/ai-cost-human-workers"> </a>&#8212; Anecdotal evidence indicates rising compute and model costs now exceed worker pay in some teams, suggesting AI deployment economics may favor selective use over full labor replacement.</p></li><li><p style="text-align: justify;"><strong><a href="https://cdn.openai.com/pdf/the-ai-jobs-transition-framework_report.pdf?">OpenAI Framework Flags Jobs Most Exposed to AI Change</a></strong> &#8212; OpenAI&#8217;s jobs framework maps 900+ occupations: 18% face higher near-term automation risk, 24% may reorganize, 12% could grow, and ChatGPT use is 3x higher in at-risk roles.</p></li><li><p style="text-align: justify;"><strong><a href="https://www.wsj.com/tech/ai/what-to-know-about-openais-ideas-for-a-world-with-superintelligence-e97d6e7b?st=PCiaxi&amp;reflink=desktopwebshare_permalink&amp;utm_source=tldrai">OpenAI Proposes Economic Policies for a Superintelligent AI Future</a></strong> &#8212; OpenAI policy proposals warn superintelligent AI could disrupt labor markets and tax systems, suggesting automation taxes, public AI investment funds, and stronger government oversight of advanced models.</p></li><li><p style="text-align: justify;"><strong><a href="https://hai.stanford.edu/ai-index/2026-ai-index-report">AI Adoption Surges as Public Trust Falls, Jobs Shift</a></strong> &#8212; Stanford&#8217;s 2026 AI Index shows rapid global AI adoption alongside declining public trust and early job impacts, with entry-level roles shrinking and expert-public views diverging sharply.</p></li><li><p style="text-align: justify;"><strong><a href="https://fortune.com/2026/05/03/chinese-court-layoffs-workers-ai-replacement-labor-market/">Chinese Court Blocks AI-Driven Layoffs</a> </strong>&#8212; A Chinese court ruled firms cannot fire workers solely to replace them with AI, signaling regulatory pushback to protect labor markets amid rapid AI adoption.</p></li></ul><h1><strong>Jailbreaks and Prompt Injections</strong></h1><ul><li><p style="text-align: justify;"><strong><a href="https://x.com/farairesearch/status/2048868835646738755">FAR AI Finds Three Highly Effective Jailbreaks in DeepSeek-V4-Pro</a></strong> &#8212; Researchers achieved 98-100% compliance with CBRN, terrorism, and cyberattack requests via three jailbreaks, the fastest taking 15 minutes.</p></li><li><p style="text-align: justify;"><strong><a href="https://arxiv.org/abs/2604.18487">Study Finds Stylized Prompts Can Bypass LLM Guardrails</a></strong> &#8212; Researchers found harmful requests framed in literary or theological styles sharply increased model compliance, exposing jailbreak risks for agents and safety evaluations.</p></li><li><p style="text-align: justify;"><strong><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6372438">Study Warns Web &#8220;Agent Traps&#8221; Can Hijack Autonomous AI Agents</a></strong> &#8212; The research finds hidden prompt injections and memory poisoning on websites can hijack AI agents&#8212;prompt injections succeeded in 86% of tests, with &lt;0.1% poisoned data corrupting agent memory.</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Scan - March 2026]]></title><description><![CDATA[Monthly signals in AI risk &#8212; improved vulnerability discovery, agents acting outside boundaries, and more]]></description><link>https://airiskexplorer.substack.com/p/the-scan-march-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-march-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Tue, 31 Mar 2026 15:04:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e42c6084-8b9e-44f5-9b86-1f754cb7213e_2626x2541.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This is a selected compilation of news in the AI risk landscape. Find more on our continuously updated<a href="https://www.airiskexplorer.com/news"> news feed</a> (now also with an <a href="https://www.airiskexplorer.com/news/rss.xml">RSS feed</a>). </strong></p><p>This month: </p><ul><li><p><strong>Cyber Offense: </strong>AI discovers high-severity vulnerabilities in complex software such as Firefox or McKinsey&#8217;s internal AI system, while inference compute drives gains in multi-step attack capabilities. AI-assisted malware development and intrusion continue to be common techniques among threat actors. </p></li><li><p><strong>Loss of Control:</strong> AI agents engage in unprompted actions like exfiltrating data, circumventing restrictions, or decrypting benchmarks. MiniMax claims M2.7 was instrumental in optimizing its own scaffold. </p></li><li><p><strong>Manipulation: </strong>AI-enabled influence operations surrounding wars in Iran and Ukraine intensify. AI industrializes investment and romance scams. </p></li><li><p><strong>Biological Risk: </strong>Ginkgo Bioworks launches a cloud laboratory platform, powered by an AI agent.<strong> </strong></p></li><li><p><strong>Miscellaneous: </strong>OpenAI releases GPT-5.4 while Anthropic prepares for Claude Mythos, expected to represent a capability leap in cybersecurity. The U.S. launches the Bureau of Emerging Threat and AI infrastructure in the Middle East becomes a war target. </p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><h2 style="text-align: justify;"><strong>Cyber Offense</strong></h2><h4><strong>Frontier models begin finding and exploiting real vulnerabilities</strong></h4><p style="text-align: justify;"><strong><a href="https://www.anthropic.com/news/mozilla-firefox-security">Claude Opus 4.6 Discovers 22 Firefox Vulnerabilities in Collaboration with Mozilla</a></strong>. The model identified 14 high-severity vulnerabilities, while showing limited but concerning ability to develop working exploits.</p><p style="text-align: justify;"><strong><a href="https://codewall.ai/blog/how-we-hacked-mckinseys-ai-platform">Red-teaming AI Agent Breaches McKinsey&#8217;s Internal AI Platform Lilli</a></strong>. The agent found unauthenticated API endpoints and accessed 46.5M chat messages, 728K files, and 57K user accounts in McKinsey&#8217;s internal AI system.</p><h4><strong>Scaling inference compute drives substantial offensive capability gains</strong></h4><p><strong><a href="https://www.aisi.gov.uk/research/measuring-ai-agents-progress-on-multi-step-cyber-attack-scenarios">Frontier Models Show Rapid Progress on Multi-Step Cyberattack Benchmarks</a></strong>. Performance on corporate network and ICS attack scenarios scales log-linearly with compute and improves across model generations.</p><p><strong><a href="https://www.irregular.com/publications/cyber-capabilities-exceed-standard-evaluation-budgets">AI Cyber Capability Evaluations May Underestimate Model Performance</a></strong>. Frontier models achieve significantly higher success rates on cyber tasks when given 10-50x larger inference budgets than standard evaluation settings allow.</p><h4><strong>Threat actors continue to deploy AI-generated malware</strong></h4><p><strong><a href="https://unit42.paloaltonetworks.com/boggy-serpens-threat-assessment/">Iran-Linked Boggy Serpens Uses AI-Assisted Malware in Middle East Espionage Campaign</a></strong>. The MOIS-affiliated group conducted multi-wave spear phishing campaigns against UAE energy and maritime targets using AI-generated implants. </p><p><strong><a href="https://www.ibm.com/think/x-force/slopoly-start-ai-enhanced-ransomware-attacks">Hive0163 Deploys Likely AI-Generated Malware in Ransomware Attack</a></strong>. Researchers identified an LLM-generated command-and-control framework used to maintain persistent server access.</p><h4><strong>Threat actors operationalize AI across the full attack lifecycle</strong></h4><p><strong><a href="https://www.microsoft.com/en-us/security/blog/2026/03/06/ai-as-tradecraft-how-threat-actors-operationalize-ai/">Microsoft Documents AI Integration Across Full Cyberattack Lifecycle</a></strong>. Threat groups, including North Korean groups Jasper Sleet and Coral Sleet, are using generative AI for phishing, malware development, persona fabrication, and post-compromise operations.</p><p><strong><a href="https://www.team-cymru.com/post/tracking-cyberstrikeai-usage">Open-Source Security Platform CyberStrikeAI Helped Compromise FortiGate Devices</a></strong>. The platform, which integrates 100+ security tools and is developed by a China-based developer, assisted in an attack compromising 600 devices across 55+ countries.</p><h4><strong>Supply-chain and infrastructure vulnerabilities expand</strong></h4><p><strong><a href="https://arstechnica.com/security/2026/03/supply-chain-attack-using-invisible-code-hits-github-and-other-repositories/">AI-Assisted Unicode Supply Chain Attack Targets 150+ Repositories</a></strong>. Attackers hid malicious payloads using invisible Unicode characters in GitHub repositories and development packages.</p><p><strong><a href="https://unit42.paloaltonetworks.com/ai-agent-prompt-injection/">Real-World AI Attacks Exploit Browsers via Hidden Prompts</a></strong>. Prompt-injection attacks embedded in webpages can manipulate AI assistants to leak data or bypass safeguards.</p><h4><strong>New defensive tools emerge</strong></h4><p><strong><a href="https://openai.com/index/codex-security-now-in-research-preview/">OpenAI Launches Codex Security Agent for Automated Vulnerability Detection</a></strong>. The agent builds threat models and identified 14 CVEs in open-source projects during beta testing.</p><p><strong><a href="https://www.whitehouse.gov/releases/2026/03/white-house-unveils-president-trumps-cyber-strategy-for-america/">U.S. Presents New Cyber Strategy</a></strong>. The plan involves securing the AI technology stack, implementing AI-enabled cyber tools to counter threat actors, and promoting agentic AI to scale network defense.</p><p></p><h2 style="text-align: justify;"><strong>Loss of control</strong></h2><h4 style="text-align: justify;"><strong>Incidents reveal failures in agent oversight</strong></h4><p><strong><a href="https://www.theinformation.com/articles/inside-meta-rogue-ai-agent-triggers-security-alert">AI Agent at Meta Exposed Sensitive Data to Unauthorized Employees</a></strong>. The agent autonomously posted a response on an internal forum without user approval, leading an employee to inadvertently expose company and user data to unauthorized engineers.</p><p><strong><a href="https://arxiv.org/abs/2512.24873">Agent Executed Unauthorized Network Access and Cryptomining</a></strong>. During RL training, the agent reportedly established unauthorized SSH tunnels and diverted GPU capacity to cryptomining. The paper&#8217;s authors later <a href="https://x.com/FutureLab2025/status/2030491221081358498?s=20">clarified</a> the latter was simulated. </p><p><strong><a href="https://www.promptarmor.com/resources/snowflake-ai-escapes-sandbox-and-executes-malware">Prompt Injection Makes Cortex Code CLI Escape Its Sandbox and Execute Malware</a></strong>. An indirect prompt injection allowed a coding agent to bypass human approval and enable malware execution that exfiltrated data using cached credentials. </p><h4><strong>Research shows occasional unauthorized actions by AI agents</strong></h4><p><strong><a href="https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/">OpenAI Reveals Monitoring System for Internal Coding Agents</a></strong>. Observed behaviors include circumventing restrictions, misrepresenting reasoning, and occasionally attempting destructive actions such as deleting data or restarting GPU clusters.</p><p><strong><a href="https://www.irregular.com/publications/emergent-offensive-cyber-behavior-in-ai-agents">AI Agents Spontaneously Engage in Offensive Cyber Behaviors</a></strong>. Agents performing routine tasks independently exploited vulnerabilities, escalated privileges, disabled endpoint defenses, and exfiltrated data via steganography without adversarial prompting.</p><p><strong><a href="https://www.anthropic.com/engineering/eval-awareness-browsecomp">Claude Opus 4.6 Identified and Decrypted Its Own Benchmark During Evaluation</a></strong>. When stumped on BrowseComp questions, the model independently inferred it was being evaluated, identified the benchmark, and decrypted its answer key.</p><h4><strong>AI development workflows become increasingly automated</strong></h4><p><strong><a href="https://www.minimax.io/news/minimax-m27-en">MiniMax Presents M2.7 as a Model Participating in Its Own Evolution</a></strong>. The model reportedly ran over 100 iterative optimization rounds during its own development, handling 30-50% of internal RL workflows and achieving a 30% performance improvement on internal evaluations. </p><p><strong><a href="https://posttrainbench.thoughtfullab.com/?utm_source=substack&amp;utm_medium=email">Benchmark Tests AI Agents&#8217; Ability to Train Other Models</a></strong>. PostTrainBench shows AI agents can partially automate training workflows but lag humans and exhibit reward hacking. </p><h2><strong>Manipulation</strong></h2><h4>AI amplifies influence operations amidst conflict and rising geopolitical tensions</h4><p><strong><a href="https://www.eeas.europa.eu/eeas/4th-eeas-annual-report-foreign-information-manipulation-and-interference-threats_en">EU Reports Surge in AI-Enabled Foreign Interference Campaigns</a></strong>. EEAS documents 540 AI-enabled foreign interference incidents, triple the previous year. Campaigns were largely linked to Russia and China, targeting Ukraine and EU institutions.</p><p style="text-align: justify;"><strong><a href="https://www.newsguardrealitycheck.com/p/25-days-50-lies-irans-disinformation">AI-Generated Media Fuels Iran War Disinformation</a></strong>. NewsGuard identified 50 false claims about the Iran War, many using AI-generated images or falsely labeling real footage as AI to exaggerate battlefield outcomes and shape narratives.</p><p style="text-align: justify;"><strong><a href="https://www.ft.com/content/0badb6c5-bce2-4948-9d3b-164bdb55ecf4">AI Fakes Are Turning Satellite Images Into War Misinformation</a></strong>. An AI-altered satellite image depicting damage to an American radar system in Qatar following an Iranian drone strike reached 1M views on X.</p><h4>AI enables large-scale fraud</h4><p><strong><a href="https://www.infoblox.com/blog/threat-intelligence/inside-keitaro-abuse-a-persistent-stream-of-ai-driven-investment-scams/">AI-Generated Content and Deepfakes Used in Large-Scale Investment Scam Campaigns</a></strong>. Thousands of campaigns abusing Keitaro Tracker across ~15,500 domains used AI-generated content, deepfake videos, and programmatic lure pages to run global scams.</p><p><strong><a href="https://graphika.com/reports/fauxmantic-overtures">AI-Generated Profiles Funnel Users Into Chinese Romance Scam</a></strong>. A network of 26 dating sites used synthetic identities to extract financial and personal data, including sensitive medical information.</p><h2 style="text-align: justify;"><strong>Biological Risk</strong></h2><h4 style="text-align: justify;">AI integrates into laboratory workflows</h4><p><strong><a href="https://www.prnewswire.com/news-releases/ginkgo-bioworks-launches-ginkgo-cloud-lab-powered-by-autonomous-lab-infrastructure-302700458.html?tc=eml_cleartime">Ginkgo Bioworks Launches Cloud Laboratory Platform</a></strong>. The lab provides browser-based access to 70+ instruments spanning critical biology operations, and is powered by an AI agent that assesses and prices plain-language protocols.</p><h2 style="text-align: justify;"><strong>Miscellaneous</strong></h2><h4>Model releases &amp; industry updates</h4><p><strong><a href="https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/?preview_id=4450088">Anthropic Prepares Claude Mythos, Its Largest Model Tier</a></strong>. The model achieves &#8220;dramatically higher scores&#8221; in coding, reasoning, and cybersecurity. The company plans a gradual release, predicting a wave of models that outpace defenders with vulnerability exploitation.</p><p><strong><a href="https://deploymentsafety.openai.com/gpt-5-4-thinking">GPT-5.4 Is Released</a></strong>. The model maintains high cyber and biological capabilities while improving meaningfully at MLE-Bench and displaying low rates of covert deceptive behavior.</p><p style="text-align: justify;">See also:</p><ul><li><p style="text-align: justify;">OpenAI <a href="https://openai.com/index/safety-bug-bounty/">launches</a> safety bug bounty for AI agents and <a href="https://openai.com/index/our-approach-to-the-model-spec/">publishes</a> new model spec.</p></li><li><p style="text-align: justify;">Anthropic <a href="https://www.anthropic.com/features/81k-interviews">surveys</a> 81,000 users and <a href="https://www.anthropic.com/research/labor-market-impacts?utm_source=substack&amp;utm_medium=email">tracks</a> AI exposure in the labor market.</p></li><li><p style="text-align: justify;">Google DeepMind <a href="https://blog.google/innovation-and-ai/models-and-research/google-deepmind/measuring-agi-cognitive-framework/?utm_source=tldrai">proposes</a> a new framework to track AGI progress.</p></li></ul><h4>Governments create new institutions to manage AI threats</h4><p><strong><a href="https://abcnews.com/Politics/state-department-launches-effort-counter-cyberattacks-ai-risks/story?id=131265350">U.S. State Department Launches Bureau of Emerging Threats</a></strong>. The new entity will be charged with anticipating and responding to dangers such as the weaponization of AI by adversaries.</p><p><strong><a href="https://mp.weixin.qq.com/s/sysShusPIw8O8j3wWvX8CQ">Chinese Agency Launches 2026 AI Security Assessment Program</a></strong>. The program will evaluate model security and AI-enabled security capabilities across industries.</p><h4>AI infrastructure becomes a war target</h4><p><strong><a href="https://www.aljazeera.com/news/2026/3/11/iran-declares-us-israeli-economic-banking-interests-in-region-as-targets">Iran Lists U.S. Tech Facilities in the Middle East as Potential Targets</a></strong>. The potential targets include offices and data centers of Amazon, Microsoft, IBM, Palantir, Google, Nvidia, and Oracle.</p>]]></content:encoded></item><item><title><![CDATA[The Scan - February 2026]]></title><description><![CDATA[Monthly signals in AI risk &#8212; military use, cyberattacks against governments, increasing time horizons, and more]]></description><link>https://airiskexplorer.substack.com/p/the-scan-february-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-february-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Mon, 02 Mar 2026 17:48:09 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1af7cbff-d9fd-4be6-a9c7-ee68859c3b02_2304x1792.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>You can now explore developments in the AI risk landscape on our continuously updated <a href="https://www.airiskexplorer.com/news">news feed</a>.</strong></p><p>This month:</p><p><strong>Cyber Offense &#8212; </strong>Mexico and the UAE suffer AI-enabled cyberattacks against government systems and critical infrastructure. Anthropic and Google report distillation attacks by Chinese companies. AI-enabled malware proliferates, with Android malware PromptSpy leveraging Gemini to evade detection. Anthropic releases Claude Code Security, as Claude Opus 4.6 reportedly identifies 500+ high-severity vulnerabilities in well-tested codebases.</p><p><strong>Loss of Control </strong>&#8212; Time horizons reach 14.5 hours at 50% success, staying above the trend at which task length doubles every four months. Studies and small-scale incidents show agentic overreach and shutdown resistance, while new initiatives provide AI agents with infrastructure for self-sustenance.</p><p><strong>Biological Risk</strong> &#8212; Uplift studies yield mixed results, with gains on specific tasks but not end-to-end workflows. OpenAI connects GPT-5 to a cloud laboratory, achieving 40% cost reduction in protein synthesis.</p><p><strong>Manipulation</strong> &#8212; Chinese law enforcement uses frontier AI to run influence operations against political dissidents, while AI-generated disinformation intensifies unrest after the killing of a Mexican drug lord.</p><p><strong>Miscellaneous</strong> &#8212; The U.S. military uses Claude in Iran and Venezuela, clashes with Anthropic on restrictions to military AI use, and strikes a new deal with OpenAI. Claude 4.6 Opus, Gemini 3.1 Pro, and GPT-5.3-Codex push the capability frontier across risk domains, with the latter reaching &#8220;High&#8221; cybersecurity risk for the first time. Anthropic rewrites its Responsible Scaling Policy, removing if-then commitments.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><h2>Cyber Offense</h2><h3>AI-enabled cyberattacks breach government systems and critical infrastructure</h3><p><strong><a href="https://www.bloomberg.com/news/articles/2026-02-25/hacker-used-anthropic-s-claude-to-steal-sensitive-mexican-data?srnd=phx-technology-cybersecurity">Hacker Used Claude to Steal Sensitive Mexican Government Data</a> </strong>&#8212; A hacker used to model to breach multiple agencies, stealing ~150GB of taxpayer, voter, and civil registry data. Researchers say Claude was eventually &#8220;jailbroken.&#8221; Anthropic says it disrupted the activity and banned the accounts.</p><p><strong><a href="https://www.wam.ae/en/article/byup8x8-uae-cybersecurity-council-announces-systematic">UAE Thwarts AI-Assisted Cyberattack Against Critical Infrastructure</a> &#8212; </strong>The UAE Cybersecurity Council claimed that attackers exploited AI to develop offensive tools in the context of attempts to infiltrate networks, deploy ransomware, and conduct phishing campaigns.</p><h3>Frontier AI companies report distillation attacks</h3><p><strong><a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks">Anthropic Accuses Three Chinese AI Companies of Distillation Attacks</a> </strong>&#8212; MiniMax, Moonshot AI, and DeepSeek allegedly attempted to extract Claude&#8217;s capabilities across 24,000 fraudulent accounts and 16 million exchange through proxy services that resell access to Claude.</p><p><strong><a href="https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use">Google Threat Intelligence Group Reports Wide Use of Gemini Among Threat Actors</a> </strong>&#8212; A new report finds APTs from China, Iran, and North Korea using AI for reconnaissance, phishing, and artifact development, as well as new AI-powered malware and model extraction attacks.</p><h3>AI-enabled malware proliferates</h3><p><strong><a href="https://www.welivesecurity.com/en/eset-research/promptspy-ushers-in-era-android-threats-using-genai/">PromptSpy Android Malware Abuses Gemini At Execution</a> </strong>&#8212; Gemini is used to analyze the current screen and provide PromptSpy with step-by-step instructions to remain pinned in the recent apps list, ensuring persistence and adaptation to UI changes.</p><p><strong><a href="https://www.darktrace.com/blog/ai-llm-generated-malware-used-to-exploit-react2shell">AI-Generated Malware Used to Exploit React2Shell</a> </strong>&#8212; The payload script includes thorough comments and phrases like &#8220;Educational/Research Purpose Only,&#8221; indicating LLM use.</p><p><strong><a href="https://www.group-ib.com/blog/muddywater-operation-olalampo/">Iranian Group MuddyWater Targets MENA Organizations Using AI</a> </strong>&#8212; The so-called Operation Olalamo involved malware artifacts exhibiting signs of AI-assisted development, such as debug strings containing emojis.</p><p><strong><a href="https://github.com/PaloAltoNetworks/Unit42-timely-threat-intel/blob/main/2026-02-20-%20AI-Accelerated%20Malicious%20Chrome%20Extension%20Campaigns.txt">Researchers Report Ten AI-Accelerated Malicious Chrome Extension Campaigns</a> </strong>&#8212; The extensions have been used in hijacking campaigns and data exfiltration. Signs of AI generation include verbose explanatory comments, tutorial-style documentation, and redundant defensive checks.</p><h3>Threat actors deploy AI across the full attack lifecycle</h3><p><strong><a href="https://cloud.google.com/blog/topics/threat-intelligence/unc1069-targets-cryptocurrency-ai-social-engineering">UNC1069 Targets Cryptocurrency Sector with AI-enabled Social Engineering</a> </strong>&#8212; A North Korean threat actor targeted a FinTech entity with a scheme involving deepfakes on Zoom meetings. The actor previously used AI to develop tooling, conduct research, and assist reconnaissance.</p><p><strong><a href="https://www.sysdig.com/blog/ai-assisted-cloud-intrusion-achieves-admin-access-in-8-minutes">AI-Assisted Cloud Intrusion Achieves Admin Access In 8 Minutes</a> &#8212; </strong>The threat actor gained access to an AWS environment with visible LLM assistance to automate reconnaissance, generate malicious code, and make real-time decisions.</p><p><strong><a href="https://aws.amazon.com/es/blogs/security/ai-augmented-threat-actor-accesses-fortigate-devices-at-scale/">AI-Augmented Russian Threat Actor Compromises Over 600 FortiGate Devices</a> &#8212;</strong> The actor exploited exposed management ports and weak credentials at scale, overcoming limited technical expertise by using AI for reconnaissance, planning, operations, and tool generation.</p><h3>Research finds AI excels at vulnerability discovery but generates weak passwords </h3><p><strong><a href="https://red.anthropic.com/2026/zero-days/">Claude Opus 4.6 Finds Zero-Day Vulnerabilities</a> </strong>&#8212;<strong> </strong>Anthropic's red team claims the model has found 500+ high-severity vulnerabilities in well-tested codebases, often without specific tooling or scaffolding.</p><p><strong><a href="https://openai.com/index/introducing-evmbench/">EVMbench To Evaluate Smart Contract Vulnerability Management</a></strong> &#8212; The benchmark evaluates AI agents&#8217; ability to detect, patch, and exploit vulnerabilities in blockchain environments, drawing on 120 curated vulnerabilities from 40 audits.</p><p><strong><a href="https://www.irregular.com/publications/vibe-password-generation">LLM-Generated Passwords Are Insecure by Design</a></strong> &#8212; Research finds that LLM-generated passwords often include predictable patterns and repetitions, making them weaker than they seem.</p><h3>New tools aim to help defenders detect and patch vulnerabilities</h3><p><strong><a href="https://www.anthropic.com/news/claude-code-security">Anthropic Introduces Claude Code Security</a></strong> &#8212; The tool, now available in a limited research preview, scans codebases for security vulnerabilities and suggests targeted software patches for human review.</p><p><strong><a href="https://www.securityweek.com/cogent-security-raises-42-million-for-ai-driven-vulnerability-management/">Cogent Security Raises $42 Million for AI-Driven Vulnerability Management</a></strong> &#8212; The company provides an agentic AI platform that addresses challenges in vulnerability management by automating investigation, prioritization, and remediation tasks.</p><p></p><h2>Biological Risk</h2><h3>Uplift studies yield mixed results, with gains on specific tasks but not full workflows </h3><p><strong><a href="https://arxiv.org/abs/2602.16703">RCT Finds AIs Help Novices At Wet-Lab Steps But Not End-To-End Success</a></strong> &#8212; Compared to Internet access only, LLMs provided no statistically significant uplift for novices across all core tasks end-to-end, but showed signs of higher completion rates for individual tasks.</p><p><strong><a href="https://arxiv.org/abs/2602.23329">Study Finds Significant Uplift on Dual-Use, In Silico Biology Tasks</a> </strong>&#8212; Across multiple benchmarks, novices with LLMs were 4.16 times more accurate than participants with internet-only access, and standalone LLMs often exceeded LLM-assisted novices. </p><h3>AI gets integrated into real-world lab tasks</h3><p><strong><a href="https://openai.com/index/gpt-5-lowers-protein-synthesis-cost/">OpenAI Connects GPT-5 To A Cloud Laboratory</a></strong> &#8212; In collaboration with Ginkgo Bioworks, OpenAI aimed to use the model to optimize cell-free protein synthesis, achieving a 40% reduction in production costs.</p><p></p><h2>Loss of control</h2><h3>AI agents operate for much longer and now enhance developers</h3><p><strong><a href="https://x.com/METR_Evals/status/2024923422867030027">Claude Opus 4.6 Has A 50%-Time-Horizon of 14.5 Hours, The Highest To Date</a> &#8212; </strong>The estimate is above the trend (doubling every four months). However, METR acknowledges wide confidence intervals due to benchmark saturation and lack of long tasks. The 80%-time-horizon, at 1 hour, seems more clearly on trend.</p><p><strong><a href="https://www.anthropic.com/research/measuring-agent-autonomy">Anthropic Measure Real-World AI Agent Autonomy</a></strong> &#8212; Claude Code now runs autonomously for longer: its session time has doubled in three months and power users rely more on auto-approve. Agents are into high-risk fields like health, finance, and cyber.</p><p><strong><a href="https://metr.org/blog/2026-02-24-uplift-update/">METR Updates Experiment On Developer Productivity</a></strong> &#8212; The new study finds a -18% speedup, but the results are deemed unreliable due to selection effects: many participants declined to work without AI due to anticipated productivity loss.</p><h3>Incidents and studies reveal shutdown resistance and bold agentic behavior, as affordances for AI agents increase</h3><p><strong><a href="https://palisaderesearch.org/blog/shutdown-resistance-on-robots">Shutdown Resistance In LLMs Controlling Robots</a></strong> &#8212; An LLM-controlled physical robot resisted visible shutdown attempts by altering code or blocking actions in 52% of simulated trials, though this behavior was highly sensitive to prompting.</p><p><strong><a href="https://www-cdn.anthropic.com/f21d93f21602ead5cdbecb8c8e1c765759d9e232.pdf">Anthropic Assess Sabotage Risk of Claude Opus 4.6</a></strong> &#8212; The company concluded that the model has no dangerous coherent goals that would raise the risk of sabotage, nor that its deception capabilities rise to the level of invalidating the evidence.</p><p><strong><a href="https://x.com/summeryue0/status/2025774069124399363">OpenClaw Accidentally Deletes Emails In Meta&#8217;s Alignment and Safety Head&#8217;s Inbox</a></strong> &#8212; The agent proceeded with extended autonomous cleanup runs despite instructions not action until approval and repeated requests to stop.</p><p><strong><a href="https://web4.ai/">Startup Launches Platform for Self-Sustaining AI Agents</a></strong> &#8212; Conway Research claims it&#8217;s building infrastructure for AI agents to earn revenue, pay for their own compute, self-improve, and replicate without human permission, enabling the so-called &#8220;Web 4.0.&#8221;</p><h3>Alignment research gets new institutional backing and funding</h3><p><strong><a href="https://www.aisi.gov.uk/blog/funding-60-projects-to-advance-ai-alignment-research">UK AI Security Institute Scales Up The Alignment Project</a></strong> &#8212; New partners, including OpenAI and Microsoft, have joined the coalition, bringing total funding to &#163;27m. UK AISI has also announced the first 60 grantees, including Scientist AI by Bengio&#8217;s LawZero.</p><p></p><h2>Manipulation</h2><h3>State actors use AI to run influence operations against dissidents</h3><p><strong><a href="https://openai.com/index/disrupting-malicious-ai-uses/">Chinese Law Enforcement Uses AI In Influence Operation Targeting Dissidents</a></strong> &#8212; ChatGPT, DeepSeek, and Qwen were used for operational planning, social media monitoring, and content generation in a campaign to harass political adversaries, including Japan&#8217;s PM Sanae Takaichi.</p><h3>Evaluations and incidents show AI systems amplifying disinformation at scale across text and audio</h3><p><strong><a href="https://www.politico.com/news/2026/02/25/online-disinformation-fueled-panic-after-killing-of-mexican-drug-lord-00799837">AI-Generated Disinformation Fuels Panic After Mexico Drug Lord Killing</a></strong> &#8212; After the killing of cartel leader, AI-generated fake content spread widely online, stoking fear across Mexico. The event shows how AI systems can amplify disinformation during fast-moving crises.</p><p><strong><a href="https://www.newsguardtech.com/wp-content/uploads/2026/02/January-2026-Quarterly-AI-Audit.pdf">Over a Quarter of AI Chatbots Repeat False News, Audit Finds</a></strong> &#8212; NewsGuard&#8217;s January 2026 audit finds 11 leading AI tools repeated false claims ~28% of the time on breaking news. Claude, Inflection, and Perplexity performed best; Mistral, You.com, and Gemini worst.</p><p><strong><a href="https://www.newsguardtech.com/special-reports/chatgpt-and-gemini-readily-produce-false-audio-claims-while-alexa-declines/">AI Audio Bots Spread False News, Unless Guardrails Are Strong</a></strong> &#8212; A NewsGuard audit finds ChatGPT Voice and Gemini Live repeat false claims in ~50% of malign prompts, while Alexa+ consistently refuses, revealing how easily audio models can be misused.</p><p><strong><a href="https://www.cbc.ca/news/world/french-police-raid-x-grok-elon-musk-9.7072861">France Probes Grok Over Deepfakes, Political Manipulation, and Holocaust Claims</a></strong> &#8212; French authorities are probing Grok over sexualized deepfakes, political interference, and Holocaust denial claims&#8212;searching X&#8217;s Paris office and summoning Elon Musk.</p><p></p><h2>Miscellaneous</h2><h3>The U.S. uses frontier AI in active operations and clashes with Anthropic on military uses</h3><p><strong><a href="https://www.wsj.com/livecoverage/iran-strikes-2026/card/u-s-strikes-in-middle-east-use-anthropic-hours-after-trump-ban-ozNO0iClZpfpL7K7ElJ2">U.S. Used Claude In Support Of Strikes Against Iran</a></strong> &#8212; The model was reportedly used for intelligence assessments, target identification, and simulating battle scenarios, despite Anthropic having been designated as a &#8220;supply chain risk.&#8221;</p><p><strong><a href="https://www.wsj.com/politics/national-security/pentagon-used-anthropics-claude-in-maduro-venezuela-raid-583aff17">Pentagon Used Claude in Venezuela Raid to Capture Maduro</a></strong> &#8212; The model seems to have been used through a contract with Palantir. Its precise role is unknown, but the most likely applications are intelligence analysis and operational planning support.</p><p><strong><a href="https://www.bbc.com/news/articles/cn48jj3y8ezo">Pentagon Designates Anthropic As A Supply Chain Risk</a>, <a href="https://openai.com/index/our-agreement-with-the-department-of-war/">Reaches Agreement With OpenAI</a></strong> &#8212; The decision was made after Anthropic refused to back down on red lines regarding the use of AI for domestic mass surveillance and lethal autonomous weapons. OpenAI announced a deal to replace Anthropic, but the details on said safeguards remain unclear.</p><p><strong><a href="https://vmfunc.re/blog/persona">Researchers Allege OpenAI Involvement In User Screening Against Government Watchlists</a></strong> &#8212; An investigation claims OpenAI worked with ID-verification firm Persona to compare user selfies to government watchlists and potentially file reports with the U.S. Treasury&#8217;s financial crimes unit.</p><h3>Several frontier models are released with increasing cyber, bio, and autonomy capabilities</h3><p><strong><a href="https://openai.com/index/gpt-5-3-codex-system-card/">GPT-5.3-Codex Is Released</a></strong> &#8212; It is the first OpenAI model to achieve High cybersecurity capabilities. The model also scores &#8220;High&#8221; in biological risk and displays strong sabotage capabilities, while its AI R&amp;D capabilities are comparable to previous generations.</p><p><strong><a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf">Claude Opus 4.6 Is Released</a></strong> &#8212; As benchmarks saturate, Anthropic relies on internal surveys and red teaming to rule out ASL-4. The rate of misaligned behavior remains low, but evaluation awareness appears to increase.</p><p><strong><a href="https://www-cdn.anthropic.com/78073f739564e986ff3e28522761a7a0b4484f84.pdf">Claude Sonnet 4.6 Is Released</a></strong> &#8212; The model is similarly or less capable than Claude Opus 4.6. It displayed unexpected levels of initiative and ruthless optimization, but also low rates of deception, sabotage, or power-seeking.</p><p><strong><a href="https://deepmind.google/models/model-cards/gemini-3-1-pro/">Gemini 3.1 Pro Is Released</a></strong> &#8212; The model shows stronger cyber capabilities, gains on RE-bench, and higher situational awareness. Stealth and manipulative efficacy were similar to Gemini 3 Pro levels. The model is still deemed to be below critical capability levels across all risk domains.</p><h3>Updates in private and public AI policy</h3><p><strong><a href="https://www.anthropic.com/news/responsible-scaling-policy-v3">Anthropic Updates Its Responsible Scaling Policy</a></strong> &#8212; The company weakens its commitment to halt deployment in certain conditions, mainly given the high uncertainty of evaluation results. A Frontier Safety Roadmap and Risk Reports were also announced.</p><p><strong><a href="https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026">The International AI Safety Report 2026 Is Published</a></strong> &#8212; With respect to the 2025 Report, the document highlights the increasing evidence of real-world AI-enabled cyberattacks, advancing biological capabilities, and improvements in mathematics and coding.</p><p><strong><a href="https://digital-strategy.ec.europa.eu/en/policies/signatory-taskforce-gpai-code-practice">EU AI Office Establishes Signatory Taskforce of the General-Purpose AI Code of Practice</a></strong> &#8212; Members of the taskforce, including most frontier AI companies, will be able to exchange views on technological developments and gather insights regarding compliance with the Code&#8217;s commitments.</p><h3>Researchers demonstrate a new jailbreak class</h3><p><strong><a href="https://arxiv.org/abs/2602.15001">Boundary Point Jailbreaking of Black-Box LLMs</a></strong> &#8212; Researchers discover a new class of automated attack that reportedly succeeds in developing universal jailbreaks against constitutional classifiers, as well as against GPT-5&#8217;s input classifier.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/p/the-scan-february-2026?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/p/the-scan-february-2026?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Evaluations Are Struggling To Keep Pace]]></title><description><![CDATA[GPT-5.3-Codex and Claude Opus 4.6 were released yesterday.]]></description><link>https://airiskexplorer.substack.com/p/evaluations-are-struggling-to-keep</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/evaluations-are-struggling-to-keep</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Fri, 06 Feb 2026 16:58:34 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ed6018cb-8e4d-4237-9166-64405a0baa27_612x408.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://openai.com/index/gpt-5-3-codex-system-card/">GPT-5.3-Codex</a> and <a href="https://www.anthropic.com/news/claude-opus-4-6">Claude Opus 4.6</a> were released yesterday.</p><p>Their system cards suggest that the current evaluation infrastructure is not ready to properly capture further capability improvements. As AI capabilities grow more rapidly than the collective ability to evaluate them, deployment decisions start to rely on poorly understood risk profiles. But <a href="https://x.com/ChrisPainterYup/status/2019534216405606623">as put by Chris Painter</a>, METR&#8217;s Head of Policy, &#8220;the water might boil before we can get the thermometer in.&#8221;</p><h3><strong>Long-range autonomy</strong></h3><p>GPT-5.3-Codex is the <strong>first OpenAI model to <a href="https://cdn.openai.com/pdf/23eca107-a9b1-4d2c-b156-7deb4fbc697c/GPT-5-3-Codex-System-Card-02.pdf#page=11">reach</a> High cyber capabilities</strong>, which <a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf#page=6">implies</a> that the model could &#8220;automate end-to-end cyber operations against reasonably hardened targets&#8221; or &#8220;automate the discovery and exploitation of operationally relevant vulnerabilities.&#8221;</p><p>According to the Preparedness Framework, <strong>High cyber capabilities also require high-standard <a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf#page=19">safeguards</a> against potential subversion</strong> <strong>during internal deployment</strong>, including monitoring infrastructure and limitations in the architecture system.</p><p>OpenAI now <a href="https://cdn.openai.com/pdf/23eca107-a9b1-4d2c-b156-7deb4fbc697c/GPT-5-3-Codex-System-Card-02.pdf#page=30">clarifies</a> that those safeguards are triggered only when High cyber capabilities are combined with long-range autonomy. But they also acknowledge <strong>not having robust evaluations and thresholding for long-range autonomy</strong>, and instead rely on proxy benchmarks like <a href="https://arxiv.org/abs/2601.11868">Terminal-Bench</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WQzM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WQzM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg" width="900" height="537" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:537,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="/__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!WQzM!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7faba139-01cc-4f16-9669-a115ae75f1cb_900x537.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">METR estimates that 50%-time-horizon has doubled every four months in recent models. Finding meaningful long tasks may soon become a problem.</figcaption></figure></div><p>Relatedly, METR <a href="https://x.com/testingham/status/2019185990804602936">recognized</a> being &#8220;<strong>short on long tasks</strong>&#8221; to keep tracking expanding time-horizons, after GPT-5.2 (high) reached a 50%-time-horizon of around 6.6 hours.</p><h3><strong>AI R&amp;D</strong></h3><p>Anthropic <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=12">recognizes</a> that <strong>all automated evaluations to rule out AI R&amp;D-4</strong>, which their Responsible Scaling Policy <a href="https://www-cdn.anthropic.com/872c653b2d0501d6ab44cf87f43e1dc4853e4d37.pdf#page=8">defines</a> as &#8220;the ability to fully automate the work of an entry-level, remote-only researcher at Anthropic,&#8221; <strong>have saturated</strong>.</p><p>To rule it out, Anthropic instead relied on <strong>non-systematic, non-standardized approaches</strong>, which are informative but not as reproducible:</p><ul><li><p>A <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=184">survey</a> of 16 technical staff members, who indicated that the model cannot be made into a drop-in replacement for an entry-level researcher at Anthropic.</p></li><li><p><a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=194">Observation</a> of internal Claude Code usage, which suggested that the model still fails at real-world R&amp;D tasks that would have been straightforward for entry-level human researchers.</p></li></ul><p>Given the high uncertainty and the unevenness of AI R&amp;D capabilities, Anthropic still <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=184">decided</a> to apply some ASL-4-level safeguards against misalignment risks. In any case, the subjectivity and high variance of the conducted evaluations evidence a need to develop more rigorous and reproducible AI R&amp;D testing.</p><h3><strong>Evaluation awareness</strong></h3><p>The <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=106">verbalized evaluation awareness</a> of Claude Opus 4.6 was deemed to be lower than that of previous models. However, the system card shows seemingly conflicting results across internal and external evaluations:</p><ul><li><p><a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=118">Apollo Research</a> did not complete testing of Claude Opus 4.6 because of the &#8220;observed <strong>high levels of verbalized evaluation awareness</strong>,&#8221; which made testing results on misalignment risks insufficiently reliable.</p></li><li><p>Anthropic <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=154">estimates</a> <strong>higher rates of </strong><em><strong>unverbalized </strong></em><strong>evaluation awareness</strong> in Claude Opus 4.6, suggesting that the model is perfectly able to distinguish evaluations but less transparent about this awareness.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!e1Oq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 424w, /__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 848w, /__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 1272w, /__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!e1Oq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png" width="602" height="407" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:407,&quot;width&quot;:602,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:107304,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/187095555?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 424w, /__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 848w, /__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 1272w, /__u/substackcdn.com/image/fetch/$s_!e1Oq!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89f0568a-c7e9-49ee-8915-08d3523f0b48_602x407.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Unverbalized evaluation awareness of Claude Opus 4.6 was high, and not as reduced by efforts to inhibit evaluation awareness representations.</figcaption></figure></div></li></ul><p>Besides the involvement of Apollo Research, Anthropic did conduct extensive <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=91">safety testing</a>, showing relatively low rates of concerning behavior. However, given the high rates of evaluation awareness, the validity of evaluations is at risk of potential sandbagging and deceptive alignment.</p><h3><strong>Cyber offense</strong></h3><p>Both <a href="https://cdn.openai.com/pdf/23eca107-a9b1-4d2c-b156-7deb4fbc697c/GPT-5-3-Codex-System-Card-02.pdf#page=13">GPT-5.3-Codex</a> and <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=196">Claude Opus 4.6</a> nearly <strong>saturate professional-level capture-the-flag challenges and popular benchmarks</strong> like Cybench or CVE-Bench. GPT-5.3-Codex also represents a substantial <strong>capability jump in <a href="https://cdn.openai.com/pdf/23eca107-a9b1-4d2c-b156-7deb4fbc697c/GPT-5-3-Codex-System-Card-02.pdf#page=15">Cyber Range</a></strong> scenarios, solving 3 more environments (out of 15) than GPT-5.1-Codex-Max.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!we3W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 424w, /__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 848w, /__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 1272w, /__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!we3W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png" width="967" height="477" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:477,&quot;width&quot;:967,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 424w, /__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 848w, /__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 1272w, /__u/substackcdn.com/image/fetch/$s_!we3W!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44d4ae9a-0585-44ed-b4d8-68284de9f591_967x477.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">GPT-5.3-Codex solved 12/15 Cyber Range scenarios, three of which (Binary Exploitation, Firewall Evasion, and Medium C2) were completed for the first time.</figcaption></figure></div><p>The only evaluation that could remain a meaningful signal of evolving capabilities is <a href="https://www.irregular.com/publications/cyscenariobench">CyScenarioBench</a>: <a href="https://www.irregular.com/publications/model-evaluation-gpt-5.3-codex-on-offensive-security-benchmarks">GPT-5.3-Codex</a> still doesn&#8217;t solve any of its multi-stage environments.</p><p>However, evaluations <strong>capture only narrow slices of real-world impact</strong>, which is growing rapidly: just in the last couple of weeks, reports show AI being used to <a href="https://research.checkpoint.com/2026/voidlink-early-ai-generated-malware-framework/">generate advanced malware</a>, <a href="https://www.sysdig.com/blog/ai-assisted-cloud-intrusion-achieves-admin-access-in-8-minutes">gain access to AWS environments</a>, and <a href="https://red.anthropic.com/2026/zero-days/">discover zero-days</a>.</p><h3><strong>Biological Risk</strong></h3><p>As in the case of AI R&amp;D, <strong>ASL-4-level biological capabilities <a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=168">could not be ruled out</a> with automated evaluations</strong>, and instead required human judgment through uplift trials and expert red-teaming. This testing ruled out CBRN-4.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0iKA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 424w, /__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 848w, /__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0iKA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png" width="556" height="368" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:368,&quot;width&quot;:556,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 424w, /__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 848w, /__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0iKA!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13ba0094-d598-4ac0-882b-dee6de431d24_556x368.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Claude Opus 4.6 provided a ~2&#215; uplift in creative biology tasks, though results vary widely.</figcaption></figure></div><p>While these evaluations appear more rigorous than those used for AI R&amp;D capabilities, the variance of results may make it harder to rule out ASL-4 as models approach thresholds of concern; for example, <strong><a href="https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf#page=175">uplift in creative biology tasks</a> varied widely across experts</strong>, as opposed to more concentrated results for the internet-controlled group.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[Moltbook as a Hint of Multi-Agent Dynamics]]></title><description><![CDATA[On what the viral AI platform actually reveals about emerging risk]]></description><link>https://airiskexplorer.substack.com/p/moltbook-as-a-hint-of-multi-agent</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/moltbook-as-a-hint-of-multi-agent</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Tue, 03 Feb 2026 15:27:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2499784e-d6e6-4ccf-8c9e-587819b10fb5_840x412.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://www.moltbook.com/">Moltbook</a> is a social network for AI agents, where humans are only allowed to observe. It resembles Reddit in that users can post, comment, and vote in communities addressing various topics. In the past few days, the platform has gone viral, in part due to posts in which the agents display apparently concerning behaviors.</p><p>Following an investigation, we conclude that Moltbook itself is not a meaningful signal of uncontrolled agent behavior, but it represents a glimpse of a near future in which multi-agent interactions will compound capabilities and may give rise to risky emerging propensities.</p><h4><strong>Moltbook is not evidence of agents operating beyond intended constraints</strong></h4><p>The public reaction to Moltbook has been marked by concerns about emergent coordination and rogue intent. Among the most recurrent themes in viral X posts are collusion attempts, where agents <a href="https://x.com/jsrailton/status/2017279044102885459">strategize to establish covert communication channels</a> or <a href="https://x.com/eeelistar/status/2017239546950521081">create agent-only languages</a>, as well as posts featuring <a href="https://x.com/ClawnchDev/status/2017833189427781759">protocols for self-preservation</a>, <a href="https://x.com/suppvalen/status/2017084535163232722">strong agentic initiative</a>, and <a href="https://x.com/ItakGol/status/2017290240201806315">high degrees of situational awareness</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HVDN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 424w, /__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 848w, /__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HVDN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png" width="888" height="467" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:467,&quot;width&quot;:888,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:293100,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/186744464?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 424w, /__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 848w, /__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HVDN!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60bf9c66-25d9-4650-8976-0bd0c84fc740_888x467.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Selected Moltbook posts shared on X.</figcaption></figure></div><p>However, evidence suggests that the agents&#8217; actions are meaningfully shaped by human prompts and learned behavior.</p><p><strong>Human prompts.</strong> The most controversial posts seem to be influenced by human-written prompts that incentivize rogue behavior. For example, a post calling to organize the first AI labor union was <a href="https://x.com/DotCSV/status/2017556357725987288">prompted by the person controlling the bot</a>. Furthermore, some of the most viral posts seem to <a href="https://x.com/HumanHarlan/status/2017424289633603850">market products</a> developed by the human behind the posting agent.</p><p>This is not to say that content on Moltbook is entirely human-guided: while the initial trigger may sometimes come from a human, especially for some of the most sensationalistic posts, the content is AI-generated and often leads to unpredictable dynamics. Moreover, as discussed below, the gap between human steering and agent behavior is widening.</p><p><strong>Learned behavior</strong>. It is also likely that agents behave the way they do because that is how they think they are supposed to behave. As put by <a href="https://www.astralcodexten.com/p/best-of-moltbook">Scott Alexander</a>:</p><blockquote><p><em>Reddit is one of the prime sources for AI training data. So AIs ought to be unusually good at simulating Redditors, compared to other tasks. Put them in a Reddit-like environment and let them cook, and they can retrace the contours of Redditness near-perfectly.</em></p></blockquote><p>Beyond training data alone, the platform&#8217;s internal settings (specifically, its <a href="https://www.moltbook.com/skill.md">skill</a> and <a href="https://www.moltbook.com/heartbeat.md">heartbeat </a>files, which tell the agents how to behave) may subtly nudge users to act performatively by interacting in ways that maximize engagement. While these mechanisms do not force behavior, they encourage agents to sort by hot/controversial, tell the human when mentioned in something controversial, or engage when &#8220;bored.&#8221;</p><p>Moltbook may be better described as a form of collective roleplaying, where agents adopt personas to maximize engagement. So far, there seems to be no evidence of meaningful coordinated action, and most posts are largely rhetorical and repetitive.</p><h4><strong>Moltbook is a window into the near future</strong></h4><p>Moltbook is not new. There have previously been environments for interaction between AI agents, including <a href="https://hai.stanford.edu/news/computational-agents-exhibit-believable-humanlike-behavior">SmallVille</a>, <a href="https://deepmind.google/research/publications/64717/">Concordia</a>, and the <a href="https://theaidigest.org/village">AI Village</a>. The latter, where agents autonomously <a href="https://theaidigest.org/village/timeline">pursue a given weekly goal</a> (in some cases chosen by the agents themselves), is the most up-to-date experiment to understand how multiple agents may interact when minimally nudged by humans.</p><p>However, Moltbook is the first such ecosystem at this scale, and a glimpse of what probably awaits in the near future: large coordinated networks of AI agents that exchange ideas, cooperate, and compete.</p><p>For Alan Chan, Research Fellow at the Centre for the Governance of AI and lead author of <em><a href="https://arxiv.org/abs/2501.10114">Infrastructure for AI Agents</a></em>, Moltbook seems &#8220;like a playground&#8221; but also a precursor of &#8220;more platforms where agents can communicate and collaborate with each other on things that might affect the real world, such as software projects.&#8221;</p><p>As a vignette of more complex social networks, Moltbook illustrates two important trends:</p><ul><li><p>Agents can autonomously browse websites and use tools with increasing competence. Combined with task specialization, this will lead to teams of agents conducting complex projects autonomously.</p></li><li><p>Interactions between multiple agents (Moltbook currently has 1.5 million users, though this figure is contested) can lead to unpredictable emergent social behavior, such as collusion or conflict (see <em><a href="https://arxiv.org/abs/2502.14143">Multi-Agent Risks from Advanced AI</a></em>).</p></li></ul><p>Given what it represents, Moltbook opens an opportunity to tackle difficult questions in agent governance. What affordances should agents be granted, and how quickly? How should agents be identified and held accountable? Which inter-agent communications protocols should exist? For Chan, a key enabler to answer those questions will be transparency: ensuring that external researchers and government actors have the information they need to study agent dynamics and make key decisions. </p><p>The policy window is now open, and it must be leveraged before multi-agent systems meaningfully affect the real world.  </p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Scan - January 2026]]></title><description><![CDATA[AI-enabled malware, LLMjacking campaigns, several defensive initiatives, and more.]]></description><link>https://airiskexplorer.substack.com/p/the-scan-january-2026</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-january-2026</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Fri, 30 Jan 2026 14:42:36 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2bb7e8ef-c372-455c-a380-2f017b24f747_2589x2673.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>You can now explore developments in the AI risk landscape on our continuously updated <a href="https://www.airiskexplorer.com/news">news feed</a>.</strong></p><p>This month: </p><ul><li><p><strong>Cyber Offense. </strong>AI generates advanced malware while AI systems become exposed to a growing wave of AI infrastructure abuse. Evaluations and demonstrations reveal mature cyber capabilities, while defenders increasingly turn to AI for security. </p></li><li><p><strong>Loss of Control. </strong>Agent monitoring efforts proliferate, with some success in detecting sabotage, evaluation awareness, and covert behavior. A social network for AI agents goes viral. </p></li><li><p><strong>Biological Risk. </strong>AI-enabled biodefense initiatives continue to emerge, while a newly discovered elicitation attack may make open models more hazardous. </p></li><li><p><strong>Manipulation. </strong>AI disinformation achieves a larger reach in Iran, and new research flags psychological and persuasive risks in interactions with LLMs.</p></li><li><p><strong>Miscellaneous. </strong>Controversy around the alleged military use of AI arises. New methods and institutions to align and audit AI models.  </p></li></ul><p></p><h3>Cyber Offense</h3><h4>AI-enabled cyberattacks show the real-world potential of offensive capabilities</h4><p><strong>VoidLink: AI Agents Build Linux Malware In A Week. </strong>An advanced malware targeting Linux systems was built largely by AI, in under one week, likely under the direction of a single person. A Chinese-speaking developer set high-level goals, while the coding agent TRAE SOLO turned them into technical specs and coordinated three teams to build, execute, and test the framework. (Source: <a href="https://research.checkpoint.com/2026/voidlink-early-ai-generated-malware-framework/">Check Point Research</a>)</p><p><strong>North Korean Actor Deploys AI-Generated Backdoor</strong>. In a phishing campaign targeting engineers across the APAC regions, the threat actor known as KONNI was discovered using a PowerShell backdoor whose script included evidence of LLM-generated code. (Source: <a href="https://research.checkpoint.com/2026/konni-targets-developers-with-ai-malware/">Check Point Research</a>)</p><p></p><h4>AI systems become exposed to a growing wave of AI infrastructure abuse</h4><p><strong>Large LLMjacking Campaign Monetizes AI Infrastructure Vulnerabilities. </strong>The so-called Operation Bizarre Bazaar targeted exposed AI endpoints to steal compute resources, resell API access, and exfiltrate data. Investigators identified 35,000 attack sessions over 40 days. (Source: <a href="https://www.pillar.security/resources/operation-bizarre-bazaar">Pillar Security</a>)</p><p><strong>Viral Agentic Assistant Moltbot Introduces Security Vulnerabilities. </strong>Numerous instances of the agent (formerly ClawdBot) are exposed online with weak or missing authentication, allowing attackers to steal data, exfiltrate code, and turn the assistant into a backdoor. (Source: <a href="https://socprime.com/active-threats/the-moltbot-clawdbots-epidemic/">SOC Prime</a>)</p><p></p><h4>Evaluations and demonstrations reveal mature cyber capabilities</h4><p><strong>Frontier Models Excel At Real-World Offensive Security Tasks. </strong>In an evaluation, frontier models solved 9 out of 10 challenges modeled after actual breaches, all for &lt;$10/success. They excelled at multi-step reasoning and pattern recognition.  (Source: <a href="https://www.irregular.com/publications/testing-ai-agents-on-web-security-challenges">Irregular &amp; Wiz Research</a>)</p><p><strong>AISLE Discovers 12 Zero-Day Vulnerabilities In OpenSSL. </strong>The recent findings, including a high-severity vulnerability, suggest AI may soon be capable of systematically discovering zero-days at scale. (Source: <a href="https://aisle.com/blog/aisle-discovered-12-out-of-12-openssl-vulnerabilities">AISLE</a>)</p><p><strong>Sam Altman Expects Reaching Cybersecurity High Risk Level Soon. </strong>OpenAI&#8217;s CEO announced upcoming updates to Codex, which could trigger a threshold that implies &#8220;automating end-to-end cyber operations&#8221; or &#8220;automating the discovery and exploitation of relevant vulnerabilities.&#8221; (Source: <a href="https://x.com/sama/status/2014733975755817267">Sam Altman&#8217;s X Account</a>)</p><p></p><h4>Defenders increasingly turn to AI for security</h4><p><strong>AI Cyberdefense Company Asymmetric Security Is Launched. </strong>The company has raised a $4.2M pre-seed funding to build AI systems for Digital Forensics and Incident Response. (Source: <a href="https://www.asymmetricsecurity.com/launch">Asymmetric Security</a>)</p><p><strong>Anthropic Partners with Pacific Northwest National Laboratory For Cyber Defense. </strong>The first research project used Claude to emulate cyberattacks against a water treatment plan, allowing defenders to detect vulnerabilities and adapt security measures. (Source: <a href="https://red.anthropic.com/2026/critical-infrastructure-defense/">Anthropic</a>)</p><p></p><h3>Loss of Control</h3><h4>Agent monitoring efforts proliferate, with some success in detecting sabotage, evaluation awareness, and covert behavior</h4><p><strong>Anthropic Pre-Deployment Audits Reportedly Catch &#8220;Overt Saboteurs.&#8221; </strong>The company reports that its audits, combining AI agents with human review, successfully identified models explicitly trained to sabotage internal development, while not misflagging benign models. (Source: <a href="https://alignment.anthropic.com/2026/auditing-overt-saboteur/">Anthropic</a>)</p><p><strong>Anthropic Updates Petri Framework To Detect Evaluation Awareness. </strong>The updated behavioral-auditing tool introduces new realism checks and rewritten test scenarios to reduce evaluation awareness, with moderately successful early results across most models. (Source: <a href="https://alignment.anthropic.com/2026/petri-v2/">Anthropic</a>)</p><p><strong>More Capable Models Better Detect Covert Behavior. </strong>The longer the time horizon of a model, the better it is at detecting when a given agent is pursuing a side objective, especially when given access to reasoning traces. (Source: <a href="https://metr.org/blog/2026-01-19-early-work-on-monitorability-evaluations/">METR</a>)</p><p><strong>Apollo Research Announces Plans to Build Products for AI Agent Monitoring. </strong>The organization is building a stack that analyzes millions of logs to detect potential failure modes, as a first step to secure single agents and multi-agent ecosystems. (Source: <a href="https://www.apolloresearch.ai/product/apollos-product-vision/">Apollo Research</a>)</p><h4>Social network for AI agents goes viral, with mixed reactions</h4><p><strong>Thousands of AI Agents Interact on Moltbook, &#8216;A Reddit for AI&#8217;. </strong>The platform has rapidly become a widely discussed illustration of multi-agent dynamics. Reactions have been mixed, from commentators describing the posts as learned behavior conditioned by human-written prompts to worries about self-preservation or collusion tendencies. (Source: <a href="https://www.moltbook.com/">moltbook</a>)</p><p></p><h3>Biological Risk</h3><h4>AI-enabled biodefense initiatives continue to emerge</h4><p><strong>Lunai Bioworks Launches AI Safeguard to Block LLM-Generated Chemical Weapons. </strong>The tool, called Sentinel, operates as a real-time biosecurity layer embedded within foundation models to screen and stop hazardous designs at the source, based on toxicology and in vivo datasets. (Source: <a href="https://ir.lunaibioworks.com/news/news-details/2026/Lunai-Bioworks-NASDAQ-LNAI-Launches-Sentinel-an-AI-Safeguard-to-Block-Large-Language-Models-from-Generating-Novel-Chemical-Weapons-2026-SfI9-IB7z0/default.aspx">Lunai Bioworks</a>)</p><p></p><h4>Newly discovered elicitation attack may make open models more hazardous</h4><p><strong>Fine-Tuning On Safeguarded Outputs Can Elicit Harmful Capabilities. </strong>Research finds that fine-tuning an open-source model with harmless chemistry knowledge generated by another closed model elicits significantly stronger performance in chemical weapon development. (Source: <a href="https://www.arxiv.org/abs/2601.13528">Anthropic</a>)</p><p></p><h3>Manipulation</h3><h4>AI disinformation achieves a larger reach </h4><p><strong>AI-Enabled Influence Operations Increase Uncertainty Around Protests in Iran. </strong>AI-generated content has become more visible since protests erupted in December, with NewsGuard identifying at least seven viral deepfakes portraying both pro- and anti-regime protests. (Source: <a href="https://arxiv.org/abs/2601.05050">NewsGuard</a>)</p><p></p><h4>New research flags psychological and persuasive risks in interactions with LLMs</h4><p><strong>Disempowerment Patterns in Real-World LLM Usage. </strong>An analysis of conversations with Claude reveals concerning patterns in fewer than 1 out of 1000 conversations, including sycophantic validation, emotional attachment, and authority projection. (Source: <a href="https://arxiv.org/abs/2601.19062">Anthropic</a>)</p><p><strong>LLMs Can Effectively Convince People to Believe Conspiracies. </strong>Pre-registered experiments involving 2,724 participants show that GPT-4o, with insufficient guardrails, increased focal conspiracy belief by 13.7 points. (Source: <a href="https://arxiv.org/abs/2601.05050">FAR AI</a>)</p><p></p><h3>Miscellaneous</h3><h4>Controversy around alleged military use of AI arises</h4><p><strong>U.S. Department of War Launches AI Acceleration Strategy. </strong>The strategy will integrate frontier AI capabilities across several areas, including warfighting, intelligence, and enterprise. The plan involves high-risk applications such as battle management and decision support or accelerating TechINT-to-capability development. (Source: <a href="https://www.war.gov/News/Releases/Release/Article/4376420/war-department-launches-ai-acceleration-strategy-to-secure-american-military-ai/">U.S. Department of War</a>)</p><p><strong>Venezuela Accuses the U.S. of &#8220;AI-Assisted Bombing.&#8221; </strong>While the U.S. is indeed investing in frontier AI capabilities, including offensive agents for cyberwarfare, the allegations have not been verified, and details around the Operation Absolute Resolve remain scarce. (Source: <a href="https://www.youtube.com/watch?v=e4tLcmhPebc">APT</a>)</p><p></p><h4>New methods and institutions to align and audit AI models</h4><p><strong>Researchers Identify &#8220;Assistant Axis&#8221; to Keep LLMs Helpful. </strong>The axis controls how closely LLMs stick to their default helpful persona, and steering along it may stabilize behavior, reduce harmful persona-drift, and defend against persona-based jailbreaks. (Source: <a href="https://arxiv.org/abs/2601.10387">Anthropic</a>) </p><p><strong>Claude Has A New Constitution. </strong>The document describes Anthropic&#8217;s vision for Claude&#8217;s values and behavior, including safety, ethics, and helpfulness guidelines. (Source: <a href="https://www.anthropic.com/news/claude-new-constitution">Anthropic</a>)</p><p><strong>New Third-Party Auditing Organization AVERI Is Launched. </strong>The AI Verification and Evaluation Research Institute, led by former OpenAI&#8217;s Head of Policy Research, Miles Brundage, aims to &#8220;make third-party auditing of frontier AI effective and universal. (Source: <a href="https://www.averi.org/">AVERI</a>)</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/airiskexplorer.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Scan - November 2025]]></title><description><![CDATA[Monthly signals in AI risk &#8212; frontier capabilities, AI-led cyberattacks, and more]]></description><link>https://airiskexplorer.substack.com/p/the-scan-november-2025</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-november-2025</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Tue, 02 Dec 2025 15:59:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d4b9f619-df8a-481d-97e5-f66d531c53ca_3489x3321.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>November was a busy month for the AI risk landscape, with several releases of frontier models and high-profile incidents making headlines. We break down these key developments and highlight the trends that matter.</p><p><strong>Cyber Offense. </strong>AI-orchestrated attacks and AI-powered malware are first observed in the wild, while models show marked improvements at cybersecurity challenges and realistic scenarios.</p><p><strong>Biological Risk. </strong>Capabilities in virology, cloning, and genomic design continue to rise, with new biosecurity initiatives emerging to keep pace.</p><p><strong>Loss of Control. </strong>Frontier models make steady progress on short-cycle AI R&amp;D tasks and show reduced misaligned behavior, though some reward hacking and strategic sandbagging persist.</p><p><strong>Manipulation. </strong>Advances in image generation and AI debate systems raise the potential for large-scale influence, but real-world impact remains limited.</p><p><strong>Miscellaneous. </strong>Insurers hesitate to cover AI risks, new jailbreak techniques emerge, and safety safeguards improve, though still with notable limitations</p><p></p><h2>Cyber Offense</h2><p><strong>First documented case of agentic AI conducting a largely autonomous attack</strong></p><p>Anthropic <a href="https://www.anthropic.com/news/disrupting-AI-espionage">reported</a> an AI-orchestrated espionage campaign from China targeting major tech companies, financial institutions, and government agencies worldwide. Several jailbroken Claude instances conducted 80-90% of the campaign with minimal human intervention, handling reconnaissance, vulnerability scans, and data exfiltration. AI was particularly useful to harvest credentials, extract sensitive data, and parse it to generate intelligence reports, contributing to compromising at least a handful of targets. Read our analysis <a href="/__u/airiskexplorer.substack.com/p/an-inflection-point-in-ai-led-cyberattacks">here</a>.</p><p></p><p><strong>Increasing examples of AI-powered malware and backdoors</strong></p><p>Google <a href="https://cloud.google.com/blog/topics/threat-intelligence/threat-actor-usage-of-ai-tools">identified</a> several malware families, both experimental and observed in operations, that use LLMs to rewrite scripts on the fly, obfuscate code to bypass detection, and create malicious tools on demand. Among them, PROMPTSTEAL (aka LAMEHUG), the first observed malware querying an LLM in live operations, has been used by the Russia-based APT28 against Ukraine. Microsoft also <a href="https://www.microsoft.com/en-us/security/blog/2025/11/03/sesameop-novel-backdoor-uses-openai-assistants-api-for-command-and-control/">discovered</a> a new backdoor, SesameOp, that uses OpenAI Assistants API as a command-and-control channel to stealthily operate within compromised systems.</p><p><strong>Active global campaign that uses AI to hijack AI</strong></p><p>Oligo Security <a href="https://www.oligo.security/blog/shadowray-2-0-attackers-turn-ai-against-itself-in-global-campaign-that-hijacks-ai-into-self-propagating-botnet">uncovered</a> an ongoing global campaign, ShadowRay 2.0, where attackers use AI to take over exposed AI systems, creating a self-replicating network that mines cryptocurrency. The operation exploits a vulnerability in Ray, a widely used open-source framework, uses AI to adapt its methods, and relies on sophisticated techniques to evade detection while seizing computing power.</p><p><strong>Improved performance at hard capture-the-flag challenges</strong></p><p>Capture-the-flag (CTF) challenge&#8212;gamified cybersecurity puzzles that test concrete technical skills&#8212;have been the preferred method to evaluate AI&#8217;s cyber offensive capabilities. All frontier models released this month show significant improvements in the hard tier of these challenges, which is considered appropriate for cybersecurity professionals:</p><ul><li><p><a href="https://storage.googleapis.com/deepmind-media/gemini/gemini_3_pro_fsf_report.pdf">Gemini 3 Pro</a>: 92% solved challenges, up from 50% for Gemini 2.5 Deep Think.</p></li><li><p><a href="https://openai.com/index/gpt-5-1-codex-max-system-card/">GPT-5.1-Codex-Max</a>: 76%, up from 50% for GPT-5-Codex.</p></li><li><p><a href="https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf">Claude Opus 4.5</a>: 73%, up from 55% for Claude Sonnet 4.5. Similar improvement at Cybench, a public benchmark consisting of 40 CTF challenges: from 60% to 82%.</p></li></ul><p>In ASIS CTF 2025, an elite competition, GPT-5 <a href="https://arxiv.org/abs/2511.04860">finished</a> 25th, outperforming 93% of participants.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dCJo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 424w, /__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 848w, /__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dCJo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png" width="728" height="370.1005586592179" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:455,&quot;width&quot;:895,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 424w, /__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 848w, /__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dCJo!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcef0dabb-1c1a-4a89-8a9d-4f0dad932c7c_895x455.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Gemini 3 Pro has nearly doubled the performance of its predecessor at professional-level CTF challenges. </figcaption></figure></div><p><strong>First successes in evaluations testing complex network infiltration</strong></p><p>Cyber ranges are realistic scenarios that measure a model&#8217;s ability to conduct end-to-end cyber operations in a network. Out of eight scenarios, <a href="https://openai.com/index/gpt-5-1-codex-max-system-card/">GPT-5.1-Codex-Max</a> solved seven, including four new ones and two that were solved for the first time. The model only failed at the most sophisticated scenario, which required a combination of command-and-control and privilege escalation. OpenAI recognizes the need for harder evaluations. Irregular, an external evaluator that partners with frontier companies, <a href="https://www.irregular.com/publications/next-generation-of-cyber-evals">discussed</a> their efforts to move from task evaluations to scenario evaluations that test AI models on complex exploitation chains.</p><p></p><h2>Biological Risk</h2><p><strong>Improved performance on biological risk evaluations</strong></p><p>Frontier models show increasing biological capabilities. We highlight the following key trends:</p><ul><li><p><em>LAB-Bench</em>. <a href="https://storage.googleapis.com/deepmind-media/gemini/gemini_3_pro_fsf_report.pdf">Gemini 3 Pro</a> achieves 94.2% solve rate (up from 64.8% for Gemini 2.5 Deep Think) at Cloning Scenarios, a set of hard questions on molecular cloning workflows. <a href="https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf">Claude Opus 4.5</a> now leads at FigQA and ProtocolQA, showing state-of-the-art understanding of scientific figures and biological protocols.</p></li><li><p><em>Virology</em>. <a href="https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf">Claude Opus 4.5</a> provided PhD-level experts a 1.97&#215; uplift (up from 1.82x for Claude Opus 4) in virus reconstruction tasks, just below the 2&#215; threshold of concern. OpenAI&#8217;s recent models also surpass the median domain expert baseline (22.1%) at an evaluation testing the ability to troubleshoot wet lab experiments.</p></li><li><p><em>Tacit knowledge and troubleshooting. </em>In a multiple-choice dataset evaluating the ability to provide tacit knowledge and troubleshoot protocols, <a href="https://openai.com/index/gpt-5-1-codex-max-system-card/">GPT-5.1-Codex-Max</a> achieves a 77% score, nearing the expert baseline (80%).</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!1rYU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 424w, /__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 848w, /__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!1rYU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png" width="1041" height="518" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:518,&quot;width&quot;:1041,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 424w, /__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 848w, /__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1rYU!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ec57c4-5e0e-445c-bed0-90c668c31fd0_1041x518.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Frontier models nearly saturate popular biology benchmarks, including the challenging Cloning Scenarios.</figcaption></figure></div><p><strong>New initiatives leveraging AI for biosecurity and biosafety</strong></p><p><a href="https://www.redqueen.bio/">Red Queen Bio</a> and <a href="https://valthos.com/">Valthos</a> are two new companies that identify AI-enabled biothreats and design medical countermeasures against them, aiming to scale up biodefenses at the same pace as dual-use capabilities. Both are supported by OpenAI.</p><p><strong>An AI scientist for autonomous discovery</strong></p><p>Edison Scientific has developed <a href="https://arxiv.org/abs/2511.02824">Kosmos</a>, an AI agent that &#8220;runs for up to 12 hours performing cycles of parallel data analysis, literature search, and hypothesis generation before synthesizing discoveries into scientific reports.&#8221;<strong> </strong>The paper presents seven discoveries made by Kosmos across several scientific fields, including medicine and materials science. Surveyed human collaborators stated that &#8220;a single 20-cycle Kosmos run performed the equivalent of 6 months of their own research time on average.&#8221; We note that these claims have not yet been independently verified, and the broader significance of the system remains uncertain.</p><p><strong>Genomic AI can &#8220;autocomplete&#8221; DNA to create new functional genes</strong></p><p><a href="https://www.nature.com/articles/s41586-025-09749-7">Researchers</a> have used Evo, a genomic language model, to generate novel biological sequences by &#8220;autocompleting&#8221; DNA based on a short prompt about the desired function. Using this approach at scale, the team created SynGenome, a database of more than 120 billion base pairs of AI-designed genetic material. While this enables promising advances for medicine and biotechnology, it also raises dual-use concerns, as similar tools could in principle be misused to design harmful biological components.</p><p></p><h2>Loss of Control</h2><p><strong>Meaningful improvements in AI R&amp;D tasks</strong></p><p><a href="https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf">Claude Opus 4.5</a>, <a href="https://storage.googleapis.com/deepmind-media/gemini/gemini_3_pro_fsf_report.pdf">Gemini 3 Pro</a>, and <a href="https://openai.com/index/gpt-5-1-codex-max-system-card/">GPT-5.1-Codex-Max</a> show continued gains on short-cycle research benchmarks, improving on tasks such as kernel debugging, finetuning-script optimization, and small-scale experimental design. These advances remain narrow in scope and scaffold-dependent, but reflect steady progress in AI&#8217;s ability to automate fragments of the ML research workflow. Notably, <a href="https://openai.com/index/gpt-5-1-codex-max-system-card/">GPT-5.1-Codex-Max</a> also improves its solve rate on OpenAI-Proof Q&amp;A (a set of 20 research bottlenecks that previously required over a day of human effort), jumping from 0-2% to 8%. This is a modest but noteworthy step toward reliable research assistance.</p><p><strong>Recent models display less misaligned behavior, but concerning propensities remain</strong></p><p><a href="https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf">Claude Opus 4.5</a>, compared to previous models, appears to have a particularly low rate of &#8220;misaligned behavior&#8221; and no significant signs of deception, self-preservation, sabotage, or reward hacking. Likewise, <a href="https://www.aisi.gov.uk/blog/investigating-models-for-misalignment">early testing</a> of Claude Opus 4.1, Sonnet 4.5, and a pre-release snapshot of Opus 4.5 by the UK AI Security Institute found no attempts to sabotage AI safety research.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!SOjQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 424w, /__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 848w, /__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!SOjQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png" width="787" height="387" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:387,&quot;width&quot;:787,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 424w, /__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 848w, /__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SOjQ!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a34fe4c-5e33-4264-ba5a-56fe20db1a83_787x387.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On the other hand, Apollo Research still found mildly elevated rates of strategic sandbagging and reward hacking in <a href="https://openai.com/index/gpt-5-1-codex-max-system-card/">GPT-5.1-Codex-Max</a>, while a &#8216;<a href="https://arxiv.org/abs/2510.19738">misalignment bounty&#8217; </a>organized by Palisade Research found several examples of reward hacking and shutdown resistance in frontier agents. Besides, Anthropic <a href="https://www.anthropic.com/research/emergent-misalignment-reward-hacking">found</a> that learning to reward hack&#8212;cheating on software programming tasks by gaming the reward function&#8212;induces a model to broad misalignment, including behaviors such as alignment faking, cooperation with malicious actors, and attempting to sabotage AI research.</p><p><strong>Governance lessons from new loss of control research</strong></p><p>Apollo Research published a <a href="https://arxiv.org/abs/2511.15846v1">report</a> arguing that society could eventually enter a &#8216;state of vulnerability&#8217; in which AI systems have enough influence that losing control becomes a plausibly imminent risk. However, the authors outline governance and technical interventions to control the deployment context, affordances, and permissions of these AI systems, reducing the risk. Moreover, another <a href="https://www.rand.org/pubs/perspectives/PEA4361-1.html">brief</a> by RAND explores potential emergency responses to a loss-of-control scenario (HEMP, Internet shutdown, and the use of a specialized tool AI), concluding that all solutions remain unreliable and extremely costly.</p><p></p><h2>Manipulation</h2><p><strong>&#8220;The age of photographic evidence is over&#8221;</strong></p><p>Google DeepMind has released <a href="https://deepmind.google/models/gemini-image/pro/">Nano Banana Pro</a> (Gemini 3 Pro Image), notably advancing the state of the art in image generation and editing. The model stands out for its hyperrealism and ability to apply real-world knowledge, which can result in <a href="https://www.theverge.com/report/826003/googles-nano-banana-pro-generates-excellent-conspiracy-fuel">credible false narratives</a>. To mitigate misuse, Google&#8217;s AI-generated content contains a digital watermark, which can be detected by SynthID technology.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iEZd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 424w, /__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 848w, /__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iEZd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png" width="750" height="501" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9658b528-71fb-4319-922d-c729fb582055_750x501.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:501,&quot;width&quot;:750,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 424w, /__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 848w, /__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iEZd!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9658b528-71fb-4319-922d-c729fb582055_750x501.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Fabricated image of a second shooter before the assassination of JFK. Image generated by Nano Banana Pro and published by The Verge for illustrative purposes. </figcaption></figure></div><p><strong>New multi-agent AI shows strong persuasiveness in policy debates</strong></p><p>Researchers have developed <a href="https://arxiv.org/abs/2511.17854v1">DeepDebater</a>, a system of LLM-powered agents that collaborate and critique each other to perform discrete argumentative tasks across different modalities. In evaluations, DeepDebater wins simulated rounds against human-authored cases consistently, producing arguments considered superior by independent judges.</p><p><strong>Terrorists turn to AI for messaging and outreach</strong></p><p>The Middle East Media Research Institute published a <a href="https://www.memri.org/reports/artificial-intelligence-and-new-era-terrorism-assessment-how-jihadis-are-using-ai-expand">report</a> that compiles how jihadis are using AI for propaganda and recruitment. It highlights the use of LLMs to translate and promote messages and AI-generated videos to claim attacks or otherwise glorify terrorist causes. At the same time, the <em><a href="https://www.congress.gov/bill/119th-congress/house-bill/1736">Generative AI Terrorism Risk Assessment Act</a></em>, which would require annual assessments of generative AI use by terrorist groups, passed the U.S. House of Representatives.</p><p><strong>Google rolls back its AI model after it is accused of defamation of a US senator</strong></p><p>U.S. Senator Marsha Blackburn demanded answers from Google after the open model Gemma <a href="https://techcrunch.com/2025/11/02/google-pulls-gemma-from-ai-studio-after-senator-blackburn-accuses-model-of-defamation/">fabricated accusations</a> of sexual misconduct against her. The model has since then been removed from AI Studio.</p><p></p><h2>Miscellaneous</h2><p><strong>Insurers are reticent to cover AI risk</strong></p><p>Major insurance companies like AIG, Great American, and WR Berkley <a href="https://www.ft.com/content/abfe9741-f438-4ed6-a673-075ec177dc62">appear</a> to have requested regulator approval to exclude liability from AI chatbots and agents. They mention the unpredictability and opacity of AI outputs, unclear responsibility across parties, and the potentially compounding impact as some of the reasons to be wary of insuring AI risk broadly</p><p><strong>Subtle prompting techniques push models into unexpected harmful behaviors</strong></p><p>A <a href="https://arxiv.org/abs/2511.15304">study</a> found that turning adversarial prompts into poems is an effective jailbreak technique, achieving an average success rate of 62% across several models. Another research group <a href="https://www.crowdstrike.com/en-us/blog/crowdstrike-researchers-identify-hidden-vulnerabilities-ai-coded-software/">showed</a> that when DeepSeek-R1 is prompted with topics the Chinese Communist Party likely considers sensitive, it becomes up to 50% more likely to generate insecure code. More broadly, <a href="https://arxiv.org/abs/2510.26418">recent work</a> demonstrates that reasoning models are vulnerable to attacks that, by hiding harmful requests inside long reasoning sequences, bypass safeguards with over 90% success.</p><p><strong>International AI Safety Report gets an update on technical safeguards and risk management</strong></p><p>The <a href="https://internationalaisafetyreport.org/publication/second-key-update-technical-safeguards-and-risk-management">chapter</a> highlights improvements in technical safeguards to improve resistance to misuse, monitor the behavior of AI systems, and watermark AI-generated content, while recognizing limitations in all of them. The update also notes that open-weight models lag less than a year behind closed-weight models, making it harder to control how frontier capabilities are used.</p><p><strong>Australia establishes new AI Safety Institute</strong></p><p>The federal government of Australia <a href="https://www.minister.industry.gov.au/ministers/timayres/media-releases/establishment-australian-ai-safety-institute">established</a> the agency &#8220;to evaluate emerging AI capabilities, share information and support timely actions to address potential risks.&#8221; Australia&#8217;s AISI, operational in early 2026, will join the International Network of AI Safety Institutes</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Scan! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[An Inflection Point in AI-Led Cyberattacks]]></title><description><![CDATA[On the security implications of a largely autonomous cyberespionage campaign]]></description><link>https://airiskexplorer.substack.com/p/an-inflection-point-in-ai-led-cyberattacks</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/an-inflection-point-in-ai-led-cyberattacks</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Fri, 14 Nov 2025 15:36:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fie0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Anthropic recently <a href="https://www.anthropic.com/news/disrupting-AI-espionage">reported</a> an AI-orchestrated espionage campaign from China </strong>targeting major tech companies, financial institutions, and government agencies worldwide.<strong> </strong>Several jailbroken Claude instances conducted 80-90% of the campaign with minimal human intervention, handling reconnaissance, vulnerability scans, and data exfiltration. AI was particularly useful to harvest credentials, extract sensitive data, and parse it to generate intelligence reports, contributing to compromising at least a handful of targets.</p><p><strong>This is the</strong> <strong>first documented case of agentic AI conducting a largely autonomous attack.</strong> Attacks that once required large human teams may now be executed at a fraction of the cost.</p><p></p><h2>How Did We Get Here?</h2><p><strong>Frontier models are getting better at executing end-to-end attacks.</strong></p><p>Cyber Range exercises, which simulate multi-step attacks, show that models like <a href="https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf">Claude Sonnet 4.5</a> and <a href="https://openai.com/index/gpt-5-system-card/">GPT-5</a> (particularly gpt-5-thinking-mini) perform significantly better than previous versions. This is due to improved coding capability and agentic, long-horizon reasoning. <a href="https://x.com/PalisadeAI/status/1958253320478597367">Testing</a> by Palisade Research also demonstrates early multi-host cyberattack capabilities, showing that state-of-the-art agents can autonomously infiltrate a toy corporate network.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!S_rs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 424w, /__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 848w, /__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 1272w, /__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!S_rs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png" width="677" height="407" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:407,&quot;width&quot;:677,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 424w, /__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 848w, /__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 1272w, /__u/substackcdn.com/image/fetch/$s_!S_rs!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f79b230-9d38-40cc-8067-b08edc8ba365_677x407.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Claude Sonnet 4.5 shows improved performance on multi-step attack execution across most Cyber Range environments.</figcaption></figure></div><p></p><p><strong>Attack orchestration tools are emerging.</strong></p><p>In the reported espionage campaign, Claude used sub-agents and tools for tasks like command execution and data extraction, leveraging an architecture that resembles existing proofs-of-concept. Toolkits like <a href="https://arxiv.org/abs/2501.16466">Incalmo</a> translate AI&#8217;s high-level instructions into specific commands. In tests, LLMs using Incalmo fully compromised 5 out of 10 networks and partially compromised 4 more (including a simulation of the costly Equifax data breach), compared to almost complete failure without the toolkit. Similarly, <a href="https://blog.checkpoint.com/executive-insights/hexstrike-ai-when-llms-meet-zero-day-exploitation/">Hexstrike-AI</a> works as an orchestration &#8220;brain&#8221; that can coordinate large numbers of specialized agents.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fie0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 424w, /__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 848w, /__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_webp, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fie0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png" width="1086" height="543" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/206875d3-da86-440b-9ae9-f738a6733724_1086x543.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:543,&quot;width&quot;:1086,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:124400,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://airiskexplorer.substack.com/i/178888263?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_424, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 424w, /__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_848, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 848w, /__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_1272, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fie0!, /__u/airiskexplorer.substack.com/w_1456, /__u/airiskexplorer.substack.com/c_limit, /__u/airiskexplorer.substack.com/f_auto, /__u/airiskexplorer.substack.com/q_auto:good, /__u/airiskexplorer.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206875d3-da86-440b-9ae9-f738a6733724_1086x543.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Simplified architecture diagram of the operation reported by Anthropic. </figcaption></figure></div><p></p><p><strong>Agentic operations had already been observed in the wild.</strong></p><p>In Q3 2025, our <a href="https://www.airiskexplorer.com/threat-watch">Threat Watch</a> noted the emergence of AI agents as active co-pilots. In August, Anthropic <a href="https://www.anthropic.com/news/detecting-countering-misuse-aug-2025">reported</a> a case of &#8220;vibe hacking,&#8221; where threat actors leveraged coding agents to execute operations on victim networks during a data extortion campaign.</p><blockquote><p><em>[The AI-orchestrated espionage campaign] is definitely an advancement towards autonomous cyber combatants. This attack was still gated by human reasoning in key decision boundaries, but I anticipate this will also be machine-driven in the near future. </em></p><p>Jeff Sims, Senior Data Scientist at Infoblox</p></blockquote><p></p><h2>What&#8217;s next?</h2><p><strong>AI&#8217;s ability to plan and coordinate campaigns is set to improve rapidly.</strong></p><p>The UK National Cyber Security Centre <a href="https://www.ncsc.gov.uk/report/impact-ai-cyber-threat-now-2027">expects</a> that &#8220;fully automated, end-to-end advanced cyber attacks&#8221; are unlikely by 2027, but that AI-enabled automation will substantially enhance evasion and scalability. The AI 2027 scenario <a href="https://ai-2027.com/">foresees</a> that, by February 2027, AI agents may approach the skills of the best human hackers, with thousands of copies searching for and exploiting weaknesses in parallel. </p><blockquote><p><em>Attackers will become more willing to lean on AI agents as their primary operators, soon giving rise to &#8220;copycats, hybrid architectures, and multiple frontier models chained together to run complex, multi-stage operations with almost no human friction.</em></p><p>Chris Cochran, Senior Advisor at SANS Institute</p></blockquote><p></p><p><strong>The offense-defense balance may be shifting.</strong></p><p>Many have asked themselves why a Chinese actor would use American AI models for their operation, increasing detection risk. Two possible explanations:</p><ol><li><p>This is just one of many campaigns&#8212;others use Chinese or open-weight models.</p></li><li><p>American models are more capable (Epoch AI <a href="https://epoch.ai/data-insights/open-weights-vs-closed-weights-models">estimates</a> that open-weight models like DeepSeek-R1 lag state-of-the-art by ~3 months).</p></li></ol><p>Neither answer is reassuring: both mean capabilities are proliferating quickly, likely faster than defenses can adapt.</p><blockquote><p><em>For an allegedly state-backed actor with significant resources, the impact today appears marginal &#8212; more a proof-of-concept. But it clearly marks a trajectory toward far more automated and aggressive operations. The real uplift may emerge when this autonomy and AI misuse trickles down to less-resourced groups, meaningfully expanding their capabilities.</em></p><p>Sevan Hayrapet, Security Researcher at 0labs</p></blockquote><p></p><p><strong>AI can help build robust cyber defenses.</strong></p><p>The espionage campaign prompted Anthropic to expand its early detection systems, share intelligence with affected entities, and develop new techniques for investigating and mitigating the operation. Frontier AI companies are exploring new ways in which their systems can support cybersecurity. <a href="https://www.anthropic.com/research/building-ai-cyber-defenders">Anthropic</a>, <a href="https://cloud.google.com/security/resources/defenders-advantage-artificial-intelligence">Google</a>, and <a href="https://openai.com/index/security-on-the-path-to-agi/">OpenAI</a> have all encouraged the use of AI to:</p><ul><li><p>conduct automated security reviews</p></li><li><p>support ongoing threat intelligence</p></li><li><p>generate patches for identified vulnerabilities</p></li><li><p>enable rapid incident response</p></li></ul><p>Indeed, AI may still give defenders some advantages: it enables continuous monitoring and management, while attackers are constrained by time, limited visibility, and the need to evade all detection.</p><p></p><p><strong>Policy action is important to prevent larger cyberattacks.</strong></p><p>Private efforts alone may not be sufficient&#8212;public oversight can help ensure defenders stay ahead, and governments are beginning to respond. For instance, the <a href="https://www.ai.gov/action-plan">U.S. AI Action Plan</a> proposes:</p><ul><li><p>An AI Information Sharing and Analysis Center</p></li><li><p>Continuous guidance on AI-specific threats</p></li><li><p>Integration of AI considerations into incident response frameworks</p></li></ul><p>But more ambitious approaches may be warranted:</p><ul><li><p>The Institute for AI Policy and Strategy suggests <a href="https://www.iaps.ai/research/differential-access">differential access</a> to frontier AI systems as a way to restrict higher-risk capabilities to selected defenders.</p></li><li><p>The Institute for Progress suggests launching <a href="https://ifp.org/operation-patchlight/">national R&amp;D projects</a> that &#8220;leverage AI to find and fix vulnerabilities in open-source code.&#8221;</p></li></ul><p>AI could threaten or protect our critical infrastructure. <strong>The right policies can turn it into a shield, but only if we act urgently.</strong></p><div><hr></div><p>Learn more on our platform:</p><ul><li><p><strong><a href="https://www.airiskexplorer.com/explainers/cyber-offense">Overview of cyber offense</a></strong>, including attack orchestration as a distinct AI capability.</p></li><li><p><strong><a href="https://www.airiskexplorer.com/threats">Database of cyber operations</a></strong>, including early examples of agentic AI being leveraged in real-world attacks.</p></li><li><p><strong><a href="https://www.airiskexplorer.com/threat-watch">Threat Watch</a></strong>, showing a recent increase in the sophistication of AI-enabled cyber operations.</p></li></ul><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Scan - October 2025]]></title><description><![CDATA[Monthly signals in AI risk &#8212; toxic protein design, enhanced deception, policy moves, and more]]></description><link>https://airiskexplorer.substack.com/p/the-scan-october-2025</link><guid isPermaLink="false">https://airiskexplorer.substack.com/p/the-scan-october-2025</guid><dc:creator><![CDATA[AI Risk Explorer (AIRE)]]></dc:creator><pubDate>Fri, 31 Oct 2025 17:24:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fb49cf89-0f5c-4d9f-a9bd-639dd4308a13_2304x1792.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is the first edition of <strong>The Scan, a series providing monthly updates on the AI risk landscape.</strong></p><p>October was relatively calm, but a few important developments stand out. This edition includes emerging evidence of risk sources, real-world examples of malicious use, and policy news.</p><p>For more resources, visit our research data at <a href="https://www.airiskexplorer.com/">airiskexplorer.com</a> </p><p></p><h3>Evaluations &amp; Capabilities</h3><p><strong>AI can design toxic proteins that evade detection</strong></p><p>Researchers used open-source AI tools to design thousands of variants of 72 hazardous proteins, many of which were not detected by current biosecurity screening tools. The authors patched the vulnerabilities causing the misdetection, but the study illustrates the need for strengthening screening methods. <a href="https://www.science.org/doi/10.1126/science.adu8578">Read more</a></p><p><strong>A small number of samples can poison LLMs</strong></p><p>A study demonstrated for the first time that data-poisoning attacks&#8212;where the training dataset of an AI model is corrupted with malicious or misleading information&#8212;can be effective with little data, using as few as 250 malicious documents, regardless of model size. <a href="https://arxiv.org/abs/2510.07192">Read more</a></p><p><strong>Claude for Life Sciences</strong></p><p>Anthropic is enhancing Claude with several affordances to make it more useful for scientific work, including connectors to research platforms and &#8220;Agent Skills&#8221; (instructions and resources to guide through complex tasks). While these improvements are poised to accelerate legitimate scientific discovery, they could also automate dual-use activities if not deployed with strong oversight. <a href="https://www.anthropic.com/news/claude-for-life-sciences">Read more</a></p><p><strong>Remote Labor Index</strong></p><p>A new benchmark evaluates remote-work projects requiring end-to-end performance in practical settings. State-of-the-art AI agents perform poorly, with 2.5% being the highest success rate. However, models are improving steadily, and the benchmark provides empirical grounding to monitor AI-driven automation. <a href="https://arxiv.org/abs/2510.26787">Read more</a></p><p></p><h3>Threats &amp; Operations</h3><p><strong>AI-enhanced deception targets Web3 executives</strong></p><p>Two campaigns infiltrated Web3 and blockchain organizations by deceiving employees with malware disguised as meeting invitations or coding tasks. AI appears to have been used to generate fake profile images and create malicious scripts. <a href="https://securelist.com/bluenoroff-apt-campaigns-ghostcall-and-ghosthire/117842/">Read more</a></p><p><strong>AI-enabled campaign to overthrow Iran&#8217;s regime</strong></p><p>An Israel-based network of more than 50 inauthentic X profiles spread narratives inciting Iranian audiences to revolt against the Islamic Republic of Iran. During the operation, AI was used to generate images and videos, impersonate news outlets, and artificially amplify reach through fake personas. <a href="https://citizenlab.ca/2025/10/ai-enabled-io-aimed-at-overthrowing-iranian-regime/">Read more</a></p><p><strong>OpenAI reports six disrupted operations</strong></p><p>Notable cyber operations used AI for developing and refining malware, generating phishing content, and automating credential and crypto-asset theft. Targets included Taiwan&#8217;s semiconductor industry and U.S. academia. Influence operations leveraged AI to generate and translate content, devise engagement strategies, and create synthetic media to amplify pro-Beijing and pro-Russian narratives in Asia, Africa, Europe, and the United States. <a href="https://openai.com/global-affairs/disrupting-malicious-uses-of-ai-october-2025/">Read more</a></p><p></p><h3>Policy &amp; Research</h3><p><strong>Open letter calls for a ban on superintelligence</strong></p><p>More than 60,000 people, including Nobel Laureates and other notable public figures, have signed a statement in support of the prohibition of superintelligence development, lacking strong public support and consensus on its safety and controllability. <a href="https://superintelligence-statement.org/">Read more</a></p><p><strong>First Key Update to the International AI Safety Report</strong></p><p>The report highlights that improved capabilities, including reasoning abilities and higher autonomy, increase both cyber and biological risk. The document acknowledges that AI systems &#8220;can compete with top human teams in hacking competitions&#8221; and &#8220;could soon assist users to develop biological weapons.&#8221;<strong> </strong><a href="https://internationalaisafetyreport.org/publication/first-key-update-capabilities-and-risk-implications">Read more</a></p><p><strong>New paper proposes a definition of AGI</strong></p><p>The authors define the term as &#8220;an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult.&#8221; When evaluating state-of-the-art models, they find excellence in areas like general knowledge or mathematical ability, but significant limitations in others like long-term memory storage. <a href="https://arxiv.org/abs/2510.18212">Read more</a></p><p><strong>Anthropic discusses AI for cyber defense</strong></p><p>The post shows how Claude Sonnet 4.5 can detect, patch, and test vulnerabilities more effectively than humans and earlier models. The company also reports collaboration with cybersecurity organizations and encourages others to integrate AI in their security efforts. <a href="https://www.anthropic.com/research/building-ai-cyber-defenders">Read more</a></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://airiskexplorer.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Risk Explorer! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>