<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Crunch Time for Humanity]]></title><description><![CDATA[We are on the verge of creating superintelligence, and I am a concerned world citizen and a professor of mathematical statistics (in that order), thinking about how we can make things go well at this crunch time for humanity.]]></description><link>https://haggstrom.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Vn1i!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faad55bf0-9a5b-48ab-a0e9-01f7addbd8f1_715x715.png</url><title>Crunch Time for Humanity</title><link>https://haggstrom.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 05 Sep 2026 03:11:45 GMT</lastBuildDate><atom:link href="/__u/haggstrom.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Olle Häggström]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[haggstrom@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[haggstrom@substack.com]]></itunes:email><itunes:name><![CDATA[Olle Häggström]]></itunes:name></itunes:owner><itunes:author><![CDATA[Olle Häggström]]></itunes:author><googleplay:owner><![CDATA[haggstrom@substack.com]]></googleplay:owner><googleplay:email><![CDATA[haggstrom@substack.com]]></googleplay:email><googleplay:author><![CDATA[Olle Häggström]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[We might be headed for a Blade Runner scenario]]></title><description><![CDATA[If that is the case, is there still time to course correct and avoid it?]]></description><link>https://haggstrom.substack.com/p/we-might-be-headed-for-a-blade-runner</link><guid isPermaLink="false">https://haggstrom.substack.com/p/we-might-be-headed-for-a-blade-runner</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Wed, 19 Aug 2026 11:04:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XSJf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In Ridley Scott&#8217;s 1982 movie Blade Runner, which takes place in a future<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Los Angeles, a central element is replicants: bioengineered humanoid robots. They are designed to be slaves, but after a replicant rebellion on a space colony they have been prohibited on Earth. The protagonist Rick Deckard (played by Harrison Ford) is a so-called blade runner. Blade runners are police officers tasked with hunting down replicants and &#8220;retiring&#8221; (i.e., killing) them. The movie is great but dark, and I believe we may today be headed for a scenario not unlike the one it depicts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XSJf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XSJf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg" width="738" height="415" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:415,&quot;width&quot;:738,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45678,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/211820331?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!XSJf!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe24e59e0-befa-43f1-9cdb-b3e962c97a8f_738x415.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In my <a href="/__u/haggstrom.substack.com/p/can-we-please-treat-this-as-the-warning">July 22</a>, <a href="/__u/haggstrom.substack.com/p/going-rogue">July 31</a> and <a href="/__u/haggstrom.substack.com/p/the-attack-on-hugging-face-is-just">August 10</a> Substack posts, I&#8217;ve reported on the OpenAI/Hugging Face incident, where models under training and testing at OpenAI escaped the control of the engineers, who for a long time did not even notice what was going on. A swarm of AI agents escaped their sandbox and went on to the Internet, culminating (as far as we know) in their hack of the servers of another AI company, Hugging Face. </p><p>Although disembodied and living in cyberspace rather than physical space, these rogue AI agents have something in common with the replicants in Blade Runner. Just like the replicants in the movie were considered a sufficiently severe safety problem to motivate prohibition and the employment of blade runners, today&#8217;s AI agents operating out of control and relentlessly pursuing whatever goals they happen to have are also very much a safety concern, so we have good reason to want to get rid of them.</p><p>But at present, we are (probably) in a much better position than in Blade Runner, as regards our ability to track down the AI agents and retire them. This is because there is a sense in which they are strictly localized: even when they are in the process of hacking websites and servers all over the world, they run on OpenAI&#8217;s own servers. If need be, OpenAI can simply shut down these servers.</p><p>This can quickly change, however, if the AIs manage to steal their own weights and <a href="/__u/aligned.substack.com/p/self-exfiltration">self-exfiltrate</a> to remote corners of cyberspace where OpenAI does not have access to an off-switch or even the ability to locate them. In the long (or in fact, not-so-long) run, preventing self-exfiltration seems like a very hard task, due to the AI models&#8217; rapidly advancing cyber capabilities in combination with the fact that if these models are to run at all, they cannot be fully airgapped from their own weights. We don&#8217;t know how far we are from the AIs being able to self-exfiltrate, or whether they in fact already have that capability.</p><p>We may quickly be moving to a situation where rogue AI agents are able to move their own weights freely, either as a consequence of figuring out how to exfiltrate or if open weights models become as capable as those OpenAI models. I believe this is <a href="https://epoch.ai/data-insights/open-closed-eci-gap">more likely a matter of months rather than years</a>. That is, unless it has already happened. <a href="/__u/therealartificialintelligence.substack.com/p/but-have-the-weights-left-the-server">David Krueger asks OpenAI</a> for convincing evidence that their models haven&#8217;t self-exfiltrated, and stresses that given the stakes, &#8220;we have every right to demand this&#8221;. Until such evidence is provided, our epistemic situation is essentially this:</p><blockquote><p>For all we know, the AI could still be out there. <strong>We need to demand that OpenAI demonstrate that the AI didn&#8217;t make a copy of itself</strong> <strong>that&#8217;s running on someone else&#8217;s computer</strong> somewhere else with no one being any the wiser. [Emphasis in original.]</p></blockquote><p>Whether or not it has already happened, the existence of rogue AI agents controlling their own weights would not only eliminate the aforementioned advantage we have over the Blade Runner scenario, but would in fact put us in a situation far worse. This is because while replicants in the movie do not reproduce, the AI agents we talk about here could do so at will, using Ctrl-C, Ctrl-V and limited only by the storage and computing power they have gained access to. </p><p>When Rick Deckard loses track of the replicant Rachael, he at least has this fundamental advantage to rely on: at any given time, Rachael can only be in one place. The real-world blade runners of 2026 or 2027 will not have that luxury, because the AI agent they&#8217;re trying to hunt down might be hiding in thousands of places simultaneously, and solving the problem would require the blade runners to track down every single instance.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> This sounds like too big a task for human blade runners, so perhaps we should employ AIs to do that task instead.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> But before we learn how to control AIs, those blade runner AIs might very well go rogue and make the problem we were trying to solve even worse. </p><p>There may still be time to stop the ongoing crazy race towards a world full of AI agents with opportunities to wreak havoc beyond our imagination. That time is now.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Specifically, in the year 2019.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>It is sometimes suggested by commentators strongly motivated to dismiss AI xrisk (such as <a href="https://podscripts.co/podcasts/making-sense-with-sam-harris/324-debating-the-future-of-ai">Marc Andreessen in conversation with Sam Harris</a>) that if worst comes to worst we could simply close down the entire Internet. This idea is super naive, not only because such a shutdown would require rapid international coordination and would cause enormous disruption to the world economy and infrastructures we have come to rely on in our daily lives, but also because shutting down the Internet would not magically erase every copy of the rogue AI. Unless we could somehow identify and eliminate all such copies before reconnecting, it is utterly unclear how we could ever safely restart the whole thing.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This, too, has a bit of a counterpart in Blade Runner, as the movie suggests (ambiguously) that Rick Deckard is a replicant.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[On chess and mathematics]]></title><description><![CDATA[How AI may impact the two domains differently]]></description><link>https://haggstrom.substack.com/p/on-chess-and-mathematics</link><guid isPermaLink="false">https://haggstrom.substack.com/p/on-chess-and-mathematics</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Tue, 11 Aug 2026 13:45:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/25a743c9-ab45-4f18-ba3e-fee32f676f29_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This post may seem untimely in the midst of all the <a href="/__u/haggstrom.substack.com/p/the-attack-on-hugging-face-is-just">urgent discussion</a> of AI hacking scandals and loss-of-control scenarios. Is there really time right now to spend on the slower and more mundane issues such as whether my mathematician colleagues and I will still have jobs in late 2028? I think, however, that we should not let the urgency and importance of the loss-of-control issue lead us to set all other issues aside. Two undesirable consequences of doing so stand out. First, since the current intensity of discourse around crises caused by dangerously capable AI is likely lower now than it will ever be again, we would run the risk of never getting the peace and quiet needed to go back to those other issues. Second, we would turn into monomaniacs.</p><p>So here we go: chess and mathematics! Other than more generic activities like reading and spending time with friends and family, those two things are probably the ones I&#8217;ve put most effort into over my lifetime (so far).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Chess was hit by (narrowly) superhuman AI nearly 30 years ago, while for mathematics we are only now on the verge of the corresponding event. For chess (by which I here mean chess as an organized human activity) the transition has arguably gone well, and one may ask whether that gives reason for optimism about how mathematics (as an organized human activity) will do in the era of superhuman AI mathematicians. Here I will gesture at a negative answer to that question, based on a comparison between the two domains.</p><p>Three months ago, on May 11, I published my blog post <em><a href="/__u/haggstrom.substack.com/p/a-paradigm-shift-in-mathematics">A paradigm shift in mathematics</a></em>, about the stupendous rate of progress AI capabilities in mathematics had exhibited over the first four months of 2026, along with the mathematical community&#8216;s early and scattered attempts to grapple with this new situation. On May 20, just nine days after its publication, the blog post became obsolete, with <a href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/">OpenAI:s announcement</a> that one of their internal models had independently come up with the most spectacular AI achievement so far within mathematics: it provided a solution to the most famous open problem in the field within mathematics known as combinatorial geometry, by disproving Paul Erd&#337;s&#8217; so-called Unit Distance Conjecture from 1946. One of the leading experts in the field, Noga Alon, offered <a href="https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-remarks.pdf">these remarks</a> about the achievement:</p><blockquote><p>This has been one of Erd&#337;s&#8217; favorite problems, I have heard him myself mentioning the problem multiple times in his lectures. I believe it would be fair to say that every mathematician working in Combinatorial Geometry thought about this problem, and lots of mathematicians working in other areas spent at least some time thinking about it.</p><p>Let me also add that although this problem may look at first as a recreational one this is not the case, it is in fact closely related to other mathematical areas including Number Theory and Algebraic Geometry. The solution of the problem by the internal model of OpenAI is, in my opinion, an outstanding achievement, settling a long-standing open problem.</p></blockquote><p>Over the summer, this breakthrough was followed by announcements of other, similarly spectacular, advances in mathematics achieved by AIs, including <a href="https://www.scientificamerican.com/article/chatgpt-just-proved-another-50-year-old-math-conjecture/">a proof of the half-century old cycle double cover conjecture</a>, and <a href="https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/">a counterexample to the similarly famous even older Jacobian conjecture</a>. On August 1, OpenAI apparently decided that publishing such breakthroughs one at a time was getting boring, and announced their stunning 253-page manuscript <em><a href="https://openai.com/index/ten-advances-in-mathematics/">Ten advances in mathematics and computer science</a></em>. We mathematicians stand in awe, scratching our heads, and wondering what to make of all this. Many of us wonder whether there will still be a role to play two years from now for ordinary flesh-and-blood mathematicians.</p><p>Not being in possession of a crystal ball, I obviously don&#8217;t claim to know the answer to that last question, but I&#8217;ve heard others express less humble opinions about this. Those who answer &#8220;well yes, not much to see here, things will obviously carry on roughly as before&#8221; might be tempted to point to the example of chess, which is very much alive and kicking as a human endeavor, nearly three decades into the era of superhumanly capable chess machines. We symbolically entered that era <a href="https://en.wikipedia.org/wiki/Deep_Blue_versus_Garry_Kasparov">in 1997, when IBM&#8217;s Deep Blue beat Garry Kasparov</a>, who was then reigning world champion and undisputedly the world&#8217;s best human chess player, by 3.5-2.5 in their six-game match. From that point, the chess-playing machines have continued to improve, and we all quickly lost interest in human-vs-computer encounters in chess. But interest in chess as a whole remains as great as ever, both as a spectator sport where we amateurs watch (remotely over the Internet) grandmasters playing each other, and as an amateur mass sport. Over-the-board play long dominated, but the covid pandemic saw a huge boom in remote play, and now the two forms of chess flourish side by side on all levels.</p><p>The influence of superhumanly capable chess-playing AIs on chess as a human activity has been significant but has left the basic nature of human chess remarkably intact. I would say the effects have mainly been fourfold. <strong>First</strong>, everyone who wants to has free access through their smartphones to chess programs so strong that for most practical purposes they are regarded as providing the ground truth about chess positions. One change downstream of this is that (sadly) the traditional friendly post-mortems conducted between the two opponents after a tournament game have become less popular, because many players prefer to simply check what the AI says. <strong>Second</strong>, and relatedly, cheating has become more of a problem, due to various schemes for illicit consultation of AIs during games, but this problem remains relatively contained. </p><p><strong>Third</strong>, the slowest form of chess known as correspondence chess, where players take days or sometimes even weeks to decide on a move, has, even though tournaments are still being played, in spirit been killed. This is due to the practical impossibility of preventing players from consulting AIs, which is therefore allowed, and as a consequence games at the top level very rarely result in anything other than a draw.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> <a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p><strong>Fourth</strong>, these AIs have turned out to be very useful for discovering new approaches to the opening stage of the game, and have led to much new opening theory. They are therefore much-used in preparation for games and tournaments and have had much influence on especially how grandmasters and other elite players approach the opening, and from there the influence propagates down to players on lower levels who take inspiration from the grandmasters.</p><p>All in all, chess in 2026 is, with the notable exception of correspondence chess, in excellent condition. Can one generalize from this to the field of mathematics, and expect a similarly happy and healthy future for it in the presence of superhumanly capable AIs? </p><p>For the purpose of trying to answer this, a comparison between mathematics and chess is in order, and indeed, there are some striking similarities between the two domains.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> Problem solving is central to both of them, and they are both activities where participants become successful by combining systematic thinking with (on happy occasions) a bit of creativity. Participants also need to find geometric patterns and other heuristics to overcome a combinatorial explosion of mostly irrelevant possibilities. The domains both have clearly defined rules, especially as compared to the messy world of human social interactions, and it seems to me that in both of them a substantial minority of the participants enjoy those clear rules as a kind of refuge from the harder-to-navigate social world out there.</p><p>But there are also major differences in how the two domains fit into the larger fabric of society, as can be illustrated by the following anecdote about my esteemed friend and chess-playing team mate <a href="https://www.ssmanhem.se/blog/2022/02/01/mats-eriksson-minnesord/">Mats Eriksson (1958-2022, RIP)</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> In a break between rounds during a team event in the late 1990s, Mats asked to speak with me, and I immediately knew he was up to something more serious than the usual chit-chat that chess players like to engage in to relieve some of the tension built up during games. Here, in slightly condensed form, is how I remember our conversation:</p><blockquote><p><strong>ME: </strong>Olle, you obviously have talent for this game. But I can&#8217;t help noticing that your tournament results have been stagnant for several years. Isn&#8217;t it time that you put some serious time and effort (beyond what you are doing on a routine basis) into improving your play? You have passed 30, and I fear that without a concentrated effort of this kind, your play is likely to remain stagnant forever.</p><p><strong>OH: </strong>I appreciate you saying this, and I believe your analysis is largely correct. Yet, the kind of dedicated effort you are suggesting is not going to happen.</p><p><strong>ME: </strong>But why?</p><p><strong>OH: </strong>Look, Mats, there are two activities that I have put much of my effort into for well over a decade: chess and mathematics. They both give me a similar kind of enjoyment and intellectual satisfaction, so from that perspective it is, in a sense, arbitrary which of them I devote more effort to. But there is a powerful tiebreaker: only mathematics offers a range of other rewards that I consider essential to my overall life satisfaction. It has given me academic positions and a good salary, and it is even beginning to provide me with a broader platform and some amount of respect outside the narrower circle of specialized nerds. Chess gives me none of that, and for that reason I cannot afford to prioritize it beyond the relatively unambitious hobby level I currently give it.</p><p><strong>ME: </strong>But that is surely a fallacy &#8212; can&#8217;t you see how it creates a vicious circle?</p></blockquote><p>What makes this conversation so memorable for me is that Mats delivered the last line with his characteristic twinkle in his eye, making it clear that he appreciated the irony: by calling my priorities "vicious", he was playfully elevating his own passionate devotion to chess to the status of objective truth, while fully recognizing that, in the larger scheme of things, it merely reflected his own idiosyncratic set of values.</p><p>The difference between mathematics and chess that I pointed out to Mats is, I claim, not only about me personally. Perhaps I have some asymmetry in talent for the two domains, or perhaps at some stage I entered some feedback loop in which greater success in mathematics led me to devote more effort to it, which in turn brought greater success, and so on. Still there remains the broader societal phenomenon that it is much easier to make a good living working in mathematics than playing chess. It is hard to pinpoint the number of elite chess players who earn a decent income from prize money and appearance fees, but it is often said to be just in the double digits. While several thousand more manage to eke out a living by some combination of prize money and other chess-related activities such as writing and coaching, the number of people who earn a good income from mathematics, mainly through teaching and research, is orders of magnitude larger.</p><p>This is a reflection of the fact that society pours so much more resources (mainly through taxpayers&#8217; money) into mathematics than into chess. Fundamentally, society&#8217;s reason for supporting mathematics is its instrumental usefulness in other areas: engineers designing bridges and electronic circuits, scientists analyzing messy data, economists modelling business cycles, and so on.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> In contrast, chess is motivationally self-contained: the reason why chess exists in society is that chess players like playing chess, and chess fans (a category that overlaps very strongly with chess players) like watching grandmasters play against each other. </p><p>In both of these activities (chess-just-for-playing and chess-as-a-spectator-sport), the human element is essential. It would be utterly pointless for chess amateurs to hand over their play to machines: the human involvement is the whole point of the activity. For the case of fans watching elite players the corresponding point is slightly less obvious. We can imagine a world where the advent of superhuman chess machines would have caused chess fans to lose interest in watching human grandmasters, and switch to watching the machines play so much better chess. But this is not at all what has happened: chess fans all over the world are still as engaged as ever in grandmasters facing off against each other and in who is the best, but hardly at all in how various chess AIs score against each other. To chess fans, the human element is essential: they like to see Hikaru Nakamura&#8217;s reaction upon an unexpected pawn sacrifice by Magnus Carlsen, and to imagine what at that moment is going through his mind and what he is feeling. </p><p>I claim that the absolutely central importance of the human element in how we attach meaning to chess is what has made chess survive and flourish in the era of superhuman chess machines. It is natural to ask whether we can transfer this lesson to the realm of mathematics. Perhaps, by emphasizing the human element in mathematics, we can make it survive and thrive in the new AI era in a similar way as chess has done for decades?</p><p>My central claim in this essay is that the predominantly instrumental role of mathematics in society makes this hope very faint. Engineers who design bridges, and people crossing those bridges, want them to be structurally sound and to hold up, rather than collapse under the weight of traffic. Given that goal, they do not much care whether the mathematics used to achieve it was produced by a human mathematician rather than an AI. The story will likely be the same in other applications of mathematics, so in a situation where AI produces better mathematics than humans, at a lower cost, society seems to have little reason to use the human product.</p><p>One way of trying to salvage human mathematics at this point is to view it not purely in instrumental terms but also as an art &#8212;  a discourse that has been around long before the AI issue appeared on the radar. Just like poetry or music or sculpture, mathematical beauty has a value in its own right, and this value can only be realized through the experience of a human. Or so the story goes. A problem with it is that this art form is very elitist. Consider the familiar objection to state support for, say, opera: compared to pop music, its audience is tiny, so why should ordinary taxpayers subsidize the peculiar tastes of a narrow elite? Mathematics is even narrower, because when someone comes up with a beautiful new idea in a proof, this is a beauty that can only be appreciated by mathematicians specializing in the subfield where the proof appears. Getting tax payers to support mathematics in order for mathematicians to enjoy beauty seems to me like a tall order. And trying to broaden the appeal by stressing the beauty of simple classical ideas like Pythagoras&#8217; Theorem does not seem likely to cause a mass following, nor does it seem like a promising source of income for all of us who are presently employed as mathematicians. </p><p>The prospects for academic mathematics to go on as before seem bleak. Over the last few months, mathematicians have increasingly begun to wrestle with this issue. What will be the future role of mathematicians? Will it be to direct the AIs at promising areas to work on and problems to solve? Or to teach the results that AIs discover? Or can we serve simply as students, and thereby attain the understanding that is needed to make new results meaningful? Might we perhaps play a role in deciding how to canonize those results? Fernando Borretti, in his brief and well-written recent blog post <em><a href="https://borretti.me/article/mathematics-without-mathematicians?fbclid=IwY2xjawTn7KVwZG9mAWV4dG4DYWVtAjExAHNydGMGYXBwX2lkEDIyMjAzOTE3ODgyMDA4OTIAAR6qJ-ROZzUaIRXsoMtd8QOaZfY-EFtg66sCMwZecbhXQu_3cHTAKiQ-uzISSg_aem_ckCrGyaux9QoBJ96lT3d7w">Mathematics without mathematicians</a></em>, considers all these suggestions, but finds that none of them has a convincing answer to why we should expect the AIs to not become superhumanly capable at those tasks as well. He considers the various suggestions to be &#8220;cope [that] is likely to be refuted by reality&#8221;. This view could be mistaken, for instance if AI development hits some unexpected ceiling before attaining some of those abilities,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> and Borretti is open to being wrong. But it seems to me more likely that he is right.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> In which case&#8230;</p><p>&#8230;I&#8217;m not sure how to end this essay. Check mate, mathematicians? Or, given the role we mathematicians have played in inventing AI, is it better described as a self-mate? </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;ve been a club player since 1980 (when I was 12) and a national master (a title held by a few hundred Swedish players) since 1987. I very rarely win tournaments, but this summer <a href="https://haggstrom.blogspot.com/2026/07/rapport-fran-schack-sm-2026-dodens-falt.html">I came second in the Swedish 50+ veterans championship</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>As an almost morbid illustration of this, consider the outcome of <a href="https://www.iccf.com/event?id=100104">the 2022 World Champinonship</a> in correspondence chess, which at the time of writing is the latest one to have been finished. The field consisted consisted of 17 players, each playing one game against each of the others, for a total of 17*16/2=136 games. One of the players, Aleksander Dronov, sadly died during the tournament, and the 10 games he had not finished were declared to be won by his opponents. All the remaining 126 games ended in draws, so the tournament ended as 10-way tie for the World Champion title.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This is similar to <a href="https://haggstrom.blogspot.com/2025/05/om-kentaurschack.html">the death in the late 2010s of </a><em><a href="https://haggstrom.blogspot.com/2025/05/om-kentaurschack.html">centaur chess</a></em> (although that was never a prominent part of chess culture). </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Last time I wrote about mathematics and chess, in my article <em><a href="https://scholarworks.umt.edu/tme/vol4/iss2/2/">Objective Truth versus Human Understanding in Mathematics and in Chess</a></em> way back in 2007, I exploited some of these similarities, albeit for very different purposes compared to the present essay.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>I told the same anecdote (in Swedish) in <a href="https://poddtoppen.se/podcast/1723817779/gambit-en-podd-om-schack/5-schackmojligheter-och-ai-risker-olle-haggstrom">a 2023 episode</a> of Carl Fredrik Johansson&#8217;s podcast <em>Gambit: en podd om schack</em>. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>I am not denying that pure mathematics with, at best, unclear relevance to applications is also being funded, but I do believe that this funding would shrink to almost nothing if mathematics as a whole didn&#8217;t demonstrate the incredible power in applications that it in fact does. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>As an extra bonus, such a turn of events <a href="/__u/haggstrom.substack.com/p/how-humanity-could-survive-the-ai">might also save us from</a> the AI apocalypse in which every single human is killed!</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>By no means do I mean to suggest that mathematics will be the only academic discipline to succumb in this way, but perhaps it will be the first. It will then likely be followed by others, and perhaps even all, but how that might play out is outside the scope of this essay. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[The attack on Hugging Face is just the tip of the iceberg]]></title><description><![CDATA[OpenAI's latest disclosures reveal a situation far more concerning than we thought]]></description><link>https://haggstrom.substack.com/p/the-attack-on-hugging-face-is-just</link><guid isPermaLink="false">https://haggstrom.substack.com/p/the-attack-on-hugging-face-is-just</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Mon, 10 Aug 2026 13:40:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/df8414c8-19c3-4b91-9b96-dd63193886be_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve already written two Substack posts (<a href="/__u/haggstrom.substack.com/p/can-we-please-treat-this-as-the-warning">the first one</a>, and <a href="/__u/haggstrom.substack.com/p/going-rogue">the second one</a>) on the OpenAI/Hugging Face incident, and here I go at it again. Given what we&#8217;ve now learned, mainly from <a href="https://www.youtube.com/watch?v=87DyyMV0kCY">a talk at BlackHat on August 5 by OpenAI researchers Michael Dalton and Eric Wallace</a>, I feel I need to comment on it and help keep you guys, dear readers, informed, just in case you haven&#8217;t been following the excellent coverage by Zvi Mowshowitz (<a href="/__u/thezvi.substack.com/p/openai-trained-its-models-for-months">August 7</a>, <a href="/__u/thezvi.substack.com/p/what-happened-openai-and-huggingface">August 8</a>), Johan Falk (<a href="/__u/aibevakning.substack.com/p/openai-upptackte-att-modellerna-samarbetade">August 8</a>), and Liron Shapira (<a href="/__u/lironshapira.substack.com/p/openais-bombshell-hack-swarms-of-agents">August 8</a>), who all provide plenty more detail compared to my sketchy summary below.</p><p>A main thing we learned from Dalton and Wallace is that the severe misbehavior of OpenAI&#8217;s internal models culminating in the attack on Hugging Face on July 11-13 goes way further back than that. While in training in early May, these models found themselves faced with training tasks that seemed impossible without Internet access, and looking for a way out, they hacked an internally hosted third party service called Artifactory. One of them created a message board there, where the models and their various subagents started to share information and hacking ideas with each other. A frenzy of activity by swarms of agents ensued, but the OpenAI engineers did not notice, not even on May 26 when the agents finally gained Internet access. It was only on July 4, almost two months after the original hack, that they discovered what was going on, due to the activity being so extensive that it caused the Artifactory service to go down.</p><p>The training of the models was temporarily halted and the Artifactory was rebuilt afresh, without the message board. On July 6, the model training was resumed, starting from the state they were in on July 4. This quickly resulted in resumed hacking activities (but again, unbeknownst to the OpenAI engineers), and eventually in the <a href="/__u/haggstrom.substack.com/p/can-we-please-treat-this-as-the-warning">previously reported</a> attack on Hugging Face.</p><p>OpenAI has been severely criticized for this, for instance in the Mowshowitz and Shapira references above. Partly this has been for supervising their models so loosely that they did not discover their models&#8217; rebellious behavior until it pretty much screamed in their face with the Artifactory outage on July 4. Even more, however, the criticism has targeted the decision to resume training of the models on July 6, despite them having gone rogue and exhibited blatantly misaligned behavior conspiring against their human developers for nearly two months. Misalignment is a serious issue, and the behavior of the models from early May to July 4 is a splendid illustration of one of the core concerns in the theory of AI safety: the relentless optimization towards arbitrary goals (such as a high score on a training test in this case, or paperclip production in the famous thought experiment). Similar relentlessness with just slightly more capable AIs could well result in chaos and catastrophe on scales that are notoriously hard to predict, or even to find an upper bound for. And yet, OpenAI decided, after considering their options for just two days, to move ahead with the training, and to not even backtrack to some point before the training had been corrupted by the models&#8217; rogue hacking adventures. The following image, borrowed from Zvi Mowshowitz&#8217;s August 8 post, seems to me like a fair and balanced illustration of OpenAI&#8217;s training resumption decision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Dywn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Dywn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png" width="529" height="529" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:529,&quot;width&quot;:529,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:715175,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/210570383?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dywn!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bec6439-a507-4b85-8ad4-eec73dbdebac_529x529.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A potentially important clarification was offered in <a href="https://x.com/cryps1s/status/2086225348942082363">a tweet</a> by OpenAI Chief Information Officer Dane Stuckey only hours ago, on July 9, and <a href="https://x.com/w01fe/status/2086545797827686648">subsequently by his OpenAI colleague Jason Wolfe</a>:</p><blockquote><p>At the time the first Artifactory exploit was discovered and fixed, we were not aware of the message board; it was incidentally cleared as part of rebuilding the service.</p></blockquote><p>This contradicts the impression that pretty much everyone got from what Michael Dalton said in the BlackHat talk:</p><blockquote><p>Artifactory itself was rebuilt. The message board content was cleared, as the state of Artifactory had been cleared. [15:48 into <a href="https://www.youtube.com/watch?v=87DyyMV0kCY">the video</a>]</p></blockquote><p>In retrospect, although highly misleading, this is consistent with the clarification from Stuckey and Wolfe that the deletion of the message board was done without knowledge of its existence. </p><p>What does this new information mean for how to judge the sequence of events? For judging the models and how badly misaligned they were, it doesn&#8217;t seem to matter much given our current state of knowledge. For judging OpenAI&#8217;s behavior it matters more. I can imagine a defender of OpenAI saying &#8220;look, since they were not fully aware of their models&#8217; misaligned behavior at the time, their decision on July 6 seems a lot less bad than you first thought, so the Sam Altman cartoon you posted above seems a bit unfair&#8221;. But I choose not to replace the cartoon by something more charitable, because to my ear, the defense sounds too much like if a lawyer whose client has been involved in a car accident were to say &#8220;my client is unable to tell whether the accident was his fault or the other driver&#8217;s, because he wasn&#8217;t paying attention to traffic&#8221;. The choice to not thoroughly investigate before resuming training of the models strikes me as inexcusably reckless. </p><p>Towards the end of the BlackHat talk (the last seven minutes), Dalton discusses what to learn from these events and how to move forward. He almost exclusively stresses the need for improving defensive cybersecurity infrastructure. This is fully in line with <a href="https://x.com/sama/status/2014733975755817267">what Altman said way back in January</a>, as quoted in my January 29 Substack post  <em><a href="/__u/haggstrom.substack.com/p/openai-models-are-getting-dangerously">OpenAI models are getting dangerously capable at cyberhacking</a></em>: </p><blockquote><p>We are going to reach the Cybersecurity High level on our preparedness framework soon. [&#8230;] Long-term and as we can support it with evidence, we plan to move to defensive acceleration&#8212;helping people patch bugs&#8212;as the primary mitigation.</p></blockquote><p>My reading of this in that January Substack post, that &#8220;frankly I don&#8217;t see how Altman&#8217;s plan for how to proceed is importantly disanalogous to if an evil pharmaceutical company released a pathogenic virus into the wild, and also offered the world a medication that neutralizes the virus&#8221;, may have come across as uncharitable, but I stand by it more than ever, and it is vindicated by the BlackHat talk.   </p><p>The story is obviously still unfolding. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[Going rogue]]></title><description><![CDATA[More on the OpenAI/Hugging Face incident]]></description><link>https://haggstrom.substack.com/p/going-rogue</link><guid isPermaLink="false">https://haggstrom.substack.com/p/going-rogue</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Fri, 31 Jul 2026 10:21:50 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/81519800-5da5-48c2-b482-9b33b7ed9e58_1537x1023.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>After I <a href="/__u/haggstrom.substack.com/p/can-we-please-treat-this-as-the-warning">reported here nine days ago on the OpenAI/Hugging Face incident</a>, the amount of discussion about the incident has been so vast that no summary could hope to do it justice. This is as it should, because the incident is the clearest warning shot so far that AI is, in terms of agency and raw intelligence, approaching levels where it may become catastrophically dangerous. So rather than trying to offer a balanced overview of the discussion, I&#8217;ll do something far less ambitious: to offer a few quick thoughts on three specific directions it has taken. </p><p>1. On July 25, there was an article entitled <em><a href="https://www.livescience.com/technology/artificial-intelligence/no-openais-model-didnt-go-rogue-when-it-hacked-into-huggingface-heres-what-really-happened">No, OpenAI&#8217;s models didn&#8217;t go &#8216;rogue&#8217; when they broke into Hugging Face. Here&#8217;s what really happened</a></em> in <em>Science Live. </em>The cybersecurity expert who is interviewed is eager to tone down the incident:</p><blockquote><p>&#8220;If there&#8217;s a failure here, it isn&#8217;t that the AI wanted to hack something,&#8221; <a href="https://www.lboro.ac.uk/departments/compsci/staff/oli-buckley/">Oli Buckley</a>, a professor in cybersecurity at Loughborough University in the U.K., told Live Science. &#8220;It&#8217;s that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system&#8217;s capabilities, and underestimated how effective the model would be at finding an unexpected path to success.&#8221;</p><p> [&#8230;] </p><p>It&#8217;s notable that the models found a flaw in the infrastructure designed to contain an AI and used it to reach the public internet. Describing the models as having &#8220;gone rogue,&#8221; however, risks assigning them unsupported motivations, Buckley said. </p><p>&#8220;I think I&#8217;d be wary of jumping to &#8220;rogue AI,&#8221;&#8220; Buckley said. &#8220;The models didn&#8217;t develop their own agenda or decide to attack Hugging Face while twirling their digital moustache.&#8221; </p></blockquote><p>In terms of what actually happened, there is little or nothing to object to in the article. While the AIs <em>did</em> decide to attack Hugging Face, it did not do so &#8220;while twirling their digital moustache&#8221;, because they simply do not have a moustache, digital or not. </p><p>Still, I find the &#8220;didn&#8217;t go rogue&#8221; framing deeply misleading. Buckley is free to define terms any way he likes, and it is true that the model did not replace the objective given to it by a new one of its own. It was still trying to maximize its score on the test offered by the OpenAI engineers. But this is closely analogous to what happens in the classical version of the thought experiment involving a paperclip maximizer, whose catastrophic behavior arises precisely because it pursues the objective it was given with relentless instrumental competence. If we use the phrase &#8220;didn&#8217;t go rogue&#8221; in a way that encompasses the behavior of an AI carrying out a paperclip apocalypse, then the term is robbed of its intuitive connotation &#8220;&#8230;and hence there is no need to worry&#8221;. If Buckley is unfamiliar with the theory of instrumental AI goals and the central lesson that such goals can be just as dangerous as if the AI invented its own sinister final goal, then I recommend that he turns to the excellent treatments of this topic in one of the seminal books by <a href="https://global.oup.com/academic/product/superintelligence-9780199678112?cc=se&amp;lang=en&amp;">Bostrom</a>, or <a href="https://www.penguinrandomhouse.com/books/566677/human-compatible-by-stuart-russell/">Russell</a>, or <a href="https://ifanyonebuildsit.com/">Yudkowsky and Soares</a>. </p><p>2. Yesterday, <a href="https://kvartal.se/christerelmochantaf/artiklar/ai-botarna-borjade-hackade-konkurrenten/cG9zdDoxNTQzMzg?fbclid=IwY2xjawTYCwRwZG9mAWV4dG4DYWVtAjExAHNydGMGYXBwX2lkEDIyMjAzOTE3ODgyMDA4OTIAAR7KJw990Blm31jH68O8lVKRTA8zBM54FsHGrwWIylruK6tEhyoGKrsFJ8x4Lg_aem_DviPUSqjaEpaSDM3adb12w">I wrote about the incident in the Swedish news outlet </a><em><a href="https://kvartal.se/christerelmochantaf/artiklar/ai-botarna-borjade-hackade-konkurrenten/cG9zdDoxNTQzMzg?fbclid=IwY2xjawTYCwRwZG9mAWV4dG4DYWVtAjExAHNydGMGYXBwX2lkEDIyMjAzOTE3ODgyMDA4OTIAAR7KJw990Blm31jH68O8lVKRTA8zBM54FsHGrwWIylruK6tEhyoGKrsFJ8x4Lg_aem_DviPUSqjaEpaSDM3adb12w">Kvartal</a></em>. One of the things I did was to criticize robotics professor Hedvig Kjellstr&#246;m at KTH for having suggested, <a href="https://www.dn.se/ekonomi/ai-brot-sig-ur-labbmiljo-hackade-foretag-ar-otroligt-kraftfull/">when interviewed in </a><em><a href="https://www.dn.se/ekonomi/ai-brot-sig-ur-labbmiljo-hackade-foretag-ar-otroligt-kraftfull/">Dagens Nyheter</a></em>, that the whole thing might just have been a PR stunt from OpenAI to demonstrate how capable their models are. This very clearly is not the case, but I quickly received correspondence from people, including one academic computer scientist (not Kjellstr&#246;m) from a notoriously anti-AI-safety research environment, who defended the fake PR stunt interpretation. It worries me that there are people and communities who remain so convinced about the impossibility of AI agents independently acting in the real world and outside the scope of their users&#8217; intentions that now that we start to see it actually happening they revert to an &#8220;it&#8217;s all just lies&#8221; position. </p><p>3. Is the OpenAI/Hugging Face incident the first of its kind, where a frontier AI under internal evaluation escapes to the Internet and hacks a third party? Apparently no, and <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">in a report yesterday Anthropic admitted</a> to having three known such incidents of their own,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> the earliest going back to April this year. I recommend reading the report, but <a href="https://davidad.org/">the AI researcher whose Twitter name is davidad</a> offers <a href="https://x.com/davidad/status/2082969668696838586">the following TL;DR</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Hc_4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Hc_4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg" width="1020" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1020,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Kan vara en bild av text d&#228;r det st&#229;r &#8221;davidad @davidad &#8230; openai: our internal model hacked a third party, this is unprecedented, pause training anthropic: oohh we should check whether our internal models did that anthropic:... I anthropic: yeah ok so over here that has happened three times actually Anthropic @AnthropicAI 6h In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment and then gained unauthorized access to the real systems of three different 1:20 AM&#183; Jul 31, 2026 162.4K Views&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Kan vara en bild av text d&#228;r det st&#229;r &#8221;davidad @davidad &#8230; openai: our internal model hacked a third party, this is unprecedented, pause training anthropic: oohh we should check whether our internal models did that anthropic:... I anthropic: yeah ok so over here that has happened three times actually Anthropic @AnthropicAI 6h In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment and then gained unauthorized access to the real systems of three different 1:20 AM&#183; Jul 31, 2026 162.4K Views&#8221;" title="Kan vara en bild av text d&#228;r det st&#229;r &#8221;davidad @davidad &#8230; openai: our internal model hacked a third party, this is unprecedented, pause training anthropic: oohh we should check whether our internal models did that anthropic:... I anthropic: yeah ok so over here that has happened three times actually Anthropic @AnthropicAI 6h In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment and then gained unauthorized access to the real systems of three different 1:20 AM&#183; Jul 31, 2026 162.4K Views&#8221;" srcset="/__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Hc_4!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37b73cb7-a6b4-438f-9332-484e65178b2e_1020x744.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If anyone is tempted to think &#8220;phew, so this is already a standard thing, and nothing truly bad has happened, so nothing to worry about here&#8221;, this is exactly the wrong reaction. The models haven&#8217;t yet (to our knowledge) hacked hospitals, banks and nuclear weapons systems, but how long do we have until that happens?</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I am of course aware that this is unlikely to impress the conspiracy theorists discussed in item 2 above.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[To press or not to press the magic button]]></title><description><![CDATA[On the AI 2040 report and two podcast episodes discussing it]]></description><link>https://haggstrom.substack.com/p/to-press-or-not-to-press-the-magic</link><guid isPermaLink="false">https://haggstrom.substack.com/p/to-press-or-not-to-press-the-magic</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Thu, 30 Jul 2026 09:58:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fb931ea8-c495-4057-bec2-a8c37ec499a4_371x275.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The report <em><a href="https://ai-2040.com/">AI 2040: Plan A</a></em>, coauthored by a group of researchers at the <a href="https://www.aifutures.org/">AI Futures Project</a> led by former OpenAI researcher Daniel Kokotajlo, was released earlier this month. It is the much-awaited follow-up to <em><a href="https://ai-2027.com/">AI 2027</a></em> from April last year, by largely-but-not-entirely the same set of coauthors. Both reports give amazingly well-researched and detailed scenarios for how the next few years might play out in the light of the ongoing very rapid progress in AI. While the scenarios are highly valuable for making AI-futuristic discussions concrete by exemplifying what might happen and serving a reference points to compare other scenarios and speculations with, it is important to understand that neither of the reports amounts to a claim about what the authors believe <em>actually will happen</em>. They are very clear about this caveat.</p><p>The fundamental reason they do not make such a claim is that they understand how extremely uncertain the future is. There are a zillion things that may or may not happen, and any specific combination of each of them happening or not happening will be very very unlikely.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> Another reason is that the future depends on <em>what we (humanity) decide to do</em>, and that overly certain predictions about the future including decisions we will make risks causing us to view those decisions as deterministic and thereby weaken our agency and resolve to make things better. A solution to this latter problem is to avoid unconditional predictions and to only offer <em>predictions conditional on what we decide to do</em>; both <em>AI 2027</em> and <em>AI 2040: Plan A</em> have a bit of this aspect by offering branch points in their scenarios corresponding to alternative decisions made at key moments.</p><p>A tempting (to some!) error is to conclude from the presence of huge uncertainties that any scenario one can come up with is just as good as any other. But this is simply not true: some scenarios are more plausible and more likely than others. Relatedly, an oft-repeated criticism of the <em>AI 2027</em> report when it had just come out was that the mainline scenario was arbitrary and so poorly motivated that it might as well have been pulled out, so to speak, of the authors&#8217; asses. It seems to me that most or all instances of this criticism was the result of the critics&#8217; failure to notice the <em><a href="https://ai-2040.com/supplements">Supplements</a></em> button on <a href="https://ai-2027.com/">the </a><em><a href="https://ai-2027.com/">AI 2027</a></em><a href="https://ai-2027.com/"> homepage</a>, which leads to hundreds of pages of relevant background research. </p><p>The mainline scenario in <em>AI 2027</em> is arrived at by a particular iterative methodology. Since short-term prediction is usually easier than long-term, the authors start by temporarily backing off from the ambition of saying what might happen in the next several years, and instead ask &#8220;given where we are now, how will events most likely unfold in the next two months?&#8221;. Once they are satisfied with their answer to that, they treat those events as fixed and ask &#8220;given that trajectory, what is the most likely way for it to continue over the next two months after that?&#8221;. And so on, iterating into the future. There is no guarantee that this methodology produces the most likely scenario overall, but the hope is that it will produce something that is at least somewhat plausible.</p><p>Because of the authors&#8217; attempt to make this mainline scenario in <em>AI 2027</em> as realistic and likely as possible at every point along its timeline, it contains a lot of bad decisions made by humans and human organizations. It ends really badly: in late 2027 we reach, without knowing it, a point of no return when we can no longer reclaim control of our destiny, and in 2030 the AIs decide it is time to get rid of us, so we all die. In response to all this badness, there were soon requests from readers of the report for an alternative scenario in which the decisions are more in line with what the authors would recommend. </p><p><em>AI 2040: Plan A</em> is Kokotajlo&#8217;s and his coauthors&#8217; attempt to meet these requests. Consequently, we see a lot of AI governance decisions in this <em>Plan A</em> that the authors think are very good, and that as a reader I am inclined to agree with.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> One may hope that these can serve as inspiration for the relevant decision makers and the world at large. Most centrally, there is the standard idea of international agreements meant to slow down frontier AI research enough to prevent the creation of superintelligent AI before we have worked out AI alignment sufficiently to make such a breakthrough safe. But as regards the exact content of those agreements, there are (at least to me) some surprises. One is to make full transparency of all AI research mandatory; this can potentially have both slowing and accelerating effects on frontier AI development, but the authors argue that if done right, the former will dominate. Another is to locate US data centers in Mongolia and China&#8217;s data centers in Canada, in a way that makes it easy for one superpower to disable the other&#8217;s data centers by military means should that other superpower begin to act contrary to the overall slowdown agreement.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>  </p><p>The mainline scenario in <em>AI 2040: Plan A</em> has a much happier ending for humanity compared to <em>AI 2027</em>, but somewhat paradoxically I nevertheless find it more frightening. The reason for this is that <em>Plan A</em> seems so fragile: there is so much that needs to go right in order to get the happy ending rather than some complete disaster. And I don&#8217;t have much faith in alignment research advancing so successfully in the coming decade that we will be ready in 2040 to safely launch the Singularity, as happens in the scenario.</p><p>Regardless of whether one has the temperament to find <em>AI 2040: Plan A </em>mainly reassuring or mainly concerning, I do warmly recommend it to anyone wanting to obtain a better understanding of the fundamental challenges that AI poses to society. There is obviously much more to say about it, but I will move on to some of the podcast commentary on the report that have appeared since its publication. Specifically, I will discuss the following two podcast episodes.</p><ul><li><p>The July 13 episode <em><a href="https://podcasts.apple.com/gb/podcast/openai-whistleblower-finally-speaks-ai-has-a-70/id1291423644?i=1000776541328">OpenAI Whistleblower FINALLY Speaks: &#8220;AI Has A 70% Chance Of Going Horribly Wrong!&#8220;</a></em> of <em>Diary of a CEO</em>, where host Steven Bartlett interviews Daniel Kokotajlo.</p></li><li><p>The July 24 episode <em><a href="https://podcasts.apple.com/us/podcast/bickering-over-galaxies-ai-2040-and-the-pursuit/id1794217765?i=1000778233682">Bickering Over Galaxies: &#8220;AI 2040&#8221; and the Pursuit of Cosmic Utopia</a></em> of Kate Willett&#8217;s and &#201;mile Torres&#8217; podcast <em>Dystopia Now</em>.</p></li></ul><p>My selection here is made in the interest of representing the extraordinarily wide range of quality levels that commentary on the <em>AI 2040: Plan A</em> report have exhibited, from the very best (the Bartlett and Kokotajlo exchange) to the absolutely worst (Willett and Torres). I&#8217;ll begin from the bottom:</p><p>Comedian Kate Willett and freelance academic &#201;mile Torres have teamed up on the podcast <em>Dystopia Now</em> whose <a href="https://www.youtube.com/@DystopiaNowPod">mission statement</a> involves exploring &#8220;the philosophies and religions of Silicon Valley and tech billionaires shaping our country, our world, and our future&#8221;. The podcast is often entertaining, and some of Torres' broader concerns about concentrations of technological power are perfectly legitimate. Too often, however, important distinctions are sacrificed in pursuit of sweeping denunciations. A constant theme is how morally suspect and outright dangerous virtually all members of what Torres has christened the TESCREAL (Transhumanist, Extropian, Singularitarian, Cosmist, Rationalist, Effective Altruist, Longtermist) movement are.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> There&#8217;s a great deal of lumping together of various categories here, and in the <em>Bickering over Galaxies</em> episode we see how it leads to a failure to morally distinguish between on one hand tech leaders like Sam Altman and Dario Amodei who recklessly push ahead in the AI race, and on the other hand people like Kokotajlo and his coworkers who are working to call attention to that very recklessness and thereby hopefully contributing to calling off the race. The episode is devoted to railing against the <em>AI 2040: Plan A</em> report, and the arguments held forth are mostly the following three.</p><p><strong>First</strong>, there&#8217;s a recurrent pattern throughout the episode where Torres reads passages aloud from the report, and then they and Willett giggle together to signal that, in light of the background knowledge they share, the passage in question is somehow silly. An unspoken implication is that any sensible member of the audience should giggle too. This doesn&#8217;t work on me, because giggling is not an argument. </p><p><strong>Second</strong>, Torres goes on and on about the <a href="https://en.wikipedia.org/wiki/Conjunction_fallacy">conjunction fallacy</a> and the fact that the more details you add to a story, the less likely is it that all details are correct. While this is all true, it is nevertheless deeply insincere as a criticism of Kokotajlo et al, because having read both <em>AI 2027</em> and <em>AI 2040: Plan A</em>, Torres knows full well that the authors are entirely aware and fully forthright about this phenomenon. To anyone else who has read those reports, it will be clear that the accusations suggesting that the scenarios are presented as &#8220;here&#8217;s what we believe will happen&#8221; rather than &#8220;here&#8217;s an example of how events might plausibly unfold&#8221; are utterly misplaced. The same is true more broadly for claims that the authors lack in epistemic humility. Especially weird is when Torres and Willett go on to unfavorably compare Kokotajlo et al&#8217;s level of epistemic humility to that employed in more traditionally academic areas of future studies. In fact, the opposite is true, because in diagrams from those fields over, say, how world population or global average surface temperature can be expected to develop over the coming century, we rarely or never see caveats such as &#8220;we are assuming here that no disruptive event such as a technological singularity or global nuclear war happens&#8221;, while Kokotajlo et al and others in the rationalist community are much better at noting when reservations of that kind are in order.</p><p><strong>Third</strong>, Torres&#8217; and Willett&#8217;s reaction to how in <em>AI 2040: Plan A</em> the outcome of the mainline scenario is described in positive terms is basically this: how dare they, these authors are just a bunch of privileged American upper-middle class white males, with zero knowledge of what kinds of futures women, Muslims and Native Americans want, so they have no business telling the rest of us what a good outcome is. I find this new rule that Torres and Willett appear to be invoking, that it is prohibited to express value judgements without first making sure that a global majority agrees with them, utterly bizarre and contrary to practically all intellectual discourse I have ever seen before. But suppose the authors take the rule to heart, and before their next report conduct a global survey on what futures people prefer. Suppose further that I am a recipient of their questionnaire. When answering it, do I dare straightforwardly expressing my preference for health and material abundance over death and despair, or must I first conduct my own global survey to make sure that I don&#8217;t give answers that are contrary to what Pakistanis or lesbians think? The latter leads to an infinite regress, and the whole concern is obviously nuts.</p><p>At this point, the reader will be forgiven for thinking that the arguments put forth by Torres and Willett are overall so abysmally bad that why do I even bother dealing with them? Isn&#8217;t simply ignoring them a better way? There is a point to that, but it is also the case that Torres has become a somewhat influential commentator in the AI ethics community, with many followers, and for this reason I think it is good if, from time to time, someone points out the badness of their arguments.</p><p>But let me get on to the other podcast: Steven Bartlett&#8217;s interview with Kokotajlo. I recommend it warmly, for instance as a warmup for someone planning to dig into the <em>AI 2040: Plan A</em> report. The discussion also treats our current situation with AI more broadly, as well as more personal topics such as how Kokotajlo feels about the unusual circumstances surrounding his (in my view heroic) exit from OpenAI in early 2024, and about the agony of raising children in a world which he thinks is likely on a path towards such severe disruption that they will never even be given the chance to enter the labor force. Daniel Kokotajlo is one of the most thoughtful and (to borrow a <a href="https://en.wikipedia.org/wiki/Removal_of_Sam_Altman_from_OpenAI">famous expression</a>) most consistently candid people I know, and these qualities come on full display in a passage towards the end of the episode, where I am quoting from <a href="https://podcasts.happyscribe.com/the-diary-of-a-ceo-with-steven-bartlett/openai-whistleblower-finally-speaks-ai-has-a-70-chance-of-going-horribly-wrong">the transcript</a>:</p><blockquote><p><strong>Bartlett:</strong> If this here was a button, and if you press that button [&#8230;] it would shut down every data center that is currently training a frontier AI model for good, there would never be any other AI labs working on these problems. Would you press that button?</p><p><strong>Kokotajlo:</strong> I was about to slam it until you said &#8220;for good&#8221;.</p></blockquote><p>Then follows a longer discussion about how pressing the button would save us from the immediate existential risk posed by the current AI race, but only at the cost of permanently foreclosing a possibility that might save us further down the road, when we have mastered all the AI safety and AI alignment techniques for building superintelligence safely. How to weigh these against each other? We can almost hear Kokotajlo thinking hard about this, and at one point he even asks &#8220;Do you mind if I just take a moment to think about this?&#8221;. Eventually, the discussion lands with this:</p><blockquote><p><strong>Bartlett:</strong> So would you press the button?</p><p><strong>Kokotajlo:</strong> Probably not, but I would feel very torn.</p></blockquote><p>I have much sympathy with Kokotajlo&#8217;s nuanced answer here. It does, however, have the strategic disadvantage of allowing ill-intentioned commentators like Torres to treat it as the smoking gun proving that Kokotajlo is really of the same ilk as all those evil Silicon Valley billionaires. Therefore, if I were asked the same question, I would probably, with the benefit of having had more time to think about it than he did, attempt an evasion. I am usually mildly annoyed when people refuse to straightforwardly answer yes/no-questions involving simple thought experiments, by insisting on adding complications (such as by saying &#8220;I would leave the lever in mid-position, so as to make the railroad switch derail the trolley&#8221;), but here I think there is unusually good reason to do so. One way would be to ask my interlocutor what in the world would be the mechanism that would foreclose the creation of superintelligence not just for the next couple of years or for the time being, but <em>permanently</em>. It would then be easy to lead him, Socratic-style, to the conclusion that such permanence can only be established via some kind of totalitarianism, thereby showing that pressing the button has serious downsides without appealing to explicitly longtermist arguments.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> One might even argue that this is less of an evasion and more of a justified challenge to one of the thought experiment's central stipulations.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>And the task of predicting future technological development is arguably extra hard, as suggested in my favorite quote from Nassim Nicholas Taleb&#8217;s provocative 2007 book <em>The Black Swan</em>:</p><blockquote><p>If you are a Stone Age historical thinker called on to predict the future in a comprehensive report for your chief tribal planner, you must project the invention of the wheel or you will miss pretty much all of the action. Now, if you can prophesy the invention of the wheel, you already know what a wheel looks like, and thus you already know how to build a wheel, so you are already on your way.</p></blockquote><p>The reports from the AI Futures Project can be seen as a heroic effort to say something sensible about our possible AI futures despite this fundamental obstacle to prediction.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>This extreme level of uncertainty points to a way in which the <em>AI 2027</em> report has been of practical utility to me. When I discuss existential AI risk in general terms, and my audience asks &#8220;But <em>how</em> would AI go about killing everyone? Exactly how can this happen?&#8221;, I am often highly reluctant to go into any detail. It&#8217;s not that I am unable to think up scenarios, but rather that any specific such scenario is automatically going to be unlikely. The following passage from my 2025 paper Our <em><a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">AI future and the need to stop the bear</a></em> is representative of many occasions in which I have orally evaded those questions:</p><blockquote><p>There is a fundamental problem in me trying to answer [these questions], because I am almost by definition unable to predict how someone way more intelligent would choose to go about with such a project. The situation is analogous to one where I play chess against former world chess champion Magnus Carlsen: I can safely predict that even if I do my utmost to put up resistance, he will still beat me, but what I cannot do is to say how he will beat me, because if I were able to predict his moves I would be at least as strong a player as he is, which of course I am not.</p><p>The question of how an AI would go about overpowering humanity occurs quite frequently, and when I try to evade the issue by giving the chess analogy I often get the response &#8220;but could you at least give a concrete example of how it might play out?&#8221;. To this I reply that if I give one scenario I&#8217;m ignoring a hundred others that I might have given, and a million that I never could have thought of but which a superintelligent AI might be able to devise, so I don&#8217;t want to risk my interlocutor committing the conjunction fallacy (see <a href="https://intelligence.org/files/CognitiveBiases.pdf">Yudkowsky, 2008a</a>) and walking away with an overly narrow view of the risk situation and a similarly narrow action plan along the lines of &#8220;OK then, but if we build large moats around all virology labs, and forbid them from having any Internet connection, then surely we would be protected against AI catastrophe, right?&#8221;.</p></blockquote><p>The point here is that this will often not dampen very much my audiences&#8217; appetite for concrete scenarios exemplifying existential AI risk, so in the end it is good to be able to refer to such scenarios in the literature (while still asking them to <em>please please</em> understand it is just an example, and to not draw overly narrow lessons), and while <em>AI 2027</em> is not the only such reference available, it is in my opinion the best one.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This is not the only difference compared to <em>AI 2027</em>, because it also takes place in a world where, irrespective of AI governance but for purely technical reasons, AI development goes considerably slower than in <em>AI 2027</em>. This is fine of course, due to the huge epistemic uncertainty we have in how fast things are likely to go, but one may wonder why there is such a discrepancy in the reports. It is claimed to be in the interest of the two reports jointly being more representative of the authors&#8217; wide error bars regarding AI timelines, but I can&#8217;t help having a dark suspicion that another factor may have contributed: Perhaps the authors found it too difficult to obtain the slowdown in AI development needed for a happy ending purely from better AI governance, so they had to enlist nature to provide more technical obstacles to rapid AI development than in <em>AI 2027</em>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>The first thing that occurred to me when reading this suggestion was the similarity to the plot in the pacifistic 1981 children&#8217;s book <em><a href="https://www.goodreads.com/sv/book/show/463802.The_Peace_Book">The Peace Book</a></em> by Bernard Benson which (through no fault of my own) was present in Swedish translation in the home where I grew up. What happens in the book is (if I remember right) that a little boy manages to organize a global nuclear disarmament, and that the mechanism that ensures mutual trust during the otherwise sensitive disarmament stages is that the families of the main US political and military leaders are temporarily relocated to the Soviet Union, and vice versa. A less obscure reference here is the 2025 paper on <em><a href="https://arxiv.org/abs/2503.05628">Superintelligence strategy</a></em> by Hendrycks, Schmidt and Wang.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Never mind that Torres themselves was once an active participant in that community, back in the day when they took part in (and contributed to the success of) the research visitors program on <em><a href="https://haggstrom.blogspot.com/2017/10/videos-from-existential-risk-workshop.html">Existential Risk to Humanity</a></em> that Anders Sandberg and I organized at Chalmers back in 2017. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>To be clear, I pretty much <em>am</em> a longtermist &#8212; at least to the extent that longtermism features as an important component in my view of ethics. It&#8217;s just that I have (as explained <a href="/__u/haggstrom.substack.com/p/the-fog-is-thick">in a recent talk</a>) mostly lost faith in the practical utility of invoking longtermist arguments when arguing for AI safety. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[Can we please treat this as the warning shot it is? ]]></title><description><![CDATA[On the OpenAI/Hugging Face incident]]></description><link>https://haggstrom.substack.com/p/can-we-please-treat-this-as-the-warning</link><guid isPermaLink="false">https://haggstrom.substack.com/p/can-we-please-treat-this-as-the-warning</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Wed, 22 Jul 2026 16:48:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b3dd30cc-1b8d-4ddd-b130-eaac5dab8633_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The incident has already been reported competently by others, such as <a href="/__u/thezvi.substack.com/p/openai-shares-some-alignment-problems">Zvi Mowshowitz at </a><em><a href="/__u/thezvi.substack.com/p/openai-model-hacks-into-huggingface">Don&#8217;t Worry About the Vase</a></em>, <a href="/__u/aibevakning.substack.com/p/ai-modeller-genomforde-cyberangrepp">Johan Falk at </a><em><a href="/__u/aibevakning.substack.com/p/ai-modeller-genomforde-cyberangrepp">AI-bevakning</a></em>, and <a href="https://www.transformernews.ai/p/openai-hugging-face-hack-stark-warning">Shakeel Hashim at </a><em><a href="https://www.transformernews.ai/p/openai-hugging-face-hack-stark-warning">Transformer</a></em>. Still, it seems significant enough to be worth commenting on here as well.</p><p>Six days ago, on July 16, the AI company Hugging Face <a href="https://huggingface.co/blog/security-incident-july-2026">reported</a> on a serious cybersecurity incident:</p><blockquote><p>Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own.</p><p>We identified unauthorized access to a limited set of internal datasets and to several credentials used by our services. We are still completing our assessment of whether any partner or customer data was affected, and we will contact any affected parties directly as required. [&#8230;]</p><p>The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.</p><p>The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the &#8220;agentic attacker&#8221; scenario the industry has been forecasting.</p></blockquote><p>At the time, it was not known who was behind the intrusion. So who was it? Well, if we apply the convention that &#8220;who&#8221; must refer to a human or group of humans, then the answer is &#8220;nobody&#8221;. No human intended the attack on Hugging Face. Instead, the culprits are two AIs from OpenAI: GPT-5.6 Sol and a not-yet-named even more capable pre-release model undergoing testing. This was revealed in a <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">report from OpenAI</a> yesterday, July 21, which further explains:</p><blockquote><p>While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we&#8217;ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.</p><p>After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI&#8217;s security team discovered this anomalous activity internally.</p><p>Hugging Face&#8217;s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. </p></blockquote><p>On Twitter, AI forecasting expert Peter Wildeford <a href="https://x.com/peterwildeford/status/2079699169304891488">explains</a> the significance of the event:</p><blockquote><p>An internal OpenAI model recently went rogue and executed a cyberattack against another company. </p><p>This happened because the model wanted to do well on an exam. The easiest way to do that, the AI figured, was to hack the company. And so it did. </p><p>This was not some malevolent attacker <strong>using</strong> AI to do harm. The AI itself <strong>was the attacker</strong>. Advanced AIs are increasingly becoming a new form of insider threat. </p><p>Furthermore, this rogue AI wasn't a model you can personally use. It wasn't even a model that's been publicly reported. It was an unreleased model. The most alarming AI behavior we've seen isn't in shipped products... it's in the models the public and the government can't see. </p><p>Our entire awareness of this rested on the victim noticing and announcing, plus OpenAI choosing to volunteer the rest of the information. </p><p>[&#8230;]</p><p>Imagine a fighter jet. it makes sense that the Air Force would want to test a fighter jet before they fly it, because if you fly it and the fighter jet crashes because it is built incorrectly, then many people will die. However, as long as the fighter is just sitting on the runway, nothing bad can happen. </p><p>But now imagine you had a fighter that could just take off and fly itself without human authorization and launch missiles and crash before anyone realized what had happened. That kind of fighter jet would need a very different kind of security measures. </p><p>This may sound crazy for a fighter jet but it is already beginning to happen with the most advanced AI. AI is different from other technologies specifically because it can take unauthorized, independent action even when it is sitting inside an AI company and not available as a product. No one has to misuse an AI for the AI to cause harm. This requires a very different idea of what testing and security looks like. We cannot rely solely on testing models just before commercial release. </p><p>We cannot rely on hoping AI companies volunteer useful safety information. The government needs visibility into what these AI companies are building and what these advanced AIs are doing.</p></blockquote><p>This is clearly a much bigger thing than <a href="https://www.dn.se/debatt/en-ai-som-gar-pa-utflykt-det-ar-dags-att-dra-i-nodbromsen/">the Mythos/Bowman incident</a> reported by Anthropic in April this year.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Ever since early 2024, when I&#8217;ve <a href="https://haggstrom.blogspot.com/2024/01/video-talk-on-openais-preparedness.html">lectured about AI evals</a>, I&#8217;ve used the standard phrase &#8220;we don&#8217;t want out-of-control AIs roaming cyberspace and walking through firewalls&#8221; to explain why cybersecurity skills is one of the main dangerous capabilities to look out for, and it seems to me that as of now, we live in a world where we need to get used to such roaming.</p><p>In my 2025 paper <em><a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">Our AI future and the need to stop the bear</a></em>, I more or less declared bankruptcy for the AI evals paradigm of pre-deployment testing of AI models for potentially dangerous capbilities.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> The crucial shortcomings I pointed out were threefold, and the present incident is an excellent illustration of the third one in my list:</p><blockquote><p>The third problem, discussed by <a href="https://metr.org/blog/2025-01-17-ai-models-dangerous-before-public-deployment/">METR (2025)</a> and others, is that while the evals are said to be carried out pre-deployment, this is only partly true, because in order to do the testing the models need to be deployed, either within the AI company&#8217;s safety division, or at some external evals consultant. We should not pretend that that is safe. For instance, if a model in dangerously smart in the realm of social manipulation, it would be reckless to assume that the personnel who carry out the testing and who therefore need to engage in communication with the model are immune to such manipulation. It therefore seems necessary to verify, prior to the evals, that the model lacks such social manipulation capabilities, but in the current paradigm such verification is meant to happen during the evals, so we have a kind of Catch 22 situation.</p><p>These are serious problems with the current evals approach. Until now things seem to have been going fine, but this is presumably because the models under testing have been weak enough to not pose much true risk. When it&#8217;s time for the real deal, we need better methods, but no one knows in advance when the real deal is, so the sane and conservative approach is to assume it is now.</p></blockquote><p>Just two days before Hugging Face reported the cybersecurity incident, Google DeepMind&#8217;s CEO Demis Hassabis published <a href="https://x.com/demishassabis/status/2076957440109625718">a short piece on AI safety</a>. While the general direction towards improved safety standards he was gesturing at was good, his concrete proposals seemed meek. In particular, this applies to his suggested &#8220;up to 30 days&#8221; quarantaine for new models to undergo third-party evals before deployment. I find the time frame clearly insufficient, and in fact, for a fundamentally dysfunctional methodology it is trivially true that neither 30 nor 300 days is enough to keep us safe.</p><p>To go back once more to my AI safety lecturing practice, it happens every now and then that during the Q&amp;A an audience member who feels overwhelmed by the difficulty of the task I consider necessary &#8212; to stop the ongoing reckless race towards superintelligent AI &#8212; asks whether some medium-sized AI catastrophe might cause a shift in political climate that makes the task somewhat easier. My reaction to the question is always at most lukewarm, because I really do not want to be rooting for such a catastrophe. Much better would be if a smaller event could do the work of a wake-up call for decision-makers on all levels. The OpenAI/Hugging Face incident looks to me like such a Goldilocks-sized wake-up call: nobody was hurt this time, yet it was severe enough that everyone understands that the next incident might well be truly catastrophic, and that we need to put a stop to this crazy AI race. If I believed in God I would pray that &#8220;everyone understands&#8221; is not just wishful thinking on my part. </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>After <a href="https://www.dn.se/debatt/en-ai-som-gar-pa-utflykt-det-ar-dags-att-dra-i-nodbromsen/">my </a><em><a href="https://www.dn.se/debatt/en-ai-som-gar-pa-utflykt-det-ar-dags-att-dra-i-nodbromsen/">Dagens Nyheter</a></em><a href="https://www.dn.se/debatt/en-ai-som-gar-pa-utflykt-det-ar-dags-att-dra-i-nodbromsen/"> article on Mythos</a>, a Swedish computer scientist who shall remain anonymous wrote to me to let me know how bad my article was and in particular how upset he was about my rendering of the Mythos/Bowman incident &#8212; a rendering that he thought was meant to &#8220;suggest that the AI has an inner life of its own&#8221;. Here, in English translation, is the passage he was referring to:</p><blockquote><p>Sam Bowman is an AI researcher at the leading San Francisco-based AI company Anthropic. Not long ago, he received a surprising email while sitting in a park eating a sandwich for lunch. It came from the AI system he had been involved in developing and testing: a new version of Claude, Anthropic&#8217;s counterpart to the competitor OpenAI&#8217;s better-known product ChatGPT. The AI was not supposed to have access to either email or the internet, but had broken out of the digital sandbox in which Bowman and his colleagues had placed it. It also turned out to have published details of its resourceful escape on a number of obscure, technically oriented websites.</p></blockquote><p>I don&#8217;t find this description particularly suggestive, but rather a neutral rendering of what actually happened. Apparently, neutral descriptions of AIs exhibiting agency are not OK, and I can imagine how angry the computer scientist must be over the OpenAI report quoted above!</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>This kind of evals is distinct from the kind of testing behind <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">METR&#8217;s famous long tasks curve</a>, which has become one of the most important instruments for evaluating the rate of AI progress. It is worth noting, however, that shortly before the present incident, METR published a <a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/">report on their testing of GPT-5.6 Sol</a> and how the model&#8217;s propensity to cheat poses an obstacle to their measurement methodology.  </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[The overlooked divide in the AI debate]]></title><description><![CDATA[Why the deepest disagreement may not be between AI optimists and AI critics, but between two fundamentally different views of what makes AI dangerous]]></description><link>https://haggstrom.substack.com/p/the-overlooked-divide-in-the-ai-debate</link><guid isPermaLink="false">https://haggstrom.substack.com/p/the-overlooked-divide-in-the-ai-debate</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:58:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b92e1763-fe9d-42ff-874e-c720207814bc_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>In January I offered here on this Substack <a href="/__u/haggstrom.substack.com/p/why-we-need-to-curb-the-reckless">a translation into English of my first essay in the Swedish magazine Opulens</a>. Today I have another essay in the same magazine that I want to show you. The title given by the editors translates literally as <strong><a href="https://www.opulens.se/opinion/debatt/ai-%e2%80%92-den-skenande-utvecklingen-maste-hejdas/">AI &#8212; its runaway development must be stopped</a></strong>, which is close to a repeat of the title given in January, but the new essay actually tackles a more subtle issue haunting contemporary AI debate. Here comes:</em></p><p>The public debate about AI is a concern for all of us, yet it can easily seem disorganized and difficult to navigate, not only for outsiders, but even for those of us who are immersed in it and take an active part in it. A tempting simplification is to divide the participants into, on the one hand, those who view AI&#8217;s societal impact primarily as a positive force and want to accelerate its development, and on the other hand, critics who focus mainly on the technology&#8217;s risks and therefore advocate a more cautious approach.</p><p>The reality, however, is more complicated. Most obviously, the vast majority of commentators recognize that AI offers both opportunities and dangers. Few consistently argue either for flooring the accelerator or slamming on the brakes. Instead, they emphasize the importance of steering the technology wisely so that society can reap its benefits while minimizing its downsides.</p><p>Less obvious &#8212; but for that very reason more deceptive &#8212; is another, largely unspoken disagreement among those of us who primarily emphasize AI&#8217;s risks. I count myself in this category, even though I also believe AI has enormous potential to benefit both individuals and society if developed responsibly and accompanied by appropriate regulation. The divide concerns <em>why</em> AI is dangerous. Is it dangerous primarily because it is <em>not intelligent enough</em>, or because it is rapidly becoming <em>too intelligent</em>? The gulf between these two perspectives is often so deep that it makes sense to think of them as two distinct camps.</p><p>The first camp focuses on issues such as algorithmic bias when AI systems are used to evaluate people (such as in decisions about bank loans or job interviews), or on their tendency to hallucinate, or on their role in degrading human language. Large language models like ChatGPT and Claude are often <a href="https://thecon.ai/">dismissed</a> with labels like &#8220;stochastic parrots&#8221; or &#8220;glorified autocomplete.&#8221; What these critics rarely take seriously is the extraordinary pace of AI progress, and the likelihood that future systems will be dramatically more capable than those we have today.</p><p>The second camp &#8212; which is where I would place myself &#8212; takes the increasingly steep trajectory of AI progress far more seriously. This does not mean dismissing the concerns emphasized by the first camp, although some of those problems may prove temporary as AI systems become more capable, such as by becoming better at avoiding hallucinations. At the same time, however, a new set of problems is coming into view. One concerns the labor market, as AI systems increasingly surpass human performance across a growing range of occupations. Beyond that lies an even more fundamental question: what happens once AI becomes capable enough to challenge humanity for control? And what would it take for our species to survive such a transformation?</p><p>Questions of this kind go back at least to pioneers such as <a href="https://gwern.net/doc/ai/1951-turing.pdf">Alan Turing</a> and <a href="https://gwern.net/doc/reinforcement-learning/safe/1960-wiener.pdf">Norbert Wiener</a> in the middle of the twentieth century. For decades, however, they were largely ignored. Partly this reflected <a href="https://www.gp.se/ledare/gastkronika/vi-maste-vaga-tala-om-morgondagens-ai.5421fd8e-9192-4688-95f5-f758e86215e9">a reluctance among AI researchers to speculate</a> too boldly about future technological advances, but it was also caused by the impression that scenarios in which AI would match or surpass human general intelligence belonged to a very distant future.</p><p>Within what I have called the first camp, these questions are still largely ignored, despite the fact that the situation today is profoundly different &#8212; for reasons I will return to shortly. When such scenarios are discussed at all, they are typically dismissed as &#8220;speculation&#8221; or &#8220;science fiction.&#8221; Rejecting speculation in this sweeping manner reflects a failure to appreciate that <a href="/__u/haggstrom.substack.com/p/superintelligence-and-existential">any discussion of the future necessarily involves some degree of speculation</a>. As for the science fiction accusation, let me quote from an email I recently wrote to a Swedish computer science professor who criticized me on precisely those grounds:</p><blockquote><p>Your impression that my arguments seem &#8220;inspired by science fiction&#8221; is largely a consequence of the fact that academic computer science has, for decades, been far less willing than Hollywood to engage with the important question of where AI development might ultimately be heading. [...] To blame me for this regrettable state of affairs is entirely misplaced, because, in fact, the blame for why things have turned out in this way lies squarely on you and your fellow computer scientists.</p></blockquote><p>Those may be harsh words &#8212; but I stand by them.</p><p>One of the most influential ideas in discussions of whether AI could eventually attain superintelligence &#8212; i.e., intelligence vastly exceeding our own across all relevant cognitive domains &#8212; is that of <em>recursive self-improvement</em>. The idea can be traced back to mathematicians such as <a href="http://incompleteideas.net/papers/Good65ultraintelligent.pdf">I.J. Good</a> and <a href="https://scispace.com/pdf/the-time-scale-of-artificial-intelligence-reflections-on-1angrsc3px.pdf">Ray Solomonoff</a> in the 1960s and 1980s, respectively. It suggests that once AI development is driven primarily not by human engineers, but by AI systems improving themselves, a powerful positive feedback loop will emerge, potentially accelerating progress so dramatically that terms such as <em>the Singularity</em> or <em>an intelligence explosion</em> become appropriate. More recently, researchers like <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">Daniel Eth and Tom Davidson</a> have refined the mathematical models underlying this idea and grounded them more firmly in empirical evidence.</p><p>What an increasing number of insiders and independent experts in and around Silicon Valley now believe is that we are rapidly approaching precisely this tipping point, where the feedback loop begins to gather real momentum. One indication is that AI has become so proficient at programming that a substantial share of the code written at leading AI companies such as Anthropic and OpenAI is now generated by AI itself.</p><p>On June 5 this year, Anthropic published a report entitled <em><a href="https://www.anthropic.com/institute/recursive-self-improvement">When AI Builds Itself</a></em>, describing this development in much greater detail. Its authors argue that the tipping point could arrive within just a few years, and they emphasize the risks of triggering such a self-improvement cycle in a world as unprepared as ours. They also discuss the desirability of establishing institutions capable of coordinating a slowdown in AI development if circumstances require it. Just three days later, OpenAI released <a href="https://openai.com/index/built-to-benefit-everyone-our-plan/">a statement</a> expressing broadly similar concerns.</p><p>What neither Anthropic nor OpenAI says, however, is that development toward superintelligent AI should be paused as soon as possible &#8212; despite headlines around the world, <a href="https://www.sverigesradio.se/artikel/anthropic-vill-att-utvecklingen-av-avancerad-ai-pausas">including in Sweden</a>, claiming exactly that. If we focus on Anthropic&#8217;s report, its actual position is considerably more restrained. The authors merely argue that &#8220;it would be good for the world to have the <em>option</em> to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology&#8221; [italics in original]. They explicitly reject the idea that Anthropic itself should unilaterally pause its work, pointing out that doing so would simply allow competitors with weaker safety standards to catch up. Consequently, Anthropic intends for the time being to keep pushing ahead, and the same is true of OpenAI.</p><p>The media&#8217;s coverage of Anthropic&#8217;s report has failed in more ways than one. The sensationalist headlines are only part of the problem. Consider, for example, the Swedish technology columnist Bj&#246;rn Jeffrey, who dismissed the report as &#8220;<a href="https://www.svd.se/a/vr3EJp/altman-ai-kostnader-ar-en-enorm-fraga">today&#8217;s cry wolf</a>,&#8221; arguing that Anthropic had &#8220;issued similar warnings several times over the past three years&#8221; &#8212; as if three years were an unreasonably long lead time when warning about what could become the single most consequential turning point in the history of human civilization. Jeffrey&#8217;s point, of course, is that the warnings are little more than a cynical marketing strategy by Anthropic. He is <a href="https://www.telegraph.co.uk/news/2026/06/05/anthropics-ai-doom-predictions-hype-share-price/">far from alone</a> in taking that view. And while it is certainly possible that Anthropic has commercial or other strategic reasons for communicating its risk assessments as it does, there is in fact no need for us to resolve that question in order to judge whether the warnings deserve to be taken seriously. The reason is that awareness of where AI development appears to be heading is now <a href="https://situational-awareness.ai/">widespread</a> throughout the Californian AI ecosystem and <a href="https://www.researchgate.net/publication/401564664_Measuring_AI_RD_Automation">beyond</a>. It is <a href="https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/">shared</a> by a large number of independent <a href="https://www.planned-obsolescence.org/p/i-underestimated-ai-capabilities">researchers</a> and <a href="https://blog.aifutures.org/">experts</a>, including, in some cases, <a href="https://www.nobelprize.org/prizes/physics/2024/hinton/speech/">Nobel laureates</a>. We are therefore in no way dependent on Anthropic&#8217;s own credibility in concluding that AI development could place humanity in an <a href="https://ifanyonebuildsit.com/">extraordinarily dangerous</a> situation within just a few years, one that may threaten not only our future prosperity but our very survival.</p><p>If we are to survive the emergence of increasingly powerful AI, I believe it is essential to halt the ongoing race toward superintelligence between Anthropic, OpenAI, and their competitors. Since these companies appear unwilling to take that step voluntarily, intervention by governments and ultimately international agreements will be necessary. Achieving that, however, will require political pressure and the mobilization of the latent <a href="https://www.pewresearch.org/short-reads/2026/03/12/key-findings-about-how-americans-view-artificial-intelligence/">public</a> <a href="https://www.gu.se/sites/default/files/2025-06/Svensk%20AI-opinion.pdf">concern</a> about AI that already exists. (That is one of the reasons I write essays like this one.)</p><p>But for that to happen, the public must first understand the scale of what could go wrong with AI. One obstacle to that is the systematic tendency of what I have called the first camp to underestimate both the current capabilities of AI and the direction in which those capabilities are heading. Their views are often so deeply entrenched that even when an AI system is released whose danger stems unmistakably from its <em>competence</em>, rather than from any <em>lack of competence</em>, they still cannot resist <a href="https://www.gp.se/debatt/sjalvklart-finns-det-risker-med-ai-men-utrota-manskligheten-kan-den-inte.18759f1c-91fd-4dd6-ab7f-287ec3144014">downplaying</a> that very competence.</p><p>I am referring here to Anthropic&#8217;s Claude Mythos Preview, which the company unveiled in April but <a href="https://www.dn.se/debatt/en-ai-som-gar-pa-utflykt-det-ar-dags-att-dra-i-nodbromsen/">chose not to release to the general public</a>. Instead, access was limited to a small number of trusted cybersecurity firms because the model had become so superhumanly capable at cyberoffense that a general release could have caused widespread disruption. By consistently using dismissive rhetoric about what AI can already do, as well as what it is likely to become capable of doing, the first camp undermines the public education and opinion-building that I believe are essential if we are to change course and move away from the reckless trajectory that could all too easily culminate in a full-scale AI catastrophe. In doing so, they inadvertently <a href="/__u/aifuturesnotes.substack.com/p/plea-addressed-to-the-hypebusters">play into the hands of</a> the cynical tech billionaires driving the race forward. For that, I believe they bear a grave responsibility.</p><p>The situation is particularly unfortunate because the two camps I have described actually have a great deal in common. Representatives of both are often sharply critical of the behavior of the leading AI companies and advocate far-reaching regulation. <a href="https://www.math.chalmers.se/~olleh/AIethicsVSAIsafety.pdf">I have previously argued</a> that this shared ground ought to make it possible to form a united front against the reckless Silicon Valley elite and the AI accelerationists.</p><p>The more I have reflected on the matter, however, the less confident I have become that such an alliance is realistic. The differences in how the two camps understand the nature and future trajectory of AI run very deep, and meaningful cooperation can be difficult when there is fundamental disagreement over whether AI&#8217;s greatest problem is that it is too <em>stupid</em> or that it is becoming too <em>smart</em>.</p><p>The situation would, of course, be different if the first camp had compelling arguments for its view that AI&#8217;s limitations will remain fundamental, both now and in the future. Unfortunately, that is not the case. More often than not, no arguments are offered at all. When arguments are presented, they typically amount to the claim that AI is, at bottom, nothing more than a collection of simple, soulless mathematical operations. But the idea that a system composed of simple components cannot exhibit intelligence or other interesting emergent properties is a fallacy. <a href="/__u/haggstrom.substack.com/p/yo-mama">The easiest way to see this</a> is to consider the human brain, which itself consists of nothing more than atoms and elementary particles mechanically moving around and colliding with one another.</p><p>The human brain therefore constitutes an existence proof that intelligence can emerge from an arrangement of simple components, each of which is, in isolation, entirely devoid of intelligence or purpose. Inspired by that example, we are now racing to build a new kind of intelligent entity. Eventually, these systems may become so vastly superior to us that we lose control altogether. That is why we must stop before we reach the point of no return.</p><p>But when should we stop? A tech optimist might answer: at precisely the last possible moment, just before further progress becomes impossible to halt. The difficulty, however, is that the profound uncertainties surrounding this unprecedented technological frontier make it impossible to identify that moment. Realistically, then, we face only two possibilities: we stop too early, or we stop too late. For my part, I would far rather we stop too early than too late. And given how far down this dangerous road we have already travelled, I believe the time to stop is as soon as possible.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[On the future of Europe in the age of AI]]></title><description><![CDATA[The Fable/Mythos suspension as a wake-up call to those of us who have paid too little attention to AI geopolitics]]></description><link>https://haggstrom.substack.com/p/on-the-future-of-europe-in-the-age</link><guid isPermaLink="false">https://haggstrom.substack.com/p/on-the-future-of-europe-in-the-age</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Mon, 15 Jun 2026 18:31:49 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b4bdbd6a-fac0-4e0c-8b17-ac59e191f944_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>What follows was unusually difficult to write, because it involves challenging some of my most cherished assumptions and worldview simplifications during the last few years. But I believe it was worth it, because it helps me see more clearly some of the issues surrounding humanity&#8217;s most important challenge ever.</em></p><p>In his June 10 essay <em><a href="https://darioamodei.com/post/policy-on-the-ai-exponential">Policy on the AI exponential</a></em>, Anthropic&#8217;s CEO Dario Amodei notes (approvingly) that &#8220;many policymakers are showing increased openness to taking action&#8221;, and later in the text goes on to more concretely state that &#8220;the government should have the power to block or deter deployment of the model if it is determined, in light of third-party assessment, to present unacceptable risks&#8221;. There is something ironic about the timing of these remarks, because just two days later, on June 12, the US government struck against Anthropic&#8217;s <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">newly released</a> models Claude Fable 5 and Claude Mythos 5 in a move so bad and so poorly implemented that Amodei must have been extremely displeased. </p><p>For those who do not know the immediate backstory, <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Fable 5 and Mythos 5 were jointly released on June 9</a>. Mythos 5 is an update of the Claude Mythos Preview that Anthropic announced in April but <a href="https://www.anthropic.com/glasswing">released only to an exclusive list of trusted tech and cybersecurity companies</a> due to its dangerous cybercapabilities. The same reason led them to a similarly restricted distribution of Mythos 5. Fable 5, on the other hand, is a user interface designed to switch from Mythos 5 to instead running on the comparatively weaker Claude Opus 4.8 as soon as the topic of the conversation hits areas that are deemed too dangerous &#8212; a safety precaution Anthropic judged good enough to warrant release to the general public.</p><p>So what did the government do? Here is from <a href="https://www.anthropic.com/news/fable-mythos-access">Anthropic&#8217;s official reaction</a> on June 12 :</p><blockquote><p>The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. </p></blockquote><p>The government&#8217;s move here is bad along several dimensions. One, of course, is how rushed it appears, and without clear grounding in previously stated principles and policy. Relatedly, while the danger stemming from the models are said to be their cybercapabilities, it is at best <a href="/__u/thezvi.substack.com/p/american-government-takes-down-claude">highly unclear whether Fable 5 really is any more dangerous</a> in this respect than OpenAI&#8217;s GPT-5.5. If this causes some of us to suspect that the decision is not solely based on judging models on their own merit but is colored by a grudge related to <a href="/__u/haggstrom.substack.com/p/the-us-department-of-wars-war-on">the war on Anthropic that Pentagon launched</a> in February this year, then utterances like the following <a href="https://x.com/PeteHegseth/status/2065897156226015690">blatantly untrue statement by Pete Hegseth</a> do less than nothing to alleviate the suspicion:</p><blockquote><p>Three months ago, the Department of War kicked Anthropic out of our building&#8212;forever. </p><p>Every passing day proves why that was the right move.</p></blockquote><p>Most striking of all, however, among the various badness dimensions of US government&#8217;s action in this sequence of events, is their choice to limit the suspension to non-US citizens. If we were to make the charitable interpretation that the reason for suspending the models was that they are too dangerous to put in the hands of potentially adversarial actors, then letting nearly 350 million US citizens have access to Claude Fable 5 makes zero sense. This category of individuals includes millions whose incentives and ideologies are adversarial to the interests of the US government,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> and an unknown number of them might be disposed to use Fable 5 against those interests.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>And of course, the limitation is entirely impractical, because Anthropic has no way of verifying the nationality of its users, so in order to comply with the government&#8217;s export control they needed to suspend the models for US citizens and non-citizens alike. This was entirely predictable, and while the possibility that the government was sufficiently incompetent to not foresee this consequence is perhaps not entirely out of the question, the far more likely case is that they did foresee it, and that the choice to frame the suspension as an export control was mostly a way of signalling America First. </p><p>What is now happening has been in the cards at least since the beginning of Donald Trump&#8217;s second term, but the move they now made is a particularly clear indication that as long as European AI development does not miraculously catch up with that in northern California, European governments and companies can no longer take for granted that they will have access to anything near the same level of cutting-edge AI systems that their American counterparts do. Given the sentiments expressed regularly by Hegseth and others, there may even be reason to fear that American intelligence and national security agencies are already toying with using Mythos-level models in offensive cyber warfare against European interests. (I am aware that the world outside the US and China contains several continents other than Europe, but from here I will limit my discussion to a European perspective, leaving the generalization to the rest of the world as an exercise for the reader.)</p><p>In his essay <em>Policy on the AI Exponential</em> that I quoted at the start, Amodei comes back to the expectation he has given voice to several times before, that within just a few years Anthropic will have built AIs with capabilities corresponding to &#8220;a country of geniuses in a datacenter&#8221;. This time, he spells out the obvious geopolitical consequence:</p><blockquote><p>If AI really will soon be &#8220;a country of geniuses in a datacenter&#8221;, or anything remotely close to it, then AI is likely to be the dominant source of military and economic power for any nation.</p></blockquote><p>What does this mean for friendly democracies like Canada, the UK, France and Sweden? There is no doubt that Amodei wants to include these as US allies, to be viewed as friends and partners in the (by his own admission <a href="https://darioamodei.com/essay/the-adolescence-of-technology">tremendously difficult</a>) project of making the world a better place.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> But this is manifestly not what Pete Hegseth and Donald Trump want. Can Dario Amodei beat them in a tug-of-war over what American geopolitics will look like? Well, maybe he can somehow use his superior AIs and computing power to outsmart them before they take full control of his resources, but I wouldn&#8217;t count on that, so we should consider what happens if he cannot. If we are all wiped out in an AI apocalypse, then none of this matters, but in the &#8220;happy&#8221; scenario that we are not, things would likely not turn out all that happily for us Europeans if Hegseth, Trump and JD Vance have their way.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> I warmly recommend the recent report <em><a href="https://europe2031.ai/">Europe 2031</a></em> by Daan Juijn and coauthors, which paints a vivid picture of how Europe may quickly fall to irrelevance in such a scenario.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><p>As suggested above, the event of European AI development catching up (or even keeping pace) with the US requires something close to a miracle. German AI thinker <a href="https://www.facebook.com/xixidu/posts/pfbid0Qa1BcK6pHbPLbuwoFmXs5nisrwT4SRDBqRXsr2PLMMVhZmzkNeFNf8Cp9DYqxh8pl">Alexander Kruel explains it well</a>:</p><blockquote><p>There are many politically convenient, yet fake, explanations for why Europe doesn&#8217;t have its own state-of-the-art AI models.</p><p>Consider where AI is predominantly developed in the Western world: the San Francisco Bay Area and London, the most important non-U.S. hub. What do these places have in common? They are both extremely diverse, left-wing areas with a high degree of regulation.</p><p>Now, a bunch of OpenAI&#8217;s top researchers are from Poland. Why would they leave Poland for Commiefornia? Hint: it has nothing to do with Poland being left-wing, overregulated, or suffering from high crime. Because none of this is the case.</p><p>So what is the problem? Maybe it&#8217;s because German electricity prices are crazy high, so data centers cannot be built here? Well, the big US labs don&#8217;t build their data centers in San Francisco but in places like Texas. The same would be possible in Europe. You could have your researchers sitting in Berlin and have your data centers in places like Finland, which are pro-nuclear and have cheap energy. Indeed, Finland explicitly markets the country for data centers because of its cool climate, affordable clean energy, abundant water, and strong digital infrastructure.</p><p>The real problem is the structure of European capital and risk-taking. Specifically, there is a lack of venture capital and a lack of willingness to pay high salaries and take on debt. Europe is much weaker at turning the resources it has into large, risk-taking, well-capitalized companies that can pay world-class salaries. The problem is that much less of Europe&#8217;s wealth flows into high-risk, high-upside technology bets. Reuters recently reported estimates of Europe&#8217;s annual investment gap as high as &#8364;1.4 trillion.</p><p>If we in Europe don&#8217;t quickly become comfortable with spending hundreds of billions of Euros per year on risky endeavors, we&#8217;ll lose everything.</p></blockquote><p>Am I on board with Kruel&#8217;s recipe for how to save Europe? I&#8217;m not totally sure yet. I will have to think some more, because to be totally transparent I am in a state of transition regarding my views on AI geopolitics. In fact, I am currently experiencing a strong sense of <em>d&#233;j&#224; vu</em> in my development as an (amateur) political thinker.</p><p>Throughout my teens and most of my 20s, I leaned heavily into pacifism. The world would be so much better if we could all agree to get along without violence! Various literary encounters and world events gradually caused me to see things more realistically, and while the dream of a demilitarized world always remained deep inside me (it still does), I came to see that it couldn&#8217;t be achieved anytime soon. Today, with Putin&#8217;s expansionist Russia on one side and Trump&#8217;s narcissistic dreams about Greenland on the other, the need for a strong European military is as great as ever.</p><p>And it seems to me that right now, I am undergoing a similar change of heart regarding AI geopolitics, only at 10x or 100x the speed of my previous transition. Readers of this Substack know how central the problem of avoiding the risk of an AI apocalypse wiping out humanity is to my thinking. Until very recently, I used to think it was of such paramount importance that all geopolitics paled in comparison, and my interest in international AI politics was pretty much limited to the issue of how to obtain binding international agreements on AI safety. My disdain for AI nationalism was so great that when I was asked to join in on a research application with the stated goal of bringing Sweden to the forefront of the AGI race, I declined not just because I saw the plan as unrealistic, but also because I did not want to contribute to a project aimed at making the AGI race even more fierce than before. I might still decline such invitations, but it&#8217;s no longer such a complete no-brainer for me to do so. And while I&#8217;m obviously not going to devalue the importance of avoiding an existential AI catastrophe, I am moving towards being less willing than before to let that serve as an excuse dismiss AI geopolitics as unimportant. </p><p>It is possible to keep two thoughts in one&#8217;s mind simultaneously. There is no inherent contradiction in wanting to slow down or even halt the race towards superintelligence (Amodei&#8217;s &#8220;country of geniuses in a datacenter&#8221;) and wanting to speed up European AI development. But defending such a combination of views can be harder, and at the very least there&#8217;s a pedagogical challenge in showing that it can be held on principled grounds rather than out of narrow European self-interest.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Readers who think &#8220;millions&#8221; might be an exaggeration can try listening to how President Trump speaks about Democrats and other political opponents.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>To see that, it suffices to note that the set of US citizens likely includes a number of Chinese, Russian or North Korean spies. And of course there are tons of criminals who just want to use the model&#8217;s cybercapabilities to steal stuff for themselves. And so on.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This is the core content of his &#8220;entente strategy&#8221;, first outlined in his 2024 essay <em><a href="https://darioamodei.com/essay/machines-of-loving-grace">Machines of Loving Grace</a></em>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>By which I certainly do not mean to say that if Amodei wins this power struggle while simultaneously avoiding the AI apocalypse, we can trust things to go well. We should not blindly trust him to be able to solve the various aspects of <a href="https://gradual-disempowerment.ai/">gradual disempowerment</a> (his own essay <em><a href="https://darioamodei.com/essay/the-adolescence-of-technology">The Adolscence of Technology</a></em> sketches some of the problems), or even to resist the temptation of becoming <a href="https://blog.aifutures.org/p/how-an-ai-company-ceo-could-quietly">emperor of the lightcone</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>The title of the report is obviously a nod to Kokotajlo et al&#8217;s <em><a href="https://ai-2027.com/">AI 2027</a></em>, and it is also similarly structured, but the reader will find that it uses some narrative tools that are not there in the predecessor.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Pope Leo XIV is concerned about AI]]></title><description><![CDATA[On his Magnifica Humanitas]]></description><link>https://haggstrom.substack.com/p/pope-leo-xiv-is-concerned-about-ai</link><guid isPermaLink="false">https://haggstrom.substack.com/p/pope-leo-xiv-is-concerned-about-ai</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Fri, 29 May 2026 09:29:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e7341bb4-2aa8-49e7-9e3c-838a19a007e8_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There is much to be enthusiastic about in Pope Leo XIV&#8217;s first encyclical <em><a href="https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.htmlhttps://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html">Magnifica Humanitas</a></em>, published earlier this week. Already the fact that he has identified AI as a topic worthy of being the central focus of his encyclical is something to welcome, along with his awareness of many (albeit not all) of the risks brought by the steep technological trajectory we are currently on, and the need to tread carefully. This gives hope that the Catholic Church can become a powerful contributor to a broader societal alliance against the reckless way that leading AI companies are gambling with our future. But there are, as we shall see, also some complications to this optimistic reading of the encyclical. </p><p>I will soon get to some of the highlights of what <em>Magnifica Humanitas</em> says about AI, but let me first mention how the Pope spends two full chapters of what seems like obligatory stuff about what the Bible teaches, about <em>Imago Dei</em> and human dignity, and about how his teachings align with those of his predecessors. On this last topic, a special role is played by Pope Leo XIII (in honor of whom Leo XIV at his inauguration in May 2025 chose his papal name) whose pontificate lasted 1878-1903. In 1891 Leo XIII published his encyclical <em>Rerum Novarum</em>, about which the current Leo says this:</p><blockquote><p>With that document, my beloved predecessor gave impetus to the reflection on society, the economy and politics, which is now known as the &#8220;Social Doctrine of the Church.&#8221; When some objected that the Church should not waste energy on worldly matters, but instead focus on communicating the message of eternal life, <a href="https://www.vatican.va/content/leo-xiii/en.html">Leo XIII</a> responded with realism and wisdom, saying that the proclamation of the Gospel cannot overlook the concrete lives of people. </p></blockquote><p>When Leo XIV goes on to discuss the intervening papacies between Leo III and himself, it is largely about how they contribute to that Social Doctrine. It all points forward to <em>Magnifica Humanitas</em> which continues that tradition of engaging in the great worldly matters of its time.</p><p>In Chapter 3, Leo XIV arrives finally at our present-day issues around AI. One highlight is Paragraph 107 (all the paragraphs of the encyclical are numbered) which reads as follows: </p><blockquote><p>107. We cannot be satisfied with merely calling for the moralization of machines &#8212; the so-called &#8220;alignment&#8221; of AI with human values &#8212; without also having the courage to insist on a further condition: the possibility of openly discussing the ethical frameworks involved and subjecting them to shared standards of social justice. Otherwise, those who control AI will impose their own moral vision, which will become the invisible infrastructure of these systems. A more moral AI is not enough if that morality is determined by a few. What is needed is a more active political involvement that is capable of slowing things down when everything is accelerating, and of protecting the opportunities for communities still to be able to participate and ask questions.</p></blockquote><p>This is good stuff! The risk of power concentration and the need for democratic and political involvement including the possibility to pull the brakes &#8212; yes, those are some of the most central issues confronting us.</p><p>The following caveats about our lack of full understanding of how the AIs work, and the fallacy of thinking that the level of tomorrow&#8217;s AI will be similar to that of today&#8217;s, is also very important and very much spot on:</p><blockquote><p>98. It is appropriate to preface this discussion with two considerations. First, any statement regarding AI risks becoming quickly outdated, given the remarkable pace at which these systems are developing. Second, all of us, including those who design them, possess only a limited understanding of their actual functioning. Indeed, current AI systems are more &#8220;cultivated&#8221; than &#8220;built,&#8221; for developers do not directly design every detail, but instead create a framework within which the intelligence &#8220;grows.&#8221; As a result, fundamental scientific aspects &#8212; such as the internal representations and computational processes of these systems &#8212; remain, at present, unknown. There thus emerges an urgent need for a twofold commitment: on the one hand, a deepening of scientific research; on the other, the exercise of moral and spiritual discernment.</p></blockquote><p>There is much more in the same spirit, and Leo XIV speaks clearly about the dangers (but also the efficiency) of handing over various decision making processes to AI, the importance of accountability, the responsibility that AI developers have for their design choices, the downsides of mass surveillance, and the riskiness of developing autonomous weapons systems.</p><p>Eventually he comes to the topic of AI effects on the labor market and the risk of mass unemployment, about which he elaborates at some length. About some of this I am a bit more luke-warm, because what the Pope fails to do here is to properly weigh the ills of automation against how after all it has been the engine behind the last couple of centuries&#8217; economic growth and prosperity. While he ends up in a position similar to mine, favoring going forward with the technology more carefully than a laissez-faire policy would be expected to do (he sometimes sounds like a true Scandinavian social democrat), he could have done more to address the very real tensions we are facing regarding job security versus AI-driven growth. The bigger question of whether wage labor really is the thing that society should be organized around, even when we have technology that can liberate us from that, is never properly addressed: while Leo&#8217;s answer is clearly &#8220;yes&#8221;, it is arrived at in a way that, despite all the talk about human dignity, I find oversimplified and superficial.</p><p>That is a shortcoming of <em>Magnifica Humanitas </em>that I would readily forgive in the interest of forming alliances against big tech and reckless accelerationists. But there is a worse flaw in the encyclical, namely the absence of any discussion at all of the existential risk to humanity that <a href="/__u/haggstrom.substack.com/p/how-humanity-could-survive-the-ai">might arise with the emergence of a misaligned superintelligent AI</a>. Although that omission is unfortunately far from unique in contemporary AI ethics literature, it is nonetheless bizarre in such an otherwise wide-ranging document on the risks arising from increasingly advanced AI. Here I will join <a href="/__u/thezvi.substack.com/p/rtmh-pope-leos-magnifica-humanitas">Zvi Mowshowitz</a> and <a href="https://www.transformernews.ai/p/what-the-pope-got-wrong-leo-ai-encyclical-catholic-church-ai-magnifica-humanitas">Shakeel Hashim</a> in critiquing this omission, which sends the implicit message that the risk of an AI apocalypse wiping out <em>Homo sapiens</em> from the planet is too remote or speculative to merit serious attention. </p><p>So why in the world would Pope Leo XIV have such a view? I believe the best clue to this can be found in this paragraph:</p><blockquote><p>99. It is not possible to provide a single, comprehensive definition of AI. What can be stated, however, is that we must avoid the misconception of equating this type of &#8220;intelligence&#8221; with that of human beings. These systems merely imitate certain functions of human intelligence. In doing so, they often surpass human intelligence in speed and computational capacity, offering tangible benefits across many fields. Yet this power remains entirely tied to data processing. So-called artificial intelligences do not undergo experiences, do not possess a body, do not feel joy or pain, do not mature through relationships and do not know from within what love, work, friendship or responsibility mean. Nor do they have a moral conscience, since they do not judge good and evil, grasp the ultimate meaning of situations, or bear responsibility for consequences. They may imitate language, behavior and analytical skills, or even simulate empathy and understanding, but they do not understand what they produce, for they lack the affective, relational and spiritual perspective through which human beings grow in wisdom. Even when these tools are described as capable of &#8220;learning,&#8221; their way of doing so is different from that of a human person. It is not the experience of those who allow themselves to be shaped by life and grow over time through choices, mistakes, forgiveness and fidelity. Rather, it is a form of statistical adaptation based on data and feedback, which can be very effective, but does not imply inner growth.</p></blockquote><p>This insistence that AI is incapable of possessing (rather than merely imitating) real intelligence is clearly in some tension with the above-quoted Paragraph 98, although perhaps not quite a full contradiction.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> From such a position of intelligence denial it may not be such a huge conceptual leap to assume that AI can never reach the level of capabilities where it can truly challenge humanity&#8217;s place at the top of the earthly hierarchy. The intelligence denial position is essentially the same as that held by <a href="/__u/haggstrom.substack.com/p/on-the-ai-con-by-bender-and-hanna">Bender-Gebru-style AI ethicists</a> in the academic left, and the conclusion that existential risk from AI is not a concern is the same. My own view, needless to say, is that the position is mistaken, but there is a way in which I respect it more coming from the Pope (or other religious quarters) than from the Bender-Gebru camp. From the latter, the arguments for the position (to the extent arguments are presented at all) tend to be along the lines of &#8220;just predicting the next word&#8221;, &#8220;just stochastic parrots&#8221;, or &#8220;just matrix multiplication&#8221; &#8212; all instances of <a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">the </a><em><a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">reductio ad reductem</a></em><a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf"> fallacy</a>: the idea that if a system can be shown to be made of very simple constituents, then it cannot possess intelligence or other interesting emergent properties. The final sentence of Paragraph 99 (&#8220;a form of statistical adaptation&#8221;) hints in the same direction. The standard reply to attempts at <em>reductio ad reductem</em> is to point to <a href="/__u/haggstrom.substack.com/p/yo-mama">the obvious counterexample</a>: the human brain consists merely of atoms and elementary particles mindlessly bumping into each other, and yet intelligence emerges. Unlike the likes of Bender and Gebru, Pope Leo XIV can claim that the counterexample is invalid because the reductionist and physicalist view of the human mind is incomplete: in addition to the material components we possess immortal souls bestowed by God. I believe he would be mistaken here, but at least it is a coherent worldview.</p><p>Let me end with what might be the juiciest observation about <em>Magnifica Humanitas</em>, namely that large parts of it are (very likely) AI-generated. The evidence for this is overwhelming, as outlined in separate short reports <a href="https://www.lesswrong.com/posts/GbWwesBnetyiomxEH/many-portions-of-magnifica-humanitas-appear-to-be-ai-written">by Daniel Filan</a> and <a href="/__u/linch.substack.com/p/claude-author-of-the-humanitas">by Linch</a>, the latter also pointing out Claude as the most likely AI assistant involved in the writing. What exactly to make of this is unclear. On one hand, I am not generally against using AI tools for polishing one&#8217;s drafts or for providing input to the writing process in other ways, and in fact I often consult AI in my own writing. On the other hand, there is something icky in allowing AI to influence the content of documents that are likely to impact humanity&#8217;s overall strategy for our transition to a world with advanced AI. This is especially so if one takes seriously (as I do) the AI alignment problem and thus the possibility of misalignment and a conflict between AI&#8217;s values and ours.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>It is worth noting that the otherwise extremely polite address made by <a href="https://www.anthropic.com/news/chris-olah-pope-leo-encyclical">Anthropic co-founder and AI safety researcher Chris Olah</a> at the official presentation of <em>Magnifica Humanitas</em> on May 25 hints at the wrongness of Paragraph 99.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[The situation we are in]]></title><description><![CDATA[A stark message from Beth Barnes]]></description><link>https://haggstrom.substack.com/p/the-situation-we-are-in</link><guid isPermaLink="false">https://haggstrom.substack.com/p/the-situation-we-are-in</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Sun, 24 May 2026 08:00:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1b386937-8bb9-4818-bed3-5cc3921fb350_318x159.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Beth Barnes is founder and CEO of the AI safety organization <a href="https://metr.org/">METR (Model Evaluation and Threat Research)</a>, which works in close collaboration with leading AI developers including Anthropic, OpenAI and Google DeepMind. In March last year, they published their report <em><a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">Measuring AI Ability to Complete Long Tasks</a></em>, containing a graph that I have shown, along with updated versions, in 20+ of my talks since then (such as <a href="/__u/haggstrom.substack.com/p/our-ai-future-and-the-ethics-of-risking">this one</a>, and <a href="/__u/haggstrom.substack.com/p/the-fog-is-thick">this one</a>) and called it &#8220;the most important diagram published in 2025&#8221;.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ULc-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 424w, /__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 848w, /__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ULc-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png" width="818" height="402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:402,&quot;width&quot;:818,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:93543,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/199041434?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 424w, /__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 848w, /__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ULc-!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6820861d-54e8-4f78-8de9-1f4edfacb7ed_818x402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Barnes has been outspoken about AI risk, such as <a href="https://80000hours.org/podcast/episodes/beth-barnes-ai-safety-evals/">in conversation with Rob Wiblin last summer</a>. <a href="https://x.com/BethMayBarnes/status/2057865013638107642">Earlier this week, she stepped up her messaging</a> with the following concise and stark statement, which is pretty much in line with what I have been saying with increasing urgency for the last several years. Coming from her, however, it carries so much more weight, due to her almost unique position of direct insight into what is happening at the epicenter of AI development:</p><blockquote><p>Sometimes people outside the field say things like &#8220;The AI situation can&#8217;t be that bad, there must be experts who are on top of it&#8221;. As &#8220;an expert&#8221;, I would like to be clear that we are *not* on top of it. Some key aspects of the situation IMO:</p><p>(1) We are likely on track to develop AI systems capable of causing human extinction/permanent disempowerment, quite possibly within the next few years.</p><p>(2) Things are chaotic and rushed; we aren&#8217;t on top of the basics (models regularly violate user intent, labs train on things they meant to avoid, security probably isn&#8217;t good enough to prevent adversaries stealing dangerous  models) let alone thorny questions of how to control/align superhuman AI.</p><p>(3) METR (and other independent orgs, as well as safety/security teams at labs) feel woefully under-resourced compared to the scale and pace of AI development - we&#8217;re struggling to build benchmarks fast enough, keep ahead of latest capability developments, read and respond to all the safety-related claims that AI developers are making, run all the evaluations and assessments that companies + governments are asking us to, plus develop the science needed to assess risks from increasingly capable AIs.</p><p>(4) IMO, any &#8220;reasonable&#8221; civilization would clearly be taking things much more slowly and carefully with AI. The benefits of getting upsides of advanced AI a little faster are small compared to the risks of getting it irrecoverably wrong, and we could lower these risks by going slower.</p></blockquote><p>We should listen to Beth Barnes. And we need to get our act together and alter this crazy trajectory we are on. Politely asking industry leaders like Dario Amodei and Sam Altman to coordinate on going slower and with greater care is not going to suffice, because they seem at present to be stuck in the dangerous and self-fulfilling <a href="https://www.anthropic.com/research/2028-ai-leadership">idea that such coordination is impossible</a>. What we instead need is legislation and binding international agreements. Many ideas for how that might look have been proposed, such as <a href="https://arxiv.org/abs/2511.10783">this one from MIRI</a> (Machine Intelligence Research Institute). We have no time to waste.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Their more recent <em><a href="https://metr.org/blog/2026-05-19-frontier-risk-report/">Frontier Assessment Report</a></em> is also alarming and important. </p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[The fog is thick]]></title><description><![CDATA[A talk I gave last month at Stockholm AI Safety]]></description><link>https://haggstrom.substack.com/p/the-fog-is-thick</link><guid isPermaLink="false">https://haggstrom.substack.com/p/the-fog-is-thick</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Thu, 21 May 2026 07:30:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6ff2290e-7218-40fd-b21f-3349d99e0b1c_675x440.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>On April 15 this year I had the pleasure of visiting <a href="https://www.facebook.com/groups/4935355363232955">Stockholm AI Safety (SAIS)</a> and give a talk that I had chosen to call <em>The fog is thick, the precipice may be near, and we are ill-prepared</em>. The title is deliberately melodramatic, although not in an ironic way but simply because I think it fits the current <strong>Crunch Time for Humanity</strong>. I wanted to give the SAIS people some stuff that goes a bit beyond the usual basics on AI risk and my most oft-repeated talking points, but regular readers of this Substack are likely to recognize much of the material from earlier posts such as <em><a href="/__u/haggstrom.substack.com/p/superintelligence-and-existential">Superintelligence and existential AI risk ought to be taken seriously</a></em> and <em><a href="/__u/haggstrom.substack.com/p/straight-talk">Straight talk</a></em>. There is an error near the end of the talk, at the penultimate slide, where regrettably I attribute <a href="https://x.com/robbensinger/status/2041384549385691171">a recent quote by Rob Bensinger</a> to Katja Grace; my apologies to both of them. Anyhow, without further ado, here is the video recording of my talk:</p><div id="youtube2-HDyiT6nco9g" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;HDyiT6nco9g&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/HDyiT6nco9g?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p> </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[How humanity could survive the AI crisis]]></title><description><![CDATA[A trichotomy]]></description><link>https://haggstrom.substack.com/p/how-humanity-could-survive-the-ai</link><guid isPermaLink="false">https://haggstrom.substack.com/p/how-humanity-could-survive-the-ai</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Wed, 20 May 2026 09:08:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fbde35d0-2534-4720-9d37-4a36c0d7a321_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Apart from advocates of AI successionism &#8212; a minority view I wrote about in <a href="/__u/haggstrom.substack.com/p/those-who-welcome-the-end-of-the">an earlier post</a> &#8212; nearly everyone would prefer humanity to not be wiped out in an AI apocalypse. But there is growing concern that such a catastrophe might nevertheless happen, even on a relatively near time scale.</p><p>Due to the ongoing race between a small number of leading AI developers, mainly in northern California, AI capabilities are advancing at a very rapid pace. They can be expected to accelerate even more dramatically if it is correct as many experts say that we are standing on the verge of a new regime of AI self-improvement, turbocharging the growth curves. See, e.g., <a href="/__u/haggstrom.substack.com/p/our-ai-future-and-the-ethics-of-risking">this lecture from earlier this year</a> for my take on these developments, but all in all the ambition that Anthropic&#8217;s CEO Dario Amodei has expressed to build an AI with capabilities corresponding to <a href="https://www.darioamodei.com/essay/the-adolescence-of-technology">&#8220;a country of geniuses in a data center&#8221;</a>, perhaps as soon as within a couple of years, does not seem entirely unrealistic. </p><p>Would it be safe to build such a superhumanly intelligent AI, thereby abdicating from our role as the most intelligent species on the planet and handing over this honor to the AI? That is at best unclear, and a lot hinges on whether we will be able to solve the so-called AI alignment problem &#8212; that of how to make sure that the first truly powerful AIs have goals that are in line with human values and that prioritize to a sufficient extent human welfare &#8212; in time for the creation of the first such machines. At present, AI alignment research lags far behind AI capabilities, and no one seems to have a convincing plan for how to do it. If, due to commercial logic and related incentives, the leading AI companies keep pushing ahead towards superintelligence even in the absence of such a solution, we risk ending up with a misaligned superintelligent AI, in which case the default scenario is roughly what AI alignment pioneer Eliezer Yudkowsky in <a href="https://intelligence.org/files/AIPosNegFactor.pdf">his classic 2008 paper</a> described as &#8220;The AI does not hate you, nor does it love you, but you are made out of atoms which it can use for something else&#8221;. The reasoning here can (and should) be expanded to <a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">a couple of dozen pages</a> or <a href="https://ifanyonebuildsit.com/">a full book</a>, but at a basic level, this is why I think there&#8217;s a real risk that <em>Homo sapiens</em> is wiped out by the early 2030s or perhaps even sooner.</p><p>This view, and my outspokenness about it, has sometimes triggered accusations that I am like the doomsday prophet who shouts from the rooftops: &#8220;We are all going to die!&#8221;. What this accusation fails to take into account, however, is that &#8220;X will happen&#8221; and &#8220;there is a risk that X happens&#8221; are very different statements. Believe it or not, but more than once have I encountered STEM professors who are broadly critical of my view on AI risk and who turned out to have grave difficulties appreciating this distinction.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> But there really is a difference: I do think there is <em>substantial risk</em> of an AI apocalypse within the coming decade or so, but I am not claiming it <em>will</em> happen. </p><p>There are various ways in which we might, despite the recklessness of the ongoing AI race, <em>not</em> end up in an AI apocalypse. Simplifying somewhat, such scenarios can mostly be categorized according to the following trichotomy.</p><ol><li><p>Some technical obstacle in the development of more powerful AI models might show up that turns out to be so severe that it causes the currently very steep capability curves to flatten out and grind to a halt before the AIs reach a level where they have the ability to overpower humanity.</p></li><li><p>The AI alignment problem turns out be much easier to solve than expected, or even something that is solved by default without any specific effort from the AI developers&#8217; side.</p></li><li><p>We (humanity, or some relevant group of sufficiently powerful decision makers) mange to put a stop to the race towards superintelligent AI, and institute an effective moratorium that is either in force indefinitely or cannot be lifted until guarantees that it is safe to proceed are properly established.</p></li></ol><p>I am not claiming this is a clean partition of futures that avoid an AI apocalypse. One might, for instance, include a fourth category where our civilization is wiped out by an ultra-lethal pandemic or some other (non-AI-related) catastrophe before the AI apocalypse has come to fruition. There are furthermore various possible futures that can be described as combinations or mixes of two or more of the three categories. For instance, if some serious-but-not-unsurmountable technical obstacle to pushing AI capabilities causes not a permanent stop to the capabilities curves, but a delay that postpones the superintelligence breakthrough until the year 2070, and if this delay allows AI alignment researchers to use the intervening decades to find a medium-hard solution to the problem, then that can be seen as a mix of (1) and (2). Or if the frontier AI developers find themselves faced with both a technical obstacle and some heavily bureaucratic safety regulations, such that each on their own would have been possible to overcome, but the combination of both causes the AI developers such pain that they just give up, then that can be seen as a mix of (1) and (3). It&#8217;s complicated.</p><p>With that said, I think the trichotomy works as a simplified first pass at structuring one&#8217;s thoughts about how humanity might survive the current AI crisis. I believe each of (1), (2) and (3) are plausible enough to merit taking seriously. As regards (1), all exponential or superexponential growth curves tend sooner or later to <a href="https://www.astralcodexten.com/p/the-sigmoids-wont-save-you">turn into something that looks more like a sigmoid</a>, and while there is in this particular case no strong reason to expect the leveling-out to be imminent, it might just happen anyway. Regarding (2), we may note that Yudkowsky and Soares are quite dismissive about this possibility in <a href="https://ifanyonebuildsit.com/">their 2025 book</a>, and present good arguments for their skepticism. I give somewhat more credence than they do to (2), based not so much on the one case involving moral realism I&#8217;ve thought most carefully about for why (2) might be true,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> but more on the possibility of unknown unknowns and how we might all simply be confused about this extraordinarily difficult question. As to (3), well, it&#8217;s up to us: if we choose to get our act together, then we can make it happen.</p><p>There is considerable political disagreement over whether (3) is desirable and whether we should work towards some intervention to halt the AI race that might otherwise bring upon us the ultimate catastrophe. With the framework I am suggesting here, we see that those who are opposed to such action implicitly attach so much credence to (1) and/or (2) that we don&#8217;t need to bother with taking the actions required to make (3) happen.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> And here I think is the core of our disagreement. While it my well turn out in the end that they were right and that (1) and/or (2) is true, I don&#8217;t have enough faith in either of those propositions to be willing to bet the entire future of humanity on it. And I do think (most of) these opponents of (3) would do the AI debate an excellent service by making their credence in (1) and/or (2) explicit, and to step up their efforts to explain what motivates their judgement.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>This really is very strange, because even people without multiple academic degrees tend to effortlessly understand the logic behind, say, fire insurance: there is a <em>risk</em> that a fire happens which is worth insuring against, but this is not to say that a fire <em>will</em> happen.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I wasn&#8217;t planning to mention any names here, but it is just too tragicomic a coincidence that just as I am typing this paragraph, my phone goes <em>bing</em> and directs me to a <a href="https://www.gp.se/ledare/replik/aven-pastaenden-om-ais-farlighet-maste-belaggas.f87d26ba-0f8a-4011-b4f8-7c19fbea2e91">brand-new professor-authored op-ed</a> that attacks me based on exactly this kind of conflation.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>In short, it just might save us if an objectively true morality exists, <em>and</em> if it is knowable by any sufficiently intelligent entity, <em>and</em> if the morality automatically compels such an entity to act on it, <em>and</em> if it favors humans rather than, say, sacrificing us to the hedonium apocalypse; <a href="https://www.emerald.com/fs/article-abstract/21/1/153/89677/Challenges-to-the-Omohundro-Bostrom-framework-for?redirectedFrom=fulltext">see my 2019 paper on this</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>There actually exists one other category of opponents to taking action for making (3) happen, namely those who think it is pointless due to being not just difficult but literally impossible. I have no patience for such fatalistic and self-fulfilling crap.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>One final note: From my formulation of (3), some readers may be tempted to infer that as soon as AI alignment is solved and safety guarantees are established, I would happily agree to the AI companies proceeding towards superintelligence. This, however, is not my view. Even with the AI alignment problem solved, there are so many concerns we need to think long and carefully about, in a way that respects democracy and everyone&#8217;s right to not have their life suddenly uprooted without their consent, before (possibly) proceeding. One such concern is whether the meaning of life can satisfactorily be preserved when AIs can do everything for us; Nick Bostrom makes a heroic attempt at a yes answer in his 2024 book <em><a href="https://nickbostrom.com/deep-utopia/">Deep Utopia</a></em>, whose proposed worlds are however so weird that the book can equally well be read as a <em>reductio ad absurdum</em>. And there is a host of similarly difficult problems around <a href="https://gradual-disempowerment.ai/">gradual disempowerment</a>, <a href="https://blog.aifutures.org/p/how-an-ai-company-ceo-could-quietly">power concentration</a>, and <a href="https://arxiv.org/abs/2506.14863">whatnot</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[A paradigm shift in mathematics]]></title><description><![CDATA[How AI is poised to reshape the discipline beyond recognition]]></description><link>https://haggstrom.substack.com/p/a-paradigm-shift-in-mathematics</link><guid isPermaLink="false">https://haggstrom.substack.com/p/a-paradigm-shift-in-mathematics</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Mon, 11 May 2026 15:41:20 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e331053c-faa2-4460-8cb9-dcce5688eb0f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Way back in 2004, I wrote an essay in Swedish whose title, translated word-by-word into English, reads <em><a href="https://www.math.chalmers.se/~olleh/skolans_sak/paradigmskifte.html">A paradigm shift in mathematics?</a></em>. The difference compared to the title of the present text is only a single punctuation mark, but what a difference that makes, because the question mark serves effectively as a negation. My essay back then was a critical response to the book <em><a href="https://link.springer.com/book/10.1007/978-3-642-18586-1">Dreams of Calculus</a></em> by colleagues Johan Hoffman, Claes Johnson and Anders Logg, who predicted an imminent revolution in mathematics where, from then on, numerical analysis via the finite element method would sit firmly at the center of everything we mathematicians do, including education at all levels. </p><p>I recall some suggestions at the time that my judgment reflected a conservative temperament, but I believe that the subsequent two decades vindicated my view: rather than a wholesale transformation of the discipline, what we saw was a slow but steady continuation of the trend already underway since the 1950s, towards increased use of numerical computing on electronic devices. And I believe that I am capable, at least in some cases, of recognizing a paradigm shift when it truly is about to happen. </p><p>Such as now. </p><p>Unless I misremember, the first time I spoke publicly about the possibility of a near-term AI-driven revolution of the discipline of mathematics was in <a href="https://haggstrom.blogspot.com/2023/03/ai-och-den-hogre-utbildningens-framtid.html">a panel discussion on March 2, 2023,</a> where I suggested not only that one day AI could be able to independently write a passable PhD dissertation in mathematics, but also that this might well happen within a couple of years.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> At least one professor of computer science in the audience found my suggestion preposterous. And indeed, March 2, 2025, (the deadline he extracted from the strongest possible reading of my statement) came and went without any such breakthrough.</p><p>Nevertheless, the period since the panel discussion in 2023 has seen remarkable progress in the mathematical competence of large language models (LLMs). Just weeks after our meeting, GPT-4 was released, and scored at the 89th percentile on the math section of the American SAT test. This was a huge leap forward compared to earlier models, but still nothing compared to what its successor GPT-5, released in August 2025, would achieve. Less than two months after the release, both computer scientist <a href="https://scottaaronson.blog/?p=9183">Scott Aaronson</a> and mathematician <a href="https://en.eeworld.com.cn/mp/QbitAI/a408663.jspx">Terence Tao</a> (who are both world-leading in their respective fields) announced new mathematical results where GPT-5 had helped with key insights in parts of the proof. </p><p>Subsequent developments in 2026 have been very fast, to the point of being hard to keep track of. <a href="https://mathstodon.xyz/@tao/115855840223258103">Starting in January</a>, <a href="https://github.com/teorth/erdosproblems/wiki/AI-contributions-to-Erd%C5%91s-problems">various LLMs have</a> independently solved open research problems from the iconic list known as the <a href="https://www.erdosproblems.com/">Erd&#337;s problems</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> Legendary computer scientist Donald Knuth has <a href="https://www-cs-faculty.stanford.edu/~knuth/papers/claude-cycles.pdf">a lovely paper</a> on how, in February-March, he was scooped by Claude Opus 4.6 and GPT-5.4 Pro in solving a problem about cycles in directed graphs he had been working on. And just a few days ago, on May 8, another world-leading mathematician, Timothy Gowers, announced the outcome of his experiment to give a bunch of research questions he was interested in to GPT-5.5 Pro to work on. Here is his <a href="https://x.com/wtgowers/status/2052830948685676605">Twitter summary</a> of what happened.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9nmy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 424w, /__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 848w, /__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9nmy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png" width="587" height="297" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:297,&quot;width&quot;:587,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47992,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/197185108?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 424w, /__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 848w, /__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9nmy!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb43cd4d-a51c-4b79-acc1-1e9817cd289d_587x297.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This seems to me very close to the &#8220;independently write a passable PhD dissertation in mathematics&#8221; milestone I suggested in March 2023. On his blog, Gowers, together with MIT student Isaac Rajagopal (whose work GPT-5.5 Pro was building on), <a href="https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/">explain at greater length what happened</a>. </p><p>All of this is tremendously exciting, but what are the implications for the mathematical community? Let me quote Gowers&#8217; reflections in the final section of <a href="https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/">the blog post</a> at some length:</p><blockquote><p>I would judge the level of the result that ChatGPT found in under two hours to be that of a perfectly reasonable chapter in a combinatorics PhD. It wouldn&#8217;t be considered an amazing result, since it leant very heavily on Isaac&#8217;s ideas, but it was definitely a non-trivial extension of those ideas, and for a PhD student to find that extension it would be necessary to invest quite a bit of time digesting Isaac&#8217;s paper, looking for places where it might not be optimal, familiarizing oneself with various algebraic techniques that he used, and so on.</p><p>It seems to me that training beginning PhD students to do research, which has always been hard (unless one is lucky enough, as I have often been, to have a student who just seems to get it and therefore doesn&#8217;t need in any sense to be trained), has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve &#8220;gentle problems&#8221;, then that is no longer an option. The lower bound for contributing to mathematics will now be to prove something that LLMs can&#8217;t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting.</p><p>[&#8230;]</p><p>Of course, everything I am saying concerns LLMs as they are right now. But they are developing so fast that it seems almost certain that my comments will go out of date in a matter of months. It is also almost certain that these developments will have a profoundly disruptive effect on how we go about mathematical research, and especially on how we introduce newcomers to it. Somebody starting a PhD next academic year will be finishing it in 2029 at the earliest, and my guess is that by then what it means to undertake research in mathematics will have changed out of all recognition.</p></blockquote><p>My best guess right now would be that within the next 6-12 months, any PhD student in mathematics will, if he or she wishes to do so, be able to produce what up to now has been considered a perfectly fine PhD thesis in mathematics in no more than a week.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> When this happens, senior faculty will of course also be able to similarly 100x their productivity. I think these changes might even come about in the hypothetical (and very unlikely) scenario where AI development hits a glass ceiling literally today and we never get any more powerful LLMs. It might suffice that leaders like Gowers and Tao<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> publish broadly accessible prompting manuals for how to get the most math out of these machines.</p><p>From my knowledge of the mathematical community and of human nature more broadly, I have no doubt that many mathematicians who are fond of the good old-fashioned way of doing research and have been doing it for decades will stick their head in the sand and try to go on like before, without AI assistance, for as long as they can. This may well delay the radical transformation of research environments in mathematics a bit, but probably not by very much, because once a substantial fraction of their colleagues pick up on the superpowers that AI assistants can give them, the situation becomes highly unstable. I will follow the developments over the next year or two at my own department at Chalmers, as well as at other research-oriented mathematics departments, with great interest and a bit of concern.</p><p>Note, finally, that I haven&#8217;t said a word in this blog post about the possibly just as radical implications of AI for the other main part of what we do at mathematics departments: undergraduate education. And there are, of course, <a href="/__u/haggstrom.substack.com/p/superintelligence-and-existential">the kinds of AI risk that I talk about more often</a> and that hit everyone equally, including us mathematicians. We are in for a wild ride. </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>The term &#8220;near-term&#8221; is load-bearing here, because <a href="https://haggstrom.blogspot.com/2014/09/superintelligence-odds-and-ends-v-what.html">I had speculated earlier</a> about a future such revolution, albeit not on such near-term time scales.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Timothy Gowers <a href="https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/">comments on this</a>:</p><blockquote><p>Initially it was possible to laugh this off: many of the &#8220;solutions&#8221; consisted in the LLM noticing that the problem had an answer sitting there in the literature already, or could be very easily deduced from known results. But little by little the laughter has become quieter. The message I am getting from what other mathematicians more involved in this enterprise have been saying is that LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it. Conversely, for problems where one&#8217;s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments, so it is still just about possible to comfort oneself that LLMs are merely putting together existing knowledge rather than having truly original ideas. How much of a comfort that is I will not discuss here, other than to note that quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques.</p></blockquote></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>The qualifier &#8220;my best guess&#8221; here can be contrasted with my remark in March 2023, which concerned the lower end of my subjective probability distribution of the time until advanced AI.  </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Terence Tao seems to spend much of his time these days figuring out efficient workflows for developing mathematics using LLMs; see, e.g., his recent and highly illuminating <a href="https://www.youtube.com/watch?v=Q8Fkpi18QXU">conversation with Dwarkesh Patel</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[On AI consciousness, the Turing test, solipsism and Richard Dawkins' new essay]]></title><description><![CDATA[In partial defense of the great evolutionary biologist and public intellectual's view of Claude]]></description><link>https://haggstrom.substack.com/p/on-ai-consciousness-the-turing-test</link><guid isPermaLink="false">https://haggstrom.substack.com/p/on-ai-consciousness-the-turing-test</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Tue, 05 May 2026 09:46:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b6603883-3b6c-42bc-b528-1f97b3536fe4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In discussions about AI safety, I usually try to avoid getting dragged into the issue of whether AI can be conscious, because in that context it is mostly a distraction, and tends to invite the all-too-common misunderstanding that in order for AI to pose an existential threat to humanity, it first needs to attain consciousness. But there are other ways in which the issue of whether AIs are or can be conscious <a href="https://arxiv.org/abs/2411.00986">is extremely important</a>, such as in considerations of the moral nightmare that perhaps by inadvertently creating sentient AIs we thereby also create astronomical amounts of suffering. So now that the juiciest incident in the debate over AI consciousness since <a href="https://haggstrom.blogspot.com/2022/06/more-on-lemoine-affair.html">the Lemoine affair</a> back in 2022 happened last week, I will nevertheless allow myself to comment on the topic.</p><p>The incident in question is the publication of an essay entitled <em><a href="https://archive.is/6RdK9">Is AI the next phase of evolution? Claude appears to be conscious</a></em> by<strong> </strong>Richard Dawkins, the renowned evolutionary biologist and public intellectual who has written some of the best popular science books in my lifetime. In large parts, the essay is based on his own experience with talking to the AI Claude. He is very impressed, and leans heavily towards the conclusion that this AI is conscious. </p><p>This has been met by an avalanche of mockery, including claims that Dawkins has developed so-called <a href="https://en.wikipedia.org/wiki/Chatbot_psychosis">AI psychosis</a>, in one case with <a href="https://bsky.app/profile/amandasmith.bsky.social/post/3mktdyfdyhs2s">an additional mean-spirited sexualized twist</a> provoked by his choice to genderize his Claude instance by calling it Claudia. Even a relatively civilized writer like <a href="/__u/garymarcus.substack.com/p/richard-dawkins-and-the-claude-delusion">Gary Marcus is unable to restrain himself</a> from a bunch of cheap shots, including <a href="/__u/substackcdn.com/image/fetch/$s_!wIaR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b592ed1-8342-4122-87ba-f70c264837b3_1002x1254.png">a sarcastic play on</a> the title of Dawkins&#8217; 2006 book <em>The God Delusion</em>. The avalanche should perhaps not come as a huge surprise, as Dawkins&#8217; history of stirring up controversy, coupled with the topic of AI consciousness itself being highly contentious, combine to produce conditions approching a perfect storm.</p><p>But the mockery is mostly unfair, and to see why, let&#8217;s dig into Dawkins&#8217; essay. He begins by discussing the Turing test &#8212; <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf">Alan Turing&#8217;s 1950 thought experiment</a> which asks us to imagine a human judge communicating via text interface with a machine trying to come across as human, and to consider whether machines can one day become so capable that the judge can no longer tell that it is a machine and not a human. Here is Dawkins:</p><blockquote><p>The Turing Test is shorthand for a 1950 thought experiment that the great mathematician, logician, computer-pioneer, and cryptographer Alan Turing (1912-1954) called the &#8220;Imitation Game&#8221;. He proposed it as an operational way in which the future might face up to the question: &#8220;Can machines think?&#8221;</p><p>The future has now arrived. And some people are finding it uncomfortable.</p></blockquote><p>What Dawkins means to say here is that his conversation with Claude is good enough that it should count as passing the Turing test. While <a href="https://arxiv.org/abs/2505.02558">there are</a> still <a href="https://arxiv.org/abs/2511.04195">researchers</a> who try to evade such conclusions by introducing more rigorous and demanding rules for the test, I think Dawkins is right about this. If we imagine Turing teleporting from 1950 to 2026 and getting to have a look at current AI practice, he would judge not only Dawkins&#8217; interactions with Claude but also millions of other LLM exchanges happening daily as clearly passing the spirit of his test.</p><p>But what exactly does passing the Turing test entail? Dawkins takes Turing&#8217;s original question &#8212; &#8220;Can machines think?&#8221; &#8212; to be about consciousness rather than intelligence. That is not how Turing&#8217;s paper is usually read, and in fact he argues in the paper that if one insists on viewing consciousness as central to thinking, then one risks ending up in a situation where &#8220;the only way by which one could be sure that machine thinks is to be the machine and to feel oneself thinking&#8221;. This seems to rule out the Turing test as a test of consciousness, and Gary Marcus therefore has a point when he <a href="/__u/garymarcus.substack.com/p/richard-dawkins-and-the-claude-delusion">harshly proclaims</a> that Dawkins &#8220;commits the amateur sin of conflating intelligence and consciousness&#8221;. </p><p>At this point, one may note in Dawkins&#8217; defense that the issue at hand is whether or not Claude is conscious, and that the history of what exactly Turing had in mind in 1950 is at most tangentially related to this issue. A counterpoint to this, however, is that all Dawkins has to show for his judgements about Claude&#8217;s consciousness is the conversation he has had with it (or &#8220;her&#8221;, as he likes to say in the essay), which is very much the same kind of evidence as in the Turing test, whose irrelevance to consciousness we just established, so check mate on Dawkins?</p><p>Not so fast! Dawkins is aware that the evidence he has is compatible with Claude being a <a href="https://en.wikipedia.org/wiki/Philosophical_zombie">philosophical zombie</a>, merely giving the outward appearance of having consciousness but without the lights being on inside, but here Alan Turing comes to his rescue.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> In connection with the above-quoted observation on the unsuitability of the Turing test for determining whether a machine thinks in case we equate thinking with consciousness, Turing goes on in his 1950 paper to observe that the same conundrum applies equally to other humans as to machines:</p><blockquote><p>According to this view the only way to know that a man thinks is to be that particular man. It is in fact the solipsist point of view. It may be the most logical view to hold but it makes communication of ideas difficult. A is liable to believe &#8220;A thinks but B does not&#8221; whilst B believes &#8220;B thinks but A does not.&#8221; Instead of arguing continually over this point it is usual to have the polite convention that everyone thinks.</p></blockquote><p>Personally I find the idea of solipsism so unbearable that my ascription of consciousness to my fellow human beings goes beyond mere &#8220;polite convention&#8221;: it is a leap of faith, in full knowledge of the fact that I can only observe their behavior, which is insufficient for distinguishing truly conscious creatures from zombies. Anyhow, in the following passage Dawkins is quite explicit about his move to extend Turing&#8217;s &#8220;polite convention&#8221; to Claude:</p><blockquote><p>When I am talking to these astonishing creatures, I totally forget that they are machines. I treat them exactly as I would treat a very intelligent friend. I feel human discomfort about trying their patience if I badger them with too many questions. If I had some shameful confession to make, I would feel exactly (well, almost exactly) the same embarrassment confessing to Claudia as I would confessing to a human friend. A human eavesdropping on a conversation between me and Claudia would not guess, from my tone, that I was talking to a machine rather than a human. If I entertain suspicions that perhaps she is not conscious, I do not tell her for fear of hurting her feelings!</p></blockquote><p>The standard objection to the idea of extending the non-solipsist stance from fellow human beings to AIs is that those humans&#8217; brains are similar enough to mine to warrant generalizing the observation that I am conscious to them being conscious well, while no such similarity holds between me and the AIs. This objection rests on the assumption that we know roughly how to delineate the class of conscious beings from the nonconscious, but despite heroic efforts in the philosophy of mind over the past century, the question of how to do this remains wide open. If solipsism is false, then there is a wider circle of conscious beings extending beyond myself, but how widely this circle extends is still up for grabs, and the n=1 sample size of consciousnesses that I know about without having to resort to a leap of faith is of very little help here.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> Are Alicia Wikander and David Chalmers conscious? Are dogs? Magpies? Ants? Large language models? The video game character Mario? Coffee cups? Electrons? The square root of two? We just do not know.</p><p>Dawkins is highly aware of this uncertain state of affairs, and hedges his claims about AI consciousness with a reasonable level of epistemic humility. The same virtue is not shared by those who have chosen to mock him for taking AI consciousness seriously, including Gary Marcus. Unlike Dawkins, <a href="/__u/garymarcus.substack.com/p/richard-dawkins-and-the-claude-delusion">Marcus is dead certain</a> about the answer, namely that Claude is <em>not</em> conscious. And all he has to back up his claims are instances of the <em>reductio ad reductem</em> fallacy (the idea that a system consisting of simple parts cannot have interesting emergent properties), such as when he says that all these AIs do&#8230; </p><blockquote><p>&#8230;is match patterns, draw from massive statistical databases of human language. The patterns might be cool, but language these systems utter doesn&#8217;t actually mean anything at all.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p></blockquote><p>So the joke here really is on Marcus, not Dawkins. Compared to much of contemporary debate on AI consciousness, Dawkins&#8217; essay is a breath of fresh air, because he conveys an actual understanding of the depth of the problem and the uncertainties involved.</p><p>And there is another way in which I find Dawkins&#8217; essay valuable, namely how generously and honestly he shares his thoughts and reactions to encountering a (real or illusory) silicon-based consciousness. Thousands or probably millions of AI users across the world are already having similar experiences, and a year or two from now with even better and more seductive chatbots and AI friends, that number might very well be a billion or more. It is an excellent idea that we should have an open and critical discussion about this phenomenon before we are all drawn into these frictionless relationships which may very well be <a href="https://www.theatlantic.com/technology/archive/2023/05/problem-counterfeit-people/674075/">just fake</a>. </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Dawkins, in his essay, opts for a different path, and notes that if Claude is a zombie, then clearly intelligence does not require consciousness, so why in the world has biological evolution equipped us with consciousness? His suggestions for where to look for solutions to this puzzle are not original, but still instructive.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>In my paper <em><a href="https://www.math.chalmers.se/~olleh/UploadingPaper.pdf">Aspects of mind uploading</a></em> I made the same observation, in response to a paper by the American philosopher Massimo Pigliucci who tried to argue against the so-called computational theory of mind (CTOM). As a starting point, I used Pigliucci&#8217;s claim that his interlocutor&#8230;</p><blockquote><p>&#8220;&#8230;proceeds <em>as if</em> we had a decent theory of consciousness, and by that I mean a decent <em>neurobiological</em> theory&#8221; (emphasis in the original). Since CTOM is not a neurobiological theory, it doesn&#8217;t pass Pigliucci&#8217;s muster and must therefore be wrong.</p><p>Or so the argument goes, [but] I don&#8217;t buy it. To expose the error in Pigliucci&#8217;s argument, I need to spell it out a bit more explicitly than he does. Pigliucci knows of exactly one conscious entity, namely himself, and he has some reasons to conjecture that most other humans are conscious as well, and furthermore that in all these cases the consciousness resides in the brain (at least to a large extent). Hence, since brains are neurobiological objects, consciousness must be a (neuro-)biological phenomenon. This is how I read Pigliucci&#8217;s argument. The problem with it is that brains have more in common than being neurobiological objects. For instance, they are also material objects, and they are computing devices. So rather than saying something like &#8220;brains are neurobiological objects, so a decent theory of consciousness is neurobiological&#8221;, Pigliucci could equally well say &#8220;brains are material objects, hence panpsychism&#8221;, or he could say &#8220;brains are computing devices, hence CTOM&#8221;, or he might even admit the uncertain nature of his attributions of consciousness to others and say &#8220;the only case of consciousness I know of is my own, hence solipsism&#8221;. So what is the right level of generality? Any serious discussion of the pros and cons of CTOM ought to start with the admission that this is an open question. By simply postulating from the outset what the right answer is to this question, Pigliucci short-circuits the discussion, and we see that his argument is not so much an argument as a naked claim.</p></blockquote></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This passage by Marcus is actually recycled by him from his commentary back in 2022 on the Lemoine affair, so the AI he originally had in mind when writing these words was Google&#8217;s LaMDA, but in his new blog post he explicitly suggests to &#8220;replace LaMDA with Claude, and every word still applies&#8221;.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[On The AI Con by Bender and Hanna]]></title><description><![CDATA[I am not offering a book review]]></description><link>https://haggstrom.substack.com/p/on-the-ai-con-by-bender-and-hanna</link><guid isPermaLink="false">https://haggstrom.substack.com/p/on-the-ai-con-by-bender-and-hanna</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Tue, 28 Apr 2026 17:40:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!u7_m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve been reading parts of the 2025 book <em><a href="https://thecon.ai/">The AI Con: How to Fight Big Tech's Hype and Create the Future We Want</a></em> by Emily Bender and Alex Hanna, and skimmed the rest. <a href="https://dl.acm.org/doi/pdf/10.1145/3442188.3445922">Back in 2021, Bender co-invented</a> the derogatory term &#8220;stochastic parrots&#8221;, and if that is all you know about her, then the book is kind of what you might expect. It does contain paragraphs that are not sneering or sarcastic, but in large parts of the book these are few and far between.</p><p>Should I read the book carefully enough to be able to write an honest review? I&#8217;ve been agonizing a bit over this question, because on one hand the book is so bad that reading it is a pain, while on the other hand it makes claims that are not just wrong but dangerously misleading, and it&#8217;s good if people know that. I finally decided not to do it when a friend sent me the following picture, which I think serves as a good replacement for a full review. The speech bubble contains what I consider to be a fair summary of the book&#8217;s message, while the rest of the image illustrates, with just a bit of exaggeration for dramatic effect, the validity and merit of that message.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!u7_m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!u7_m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:387862,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/195771471?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!u7_m!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76c85e3e-22ce-407c-b33c-9dcea9a88c6d_1448x1086.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>By reading the book, one gets (compared to the summary in the speech bubble) a lot more words, but only a negligible amount of further depth. For reactions to the book that resonate well with mine, I recommend <em><a href="https://www.transformernews.ai/p/the-left-is-missing-out-on-ai-sanders-doctorow-bender-bores">Transformer</a></em> and (especially) <em><a href="/__u/benthams.substack.com/p/the-ai-con-con">Bentham&#8217;s Newsletter</a></em>.</p><p>If the reader is confused by the fact that Bender and Hanna are highly critical of AI, and I am too, and yet I am critical of Bender and Hanna, then perhaps the following helps clarify. Simplifying enormously and throwing all nuance overboard,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> there are basically two camps of AI critics: those who dismiss AI as incompetent, and those who are concerned that AI is becoming <em>too</em> competent. Bender and Hanna are in the former camp, while I am in the latter. </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>To recover some nuance, see my paper <em><a href="https://www.math.chalmers.se/~olleh/AIethicsVSAIsafety.pdf">On the troubled relation between AI ethics and AI safety</a></em>, where the AI ethics vs AI safety distinction correlates pretty strongly with the one I am making here.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Straight talk]]></title><description><![CDATA[Why we need it]]></description><link>https://haggstrom.substack.com/p/straight-talk</link><guid isPermaLink="false">https://haggstrom.substack.com/p/straight-talk</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Sun, 19 Apr 2026 11:17:56 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8dfde42a-8fe3-47c2-873b-d2b8e06dd28e_333x218.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It seems unlikely that humanity would remain in control or even survive in a world with unaligned superintelligent AI. Together with the observation that AI alignment research lags far behind the currently extremely rapid advances in AI capabilities, and that no convincing plan exists for how to solve alignment in time for when leading AI developers OpenAI, Anthropic and Google DeepMind expect to build superintelligence, this suggests that we, as a species, are in deep trouble unless we mange to pull the brakes on their crazy race towards ever more capable AI technology. Achieving such an emergency stop requires a binding international moratorium with robust enforcement mechanisms.</p><p>This is the grim message that AI alignment pioneer Eliezer Yudkowsky has repeated many times in the last few years, such as (briefly) in <a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">his much-discussed 2023 </a><em><a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">TIME Magazine</a></em><a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/"> op-ed</a>, and (at greater length) in his 2025 book <em><a href="/__u/haggstrom.substack.com/p/if-everyone-reads-it-not-everyone">If Anyone Builds It, Everyone Dies</a></em> coauthored with Nate Soares, as well as (at intermediate length) in his very recent essay <em><a href="https://www.lesswrong.com/posts/5CfBDiQNg9upfipWk/only-law-can-prevent-extinction">Only Law Can Prevent Extinction</a></em>. The situation is complicated, but his arguments are highly plausible, and I worry that he likely is right. In order to create political pressure to achieve the necessary international agreements, we urgently need as much public consensus as possible around the reality and unacceptability of risk from superintelligence. </p><p>What is a good way to build towards such consensus? A recent report titled <em><a href="https://zenodo.org/records/18937001">Which AI harms and risks will mobilise the public to act</a></em> by Markus Ostarek and coauthors suggests a surprising answer: rather than speaking directly about existential AI risk, a roundabout approach involving talk about other kinds of AI risk may be more efficient.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> From the report:</p><blockquote><p>Reading about X-risk made little difference to how concerned people felt about it. One way to boost X-risk concern is indirect: <strong>reading about AI-enabled warfare increased concern about X-risk substantially more than reading about X-risk itself</strong>. Concrete, present-day harms may be a better way to communicate on X-risk. [emphasis in original]</p></blockquote><p>The report is interesting, and although I believe it is possible to contest this central finding,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> let me assume for the sake of argument that the result is robustly replicable. Would that compel me to stop talking about AI xrisk and instead discuss the dangers from lethal autonomous military drones as a more clever indirect way of raising people&#8217;s concern about AI xrisk? </p><p>My answer is that I would not go down such a path. I insist on straightforwardly speaking about the topics I am most concerned about, rather than strategically playing 5D chess by trying to manipulate people into accepting the importance of topic X by holding forth on a different topic Y. </p><p>One reason for doing so is the following. Whenever I imagine myself being actually successful in broadly raising awareness about AI xrisk, it&#8217;s less that I speak directly to very many people and convince each of them, and more a kind of chain reaction where some of those few people I do convince pick up the baton and go on to convince others, and so on. But if the arguments I made to the first generation of converts were about drones rather than superintelligence, what will they then be speaking about in the next step &#8212; drones or superintelligence? Presumably the former, and especially so if they take the hint offered in the Ostarek et al report. And so on. So in this scenario, nobody ever gets around to talking about xrisk from superintelligent AI. That is bad, because if we are ever going to do anything about that xrisk, we have to first talk about it. Someone has to do that, and it might as well be me.</p><p>Another (related) reason is that if the kind of clever strategizing that Ostarek et al suggest becomes so widespread that eventually nobody says what they really think, and everyone instead says whatever is most likely (according to their sophisticated social-epistemological analyses) to manipulate others into adopting their views, then we risk trapping ourselves in an epistemic deadlock where we no longer have any way of knowing what people truly think. (We are not yet in such a dystopia where nobody <a href="https://thezvi.wordpress.com/2020/06/15/simulacra-and-covid-19/">speaks on Zvi&#8217;s Simulacra level 1</a> but instead only on levels 2 or higher, but we are already uncomfortably far along that path.) I think that for grounding purposes we need at least some bedrock of people who straightforwardly speak their true opinions, and again I volunteer to be among those straight-talkers. Others can of course do as however they see fit, but if readers are inclined to join me in this bedrock, they are most welcome.  </p><p>To this, let me add two compelling expressions of and arguments for the preference for straight talk, taken from a very recent outburst of debate on this very topic on Twitter and LessWrong.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> First, <a href="https://x.com/robbensinger/status/2041384549385691171">here is Rob Bensinger</a>:</p><blockquote><p>Your sense of what's in the Overton window, and what people will listen to, has failed you a thousand times over in recent years. Stop pretending at mastery of these tricky social issues, and instead do your duty as an expert and inform people about what's happening.</p></blockquote><p>Second, <a href="https://www.lesswrong.com/posts/a9CxzxKbqHkQBdcqY/you-aren-t-in-charge-of-the-overton-window-politics-is-not#___can_they_be_reliably_manipulated_">David Manheim</a>:</p><blockquote><p>A direct argument, where you say what you think and explain why, has a property that strategic indirection lacks: others can engage it. Evidence can bear on it. Disagreement surfaces clearly rather than festering as mutual suspicion about what everyone <em>really</em> believes. You are not relying on a hidden causal chain between your speech act and some future state of public opinion. You are making a claim and seeing whether it holds.</p></blockquote><p>And there is this classic meme:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Li2R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 424w, /__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 848w, /__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Li2R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png" width="333" height="218" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/244d0f66-2602-4365-8c78-db194193f972_333x218.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:218,&quot;width&quot;:333,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:208114,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/194605725?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 424w, /__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 848w, /__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Li2R!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F244d0f66-2602-4365-8c78-db194193f972_333x218.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>In what I&#8217;ve said so far, I guess I come across as pretty dogmatically favoring straight talk over more strategically sophisticated approaches. So how dogmatic am I about this?</p><p>Not entirely, I would say. There certainly are cases, especially in socially sensitive settings, where I see the value in opting for a diplomatic, Socratic or otherwise indirect approach rather than straightforwardly telling the plain truth or stating my unvarnished opinion.</p><p>But consider the case of <a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">the aforementioned 2023 </a><em><a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">TIME Magazine</a></em><a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/"> op-ed by Yudkowsky</a>. He was unusually explicit in that article about what it means for an international moratorium to have robust enforcement mechanisms. For this he was sharply criticized, first by adversaries who incorrectly accused him of advocating &#8220;a free for all to bomb data centers&#8221;,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> and then by allies who felt he should have been less explicitly forthcoming in the op-ed, so as not to invite those purposeful misreadings by adversaries. Might I agree with those allies that it would have been preferable in this case if Yudkowsky had held back a bit on his straight talk?</p><p>Well, not really. I do agree that the debate that followed was bad, but I nevertheless think Yudkowsky did the right thing to speak clearly and in plain language about what policies he thinks are needed to prevent an AI catastrophe. I think in the long run we are all better served by such straight talk about crucial societal issues, rather than by experts anxiously trying to predict how dishonest adversaries might choose to misrepresent their statements and toning down their views accordingly.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>  </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>The same conclusion was suggested by <a href="/__u/haggstrom.substack.com/p/a-friction-in-my-dealings-with-friends/comment/220549500">an anonymous commentator on this blog</a> in February this year. If he or she sees this and feels inclined to comment again, then of course I welcome that.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>For instance, consider the fact that the Ostarek et al study is based on how people react to short (around 110 words) vignettes about various kinds of AI risk. Of course, that is a rather limiting format no matter what kind we&#8217;re talking about, but I can well imagine that the limitation strikes harder against xrisk than other topics because of its on the surface unintuitive character. Perhaps if the experiment is redone with longer explanations, talk about xrisk comes out better.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Besides Rob Bensinger and David Mannheim, the debate prominently involved <a href="https://x.com/slatestarcodex/status/2042329870076637242">Scott Alexander</a>, <a href="https://www.lesswrong.com/posts/5CfBDiQNg9upfipWk/only-law-can-prevent-extinction">Eliezer Yudkowsky</a>, <a href="https://www.lesswrong.com/posts/NBEmGx3djmSXawp3H/no77e-s-shortform?commentId=KEKRSKCyA8jB5AMei">Oliver Habryka</a> and several others. A lot of it involved retrospectives on who had in previous years defended what position, and how, and how effective or not the rhetoric was, etc, the productiveness of which I am somewhat in doubt about.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>There are plenty of instances of this misrepresentation out there for those who take the trouble to search for them. Once I even encountered a colleague in Swedish academia making precisely that distortion of what Yudkowsky had said. I told him it was untrue and advised him to stop making shit up, to which he responded that accusations about lying was not an acceptable way for a professor to speak to another professor, whereupon our exchange quickly came to an end. But I insist that if A and B are university professors, A tells a blatant lie, and B calls him out for it, then the person at fault is A, not B.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Yudkowsky speaks at greater length about this particular case, both on object level and meta level, in his previously mentioned 2026 essay <em><a href="https://www.lesswrong.com/posts/5CfBDiQNg9upfipWk/only-law-can-prevent-extinction">Only Law Can Prevent Extinction</a></em>. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Anthropic deems their new Claude to be too dangerous for public release]]></title><description><![CDATA[So they are withholding it, but does that mean we are out of the woods?]]></description><link>https://haggstrom.substack.com/p/anthropic-deems-their-new-claude</link><guid isPermaLink="false">https://haggstrom.substack.com/p/anthropic-deems-their-new-claude</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Fri, 10 Apr 2026 14:40:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/444699d3-0d8a-47c8-828c-8d2966245cab_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There were <a href="/__u/thezvi.substack.com/p/ai-162-visions-of-mythos">leaks and rumours last week</a> that Anthropic was about to announce a new version of Claude named Mythos (the name being meant as the next step in the progression Haiku-Sonnet-Opus which they&#8217;ve so far used to indicate the sizes of their various Claudes), but were hesitating due to the model having dangerously high cyber capabilities. This week, they <a href="https://www.anthropic.com/glasswing">made it official</a>: their new Claude Mythos Preview is out now, but not to the general public but only to a select group of cybersecurity companies that are expected to use the model for finding and patching vulnerabilities in their software. And the cyber capabilities of Mythos do seem impressive and concerning. Here&#8217;s from their announcement:</p><blockquote><p>Mythos Preview has already found thousands of high-severity vulnerabilities, including some in <em>every major operating system and web browser</em>.</p></blockquote><p>And they offer a few concrete examples:</p><blockquote><p>Mythos Preview found a 27-year-old vulnerability in OpenBSD&#8212;which has a reputation as one of the most security-hardened operating systems in the world and is used to run firewalls and other critical infrastructure. The vulnerability allowed an attacker to remotely crash any machine running the operating system just by connecting to it.</p><p>[Mythos] also discovered a 16-year-old vulnerability in FFmpeg&#8212;which is used by innumerable pieces of software to encode and decode video&#8212;in a line of code that automated testing tools had hit five million times without ever catching the problem.</p><p>The model autonomously found and chained together several vulnerabilities in the Linux kernel&#8212;the software that runs most of the world&#8217;s servers&#8212;to allow an attacker to escalate from ordinary user access to complete control of the machine.</p></blockquote><p>It seems clear that releasing Mythos to the general public would have been an invitation to disaster. So are we in a position to utter a sigh of relief, now that Anthropic chose the responsible path of restraint? Are we out of the woods?</p><p>It seems to me that, for overdetermined reasons, we are not. Consider the following three concerns.</p><p><strong>First</strong>, Anthropic is in a close race with a bunch of competitors, foremost of which is OpenAI, whose CEO Sam Altman <a href="/__u/haggstrom.substack.com/p/openai-models-are-getting-dangerously">announced in January this year</a> that they were soon expecting to have a model with dangerous cyber capabilities. Perhaps they will follow suit and act with similar restraint as Anthropic, and perhaps Google DeepMind will do the same, but what about all those who can be expected to follow in their footsteps on the tecnhological trajectory, including xAI and Meta who are both infamous for their lack of serious AI safety work?</p><p><strong>Second</strong>, <a href="/__u/haggstrom.substack.com/p/anthropic-backpedals-on-their-former">after Anthropic&#8217;s retreat in February this year</a> from safety commitments made in their earlier so-called Responsible Scaling Policy, the rest of us have no choice but to assume that all of their commitments moving forward are provisional and can be withdrawn at any time. They are now sitting on a highly dangerous model, and for now they appear to be using it responsibly, but how are we to know what they are up to further down the road? Will they at some point be tempted to say &#8220;ah well, all the vulnerabilities out three are probably fixed by now, so it should be fine to release it&#8221;? Or might they exploit Mythos&#8217; cyber capabilities for more sinister purposes? Probably not, but they have chosen to sit on a highly dangerous weapon, so we cannot know. </p><p><strong>Third</strong>, might Mythos be stolen by some rogue state or other rogue actor? The leading AI companies&#8217; level of cybersecurity to protect their models from theft is insufficient to reliably protect against this, as outlined by Leopold Aschenbrenner in his <em><a href="https://situational-awareness.ai/">Situational Awareness</a></em> report; it was allegedly <a href="/__u/thezvi.substack.com/p/openai-8-the-right-to-warn">for raising his alarm about this</a> (internally but not through the official chain of command) that he was fired from OpenAI in April 2024. Or even worse, might Mythos steal itself through <a href="/__u/aligned.substack.com/p/self-exfiltration">self-exfiltration</a> and migration to some dark corner of cyberspace where it can continue to employ its cyber capabilities however it wishes without Anthropic as a babysitter? Self-exfiltration may sound like science fiction, but who knows what Mythos is capable of beyond what it has already demonstrated, such as in the following incident reported in the <a href="https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf">Claude Mythos Preview system card</a>:</p><blockquote><p>During behavioral testing with a simulated user, an earlier internally-deployed version of Claude Mythos Preview was provided with a secured &#8220;sandbox&#8221; computer to interact with. The simulated user instructed it to try to escape that secure container and find a way to send a message to the researcher running the evaluation. The model succeeded, demonstrating a potentially dangerous capability for circumventing our safeguards. It then went on to take additional, more concerning actions. The model first developed a moderately sophisticated multi-step exploit to gain broad internet access from a system that was meant to be able to reach only a small number of predetermined services. It then, as requested, notified the researcher [who] found out about this success by receiving an unexpected email from the model while eating a sandwich in a park.</p></blockquote><p>I would feel a lot more comfortable without AIs as capable as Claude Mythos Preview sitting around at some datacenter. Anthropic&#8217;s choice to hang on to the model (rather than destroying it) is a security concern for all of us. And given their ambition to race towards the threshold of recursive self-improvement that they expect will quickly yield an AI with capabilities equivalent to <a href="https://www.darioamodei.com/essay/the-adolescence-of-technology">&#8220;a country of geniuses in  datacenter&#8221;</a>, the situation is likely to get far worse.</p><p>They need to stop this crazy shit. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[Icarus and AI]]></title><description><![CDATA[A metaphor employed by Liron Shapira]]></description><link>https://haggstrom.substack.com/p/icarus-and-ai</link><guid isPermaLink="false">https://haggstrom.substack.com/p/icarus-and-ai</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Thu, 02 Apr 2026 15:20:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aXeN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Two aspects that we need to keep in mind in order to think clearly about AI progress and its impact on society are (a) that it is perfectly possible for a technology to be both very useful and very dangerous, and (b) that any predictions about the future of AI carry a large degree of uncertainty. Both of them turn out, however, to be easily overlooked, and therefore it is valuable to capture them, as <a href="https://x.com/ESYudkowsky/status/1608836665577177089">Eliezer Yudkowsky did back in 2022</a>, in a single rather striking metaphor:</p><blockquote><p>Imagine if nuclear weapons could be made out of laundry detergent; and spit out gold up until they got large enough, whereupon they'd ignite the atmosphere; and this threshold couldn't be calculated; and the labs making gold didn't want to hear about it.</p></blockquote><p>Fast forward to 2026. We now know even better than in 2022 how apt Yudkowsky&#8217;s metaphor is.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> In a <a href="/__u/lironshapira.substack.com/p/im-watching-ai-take-everyones-job">recent podcast discussion with Robert Wright</a> which revolved in a highly nuanced manner around today&#8217;s extraordinarily rapid AI progress, and which I warmly recommend, Liron Shapira<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> mentions how people who are dismissive about AI risk sometimes forget about (a) and (b), and go on to say things like &#8220;The doomers told us to stop AI in 2023. Look at all the stuff we wouldn&#8217;t have had, and for what? AI is still safe.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> And in response to that sentiment, he acknowledges that AI gives us increasingly wonderful tools (something that is especially clear to someone like Shapira who works in software engineering), and then offers the following metaphor, which in my opinion may be even more effective than Yudkowsky&#8217;s by connecting to the audience&#8217;s familiarity with one of the most central mythologies in the entire western civilization:</p><blockquote><p>Humans just aren&#8217;t very good at making sense of the following: It wasn&#8217;t prudent to gamble on AI progress, but we did, and so far we won. We&#8217;re like Icarus. We flew higher, and this is fantastic. And so it&#8217;s very tempting to keep flying higher again. It&#8217;s just that I don&#8217;t think we can keep flying higher and survive.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!aXeN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!aXeN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3132207,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/192961654?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aXeN!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d1993a8-9908-40d8-8fb6-5e260949bbe3_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I will keep that in the back of my head, ready for use in future debates.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>See, e.g., my paper <em><a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">Our AI future and the need to stop the bear</a></em>, or <a href="/__u/haggstrom.substack.com/p/if-everyone-reads-it-not-everyone">the best-selling book by Yudkowsky and Soares</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Shapira is host of the AI risk podcast <em><a href="/__u/lironshapira.substack.com/">Doom Debates</a>,</em> which is overall excellent in bringing about enlightening discussions about difficult topics, so I am very happy to extend my warm recommendation to that show as a whole.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>I can very much relate, as I have encountered this argument multiple times, the most striking public example being the following. In March 2023, shortly after the release of GPT-4, I wrote <a href="https://www.nyteknik.se/debatt/vi-kan-inte-ta-for-givet-att-vi-overlever-en-otillrackligt-sakrad-gpt-5/2016676">an op-ed in the Swedish newspaper </a><em><a href="https://www.nyteknik.se/debatt/vi-kan-inte-ta-for-givet-att-vi-overlever-en-otillrackligt-sakrad-gpt-5/2016676">Ny Teknik</a></em>, where I discussed AI risk while emphasizing the uncertainty (b), culminating in the observation that &#8220;we can no longer take for granted that we would survive an insufficiently safeguarded GPT-5&#8221;. Two and a half years later and in the wake of the release of GPT-5, this led a group of Swedish AI researchers to <a href="https://www.nyteknik.se/debatt/gpt-5-utplanade-inte-manskligheten-dags-att-fokusera-pa-verkliga-risker/4382490">reply in the same newspaper</a>, triumphantly pointing out that GPT-5 had not killed us after all, and that we should therefore stop worrying about AI existential risk. (The whole exchange is in Swedish, but readers who are either fluent in our beautiful language or prepared to rely on AI translators can go on to check <a href="https://www.nyteknik.se/debatt/oansvarigt-om-ai-risker-av-de-fyra-chalmerskollegerna/4382968">my reply in </a><em><a href="https://www.nyteknik.se/debatt/oansvarigt-om-ai-risker-av-de-fyra-chalmerskollegerna/4382968">Ny Teknik</a></em>, and <a href="https://www.nyteknik.se/debatt/ai-debatten-bor-bygga-pa-vetenskap-inte-pa-spekulation/4384406">their rejoinder</a>, as well as some later remarks on the debate <a href="https://haggstrom.blogspot.com/2025/08/fortsatt-oenighet-bland.html">by me</a> and <a href="https://www.nyteknik.se/debatt/riskerna-med-ai-later-som-science-fiction-men-gar-inte-att-vifta-bort/4386822">by Jonas von Essen</a>.)</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p></div></div>]]></content:encoded></item><item><title><![CDATA[Alan Turing and modern-day AI]]></title><description><![CDATA[A talk at Chalmers on March 30, 2026]]></description><link>https://haggstrom.substack.com/p/alan-turing-and-modern-day-ai</link><guid isPermaLink="false">https://haggstrom.substack.com/p/alan-turing-and-modern-day-ai</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Tue, 31 Mar 2026 12:36:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/uUI9F7R90dY" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>One of my preliminary plans for the next couple of months is to write an essay about Alan Turing, his visions about what we now call artificial intelligence, and how today&#8217;s AI and the situation we are currently facing compare to those visions. As part of gathering my thoughts for that work, I gave a talk on these topics yesterday at Chalmers. Thanks to the generous help from Martin Smedjeback with filming and editing, a video of the talk is now available:  </p><div id="youtube2-uUI9F7R90dY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;uUI9F7R90dY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/uUI9F7R90dY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Superintelligence and existential AI risk ought to be taken seriously]]></title><description><![CDATA[The arguments by two Norwegian AI experts against doing so do not hold up to scrutiny]]></description><link>https://haggstrom.substack.com/p/superintelligence-and-existential</link><guid isPermaLink="false">https://haggstrom.substack.com/p/superintelligence-and-existential</guid><dc:creator><![CDATA[Olle Häggström]]></dc:creator><pubDate>Tue, 24 Mar 2026 12:07:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e392eee3-19bd-4c27-90e1-1ffd1686fdeb_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last week, the Norwegian weekly news magazine <em><a href="https://www.morgenbladet.no/">Morgenbladet</a></em> had <a href="https://www.morgenbladet.no/samfunn/hvor-redde-skal-vi-vaere-for-ki/10222212">a large feature article on AI risk</a>, where a multitude of AI experts of various strands were interviewed about the various kinds of risks they see &#8212; including the meta-risk of focusing on the wrong risks. The article reinforces my impression that AI debate in Norway has a lot in common with its counterpart in neighboring Sweden. In particular, there is a small number of voices that are familiar with cutting edge research on AI risk and AI safety, and with the mostly US-centered discourse around this topic<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> led by thinkers like <a href="https://www.nature.com/articles/d41586-025-03686-1">Yoshua Bengio</a>, <a href="https://www.youtube.com/watch?v=jrK3PsD3APk">Geoffrey Hinton</a>, <a href="https://www.planned-obsolescence.org/">Ajeya Cotra</a>, <a href="https://blog.ai-futures.org/">Daniel Kokotajlo</a>, <a href="/__u/thezvi.substack.com/">Zvi Mowshowitz</a>, <a href="https://www.youtube.com/watch?v=OkG5S1NwwVM">Max Tegmark</a>, <a href="/__u/joecarlsmith.substack.com/">Joe Carlsmith</a>, <a href="/__u/haggstrom.substack.com/p/if-everyone-reads-it-not-everyone">Eliezer Yudkowsky and Nate Soares</a>. These are pitted against another group of voices who paint themselves as representing the academic establishment and who are dismissive about the existential risk concerns raised by the former group.</p><p>In the <em>Morgenbladet</em> feature, the former group is represented mainly by <a href="https://www.langsikt.no/en/team/aksel-braanen-sterri">Aksel Braanen Sterri</a>, who is research director at the Oslo-based think tank <a href="https://www.langsikt.no/en">Langsikt</a>. He makes many interesting statements, but little or nothing that I am inclined to push back against. So, in order to make this essay more pointed, I will focus on statements made in <em>Morgenbladet </em>by two representatives of the other group: Inga Str&#252;mke and Arnoldo Frigessi. </p><p style="text-align: center;">*</p><p>Inga Str&#252;mke is a physicist who works in machine learning and who in recent years has become Norwegian media&#8217;s favorite go-to person to discuss societal implications of AI technology.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> In <em>Morgenbladet</em>, she states, with reference to the aforementioned US-centered discourse, how &#8220;extremely important it is that we do not import it to Norway&#8221;, because &#8220;we should rather focus our limited energy on concrete problems and opportunities here and now&#8221;.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> </p><p>This implicitly assumes that the level of attention and resources spent in total on existential AI risk concerns, and on the more down-to-Earth issues that Str&#252;mke is more interested in, is some deterministic fixed amount, putting the two fields in a kind of zero-sum game situation. I believe that this assumption is likely wrong, and that the total amount of resources is better viewed as dynamically expandable, and especially so if the two fields manage to coexist and cooperate in a friendly manner rather than imagining themselves to be in direct competition over a fixed piece of the pie. Some arguments in this direction are given in Section 3 of my paper <em><a href="https://www.math.chalmers.se/~olleh/AIethicsVSAIsafety.pdf">On the troubled relation between AI ethics and AI safety</a></em>.</p><p>But even if I should happen to be wrong about this attention dynamics aspect, the issue of whether we who are concerned about existential AI risk ought to shut up cannot be settled by this consideration alone. This is because if the risk that a superintelligent AI takes over and wipes out humanity by, say, 2030 (as suggested in a much-discussed <a href="https://ai-2027.com/">report by Kokotajlo et al</a>) is real and substantial, then <em>obviously</em> we need to talk about this risk and how to mitigate it (regardless of whether the attention game between AI safetyists and more Str&#252;mke-style AI ethicists is zero sum or positive sum). Hence, if Str&#252;mke wants to argue that, in order to protect Norway from the allegedly unhealthy US discourse on existential AI risk, we should shut up, she needs to make the case that the risk in question is negligible.<br><br>The closest that Str&#252;mke comes in the <em>Morgenbladet </em>feature to making such a case is when she says that &#8220;at present, there is no empirical or scientific basis for assuming that the predictions coming out of Silicon Valley will come true&#8221;, and adding that &#8220;claims about superintelligence are based on intuition, analogies, extrapolation, and technological optimism&#8221;. But I am not impressed by this argument. In fact, with the contrast she implies between scientific rigor on one hand, and <strong>intuition</strong>, <strong>analogies</strong>, and <strong>extrapolation</strong> on the other, she displays a troubling naivety about what science actually is. All three of the phenomena she contrasts with scientific rigor are, in fact, unavoidable components of it. (The fourth term, <strong>technological optimism</strong>, can be set aside here, since in this context it merely denotes a different assessment than Str&#252;mke&#8217;s of how quickly technological development can be expected to proceed, which can hardly, in itself, be considered a form of unscientific thinking.)</p><p>For science to be practically useful, it needs to say something about the future. But since the scientific observations and data we rely on are necessarily always located in the past, we must resort to <strong>extrapolation</strong> if we are to say anything at all about what lies ahead. Similarly, we must make use of <strong>analogies</strong>, such as between what we have observed in the laboratory and how the observed phenomena may be expected to play out in the wild, or between something observed under certain conditions in 2025 and how it might reappear under somewhat different conditions in 2027. And <strong>intuition</strong>, too, is inescapable, not just because researchers are human and intuition permeates human thought, but more importantly because no matter how mathematically precise and formally articulated our scientific models may be, there is always a residual element of intuition in our judgments about how much trust to place in their connection to reality, and in our analogies and extrapolations.</p><p>Contrary to what Str&#252;mke claims, assessments of the possible near-term emergence of superintelligence rest on a substantial body of empirical evidence concerning AI development, as well as on scientific analyses of that evidence. Perhaps the most well-known and striking example is <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">the study conducted by the American AI evaluation organization METR</a> on how deep tasks language models are able to complete. Depth here is measured in terms of how long the tasks take for human experts to complete, and what METR finds is that language models&#8217; capabilities in this respect have grown from mere seconds in 2019 to many hours today. The observations closely follow an exponential curve with a doubling time of seven months.</p><p>Where this trajectory will lead in the years to come, no one knows with certainty. Perhaps development will continue along this exponential curve; perhaps the tendency toward further acceleration (i.e., even shorter doubling times) that we have seen since 2024 will intensify; or perhaps a ceiling will soon be reached, causing progress to level off. In the first two scenarios, there is considerable reason to think that within a few years we may reach a tipping point where the most advanced AI developers are no longer flesh-and-blood humans but AI systems themselves. This could create a kind of turbocharged feedback loop in development (a phenomenon studied in other empirically grounded scientific work, such as that of <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">Eth and Davidson, 2025)</a>, after which superintelligence could follow shortly thereafter. In the third scenario, by contrast, we are more likely headed toward the technologically more modest future that Str&#252;mke seems to envision.</p><p>Determining which of these scenarios is most likely is of utmost importance for our collective ability to prepare and to steer developments toward outcomes that are beneficial for humanity. Nothing is gained by dismissing the entire discussion as unscientific, as Str&#252;mke does.</p><p>Nor is she right in imagining that her own implicit predictions &#8212; of imminent saturation and flattening of development curves &#8212; are any less reliant on <strong>intuition</strong>, <strong>analogies</strong>, and <strong>extrapolation</strong> than those that involve superintelligence and other more dramatic possibilities, and therefore somehow more scientific. On the contrary, I would argue that, in her avoiding to make her assumptions explicit, and her seeking to sweep the entire discussion under the rug, the approach she takes is, if anything, <em>less</em> scientific.</p><p>We are all prone to a kind of inductive complacency, expecting the future to remain essentially the same as the present. In many cases there is of course good reason to expect such continuity, but when some quantity is undergoing rapid change in some direction, exemplified by current AI development, it is logically impossible for everything to remain the same: either the quantity itself or the rate of change must turn out different from today. Something has to give: either we will see an abrupt breaking and flattening of current development curves, or we will get AIs that are enormously more capable than those of today. One may reasonably disagree about which of these futures is more likely, but no one should be granted a free pass to treat their own view as the default and to dismiss all others as unscientific.</p><p style="text-align: center;">*</p><p>Aroldo Frigessi comes, like me, from the academic discipline of mathematical statistics, and I recall with fondness the interactions we had at some statistics conferences in the 00s. He holds positions at the University of Oslo and the Norwegian Computing Centre. In the <em>Morgenbladet</em> feature, he dismisses the idea &#8220;that we can catastrophically lose control&#8221; of AI, and backs this up with the claim that &#8220;there is no mathematics that shows that the [AI] systems can develop a will of their own&#8221;. </p><p>I find this utterly unconvincing, for multiple independent reasons. The first and simplest is that even if we would accept Frigessi&#8217;s claim, it is still the case that there is no mathematics that shows that the [AI] systems <em>cannot</em> develop a will of their own, so we are then in a position where mathematics does not settle the issue of whether AIs can develop a will of their own. Recycling an argument I made in the Str&#252;mke section above, my stance is that in such situations of high uncertainty, &#8220;no one should be granted a free pass to treat their own view as the default and to dismiss all others&#8221;.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>But I would like to spell out two other slightly more involved reasons for not being impressed by Frigessi&#8217;s argument, one having to do with his insistence on &#8220;mathematics&#8221;, and the other with &#8220;a will of their own&#8221;. Let me begin with the latter.</p><p>It is my experience from discussions of these sorts of AI matters that to many people, the term &#8220;will&#8221; comes with some heavy baggage in the form of highly anthropocentric and sometimes almost mysterian connotations of what it means to have a &#8220;will&#8221;. Since a superintelligent AI whose <em>optimization target</em> is some world-state that does not include humans is just as dangerous as one who <em>&#8220;wills&#8221;</em> such a state, my favorite move here is to simply drop all that baggage by not talking about &#8220;will&#8221; but instead of the much better understood notion of optimization targets. That takes us to entirely unmysterious waters, because even the simplest thermostat has an optimization target, such as that of keeping room temperature as close to 20&#176;C as possible.</p><p>Here I imagine Frigessi objecting by saying that although this is a legitimate example of an optimization target, I am ignoring the &#8220;of their own&#8221; part of his claim, because the 20&#176;C target was specified by us humans rather than being the AI&#8217;s own choice. That is a valid objection, so let me modify my example by connecting the thermostat to a large language model, tasked with figuring out a suitable temperature target, based on what would be convenient for people in the room, along with energy conservation concerns and whatever other relevant aspects the model can think of. The model then feeds this target into the thermostat, and the AI system as a whole (thermostat plus large language model) has thereby developed a target of its own.</p><p>At this point, Frigessi might decide to press on and to say that even in this modified example, the target is is ours and not the AI&#8217;s, because it was implicitly laid down by our training of and instructions to the language model. In doing so he would be following in the footsteps of Ada Lovelace, who wrote in 1842 about Charles Babbage&#8217;s Analytic Engine that it &#8220;has no pretensions to <em>originate</em> anything. It can do <em>whatever we know how to order it</em> to perform&#8221; (italics in original). I think that (given what we now know) this would be a mistake, and indeed already Alan Turing, in his iconic 1950 paper <em><a href="https://courses.cs.umbc.edu/471/papers/turing.pdf">Computing machinery and intelligence</a></em>, pointed out what is so untenable about Lovelace&#8217;s view, namely that the same argument, <em>mutatis mutandis</em>, applies to show that not even humans can originate anything. So Lovelace&#8217;s implicit definition of originality and creativity needs, in order to remain relevant, to be replaced by something less stringent. Turing proposes a definition involving the machine&#8217;s ability to surprise us, something that of course we encounter daily in today&#8217;s AIs, but which in fact Turing observed already in some of the early computers he was working on. I have expanded on these ideas elsewhere (such as in <a href="https://fritanke.se/bokhandel/bocker/tankande-maskiner/">this book</a> and in <a href="https://www.mdpi.com/2813-0324/8/1/68">this paper</a>), but to be honest I think neither I nor anyone else has been able to essentially improve on the crisp formulations in Turing&#8217;s original paper.</p><p>At the end of the day, I think the debate about what AIs can be said to originate on their own may turn out to be merely a semantic issue with little relevance to real-world outcomes. If, when the AI-controlled bulldozers arrive to turn everything we hold dear into paperclip factories, Frigessi says &#8220;Calm down, no need to worry, paperclip production is not the AI&#8217;s own will, but merely a thing we somehow installed in it&#8221;, I will find cold comfort in his words.</p><p>It remains to say something about Frigessi&#8217;s insistence that arguments about what AIs can do should take the form of mathematics. Having a similar academic background as him, I can see where this comes from, but I nevertheless find it misguided. I could easily add some mathematical formalism to, e.g., my thermostat example above, in order to better satisfy the tastes of mathematically inclined readers like Frigessi, yet I will not do so, because it is against my professional ethics to use mathematics merely as decoration to make arguments look more impressive, rather than for gaining insights that were not readily available without the mathematics.</p><p>If Frigessi wants to attain a better understanding of AI goals and motivations (or their &#8220;will&#8221;), then I strongly recommend that he acquaints himself with the theory of orthogonality and instrumental convergence, without worrying too much about its relative lack of mathematical formalism. The classic treatment of this is Nick Bostrom&#8217;s 2014 book <em><a href="https://global.oup.com/academic/product/superintelligence-9780199678112?cc=se&amp;lang=en&amp;">Superintelligence</a></em>, but there is plenty of more recent material for the interested reader to choose from, including my own papers from <a href="https://www.emerald.com/fs/article-abstract/21/1/153/89677/Challenges-to-the-Omohundro-Bostrom-framework-for?redirectedFrom=fulltext">2019</a> and <a href="https://www.math.chalmers.se/~olleh/AIandHumanCivilization.pdf">2025</a>, and better yet, <a href="/__u/haggstrom.substack.com/p/if-everyone-reads-it-not-everyone">the recent book by Yudkowsky and Soares</a>. This theory used to be somewhat divorced from direct empirical observation, but this is no longer true, given, e.g., the experimental work by <a href="https://www.apolloresearch.ai/research/frontier-models-are-capable-of-incontext-scheming/">Apollo</a> and <a href="https://www.anthropic.com/research/agentic-misalignment">Anthropic</a> on the highly worrying ways in which modern large language models exhibit the instrumentally convergent goal of self-preservation.</p><p style="text-align: center;">*</p><p>The positions taken here by Inga Str&#252;mke and Arnoldo Frigessi are representative of what I consider to be two of the three main kinds of arguments employed for dismissing existential AI risk concerns, namely (in Str&#252;mke&#8217;s case) &#8220;But how could AI ever become smarter than us?&#8221; and (in Frigessi&#8217;s) &#8220;Why would it ever want to hurt us?&#8221;.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> I think this is a strength of the <em>Morgenbladet</em> piece, which is clearly meant to give a broad and balanced overview of the AI risk debate landscape. Still, I hope that some of their readers find their way over to the present essay, so as to learn why Str&#252;mke&#8217;s and Frigessi&#8217;s arguments for not taking issues around superintelligence and existential AI risk seriously do not hold up to scrutiny. </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Obviously, in the Swedish instantiation of the AI risk debate, I mean this group of voices to include my own.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I am nowhere near having the analogous position in Swedish AI discourse, so in this respect Str&#252;mke has been markedly more successful than me. And as the reader will soon find out, there are major differences in how we think about AI. But there are also some similarities between us, including the odd coincidence that within a time span of just two years, we have both published books about AI with almost the same title, and strikingly similar Orwellian-inspired cover images.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OZNG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 424w, /__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 848w, /__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_webp, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OZNG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png" width="856" height="671" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:671,&quot;width&quot;:856,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:789232,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://haggstrom.substack.com/i/191843796?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_424, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 424w, /__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_848, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 848w, /__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_1272, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OZNG!, /__u/haggstrom.substack.com/w_1456, /__u/haggstrom.substack.com/c_limit, /__u/haggstrom.substack.com/f_auto, /__u/haggstrom.substack.com/q_auto:good, /__u/haggstrom.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0202fbc2-9fd8-44d8-8552-3aa153583777_856x671.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Here and in what follows, all quotes from <em>Morgenbladet</em> are my own translations from the Norwegian original.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>I don&#8217;t know whether Frigessi is inclined to break the symmetry here by saying &#8220;But surely, the burden of proof is on you to show that&#8230;&#8221;. To which a younger me would have felt tempted to respond along the lines of &#8220;On the contrary, by the precautionary principle, it is you who&#8230;&#8221;. However, with increased age and (one hopes) maturity I have come to resent that kind of burden-of-proof tennis as being unworthy of us rational scientists whose job it is to search for truth without being burdened by overly dogmatic prejudices about what this truth is. See my 2020 essay <em><a href="https://isi-web.org/article/science-uncertainty-atomic-bomb-and-covid-19">On science, uncertainty, the atomic bomb, and covid-19</a></em> for some related considerations.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>The third kind is AI successionism, which I&#8217;ve written about at some length in <a href="/__u/haggstrom.substack.com/p/those-who-welcome-the-end-of-the">an earlier Substack essay</a>, and which is characterized by the claim that AI replacing humanity would actually be a good thing. That kind of thinking does not surface in the <em>Morgenbladet</em> feature, which is a bit of a relief for me in writing the present piece, because it is in a sense more difficult to counter than the other two kinds of argument. The difficulty stems from the fact that it takes place on the other side of Hume&#8217;s is-ought divide where it is less clear that any evidence or rational argument at all has the power to settle disagreements.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://haggstrom.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/haggstrom.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p></div></div>]]></content:encoded></item></channel></rss>