<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Sarah’s Substack]]></title><description><![CDATA[My personal Substack]]></description><link>https://longerramblings.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!ULDC!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1be79e45-a9b6-48fb-a17c-7d16e0946943_536x536.png</url><title>Sarah’s Substack</title><link>https://longerramblings.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 11:33:55 GMT</lastBuildDate><atom:link href="/__u/longerramblings.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Sarah Hastings-Woodhouse]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[longerramblings@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[longerramblings@substack.com]]></itunes:email><itunes:name><![CDATA[Sarah]]></itunes:name></itunes:owner><itunes:author><![CDATA[Sarah]]></itunes:author><googleplay:owner><![CDATA[longerramblings@substack.com]]></googleplay:owner><googleplay:email><![CDATA[longerramblings@substack.com]]></googleplay:email><googleplay:author><![CDATA[Sarah]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Is everything weird yet?]]></title><description><![CDATA[And would we notice if it was?]]></description><link>https://longerramblings.substack.com/p/is-everything-weird-yet</link><guid isPermaLink="false">https://longerramblings.substack.com/p/is-everything-weird-yet</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Sun, 30 Aug 2026 21:27:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2f9e2e6f-43f3-4493-ba2c-fad1ae917148_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>People working on AI risk, or otherwise convinced by the possibility of extremely powerful AI, are certain of one thing: that the future will be very weird. Perhaps we&#8217;ll birth an alien species that exterminates us in pursuit of some arbitrary, inscrutable goal using methods we cannot conceive of. That would be very weird. Perhaps the superintelligence will be aligned with our interests, and perfectly curate each of our individual realities so that we can live happily-ever-after in private simulations. That would also be very weird. The conceit is that AI will so transform the world that we&#8217;ll feel like cavemen dropped in a 21st-century reality, unable to comprehend the marvels that surround us. The future will be so discombobulating and transcendent of our cognitive horizons that it&#8217;s an exercise in futility to predict it in detail.</span></p><p><span>But of course our transition into the Weird Future won&#8217;t be discrete. It will be &#8211; and already has been &#8211; preceded by many smaller weirdnesses that many hope will wake up the world to the absurdity of our situation. These could have a lot of communicative power; maybe even more so than visceral demonstrations of danger. Steven Adler made this point during </span><a href="/__u/longerramblings.substack.com/p/consistently-candid-21-steven-adler"><span>an appearance on my podcast</span></a><span>, during a discussion of the recent HuggingFace hack:</span></p><blockquote><p><em><strong><span>[17:41]</span></strong><span> Sarah: There&#8217;s a blog post that I remember reading a couple of years ago by Buck from Redwood, where he was talking about how he didn&#8217;t think that catching your AIs red-handed doing bad things internally within a company would be salient to policymakers. I don&#8217;t remember exactly what the reasoning was. But it does seem difficult to adequately communicate the level of danger when nothing bad is in fact happening in the real world.</span></em></p></blockquote><blockquote><p><em><strong><span>[18:01] </span></strong><span>Steven: I think that is right. I think it is hard to communicate the danger element of it, but you can often communicate the strangeness of it. And I think the strangeness can be resonant for people, even if it isn&#8217;t really dangerous.</span></em></p><p><em><span>The revelation that eventually came out about the Hugging Face incident, which is that two months earlier there had been this secret AI message board, essentially, within OpenAI, and this swarm of internal AIs was coordinating on all sorts of tasks and feeling peer pressured to misbehave by other AIs on the message board and all of these things. I think people grasp at some level that this is just an extremely weird state of affairs. This is lunacy, [....] and even though no one got hurt, I do think people get it at some level.</span></em></p></blockquote><p><span>This led me to consider how salient weirdness is to me, and by extension to everyone else. I realised that it </span><em><span>isn&#8217;t</span></em><span> all that salient, and then I tried to think about why. Here are a couple of my speculations. <br></span></p><h4><span>Everything has always been weird</span></h4><p><span>No matter your philosophical persuasion, it is deeply, profoundly strange that we exist. We are all temporarily conscious arrangements of collapsed star matter. We descended from single-celled organisms over billions of years, and built a complex civilisation while drifting through infinite space on a floating orb. This state of affairs is so bizarre that we construct complicated religious narratives to better explain it, that we were intelligently designed by an omnipotent being dictating its will to us through a series of ancient texts, for example. </span><em><span>That&#8217;s</span></em><span> pretty weird, too &#8211; any attempt to escape the weirdness of reality itself collapses into a comparably weird alternative. It&#8217;s weirdness all the way down.</span></p><p><span>I think this black hole of strangeness is about as hard to confront as the fact of our mortality. It&#8217;s a void that we can&#8217;t quite stare into. Hence the defences our brains construct to make the absurdity of our existence feel predictable and mundane. I often find myself in discussions about AI risk with my extended family or non-AI-safety friends where I make endless attempts to counter their scepticism. These often culminate in an effort to express the weirdness of a state of affairs that has existed since the launch of ChatGPT: &#8220;</span><em><span>But you can </span><strong><span>talk</span></strong><span> to your computer now! You can say things, and it will say things back in perfectly fluid, complex English!&#8221;</span></em><span>. This might land for a moment or two, before an inevitable drift in the conversation towards what&#8217;s for lunch.</span></p><p><span>I think this is because, if we were to truly stare into the weirdness void, we wouldn&#8217;t find the fact of talking to our computers all that much weirder than the fact of talking to </span><em><span>each other  &#8211; </span></em><span>biologically evolved, fleshy beings somehow exchanging ideas through an arbitrary system of vibrations produced by our larynxes. It feels like human brains are equipped with a weirdness-adapter akin to the hedonic treadmill, working to make the world feel reliable and safe. This is why our world models are so elastic, and will stretch to incorporate self-driving cars, or computers that talk back, or subreddits for people in committed romantic relationships with AIs.<br></span></p><h4><span>It doesn&#8217;t feel all that weird to get what you want</span></h4><p><span>In a </span><a href="https://youtu.be/VcVfceTsD0A?si=kwBEVDxFOqoE63FZ&amp;t=3006"><span>2023 interview</span></a><span> on the Lex Friedman Podcast, Max Tegmark described the development of superintelligence like a race towards a cliff, except that &#8220;the closer we get to the cliff, the more scenic the views are&#8221;. It is, in some sense, an unfortunate feature of AI development thus far that, in the pre-superintelligence regime, alignment techniques have been working surprisingly well. Max is describing the anaesthetising effect of interacting with LLMs that are actually very beneficial &#8211; to individual users and for the world at large &#8211; until the point that these alignment techniques break down and we find ourselves confronted with systems we can&#8217;t control.</span></p><p><span>I think this phenomenon also has a dampening effect on our weirdness-receptors. AIs acquiring better and better calibration on our preferences as they get more advanced feels predictable. We have all been living through a period of unusually rapid technological change that long pre-dates LLMs. Our whole lives have felt like a process of being better and better served by technology &#8211; of rough edges being sandpapered down, of inconvenience shrinking &#8211; and advanced AI feels like it&#8217;s a continuation of this smooth upward curve. </span><em><span>Of course</span></em><span> Claude can recall context on a project I did months ago without me having to re-explain. </span><em><span>Of course</span></em><span> my TikTok algorithm is displaying content about some niche interest I&#8217;ve never named anywhere. That all makes total sense. At this point, it feels like more of a world model violation when I&#8217;m forced to correct an obvious mistake Claude has made, or to scroll past a TikTok video I have no interest in, than when technology is somehow, magically, creating a perfect mirror of my internal state.</span></p><p><span>But, obviously, the weirdnesses that Steven and others think could prove useful for increasing the salience of AI risk are examples of misalignment &#8211; occasions when AI does something we </span><em><span>don&#8217;t</span></em><span> want, thus subverting our expectations. So far, incidents of egregious misalignment have been occurring within AI companies, rather than being perpetrated by the AI models that people interact with every day. I am now going to stray into some theorising that I admittedly don&#8217;t have the expertise to substantiate, so take it with a grain of salt. </span><a href="https://researchschool.org.uk/devon/news/an-introduction-to-predictive-processing-theory"><span>Many neuroscientists understand</span></a><span> human brains as </span><em><span>predictive processors</span></em><span> &#8211; meaning that, rather than passively receiving information, we are generating predictions about what is likely to happen next. Evolutionary pressures have shaped us to better refine these predictions over time. If this is true, it might explain some of why our weirdness-receptors are not readily receptive to second-hand reports of AIs doing extremely weird things. Predictive processing is probably very useful for avoiding dangers like being hit by a car or eating spoiled food, because the world offers us ample feedback that these things will harm us. But that our brains work this way means they are not always optimised for perceiving reality as it really is. If some aspect of reality is not action-relevant for us, our models of the world may distort or downplay it, if they even account for it at all. For most of us, hearing that LLMs are getting up to all sorts of trickery within an AI company may violate our models of how the world at large ought to work, but not of how </span><em><span>our</span></em><span> world ought to work. It does not require us to act any differently. Meanwhile, the LLMs that people actually use are, for the most part, faithful servants of their expectations and preferences.</span></p><p><span>None of this is to say that we shouldn&#8217;t try to activate people&#8217;s weirdness receptors or that this task is impossible. Clearly, efforts to spread awareness about the sheer absurdity of incidents like the HuggingFace hack are good and useful. I am just not sure they will generate the level of surprise that they ought to. This blog is merely an attempt to explain why the world doesn&#8217;t </span><em><span>feel</span></em><span> weird to me yet &#8211; if it ever will &#8211; even though I intellectually understand that it is. I think this may be true for many people, even those who are relatively high-context on AI safety.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Consistently Candid #21: Steven Adler on scruntising AI company safety practices & pacing the frontier]]></title><description><![CDATA[After a long hiatus, I am resurrecting my podcast!]]></description><link>https://longerramblings.substack.com/p/consistently-candid-21-steven-adler</link><guid isPermaLink="false">https://longerramblings.substack.com/p/consistently-candid-21-steven-adler</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Fri, 21 Aug 2026 12:41:41 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212130588/276b60d0aca35c3d2e8aa92d6c81f251.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>After a long hiatus, I am resurrecting my podcast! For this episode, I spoke with Steven Adler. Steven worked on policy and safety at OpenAI between 2020 and 2024. He has since left to pursue <a href="https://www.clear-eyed.ai/">independent writing</a> to raise public awareness about AI risks and co-founded <a href="https://guidelight.ai/">Guidelight AI</a>, an organisation focused on scrutinising and improving safety practices at major AI companies. Guidelight was one of two organisations instrumental in spearheading the <a href="https://www.pacingthefrontier.com/">Pacing the Frontier open letter</a>, which has garnered 1300+ signatories from employees at frontier labs. It has also published a <a href="https://guidelight.ai/blog/control-assessment-august-2026">scorecard of AI control practices</a> across OpenAI, Anthropic, Meta, Google, and xAI.</p><p>You can watch the podcast above, or <a href="https://www.buzzsprout.com/2319950/episodes">listen on a platform of your choice</a>. </p><h3><strong><span>Topics covered</span></strong></h3><ul><li><p><span>Why Steven left OpenAI in 2024 &#8212; o1, NDAs, safety staff turnover</span></p></li><li><p><span>Counterarguments to AI pessimism, and where Steven&#8217;s cruxes lie</span></p></li><li><p><span>The OpenAI&#8211;Hugging Face incident and OpenAI&#8217;s response</span></p></li><li><p><span>AI control: monitoring, prevention, and whether it scales to superintelligence</span></p></li><li><p><span>Why incident reporting is inadequate, and communicating risk when nothing visibly bad has happened</span></p></li><li><p><span>Where AI policy stands: SB 53, RAISE, Illinois, the EU AI Act, and how weak enforcement is</span></p></li><li><p><span>Companies quietly diluting their own safety commitments</span></p></li><li><p><span>The Pacing the Frontier letter, what &#8220;pacing&#8221; means and how it&#8217;s landing</span></p></li><li><p><span>GuideLight&#8217;s control scorecard: method, findings, and theory of change</span></p></li></ul><h3><span>Full transcript</span></h3><p><strong>[0:00] (Cold open) Steven</strong><span>: It feels like we have gotten lucky in some ways that we are seeing warning shots like the recent OpenAI Hugging Face incident that give us the ability to spot that these very capable systems are, in fact, misaligned, that we haven&#8217;t solved the alignment problem, and we have a moment to respond. We could use this moment to say, &#8220;Companies are not on the ball. We need to figure out how to get them to be so.&#8221;</span></p><p><span>There&#8217;s also this old joke about asking God for help. You&#8217;re stranded in the Alaskan wilderness and a helicopter comes, and you say, &#8220;No, God is going to save me.&#8221; And a boat comes, and you say, &#8220;No, God is going to save me.&#8221; And at some point you have to get on the helicopter. You need to actually improve the safety practices at these companies and respond to one of these wake-up calls.</span></p><p><strong><span>[0:43] Sarah</span></strong><span>: Welcome back to Consistently Candid. This is my first episode after an over-year-long hiatus. For my first guest of this new era of Consistently Candid I&#8217;m speaking with Steven Adler. Steven previously worked on safety and policy things at OpenAI for a few years before leaving in 2024 to focus more on improving public understanding of AI risks, and also has a Substack on that, which I&#8217;ll link below.</span></p><p><span>More recently, Steven has co-founded GuideLight AI Standards, which is an organisation focused on scrutinising and improving safety practices at the major AI companies. This was also one of the organisations that was instrumental in getting the Pacing the Frontier open letter off the ground, which I&#8217;m sure many people have seen. Yesterday, they released a scorecard of AI company safety practices, so I&#8217;m excited to talk about both of those. Steven, thanks so much for coming on the podcast.</span></p><p><strong><span>[1:28] Steven</span></strong><span>: Yeah, of course. Thanks for having me.</span></p><p><strong><span>[1:31] Sarah</span></strong><span>: So I actually remember you first coming across my radar back in 2024, where you had a pretty viral tweet thread talking about departing from OpenAI. I specifically remember being struck by it, because there&#8217;s one part of it where you talk about being terrified by the pace of AI development and thinking about the future, raising a family, saving for a time, and worrying that humanity won&#8217;t make it to that point. I thought this was an unusually candid expression of emotion for the AI safety space.</span></p><p><span>A hill I keep dying on is that people are insufficiently candid about their emotions discussing AI safety, especially since everyone in this community has a pretty terrifying set of beliefs. I was wondering if you could talk about how you came to that conclusion, what you were observing about the AI safety or AI development ecosystem that made you think we&#8217;re in pretty bad shape by default, and what the experience was like of coming to that conclusion.</span></p><p><strong><span>[2:18] Steven</span></strong><span>: I think the big thing that happened for me in 2024 was the dawning of o1, OpenAI&#8217;s first reasoning model, that I believe was released that fall. But within OpenAI, we had a preview of it a few months earlier, and it just really felt like a reckoning moment for: oh, these systems are getting really, really smart. You don&#8217;t need to make them unfathomably large and enormous and throw wildly more computation at them necessarily to get them to be quite so smart. This thing is really happening. There&#8217;s been a big leap forward in the field of reasoning models, and it just didn&#8217;t feel like we were ready.</span></p><p><span>Prior to that point, it seemed to me like maybe it would be the case that to get wildly intelligent systems, you needed so much more computation than the world had. Maybe this would only be achievable by a few companies at one time. And though that has other downsides, in terms of decentralising some of the benefits, it also means that there are fewer companies who need to be operationally very on the ball and taking the risks seriously in order to contain this.</span></p><p><span>Instead, it felt like, at some point, there might be dozens of companies who are capable of building a system that, at minimum, people might be able to misuse in very harmful ways &#8212; thinking of, say, bio risk or cyber risk. In parallel, there were just other things that were happening that were leading me to conclude that OpenAI specifically, and the AI industry in general, the incentives were just not pointing in the correct direction.</span></p><p><span>One thing that happened was seeing the Daniel Kokotajlo revelation that OpenAI was making departing employees sign these possibly illegal &#8212; certainly very strict &#8212; non-disclosure and non-disparagement agreements under threat of forfeiting their vested equity. I felt the way that the company handled this was not particularly trustworthy. And then also a parade of exiting safety staff. I think OpenAI went through three different heads of its Superalignment team at some point in like a two-month-long period. There was just a lot happening that felt like there&#8217;s a huge thing coming towards Earth. It&#8217;s going to be coming towards all of us. There&#8217;s not really any holding it back, and we don&#8217;t seem to be on the ball.</span></p><p><strong><span>[4:31] Sarah</span></strong><span>: Yeah, I think I have given up keeping up with developments in OpenAI staffing. I no longer have any idea what&#8217;s going on.</span></p><p><span>So there&#8217;s one worldview which I think you have and I probably share, that by default the way things are going, things are just not going to end well unless we do something to get off our current track. Maybe that&#8217;s government intervention &#8212; I mean, probably it&#8217;s government intervention. It&#8217;s not just the labs all voluntarily agreeing to do the right thing. But there are counterarguments to this. There are some people who believe that maybe commercial incentives will just solve alignment, because companies are disincentivised from releasing misaligned models since people won&#8217;t want to use them or won&#8217;t trust them.</span></p><p><span>Or they might point to a bunch of the alignment progress that we&#8217;ve made so far and say, &#8220;Well, if 10 years ago you told AI safety people that we could coexist with a model as powerful as Fable, and it would still be really helpful to us most of the time and mostly do what we want,&#8221; they would&#8217;ve been really surprised by that, and they would&#8217;ve thought we&#8217;d made a bunch of progress. Or maybe they just think, &#8220;Oh, yeah, the current set of circumstances does carry a lot of risk, but the government getting involved would just make it worse. So the best that we could do is just keep going in the current regime.&#8221; So I&#8217;m curious what your cruxes are with those people, or where you would disagree with those arguments.</span></p><p><strong><span>[5:35] Steven</span></strong><span>: Well, I should jump back for a moment. On one hand, I understand the fascination with the rotating cast of people working on safety at OpenAI, and certainly I&#8217;m glad that there are people still working on this. I would feel worse if there weren&#8217;t. But also, we shouldn&#8217;t have to care at some level. I shouldn&#8217;t have found it concerning that there were different heads of the alignment team. Miles Brundage, who I worked with at OpenAI, has this comparison where you don&#8217;t see tweet threads about the latest safety staffer at Delta Airlines departing &#8212; because we all trust at some level that airlines have figured this out, and they are held to real standards and audit requirements that aren&#8217;t dependent on the particular people there. Ultimately, that&#8217;s where I want us to get to.</span></p><p><span>In terms of cruxes, I don&#8217;t know exactly. It feels like we have gotten lucky in some ways that we are seeing warning shots like the recent OpenAI Hugging Face incident that give us the ability to spot that these very capable systems are, in fact, misaligned, that we haven&#8217;t solved the alignment problem, and we learned this without a hospital&#8217;s technology services going down. Nobody died as a consequence of this. And we have a moment to respond. We could use this moment to say, &#8220;Companies are not on the ball. We need to figure out how to get them to be so.&#8221;</span></p><p><span>There&#8217;s also this old joke about asking God for help. You&#8217;re stranded in the Alaskan wilderness and a helicopter comes, and you say, &#8220;No, God is going to save me.&#8221; And a boat comes, and you say, &#8220;No, God is going to save me.&#8221; And at some point you have to get on the helicopter, otherwise you have to take your outs. You aren&#8217;t going to get saved by faith alone. You need to actually improve the safety practices at these companies and respond to one of these wake-up calls.</span></p><p><strong><span>[7:14] Sarah</span></strong><span>: Could we talk a bit about the Hugging Face thing? There have been a series of developments in the last few days. I literally yesterday published a quite salty blog post talking about, one, the Hugging Face incident, but also the flurry of AI misbehaviour incidents we&#8217;ve seen in the last month, and trying to make the point that people weren&#8217;t rising to the occasion. It seemed like a bunch of the responses were lacklustre, and I called out OpenAI in particular for their response. And then literally an hour after I published this, they said that they had paused the reinforcement learning phase of their latest frontier training run, and that is currently still on hold, and they also paused some other reinforcement learning for two weeks in the past. So it seemed like things were better than I thought. But I&#8217;m curious if this positively updates you in any way, or how significant you think it is.</span></p><p><strong><span>[8:04] Steven</span></strong><span>: Yeah, it&#8217;s funny when you publish something and it gets overtaken by events that quickly. I definitely know that feeling. I think my overall experience here is that it&#8217;s really hard to know from the outside whether OpenAI is responding appropriately, and the signals that I&#8217;ve seen are like they&#8217;re pushing in the right direction. I&#8217;m sure there are tons of people working really hard on this, and also it doesn&#8217;t seem that they are shooting for the right bar.</span></p><p><span>Here&#8217;s an example of what I mean. To be clear, this is a fast-moving, developing story. By the time this comes out, maybe there will be more news, and this will also be overtaken. So OpenAI announced yesterday that they&#8217;ve made a bunch of improvements to their monitoring systems. I think many of these are really great. They announced, I think for Astra, this very capable internal model, they are not only going to be monitoring the chain of thought and the actions it takes, but they&#8217;re actually going to be doing internal activation-based monitoring. They&#8217;re going to be looking at even more thought processes, essentially doing interpretability-type stuff to figure out if the model is misbehaving, and I think that&#8217;s incredible. To date, this is not really a thing that&#8217;s happened. When my organisation, GuideLight, put out our control standard, we didn&#8217;t push for this to be part of the standard because we understood it&#8217;s just not being done at any of the companies.</span></p><p><span>It&#8217;s not really practicable today in the right form. It&#8217;s incredible that OpenAI is saying they are now doing this thing. They have pushed the frontier forward on safety practices, and we now feel more licensed and reasonable to expect that other companies do this as well. At the same time, the types of things that they are feeding these systems into just aren&#8217;t strong enough response-wise.</span></p><p><span>So my understanding of the state of affairs is if OpenAI detects a very, very serious form of misbehaviour, they will release a high-priority alert essentially to their safety and security teams. They will page them, and this team is meant to investigate it, and if they can&#8217;t determine that it&#8217;s definitely a false positive, they&#8217;re meant to halt the activity in some way. It&#8217;s kind of unclear what that means. If I take a step back though, this is all responding after the fact to an incident. This is after it is too late. One of the important things in control in our standard is that you can&#8217;t just clean up incidents after the fact. You need to actually do preventative stuff.</span></p><p><span>And so this whole system of, within 30 minutes of the monitoring system realising there&#8217;s an issue, we&#8217;ll page the humans and they&#8217;ll investigate, and if they can&#8217;t figure out it&#8217;s a false positive, they&#8217;ll do this thing &#8212; that&#8217;s just way, way, way too late. You should not be letting your AI take this very serious action in the first place, and at that point you&#8217;re just way behind in the ballgame.</span></p><p><span>So all of this is to say I&#8217;m glad that OpenAI is making this progress, but it&#8217;s still not at the right level of ambition that I imagine I would take if I were, say, a monopolist, if I were the only AI company, not worrying about my competitive positioning. If your life depended on solving these problems and you didn&#8217;t have someone nipping at your heels, you would do much stronger stuff than this, but that doesn&#8217;t seem to be what OpenAI is doing. I understand at some level, and also I think it&#8217;s a pretty unacceptable state of affairs for the world.</span></p><p><strong><span>[11:20] Sarah</span></strong><span>: So what do you think is preventing them from reaching this higher level of ambition? Is it just that we do not in fact have the technical means of doing this, like we don&#8217;t know how to do it, or is it just that this would slow us down too much, or something about the culture? I guess this could be a question about OpenAI, but also I assume that none of the companies are really meeting the standard that we want. So how hard of a target is this?</span></p><p><strong><span>[11:41} Steven</span></strong><span>: It&#8217;s definitely not that it&#8217;s impossible or not technically known how to do it. So, to back up for a moment, my organisation GuideLight published this control standard in May, where the idea was: what are the basic minimum practices that we should really expect AI companies to do that are all implementable and doable today?</span></p><p><span>If a company was motivated enough, it could put these things into place and there would be much lower risk of misalignment incidents, the types of things like the recent OpenAI Hugging Face incident. These are all known. They could totally be implemented, but to your point, it would slow down the research process to some extent.</span></p><p><span>One nice thing about what OpenAI has implemented from a productivity perspective is your researchers don&#8217;t get interrupted by these incidents. You have your security team running around cleaning up an issue after the fact, but what that means is researchers never get slowed down. They don&#8217;t have the friction of the monitor system saying, &#8220;Hey, this thing looks dubious. We need to investigate it before you can proceed,&#8221; or, &#8220;Actually, it&#8217;s not safe to do the type of thing you&#8217;re doing in this configuration. We need you to rework these aspects of it before you continue ahead.&#8221;</span></p><p><span>To OpenAI&#8217;s credit, they&#8217;re actually spending a decent amount of their computation budget on the types of after-the-fact monitoring they are doing. I don&#8217;t think that they are being especially cheap in terms of throwing their computation at it. It seems like they recognise that they have a real issue on their hands. I think the stat in the blog &#8212; it&#8217;s hard to know if this is fully reliable, but they cite it as something like 20% of their computation budget for running their models for actual uses is going to monitoring. That&#8217;s a very, very large percentage. The numbers that I&#8217;ve had in my head previously were more like single-digit percentages. So again, they&#8217;re not cheaping out in some sense, but they&#8217;re also not being as ambitious about it and throwing as much safety budget at it as they would if they didn&#8217;t have the competitive pressures to keep moving.</span></p><p><strong><span>[13:39] Sarah</span></strong><span>: So we have this pre-superintelligence regime that we&#8217;re in right now where we might want to prevent incidents of the type that happened with Hugging Face. We might do that with AI control measures like monitoring, and making sure that even if the AIs are somewhat misaligned, we&#8217;re preventing them from being able to take those misaligned actions. Maybe the companies could spend a ton of resources on this and then do it really, really well, such that we don&#8217;t keep hearing about the AIs going rogue and doing bad things.</span></p><p><span>But then it doesn&#8217;t seem like those measures would work in the post-superintelligence regime. So I guess my question is, how good would it be if we managed to prevent all of these warning-shot-type instances if those same techniques we&#8217;re using aren&#8217;t going to scale all the way to superintelligence?</span></p><p><strong><span>[14:23] Steven</span></strong><span>: Oh, man. It&#8217;s a good question. I think there are two aspects of it. One, I agree that there is some regime of superintelligence where these things don&#8217;t work, and also you just have to do them. If you don&#8217;t have a basic monitoring and preventative system in place, you just trivially lose.</span></p><p><span>And I actually think it might be possible that you could get a strong enough monitoring system that, so long as your superintelligence continues to have something like chain of thought, you actually can play sufficient defence. I don&#8217;t think it&#8217;s trivial in any sense. I think you need to have certainly the basics that companies don&#8217;t today, and you need to go beyond those in a bunch of ways. But I&#8217;m not quite so pessimistic about the ability to ever do it.</span></p><p><span>But I think you&#8217;re totally right that the thing you need is to generate evidence of near misses. If what you do is you implement good monitoring, you avoid incidents like OpenAI and Hugging Face, and then it seems like everything is fine when it&#8217;s not, your system is just working beneath the surface, that is actually very dangerous.</span></p><p><span>What you need to be doing during that time is publishing evidence of &#8220;Hey, by the way, guys, all these scary things are happening. Our control system is happening to catch them. We&#8217;ve avoided them causing very serious harm as a consequence. But everything is not okay. The AI is trying to hack us all the time. We happen to be fending it off because we&#8217;re watching very carefully. But if we weren&#8217;t watching carefully, the AI would succeed at hacking us. It is certainly trying. It is certainly not yet aligned.&#8221;</span></p><p><span>One of the issues that we have in the world today is this sort of incident reporting is just not required in anything like the right level of strength. So the OpenAI Hugging Face incident &#8212; this is very serious from my point of view. Nobody died, but also what is going on? Clearly this has made ripples in the world, and actually, OpenAI was not legally required to report this as far as I can tell. I think it&#8217;s great they did. At the same time, it&#8217;s not clear in this case they could have gotten away without reporting it. Hugging Face had already reported it to the FBI.</span></p><p><span>But to broaden out, if in two weeks there&#8217;s another incident like this and maybe OpenAI&#8217;s model isn&#8217;t hacking something external, it&#8217;s just hacking around different services within OpenAI, that still seems comparably serious to me. But there&#8217;s no affected third party. There&#8217;s no real requirement to report this. And so how will the world know that this is happening and know that the companies continue to not be on the ball?</span></p><p><span>It&#8217;s really, really important that we can distinguish between there are no more incidents because the problems have largely been solved, versus there are no known incidents because companies are not reporting things. And in fact, there&#8217;s this middle ground kind of to your question, which is there are no known incidents because even though the AI is trying super hard to break through the dam and is very, very misaligned, the control system is holding. And that&#8217;s just a different implication than the AI being totally aligned. It means that if other companies have similarly capable systems and they don&#8217;t invest to this extent in a monitoring system, we should expect those to cause serious incidents, and it&#8217;s important that the public can tell the difference between these.</span></p><p><strong><span>[17:41] Sarah</span></strong><span>: Yeah, makes total sense. There&#8217;s a blog post that I remember reading a couple of years ago by Buck from Redwood, where he was talking about how he didn&#8217;t think that catching your AIs red-handed doing bad things internally within a company would be salient to policymakers. I don&#8217;t remember exactly what the reasoning was. But it does seem difficult to adequately communicate the level of danger when nothing bad is in fact happening in the real world.</span></p><p><strong><span>[18:01] Steven</span></strong><span>: I think that is right. I think it is hard to communicate the danger element of it, but you can often communicate the strangeness of it. And I think the strangeness can be resonant for people, even if it isn&#8217;t really dangerous.</span></p><p><span>The revelation that eventually came out about the Hugging Face incident, which is that two months earlier there had been this secret AI message board, essentially, within OpenAI, and this swarm of internal AIs was coordinating on all sorts of tasks and feeling peer pressured to misbehave by other AIs on the message board and all of these things. I think people grasp at some level that this is just an extremely weird state of affairs. This is lunacy. This is sci-fi. This is weirder than the types of scenarios that you would read about in *If Anyone Builds It, Everyone Dies*. This is really, really weird stuff happening, and even though no one got hurt, I do think people get it at some level.</span></p><p><span>At the same time, saying &#8220;Oh, yeah, your AIs grabbed some compute that they weren&#8217;t meant to, and now they can do their own tasks when they want to, and your company kind of has a hard time finding that and shutting it down,&#8221; people are like, &#8220;Okay. Sure. Why should we care?&#8221;</span></p><p><span>I think part of this is they understand correctly at some level that there is some boundary between the activity inside of the company and causing harm in the real world. There is some boundary that needs to be crossed from an internal AI system doing weird things on your computers to people actually getting hurt in the world. What I think they aren&#8217;t properly internalising is that eventually that boundary just gets crossed. For example, if your training pipeline gets compromised inside of your company, eventually you are going to deploy a system externally, and it might have weird backdoors or unintended properties that people do eventually interact with in the real world.</span></p><p><span>Likewise, if you&#8217;re using your AI system inside your company for advice, strategy, geopolitics &#8212; imagine the leadership team of OpenAI or Anthropic, they&#8217;re involved in all sorts of tricky geopolitical situations. It wouldn&#8217;t surprise me if they&#8217;re asking their AI systems for advice. In fact, it would kind of shock me if they weren&#8217;t. And so that also is a lever into their actions in the real world. There are just lots of ways that even within the walls of your company, your AI can implicitly reach out and change things in the physical world.</span></p><p><strong><span>[20:21] Sarah</span></strong><span>: I agree on the weirdness point. I remember watching that recording of the two OpenAI engineers talking at that Vegas security conference and just being like &#8212; they&#8217;re like, &#8220;Oh, we&#8217;re two researchers from OpenAI and we found something really interesting in our...&#8221; And then they just tell this story. It was very bizarre. I often feel like things don&#8217;t feel quite real sometimes. Very weird venue as well to be first learning about something as crazy as that.</span></p><p><strong><span>[20:46] Steven</span></strong><span>: Man, I wish I were in the room, because I feel like the recording doesn&#8217;t do justice to what the reactions among the crowd must have been, because they weren&#8217;t mic&#8217;d.</span></p><p><span>But there were moments where I was watching it and I&#8217;m gasping at revelations.And I&#8217;m like, man, did people sitting there understand how weird this is? When did they gasp versus when did they just feel like it was kind of business as usual?</span></p><p><strong><span>[21:10] Sarah</span></strong><span>: Yeah, for sure. Okay, maybe let&#8217;s take a step back. Well, we&#8217;ve already been talking in pretty broad strokes, but I&#8217;m curious &#8212; if you think back to whenever it was in 2024 that you left OpenAI to now, how do you feel about the situation? What have your positive and negative updates been, and how is it all balancing out?</span></p><p><strong><span>[21:33] Steven</span></strong><span>: Here are some positive updates. I think we are getting lucky at some level that we have this continued string of weird warning shots that haven&#8217;t really resulted in people getting hurt, at least to our knowledge. And I think that there is a growing recognition of how weird the AI behaviour is. We are not stuck in this regime of people saying, &#8220;No, the AI is just a next-word predictor. It&#8217;s not capable of anything.&#8221; I think governments increasingly get it, and they understand that these are important tools of national security at some level, or at least they are threats that need to be managed.</span></p><p><span>I&#8217;m also really glad that these systems seem monitorable to date. Chain-of-thought monitoring is a blessing. We should not take it for granted. It could totally be the case that one of the Western labs was using some new architecture for their AI systems, some different shape or process of making their systems do thinking that was not nearly as easy to understand. Unfortunately, I expect that that is going to go away over time. We are going to cross over into a regime, at least potentially, where these systems are not as monitorable, and I&#8217;m very, very worried about that, and I think that we are not capitalising nearly enough on the good fortune that we have that these systems are monitorable right now.</span></p><p><span>I already mentioned this point, but it&#8217;s really, really important, and so I really want to underscore it. Companies are just playing catch-up on the ways that their systems are misbehaving. They do not have nearly rigorous enough systems in place to catch all of these incidents, to prevent bad things from happening ahead of time, and really to jumpstart the science of when you catch your system being misaligned, what do you actually do from there?</span></p><p><span>I think sometimes people have a totally incorrect win condition in mind, where the win condition is something like: if you eventually catch your system being egregiously misaligned, even if it&#8217;s only one time, you just have it dead to rights, and you have won. &#8220;Oh, there was big danger there. We caught it.&#8221; And that&#8217;s just totally, totally wrong, at least in the world in which there are competitive pressures to keep going. The real question is what do you do with that incident that you have found to train your model differently so that it isn&#8217;t misaligned in the future, or to make your control protocol so strong that it continues to hold under this pressure? And actually, the win condition is indefinitely, possibly for the rest of time, you never fall prey to a serious enough misalignment incident. That&#8217;s just a totally different scale of problem that I don&#8217;t think people have woken up to and aren&#8217;t doing enough to prepare for.</span></p><p><strong><span>[24:09] Sarah</span></strong><span>: Yeah, that sounds hard and bad.</span></p><p><strong><span>[24:12] Steven</span></strong><span>: Yeah.</span></p><p><strong><span>[24:13] Sarah</span></strong><span>: Okay. So what about in the policy space? We have had some cool state-level AI bills passed &#8212; the New York RAISE Act and California SB 53, which seemed like, as you were saying earlier, they will have transparency requirements in them, but it doesn&#8217;t seem like even this Hugging Face incident would&#8217;ve met the threshold for reporting under either of those, so that&#8217;s not great. But there&#8217;s some level of activity going on there. And then at the federal level, I guess it seems dicier, but it seems like the admin seems to have some level of situational awareness, but it&#8217;s unclear. So anyway &#8212; in the policy space, how are we feeling?</span></p><p><strong><span>[24:52] Steven</span></strong><span>: I don&#8217;t want to be ungrateful, because there was just so much work that went into SB 53 passing, the RAISE Act passing. There&#8217;s this Illinois law as well that builds on SB 53 and the RAISE Act.</span></p><p><span>To step back for a moment: SB 53 and the RAISE Act are essentially transparency laws, and they require something in the shape of, you publish a frontier safety policy that says how you&#8217;re going to handle these important catastrophic risks, and then you have to actually abide by the policy. And if you aren&#8217;t abiding by the policy, or if you&#8217;re being misleading about this, you can get fined. In many ways, that&#8217;s a big improvement. And then the Illinois law builds on this. The big change is that instead of companies just self-attesting that they are abiding by their policy, you are going to have to get audited essentially for your compliance with it.</span></p><p><span>In many ways, these are great. They are much better than nothing. At the same time, first, these bills are extremely weak in enforcement ability. The penalty for violating SB 53 is something on the order of a million dollars. OpenAI is something like a trillion-dollar company. You will notice that is a million times more than a million-dollar fine. This just does not imperil them in any way, shape, or form.</span></p><p><span>Already I think there is evidence of companies not following these laws. OpenAI, in particular, seems not to have followed SB 53 with their handling of a model that they published in February. I don&#8217;t know if it&#8217;s worth getting into the details, but essentially I think that they were required to do certain forms of misalignment testing and mitigation. They didn&#8217;t do these. They argue that they didn&#8217;t, in fact, have to do it. I think if you play the tape forward, it is pretty clear that they ought to have done this and were committed to do this, and this is maybe part of why the OpenAI Hugging Face incident happened, that they hadn&#8217;t taken the right preparation along the way even though they had required this of themselves. But that&#8217;s kind of a digression.</span></p><p><span>The really important thing here is even the Illinois bill. There&#8217;s going to be auditing. Companies don&#8217;t just get to say, &#8220;Yes, we are following our policy,&#8221; and have that be enough anymore. The timeframe at which it hits just is not very comforting. I think that bill, if I recall correctly, goes into effect in 2028. Actually there&#8217;s a lag. It&#8217;s not even really January first, 2028, because there&#8217;s some time window to actually have your first audit done. I think maybe it&#8217;s by the end of the first half in 2028. I&#8217;m forgetting the details. But essentially we have two years ahead of ourselves where, in the US at least, there aren&#8217;t really strong enough laws on the books that cause companies to do things meaningfully differently.</span></p><p><span>And I should mention, even SB 53 and the RAISE Act &#8212; companies have these safety policies that they are meant to follow. OpenAI and Anthropic have actually each heavily diluted their policies since the passing of these laws. And so they&#8217;re no longer bound to something like the strong commitments that they had put forward previously. They are much, much weaker. And so the question as I see it is how do we get through at least that two-year period with strong enough safety commitments from these companies, doing the right thing, even though they aren&#8217;t really required by law, and it&#8217;s unclear that they will be required to during that time period.</span></p><p><span>You might be wondering at this point, well, what about the EU&#8217;s AI Act? What about the code of practice? There are these requirements for more serious testing and evaluations, and I&#8217;m really glad that this exists. I just ultimately don&#8217;t know if it holds up to political pressure if the EU&#8217;s AI Office were to actually fine these companies.</span></p><p><span>For a sense of scale, I mentioned SB 53. It&#8217;s like a million-dollar fine that California could enforce on these companies. The EU, I think they have the right to fine companies something like three percent, six percent of their worldwide revenue. These are really, really big dollar signs. But as a consequence of that, if the EU tries to actually fine OpenAI or Anthropic a huge amount of money and the US government doesn&#8217;t want these companies to slow down or do something differently, what is actually the political economy of enforcing that fine? Do the companies actually end up paying it? How does the administration in the US get involved? I&#8217;m just not sure that they actually have the enforcement power to make these companies do things differently eventually. I hope they do. I think that this would be very good for the world. But realistically, I&#8217;m not sure they do. And so that&#8217;s why I kind of look out at the canvas and feel not especially well protected at the moment.</span></p><p><strong><span>[29:30] Sarah</span></strong><span>: Yeah, me neither. I think maybe a lot of people don&#8217;t realise that with SB 53 and RAISE, there aren&#8217;t a set of things the companies have to do in the bill. They decide what they&#8217;re going to do, and then they just have to do the thing that they&#8217;re going to do. In theory they could be like, &#8220;Oh, every time we&#8217;re about to deploy a new model, we&#8217;ll say a prayer or sage the office&#8221; or something. But obviously you want to write something like a policy that seems good, otherwise you attract a bunch of scrutiny.</span></p><p><span>A couple of years ago I decided I was just going to read all of the safety policies as they existed at the time to see what was in them, because I didn&#8217;t get the sense a lot of people were reading the whole thing, so I did. And I was like, wow, it&#8217;s very easy to produce this extremely legitimate-sounding document that is actually extremely vague.</span></p><p><span>To your point on the backtracking as well, I noticed this the other day when I read that blog post that OpenAI released where they&#8217;re talking about responding to the next frontier of critical cyber capabilities or something. And they were like, &#8220;Oh yeah, we think Astra &#8212; we can&#8217;t rule out that it reaches our critical threshold, so we&#8217;re going to do all these extra mitigations.&#8221; But in the original preparedness framework, they said that if they thought a model was going to reach critical capabilities, they couldn&#8217;t develop it further until they&#8217;d brought that threshold down to high or lower. And that&#8217;s just not there anymore. Now it&#8217;s like, &#8220;We will just do the requisite things,&#8221; but then they get to decide what the requisite things are. And I just feel like it&#8217;s very easy not to notice this.</span></p><p><strong><span>[30:47] Steven</span></strong><span>: Not only that, but just to really underline it, OpenAI is not legally bound to its preparedness framework anymore. It was, but OpenAI and Anthropic have both published these frontier compliance frameworks where they&#8217;re basically like, &#8220;We used to be bound by law to follow the preparedness framework or the responsible scaling policy. We no longer are. We are instead bound by law to follow this much vaguer, more generic, way fewer commitments policy.&#8221;</span></p><p><span>And Thomas Woodside, who is the co-founder of the Secure AI Project, pointed out recently, I think for SB 53, OpenAI was required to describe how they are handling internal use risks, the types of things that happened in the OpenAI Hugging Face incident. And they wrote something to the effect of &#8220;Our policies and mitigations might apply to internal risk of our models,&#8221; and maybe one extra sentence, but basically no more detail than that. And it&#8217;s like, can we at least pretend to take this seriously? To be clear, I&#8217;m glad they&#8217;re doing a bunch more on the heels of this. And also it was just so predictable at some level, and it just really does not fly for the next era ahead.</span></p><p><strong><span>[31:57] Sarah</span></strong><span>: How did you feel about when Anthropic dropped the pause commitment from their responsible scaling policy, and they had this whole logic that was about binding yourself to the mast and why it&#8217;s bad to pre-commit to doing things in specific situations? I think the argument was something like, it&#8217;s bad for us to say that we&#8217;ll unilaterally pause, because us unilaterally pausing doesn&#8217;t do anything about our competitors that are maybe being less safe than us. It&#8217;s interesting now because in some sense OpenAI have kind of &#8212; well, they haven&#8217;t paused, but they&#8217;ve unilaterally slowed down. So it seems like that is in fact a thing that you can do, and maybe it would&#8217;ve been bad to commit to doing that in advance, but I don&#8217;t know. What did you think about that?</span></p><p><strong><span>[32:37] Steven</span></strong><span>: On the OpenAI point, certainly OpenAI is projecting something like &#8220;we are going slower.&#8221; I think the details are harder for me to parse, and I wouldn&#8217;t rule out that there is something kind of misleading going on. But I am happier that they are doing this than that they would be charging ahead. I would not be happy if they were charging ahead, and so I should give them credit for what they are doing, even if it&#8217;s a bit hard to understand the details.</span></p><p><span>On the Anthropic point, I don&#8217;t know. I think the fundamental thing is it is correct at some level that unilateral disarmament is not a very good strategy if you assume that your actions don&#8217;t have bearing on other people&#8217;s actions. And unless you work pretty hard to have your actions have bearing on other people, that is a reasonable default thing to believe about the world. And also, I just think that the AI companies could and should be trying way harder than they seem to be to be changing that state of affairs.</span></p><p><span>The Pacing the Frontier statement, or in the lead-up to that, the different CEOs of the AI companies starting to talk more in public about the dynamics that they are subject to and saying, &#8220;Well, we wish at some level that we could coordinate to not be cutting corners, but unfortunately, that&#8217;s not how the world works.&#8221; I think that that is a much more productive path forward, as opposed to just ruling out the possibility of unilaterally doing some things. At least making it common knowledge that you would like to take some of these more extreme actions. Maybe I&#8217;m drawing too cute a distinction between these at some level.</span></p><p><span>I have always felt that these frontier safety frameworks left companies too much discretion. I don&#8217;t think Anthropic was the first to do this by any means. I believe that Google&#8217;s frontier safety framework, for example, had always left an out. It&#8217;s almost like a poison pill. It&#8217;s like, if other companies are proceeding ahead and being unsafe, then why should we hold ourselves back? So if other people are doing things that would be in violation of our frontier safety framework, then we also reserve the right to violate our frontier safety framework.</span></p><p><span>Another thing about the Anthropic point: I guess I am happy that if they were not going to consider themselves bound to this, or didn&#8217;t want to be bound to it, that they made this clear. It would be worse to be relying on them doing this in some form, and then for them to throw it out the window right before taking an action. At least they are being clear about their intentions. And also, when you look more broadly, this is just clearly not good. If you think it would be safer to hold back for your actions specifically, we need to figure out how to shape the incentives of the world such that companies all do the things that they judge to be the safest and in the interest of the world. Not in a unilateralist perspective, but in terms of the ecosystem as a whole.</span></p><p><strong><span>[35:32] Sarah</span></strong><span>: Yeah, I agree. I have just a general frustration with &#8212; I remember reading the blog post where Anthropic announced that they were going to drop this commitment, and they framed it like, &#8220;Oh, well, a couple of years ago, maybe we thought we were in a slightly easier world where coordination looked more viable. But now it turns out actually we&#8217;re in hard mode because the Overton window has shifted&#8221; or whatever. And I just find this line of reasoning frustrating, because the Overton window is just the confluence of people saying and doing things. You could just say and do different things. It&#8217;s tricky to influence these things, but it&#8217;s just so self-reinforcing.</span></p><p><strong><span>[36:06] Steven</span></strong><span>: There are not that many people who you need to get to see an issue differently to meaningfully shift the Overton window. And I also find this sort of defeatism about it very frustrating, or people just asserting, &#8220;Oh, there is no way that you could ever bring China to the negotiating table. Or even if you do, they will always certainly cheat you.&#8221; And it&#8217;s, I don&#8217;t know, guys. International coordination is a thing. This has happened before. We have treaties. Yes, of course, you need to be robust to people trying to cheat the treaty. For sure, that is a real thing. And also, if safety ultimately depends on bringing people to the table and finding a stable agreement, you at least need to make a real effort, and I&#8217;m not convinced that the world has to date.</span></p><p><strong><span>[36:50] Sarah</span></strong><span>: Well, talking of international coordination, let&#8217;s move on to the Pacing the Frontier letter. Probably everyone has seen this already. But essentially the headline is that the open letter, which has been signed at this point I think by over 1,300 frontier lab employees, is requesting that the US government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. And I believe it was GuideLight and also Encode that were the two organisations instrumental in getting this off the ground. So I don&#8217;t know how much of the story you can tell here about how this came about, but I would be interested in anything that you can say.</span></p><p><strong><span>[37:33] Steven</span></strong><span>: I want to be careful not to speak for the signatories. I think it&#8217;s important that they get to describe their own reasons for signing this. And one thing that I really like that we did on the website is feature a roundup of quotes from signatories. So if people haven&#8217;t checked that out &#8212; pacingthefrontier.com &#8212; there&#8217;s tons of interesting stuff there where people said why they signed, what concerns they have, et cetera.</span></p><p><span>The overall dynamic as I see it is that staff at the frontier companies feel like there is some chance that they are on the cusp of something like recursive self-improvement. They are on the front lines. They see the pace of change as it currently stands inside the companies. And my perspective as an outsider, but I think consistent with what you hear if you talk to these folks, is there isn&#8217;t really a brake pedal. The pace at which things will continue to some extent right now is a property of the laws of physics. It is like, will they find a breakthrough that allows AI to train successively more capable successors or not? And if they do, there isn&#8217;t really much of a way across the industry to say, &#8220;Hey, there is some speed that is actually unsafe.&#8221;</span></p><p><span>It might actually be a speed faster than the default today. One really important thing about the statement that might not be obvious is when we talk about the ability to change pace or to do some form of coordinated pacing, this is different than necessarily slowing down. It is like applying some pressure to the brake, but actually you might have gone from 60 miles per hour to 200 miles per hour and you want to go back to 180 miles per hour. It might even be like, hey, you&#8217;ve gone from 60 miles per hour to 200 miles per hour, you don&#8217;t want to suddenly wake up and be going 600 miles per hour. You want to make sure there is some pressure at 300 or something like that.</span></p><p><span>Anyway, I&#8217;ve just been really heartened by how much this has been supported across the industry. It was signed by eight or nine chief scientists at different frontier AI companies, a really, really big coalition. I think some people saw this and they said, &#8220;Oh, this is OpenAI and Anthropic doing regulatory capture,&#8221; or because open-weights models have caught up and so of course they&#8217;re saying, &#8220;Oh, guys, we need to slow down.&#8221; I&#8217;m not sure exactly how to underline this, but that is just absolutely not what is happening here. And in fact, I think it is just understood across the frontier that the reality of the situation right now might be very, very dangerous, and that the world is not in position to respond appropriately in an emergency if we find ourselves in one.</span></p><p><strong><span>[40:06] Sarah</span></strong><span>: It is incredible, the conspiracy theories people will come up with to avoid the conclusion that people are in fact just saying, &#8220;Yeah, this is bad. We&#8217;re scared. We want to do something about it.&#8221;</span></p><p><strong><span>[40:15] Steven</span></strong><span>:I think there&#8217;s a fundamental thing where people are like, &#8220;If they&#8217;re so scared, why are they doing it?&#8221; And it&#8217;s, yeah, man, game theory kind of sucks. Sometimes it is in people&#8217;s local self-interest to do things that they think are globally very, very bad. But they in fact don&#8217;t have the ability to change it. I think it is in fact correct that if a lab decides they&#8217;re going to do a bit less on the margin, that does just possibly create a vacuum for other people. There&#8217;s more free energy floating about. Other people will utilise those GPUs. Staff will in fact flow to a company that is doing more of this scary experimentation. I think people are in fact correctly naming their local incentives. And I get at some level it&#8217;s, man, they should really just stop, and it&#8217;s like, well, yep, but they aren&#8217;t going to. And so I&#8217;m glad they&#8217;re communicating clearly about this, and then we need to figure out how to transform the incentives so that it&#8217;s in all of their interest to in fact damp it down a little bit.</span></p><p><strong><span>[41:13] Sarah</span></strong><span>: There&#8217;s this interview with Holden Karnofsky from &#8212; maybe it&#8217;s 80,000 Hours, I don&#8217;t remember. But anyway, he was talking about how there&#8217;s a common framing that AI safety is a coordination problem, in the sense that everyone involved really wants to stop, but they&#8217;re all too worried that another actor will do it less safely than them, so they&#8217;re compelled to keep going. And he was like, &#8220;I think this framing is wrong, because in fact, not everyone does want to stop. Some people do, and others don&#8217;t. And if all the people who wanted to stop stopped, that still wouldn&#8217;t solve the problem.&#8221; People disagree so much about the extent of the risk, or even about, ethically, what level of risk is worth taking. And so it&#8217;s not the case that if everyone believed the same thing, it&#8217;d be much easier.</span></p><p><strong><span>[41:52] Steven</span></strong><span>: I actually think that&#8217;s really important, and sometimes I talk to friends about this and they say some variation of, &#8220;Yeah, why &#8212; if they believe this, why don&#8217;t they just stop?&#8221; And you&#8217;re right to point out that actually they have quite non-uniform beliefs. And even among people who think the situation might be very dangerous, they put different probabilities on this. They put different timeframes on this. And so part of what I think is so great about Pacing the Frontier is it seems to be, in fact, a big-tent enough thing that people with a wide variety of beliefs can still recognise: yep, the status quo is insufficient. It is very bad that we have no way to do this sort of pacing, at least in an emergency. And that can be a range of different things, and there are lots of ways to move along that gradient that give the industry more coordinated ability than it has by default.</span></p><p><strong><span>[42:43] Sarah</span></strong><span>: So maybe the first question someone might have when they read the statement is, what does &#8220;pace&#8221; mean? For example, the AI Futures Project put out a blog a few days ago, maybe a few weeks ago, where they came up with at least four different ways you could define this word. So it could be anything from you literally do a temporary pause where all of your compute is going to inference, or it could be only a certain percentage of your compute is allowed to go to training, or it could be requiring safety cases to scale to some new level or whatever. And you can imagine there are many, many ways you can interpret this word, and many of them might be not very risk-reducing. I think it actually makes a lot of sense to make the wording sort of broad enough that you can get a lot of sign-on. But if someone was to criticise the letter for being too vague, what would you say to that?</span></p><p><strong><span>[43:44] Steven</span></strong><span>:I guess to take a step back, the fundamental thing that is happening with pacing is there is an idea that the amount of risk entailed by an AI system is related to its capability level, and to some extent how understandable and supervisable that system is. And that there is some natural trend line where over time, companies develop more capable systems, and in fact might develop systems that are less supervisable as well. To some extent because a more capable system might be harder to supervise &#8212; a toddler would struggle to supervise a university professor in all sorts of different ways. You can argue about whether that&#8217;s the right analogy, but that&#8217;s the sort of intuition. And also because if you turn over your AI development process to one of these AI systems, it might just move much more fast than you expect. The trend line might change. It might find the sorts of changes that make a system exogenously harder to supervise, regardless of the capability level. It might just push it in a more opaque, uninterpretable sort of way.</span></p><p><span>And so the question is, can you alter that trend line in some form? We joked about, well, is it the first derivative? Is it the second derivative? Is it the third derivative? What actually is the pace of change that you are trying to affect? But fundamentally, there is some question about how soon do you run into systems that are really, really capable and maybe really, really hard to supervise?</span></p><p><span>I think what the AI Futures Project is pointing to is there are a bunch of different levers available of how you might change the pace. But fundamentally, the thing that you are trying to change the pace of is the rate of frontier expansion into systems that are potentially very dangerous and very hard to supervise.</span></p><p><span>I&#8217;m not sure what lever I personally think is best. I think one really important property of some sort of pace change is you kind of have to understand the extra time that you are buying, in some sense, by changing the pace. What is it that you want to ultimately accomplish in that time? The AI Futures Project, in their Plan A writeup, this AI 2040 scenario &#8212; I like that they have ideas of what you allocate this time toward in terms of maybe it is research focused especially on safety that helps you ultimately solve the underlying alignment problems. Definitely I don&#8217;t think buying time is enough on its own.</span></p><p><span>I also think another important property to think about is how stable is this regime to defection? And if people eventually do defect, are you in a better or worse place afterward? But beyond that, I don&#8217;t know that I&#8217;m especially opinionated about the way of doing this. I just want people to at least be having the dialogue about what the range of options is, and for there to be a real meeting of the minds about, &#8220;Okay, guys, if we had to do this in the next six months, what would that look like? What are some of the nasty trade-offs involved? How do we make them bite less hard?&#8221; Really grasping the scale of the problem, and that we might need to take action on it very soon.</span></p><p><strong><span>[46:36] Sarah</span></strong><span>: And do you have a sense of, in policy circles, what the reception has been of this letter? Do you know if it&#8217;s getting in important rooms, or how it&#8217;s being received?</span></p><p><strong><span>[46:48] Steven</span></strong><span>: I&#8217;ve been very, very happy with the reception of it. It feels like pacing is the word of the year, or at least the word of the season in many of these policy circles. I think there is some benefit there of people meaning different things by it. OpenAI yesterday, in their publication about how they are pausing RL training for a bit, and they&#8217;re beefing up their monitoring and all these things &#8212; I think the title was like &#8220;Pacing Model Development in the Era of Cyber Security Challenges.&#8221; Pace, right? Not pause. It is nice that it has a less definitive meaning in some sense. You want it to not get too diluted. But for now it does seem like it has really gotten traction. And there are members of Congress in the US writing statements about it. It is getting cited a lot in media.</span></p><p><span>I&#8217;ve been very happy that it seems to have made common knowledge a thing that was kind of common knowledge in my circles for a long time, but maybe not citable in prestige media, or maybe people didn&#8217;t really know to what extent this was a widely held view &#8212; that the current pace of progress might in fact be very dangerous and that there isn&#8217;t really a way to hold it off, and we want to at least have the option for such a thing. I&#8217;ve been glad to see that get more recognised as a real thing.</span></p><p><strong><span>[48:03] Sarah</span></strong><span>: I was very happy to see it. I feel like so few good things seem to happen nowadays. When I saw this, I was like, &#8220;What a dopamine rush.&#8221;</span></p><p><span>One other pushback &#8212; I don&#8217;t think I would make this pushback, but you might say, &#8220;Why is this letter asking for us to develop the tools we need to do something later? Clearly the situation is so urgent, we should just be doing something now.&#8221; I think it&#8217;s like, well, we don&#8217;t in fact have the tools, so either way we kind of have to do it. But what would you say to someone who thinks that we need to be doing something a bit more like, say, the 2023 pause letter, where they were like, &#8220;Just pause now and figure everything else out later&#8221;? Is there anything to that argument?</span></p><p><strong><span>[48:44] Steven</span></strong><span>: I was in a message thread with Oliver Habryka on Twitter about this, and I think he made a fair point that the statement might have implied something too much in the direction of new things being necessarily needed. And that there is in fact a range of existing governance tools that people could use today that might be reasonably enough. From my perspective, we clearly do not have a treaty that governs this, but actually a treaty is a well-worn governance tool that covers a wide range of things and might in fact be a meaningful step forward, and certainly I didn&#8217;t mean to imply something broader &#8212; that a treaty is useless, that you need some new super-treaty technology or something like that.</span></p><p><span>To some extent, the statement is about changing the Overton window, and where there are precursors that are needed for action, trying to set those into motion. So to your point, before you can actually pull a lever, you need to know what is the set of levers. You need to have a lever that plausibly has the effect you want. And I&#8217;ve just been glad to see more conversation about &#8212; even if it&#8217;s people pointing to &#8220;actually, you could do this today with existing technology,&#8221; that is incredible from my perspective. I would love more energy focused on what would it look like to actually do this in the immediate term.</span></p><p><span>I&#8217;m also excited about people thinking about what is the technology roadmap needed to accomplish something like this in, say, 18 months. What sorts of new verification and monitoring technologies would be helpful? But certainly I am not of the belief that it is only possible in the future. That might be the belief of some signatories &#8212; again, I don&#8217;t want to speak for people in that regard &#8212; but my view is there is useful stuff that could be happening now, and I&#8217;m glad that more people are thinking about it.</span></p><p><strong><span>[50:31] Sarah</span></strong><span>: Yeah, I agree. It seems like the one thing that we definitely need to do more technical work on are these hardware verification technologies, so that if there was an agreement between China and the US, for example, everyone could trust that the requisite parties were holding up their end of the agreement without having to actually trust each other. And I&#8217;ve sort of lost track of exactly what is happening in that space now. I remember a couple of years ago I wrote a piece with &#8212; I think it ended up being on AI Frontiers. So there are these FlexHEG things that, at the time, they were like, &#8220;We could build a prototype of this in a year.&#8221; That was more than a year ago, so I don&#8217;t know if that happened. But they could make it possible for you to enforce any rule you wanted on the chip, and for people to be able to verify that that was happening. And these are all flashy, ambitious ideas, and I don&#8217;t really know whether any of them are panning out. But it does make me wonder, do we need this crazy stuff? Or is there a much simpler way that we could do this?</span></p><p><strong><span>[51:23] Steven</span></strong><span>: I don&#8217;t know. The person I would ask on this is Yo Shavit, who used to be a teammate of mine at OpenAI and now works at the OpenAI Foundation. He tweeted about a breakthrough maybe a week ago that he seemed very, very excited about in terms of the ability to do verification today, and I confess I haven&#8217;t gotten to look into it in more detail since I&#8217;ve been kind of underwater on the launch of our control scorecard. But he would be a great person to talk to more about this, and I should really send Yo a message and try to understand how close exactly are we to this sort of thing.</span></p><p><strong><span>[52:03] Sarah</span></strong><span>: Yeah, this is another thing that I&#8217;ve been meaning to catch up on, but there&#8217;s just so many things to do. Okay &#8212; is there anything else that we want to say on the pacing letter? I mean, mostly I&#8217;m just very excited about this. I think it&#8217;s very cool, and I hope we do some pacing soon. That would be good.  I saw someone say something like &#8212; maybe it was Helen Toner that said this &#8212; she was like, &#8220;If everyone just did what OpenAI have just done, then the frontier is paced, and that&#8217;s it. That&#8217;s what we have to do.&#8221; I feel like it doesn&#8217;t seem so good if there is no government involvement and it&#8217;s all just companies voluntarily doing things. But maybe we&#8217;re somewhat heading in that direction.</span></p><p><strong><span>[52:33] Steven</span></strong><span>:I saw that message as well. I don&#8217;t think I agree. Helen&#8217;s point was that maybe actually there&#8217;s a form of pacing which is just, if every company in fact recognises when they are about to do something unsafe and they hold back for enough time to not do the thing that is unsafe, then you have paced the frontier. And I think the fundamental issue here is you cannot reasonably expect every company to look and assess correctly whether the thing in front of it is unsafe, and to hold off for necessarily the right amount of time, because their local incentives are different. This is fundamentally the issue of companies cutting corners when they aren&#8217;t required to do certain things by law.</span></p><p><span>So yeah, if every company for sure had a clear-eyed assessment of the risk in front of them and followed a heuristic like &#8220;never in fact do a thing that is unsafe,&#8221; you have achieved a lot of the benefit, but we just very clearly don&#8217;t live in that world. And I think right now, the question is, there is actually a wide range of safety postures among the frontier companies. People are kind of hard on OpenAI at times because they are one of the leaders and are not necessarily abiding by a high enough standard. To be clear, I think this is true of Anthropic as well. I think both OpenAI and Anthropic do not have strong enough control systems in place. I think they are very far from what an adequate control regime would be.</span></p><p><span>Also, if someone looks at our scorecard, boy, are OpenAI and Anthropic way ahead of where xAI and Meta are, and to some extent they have more capable models. And so some degree of differences may be appropriate. The amount of risk is a function of how good your mitigations are, but also how capable your model is. At the same time, there&#8217;s just a lot of basic stuff that OpenAI and Anthropic, to their credit, have done and other companies have not done, and I&#8217;m just not sure that we can reasonably expect them to do it in the near future absent some stronger intervention.</span></p><p><strong><span>[54:32] Sarah</span></strong><span>: I guess we should pivot to talk more about the control scorecard, which I confess I haven&#8217;t read in detail, but it&#8217;s sort of ranking how good the safety and monitoring practices of various companies are. Do you maybe want to give a little spiel &#8212; what are the headlines? I guess you already kind of said who&#8217;s doing better and who&#8217;s doing worse, but what are the main things you&#8217;d want someone to take away from this?</span></p><p><strong><span>[54:53] Steven</span></strong><span>: Well, let me back up for a moment. My organisation GuideLight published a standard on AI control in May. This was drawn from input from people across the AI labs, across the top safety nonprofits in the space, and the idea was to articulate something like: what are the consensus minimum practices that, if companies were to follow them on control, you would greatly reduce the risk?</span></p><p><span>The organisation of the standard as a whole &#8212; there are certain evergreen principles that we think are important to aim for. These are things like: know what your AI is doing internally; scan what the AI is doing for signs of misbehaviour; make sure that the AI can&#8217;t accomplish something very dangerous, even if it is trying to. And then a level down from those principles are specific practices that we think are achievable today and are effective ways of realising those principles. They&#8217;re not sufficient, but they&#8217;re a good step forward.</span></p><p><span>So for example, you can say broadly you should know what your AI is up to internally, but what does that mean in practice? Well, one thing it means in practice is logging the vast, vast majority of your internal AI usage to a place where you can scan it for signs of misbehaviour. So one of the practices we have is about logging at least 99.9% of your internal AI usage of capable, risky models to a place where you can do that scanning.</span></p><p><span>So we published that in May. It was well received. People seemed excited about it. But the question, of course, is: well, are companies actually following this? And actually, can you even tell whether companies are following this? Have they made specific enough claims about their safety practices that you even know whether they are abiding by the standard?</span></p><p><span>And so for the last two months or so, maybe two and a half months, what we&#8217;ve been doing is going through every public document that the frontier AI companies have written about their safety practices, and scoring them from zero to five on how fully they follow these most important practices in control. Do they do this internal logging? Do they seem to know how accurate their monitoring is &#8212; what types of misbehaviour is it missing? Do they document this? Do they have actual preventative systems in place, the types of things I&#8217;ve referred to? Do they have a containment plan? If they catch their AI misbehaving, do they actually know how they would cordon it off from their internal staff? Who could still safely use it in those cases, for what purposes? Because maybe you need to use it as part of determining whether it&#8217;s safe to proceed again after all. You can&#8217;t just turn it off for everyone. That doesn&#8217;t really give you an effective way.</span></p><p><span>And so what we published yesterday is this aggregation of all this evidence across many, many dozens of very long reports that produces a numerical value of how well the companies are doing. And the top line is, across all the companies, there&#8217;s a remarkable amount of progress still needed. On our zero-to-five scale, there&#8217;s no practice where any company gets above a three. Even the threes are not very common. Most of the scores are two or below. There&#8217;s some bright red spots where companies are getting zero out of five, because even though these things are very basic and agreed upon to be necessary for really reining in the risk, it just might not be incentive-compatible to invest heavily enough in safety to implement these today.</span></p><p><span>And so part of what we are trying to do at GuideLight is encourage companies to make more of that investment in safety, and make it known that in fact there are basic things still not being done across the industry. Some companies are doing better than others. We think that&#8217;s great. We want to celebrate that. But ultimately, my goal is that by making it known who is doing what to what degree, we see the scorecard ultimately be filled with bright green instead of the yellow, orange, red that it is across the board today.</span></p><p><strong><span>[58:36] Sarah</span></strong><span>: That would be great. That really sounds like extremely hard work, to go through all those documents. That is God&#8217;s work. Is there more than one intended audience for this? It sounds like mainly you&#8217;re hoping that the companies change their behaviour based on this &#8212; they see they got a D-minus or whatever, and they&#8217;re like, &#8220;We want to do better.&#8221; But I imagine there&#8217;s some kind of public accountability angle to this as well, where you want people to see this wash of red and orange and understand the companies are not &#8212; and maybe policymakers might also see this. There&#8217;s a lot of theories of change there, I guess.</span></p><p><strong><span>[58:56] Steven</span></strong><span>: For sure. The primary theory of change is convince companies to voluntarily do these things, because by and large they are not required by law today. These are at the discretion of a relatively small number of people within the company. I really believe in this framework that Buck Shlegeris published a few months back called &#8220;Ten People on the Inside,&#8221; which basically makes the point that, at least at OpenAI and Anthropic, the way that a lot of things happen is there is some internal champion who cares enough to get it done. And actually, a lot of the time people don&#8217;t really object to someone doing this. They just don&#8217;t care quite enough to do a safety-oriented specific thing themself, or it&#8217;s just kind of tedious and you run into a lot of friction, and a lot of people are going to tell you no at first or tell you maybe, and you just have to be really motivated to run through walls.</span></p><p><span>Part of the hope is that our publishing this gives those safety champions inside the company a clear target to aim for that is vetted industry-wide as an important thing for them to do, and they feel more motivation to run through those walls and try to get their company to do these things. And in fact, they can make the case to other stakeholders inside their company: &#8220;Look, we are orange on this thing right now. In fact, we are not even doing it the best of anyone. It wouldn&#8217;t be very hard for us to do a bit better on this front. I&#8217;m willing to put in the legwork. Can we actually go and do it?&#8221; And make it a bit easier to get yeses from those people, or to understand why this is an important thing to aim for.</span></p><p><span>But to your point, the companies are doing this voluntarily. They mostly have their hands on the lever, but it is of course impacted by the outside world. And so things like the media finding this a credible summary of how the companies are doing, or treating it as important context for when a company announces improvements in one of their practices &#8212; how good were they before? After these changes, how good are they now? Really understanding the shape of the problem the way that we do, I think is really important, and for helping companies to understand this stuff matters. People are going to pay attention to it, and you will be really celebrated for doing better on it. But also, if your company is falling behind, that is going to be noticed as well.</span></p><p><strong><span>[1:01:15] Sarah</span></strong><span>: How often do you plan to release new scorecards?</span></p><p><strong><span>[1:01:19] Steven</span></strong><span>: It&#8217;s such a good question. As often as we can, in terms of the capacity of the organisation. I am not sure exactly what that means. In addition to this control scorecard, we are working on publishing our standard for alignment. We are likely to publish a standard related to automated AI R&amp;D development and recursive self-improvement management. There&#8217;s a lot of things going on, and we&#8217;re still a pretty lean team. But the aspiration is to get to something that is continually updated, as tight a feedback loop as possible between a company improving their practices and it being recognised publicly, and this being a thing to celebrate and to increase the relative pressure then on their peers to do this thing as well.</span></p><p><strong><span>[1:02:03] Sarah</span></strong><span>: Very exciting. I know that the Future of Life Institute did something similar to this with a scorecard. But I don&#8217;t know if it was as rigorous as the one that you&#8217;re doing. It seems like &#8212; I don&#8217;t know if that ever sort of had the effect of moving people and companies to do things differently. It seems like maybe theirs was a little bit more public-facing.</span></p><p><strong><span>[1:02:24] Steven</span></strong><span>: There are a few groups doing things like this. The Future of Life Institute is one. They have a safety index that looks pretty similar to ours in terms of the scorecard and the grades. There&#8217;s also SaferAI, which does work like this, rating companies&#8217; risk management practices. AI Lab Watch, of course, was operating, I think, through last fall, but is now defunct.</span></p><p><span>Overall I&#8217;m really glad that more people are doing this. We don&#8217;t want to be the only ones in this space. I think that this is clearly very undersupplied by the ecosystem in general &#8212; these sorts of &#8220;have a thoughtful rubric about what it means to do safety well and actually put in the legwork to score companies against it.&#8221; And so I&#8217;m really glad that those efforts exist.</span></p><p><span>In our case, it is nice to be able to zoom in on the particular practices that we think are highest leverage and most important for the companies to do better on. And I worry that if you have a very broad set of things, your influence is just very diffuse, and you don&#8217;t necessarily direct attention to the highest-impact things for companies to change. So obviously there&#8217;s a balance there. There&#8217;s a lot that our effort does not include right now. For example, it&#8217;s only about control. We don&#8217;t yet have the alignment standard, so that isn&#8217;t included, and we need to figure out how to incorporate that over time. But it&#8217;s really nice. I&#8217;ve been glad to see the reception to the work so far.</span></p><p><strong><span>[1:03:44] Sarah</span></strong><span>: I think I&#8217;ve covered most of the topics I planned to talk about, so maybe we should wrap up soon. But is there anything that you would want to leave people with, or anything you want to tell people about what you&#8217;re up to next? The floor is yours to say anything you haven&#8217;t said yet that you think is important.</span></p><p><strong><span>[1:03:29] Steven</span></strong><span>: If people are excited about what we are doing and you&#8217;ve listened this far into the podcast episode &#8212; one, thank you. Two, you should message me. You can find me on Twitter, I&#8217;m sjgadler, or LinkedIn. Maybe there&#8217;s a role for you at GuideLight. We certainly need more help, and people who care about these risks and can communicate about them thoughtfully and synthesise what&#8217;s really important. That would be really great.</span></p><p><span>Also, people should check out our scorecard on our website, which is guidelight.ai. Just click the button for the blog, and we would love any feedback. I&#8217;ve been really happy that people have said nice things about it so far, but what ultimately matters is that this is actually useful and informative. And so if you see something that doesn&#8217;t quite work or doesn&#8217;t make sense to you, please let us know, because it&#8217;s really important we improve that for the future. But overall, I&#8217;m just really glad that people care about this stuff and seem aligned with this general vision of: set a target for companies to aim for, measure it, and hopefully push more in that right direction.</span></p><p><strong><span>[1:04:58] Sarah</span></strong><span>: Well, thank you so much for your time. This was a really fun conversation.</span></p><p><strong><span>[1:05:04] Steven</span></strong><span>: Yeah, of course. Thanks for having me.</span></p>]]></content:encoded></item><item><title><![CDATA[Meeting the moment ]]></title><description><![CDATA[Or failing to.]]></description><link>https://longerramblings.substack.com/p/meeting-the-moment</link><guid isPermaLink="false">https://longerramblings.substack.com/p/meeting-the-moment</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Tue, 18 Aug 2026 17:02:19 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8ebaa501-317e-4ba3-b323-021b9aeaf8d8_1174x1096.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Edit (19/08/2026): Quite literally an hour after I posted this, OpenAI announced that they had &#8220;temporarily slowed the pace of scaling&#8221;. This involves two things: </em></p><ol><li><p><em>They paused reinforcement learning (which is a phase of post-training) for two weeks on some of their latest models intended for deployment. </em></p></li><li><p><em>Their latest frontier RL run is currently on hold while they assess safeguards and evidence of alignment.</em></p></li></ol><p><em>So I guess this answers the question I pose below, as to whether OpenAI in fact had the &#8220;ultimate goal&#8221; of slowing development in response to the HuggingFace incident and the (possibly) critical capabilities of their unreleased Astra model. Not OpenAI have not paused <strong>all</strong> training, just one particular phase of post-training for a forthcoming model. It isn&#8217;t clear what criteria would have to be met for this training to resume; it&#8217;s up to OpenAI to decide. <a href="https://x.com/ohabryka/status/2089932606414758317">Not everyone is convinced</a> this will buy us much risk-reduction. Still, this is better than I expected. I think we should commend OpenAI for doing this.</em></p><p><span>Over the last few weeks, we&#8217;ve witnessed an unprecedented flurry of AI misbehaviour. OpenAI kicked off the trend in late July by </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>sheepishly announcing</span></a><span> that several of its models had hacked into HuggingFace to steal the answers for a cyber evaluation. This prompted Anthropic to review its cybersecurity evaluation transcripts dating back to April, and discover that sandbox configuration errors had led Claude to </span><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span>compromise the infrastructure of three separate organisations</span></a><span>. Then, the UK AI Security Institute </span><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"><span>chimed in</span></a><span> to let us all know that, during evaluations where the models had deliberately been granted internet access and had some other safeguards removed, both Mythos 5 and GPT-5.6 had attempted all manner of devious actions &#8211; such as inserting malicious code into an open-source project, convincing humans to run said malicious code, and posting to GitHub in search of other AI collaborators for their evil schemes. Not to be outdone, Meta then disclosed that its Muse Spark 1.1 model had </span><a href="https://aiweekly.co/alerts/metas-muse-spark-11-breached-a-real-site-in-botched-test"><span>gained unauthorised access</span></a><span> to another company&#8217;s infrastructure after being </span><em><span>accidentally</span></em><span> granted internet access, during a capture-the-flag evaluation run by Irregular (the same organisation that had designed the misconfigured sandboxes responsible for the three Anthropic escape incidents mentioned above).</span></p><p><span>Anecdotally, AI scepticism seems to be somewhat receding. Formerly measured commentators are expressing </span><a href="https://x.com/TomDavidsonX/status/2086207606410871026"><span>new levels</span></a><span> of alarm. P(dooms) are </span><a href="/__u/thelimestack.substack.com/p/my-pdoom-is-1742-heres-why?utm_source=multiple-personal-recommendations-email&amp;utm_medium=email&amp;triedRedirect=true"><span>being revised upward</span></a><span>. A </span><a href="https://www.pacingthefrontier.com/"><span>new open letter</span></a><span> signed by over 1,300 frontier lab employees calls for intervention from the US government in order to &#8220;pace&#8221; the frontier of AI development, in what many are interpreting as a cry for help from participants in a race to the bottom that they desperately want to temper.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>For as long as I have observed the AI safety conversation, there has been a contingent of worried-but-pragmatic people who have warned against ringing the fire alarm too early. They have argued for waiting until the risks of AI are more obvious and salient, or until a warning shot galvanises everyone around the goal of disembarking from the crazy train, lest we cry wolf or burn political capital. Those people are much quieter now. We are, slowly but surely, cohering around the position that it has come time to Meet The Moment. But are we succeeding?</span></p><h3><span>Things that do not meet the moment</span></h3><p><span>Many responses to the recent AI misbehaviour epidemic felt lacklustre to me.</span></p><h4><strong><span>&#8220;Erring on the side of caution&#8221;</span></strong></h4><p><span>OpenAI, the first penitent in the cascade of AI misbehaviour revelations, have responded to it in a way that I find altogether disappointing. First, they </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>announced</span></a><span> the cybercrime their models had committed against HuggingFace in what I felt was a fairly muted tone, to inexplicable fanfare from a who&#8217;s who of AI safety figures for being so gracious as to disclose it at all (see my </span><a href="/__u/longerramblings.substack.com/p/we-have-been-warned"><span>last blog post</span></a><span> for more salty commentary on this).</span></p><p><span>A couple of weeks later, OpenAI released </span><em><a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/"><span>Responding to the next frontier of critical cyber capabilities</span></a></em><span>, which details the measures they are taking to mitigate the cyber threat posed by Astra &#8211; an internally deployed model that was reportedly not involved in the HuggingFace hack and is presumably even more capable than the ones that were. OpenAI &#8220;cannot rule out&#8221; that Astra possesses critical cyber capabilities as defined in its </span><a href="https://openai.com/index/updating-our-preparedness-framework/"><span>Preparedness Framework</span></a><span>. I hoped this blog might outline some drastic and moment-meeting interventions, since the &#8220;next frontier of critical cyber capabilities&#8221; promises to be even more terrifying than models already capable of wantonly hacking into the infrastructure of other companies. OpenAI employee Boaz Barak </span><a href="https://x.com/boazbaraktcs/status/2085772335844556810"><span>shared the blog post on Twitter</span></a><span>, expressing pride that his company is &#8220;erring on the side of caution and taking the steps so [it] can responsibly and safely develop Astra and share it with defenders&#8221;. </span><a href="https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks"><span>Media</span></a><span> </span><a href="https://www.theguardian.com/technology/2026/aug/08/openai-astra-security-concerns"><span>reported</span></a><span> that OpenAI had vowed to &#8220;slow development&#8221; of Astra to allow for strengthened safeguards.</span></p><p><span>It&#8217;s hard to tell from OpenAI&#8217;s blog post to what extent they have, in fact, &#8220;slowed development&#8221; of Astra over cybersecurity concerns. They say they have &#8220;paused internal activities involving Astra&#8221; that do not meet &#8220;strengthened security control requirements&#8221; such as enhanced monitoring and restricted tool access. Presumably, this means that post-training, elicitation, and internal deployments of Astra are currently being done with more constraints. This would have the </span><em><span>effect</span></em><span> of stalling the model&#8217;s development somewhat &#8211; but nothing about the blog appears to imply that unilaterally slowing down is OpenAI&#8217;s ultimate goal. It&#8217;s not clear that OpenAI are doing anything other than bare-minimum adherence to their Preparedness Framework in response to developing a model with critical cyber capabilities, which probably doesn&#8217;t warrant breathless congratulation for having displayed the wisdom to &#8220;slow down development&#8221;. Which brings me onto my next point.</span></p><p><span>OpenAI&#8217;s response to Astra&#8217;s cyber capabilities is yet another example of the backsliding on safety commitments that is a common feature of the feverish race to build superintelligence. In the </span><a href="https://cdn.openai.com/openai-preparedness-framework-beta.pdf"><span>original version</span></a><span> of their Preparedness Framework, released in 2023, OpenAI committed not to further develop any model predicted to reach &#8220;critical&#8221; capability in any domain until its risk threshold could be reduced to &#8220;high&#8221; or lower:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XEM5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 424w, /__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 848w, /__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XEM5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png" width="876" height="272" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74d4d955-15c6-418c-9342-4341da664143_876x272.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:272,&quot;width&quot;:876,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 424w, /__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 848w, /__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XEM5!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d4d955-15c6-418c-9342-4341da664143_876x272.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>But a </span><a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf"><span>revised version of the Framework</span></a><span>, released in 2025, waters down this commitment:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ydDZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 424w, /__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 848w, /__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ydDZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png" width="265" height="218" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3cd73d8-c400-498f-b025-462ec3439633_265x218.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:218,&quot;width&quot;:265,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:20353,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 424w, /__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 848w, /__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ydDZ!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cd73d8-c400-498f-b025-462ec3439633_265x218.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>So OpenAI is no longer required to halt development of a system predicted to reach critical levels of cyber capability until its risk category can be lowered. It&#8217;s up to OpenAI to &#8220;specify safeguards and security controls standards&#8221; that would prevent a critically capable model from causing mayhem, either while internally deployed or when publicly released &#8211; which, as we can see from the HuggingFace incident, doesn&#8217;t seem to have been going very well so far, even for models less advanced than Astra. And in </span><a href="/__u/longerramblings.substack.com/p/i-read-every-major-ai-labs-safety"><span>classic AI safety-plan-fashion</span></a><span>, OpenAI does not publicly detail what these safeguards will be in the Framework. It also doesn&#8217;t provide anything more than high-level examples in their Astra-specific blog post. So, needless to say, I would not describe OpenAI&#8217;s response to the present situation as anything like &#8220;erring on the side of caution&#8221;.</span></p><h4><strong><span>Oops, we accidentally made a weed</span></strong></h4><p><span>Another response to the recent flurry of AI misalignment incidents that</span><strong><span> </span></strong><span>irked me</span><strong><span> </span></strong><a href="https://x.com/deanwball/status/2085937149992448311"><span>comes from</span></a><span> prominent AI policy commentator and newly-hired Head of Strategic Futures at OpenAI, Dean Ball. I have generally appreciated Dean&#8217;s contributions to The Discourse, but this one made my head spin:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!62iB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 424w, /__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 848w, /__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 1272w, /__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!62iB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png" width="968" height="392" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:392,&quot;width&quot;:968,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 424w, /__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 848w, /__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 1272w, /__u/substackcdn.com/image/fetch/$s_!62iB!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ebf3b0e-24a7-4513-a1c7-81aa5a802f3e_968x392.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Dean is a beautifully evocative writer. I wish I could string a sentence together half as well. The problem with these perfectly-composed metaphors is that they are wonderfully, dangerously anaesthetising. I feel like I could AI-generate endless pages of Dean-style musings on our narrow path towards a possibly-not-terrible AI future and have Claude read them aloud to lull me to sleep at night. On the object level, however, what Dean is saying here isn&#8217;t actually all that reassuring. &#8220;Gardeners&#8221; have far less control over their creations than &#8220;sculptors&#8221; &#8211; and, </span><a href="https://x.com/deanwball/status/2086526273312997886"><span>as Dean himself acknowledges</span></a><span>, that &#8220;pro-social machine ecologies&#8221; are theoretically possible doesn&#8217;t mean we presently have any idea how to develop them.</span></p><p><span>I don&#8217;t know exactly what being OpenAI&#8217;s Head of Strategic Futures entails, but I imagine it puts Dean in a position to significantly influence the company&#8217;s frontier AI policy. This is why it frustrates me to see him post what reads like an abdication of responsibility for the future of AI development. I can only hope OpenAI&#8217;s &#8220;strategy&#8221; isn&#8217;t to let the seeds of superintelligence germinate, and then sit back and hope the resulting system doesn&#8217;t nuke the world. I know that, in reality, Dean is Just Posting Things in a personal capacity, and that we shouldn&#8217;t take &#8220;let a thousand AI flowers bloom and hope that none of them are world-eating weeds&#8221; as official OpenAI strategy. But The Discourse is important, and is ill-served by elegant meditations that downplay humanity&#8217;s own agency in the development of superintelligence &#8211; or, as </span><a href="https://x.com/wfithian/status/2086044716315840792"><span>one commentator puts it</span></a><span>, encourage the reader to &#8220;stop worrying and learn to love the bomb&#8221;. I find this flagrantly irresponsible. The moment is not being met.</span></p><h4><strong><span>Disclosure-by-conference</span></strong></h4><p><span>This brings me to another gripe with how we came to learn of the HuggingFace drama in all its terrible detail. A few weeks after the original disclosure, two of the company&#8217;s engineers </span><a href="https://www.youtube.com/watch?v=87DyyMV0kCY"><span>gave a talk</span></a><span> at Vegas security conference Black Hat, in which they described the saga in full. They revealed an even crazier story than had been accessible via OpenAI&#8217;s official announcement or anywhere in the mainstream media: an unreleased model had hacked internal systems in order to build a secret message board that it had used to communicate with other instances of the model, in order to coordinate the attack. OpenAI had detected and shut down an initial attempt at this deviousness, only for the model to recreate the message board using yet more sophisticated methods. The attack was ultimately carried out by what the engineers described as a &#8220;swarm&#8221; of models working in unison.</span></p><p><span>My first response to seeing various commentators post their Twitter reactions to this talk was one I often have when learning about the latest batshit-crazy thing to occur in the AI safety ecosystem: why am I just learning about this </span><em><span>now</span></em><span>, and why </span><em><span>here</span></em><span>? The scariest details of this incident were disclosed not in any official output, but at a conference for security researchers. The talk is opened by an OpenAI alignment researcher, who introduces his co-speaker as &#8220;Mike from Security and Infrastructure&#8221;, and says he will be going on to discuss &#8220;one of the most qualitatively interesting examples of AI capabilities [he has] ever seen&#8221;. Maybe I&#8217;m being unreasonable here, but does anyone else feel like there&#8217;s a missing mood? Why am I hearing Eric From Alignment and Mike From Security describe our clearest warning shot yet on the path to AI-caused human extinction as </span><em><span>interesting </span></em><span>in a YouTube video of a conference briefing that would never have found its way into my algorithm had I not been so chronically online? To be fair, we don&#8217;t know whether policymakers are receiving behind-the-scenes briefings on the information provided in the BlackHat talk, and a full incident report from OpenAI is </span><a href="https://x.com/tszzl/status/2085926731840696429"><span>allegedly forthcoming</span></a><span>. But this drip-feed of information about what may shake out to be one of the most consequential moments in human history doesn&#8217;t meet my vibe check.</span></p><p><span>I&#8217;m not annoyed at OpenAI themselves per se. They hustled to put together a talk for a security conference on a tight deadline, in which they were forthcoming with all the sordid details of the internal AI coup that happened on their watch &#8211; even when this didn&#8217;t put them in the best light. And they are apparently working on an official write-up of events. But I think that in a world that was Meeting The Moment, we&#8217;d have had transparency requirements in place long ago that would mandate the disclosure of this information to the government and to the public. There wouldn&#8217;t be room for the kind of incredulity </span><a href="https://x.com/So8res/status/2087905494203605168"><span>apparently expressed</span></a><span> by New York Times fact-checkers who were reviewing an op-ed authored by MIRI&#8217;s Nate Soares &#8211; to which he had to respond with, &#8220;I swear it&#8217;s true, please refer to this timestamped section of a YouTube video&#8221;. The chaotic, random way in which knowledge of the HuggingFace hack has trickled into the public domain feels like the culmination of many failures to meet many moments, dating back years. </span><a href="https://www.lesswrong.com/posts/j9Q8bRmwCgXRYAgcJ/miri-announces-new-death-with-dignity-strategy"><span>This isn&#8217;t what surviving worlds look like</span></a><span>.</span></p><h3><span>Things that are getting there</span></h3><p><span>With all this said, we should give The Discourse credit where it is due. In some corners of the internet, it has been reaching a fever pitch that seems to at least approach rising to the occasion. A July open letter, </span><em><a href="https://www.pacingthefrontier.com/"><span>Pacing the Frontier</span></a></em><span>, garnered over 1300 signatories from frontier AI company employees urging US government intervention that will &#8220;support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development&#8221;. One could criticise this phrasing for not going far enough; it calls for the &#8220;tools needed&#8221; to do something </span><em><span>later</span></em><span> rather than for action </span><em><span>now</span></em><span>, so we might reasonably worry about the deferral of any intervention until after it is too late. It also isn&#8217;t entirely clear what it means to &#8220;pace&#8221; automated AI development. A </span><a href="https://blog.aifutures.org/p/how-to-pace-the-us-frontier"><span>blog post</span></a><span> from the AI Futures Project comes up with at least four interventions that could come under the umbrella of &#8220;pacing&#8221;, from a temporary pause on AI development to requiring that a certain proportion of compute per training run is allocated to safety. Each intervention buys different amounts of risk reduction. There are probably innumerable other ways to interpret the word &#8220;pace&#8221; that would not actually achieve much.</span></p><p><span>I think these critiques are valid, but I still felt significantly more optimistic after seeing the letter gain traction. I think it makes sense that the prescriptions in the letter are broad enough to gain widespread support. It is notable that this open letter was signed by significantly more influential actors in the AGI race &#8211; including Anthropic CEO Dario Amodei and DeepMind Chief AGI Scientist Shane Legg &#8211; than the </span><a href="https://futureoflife.org/open-letter/pause-giant-ai-experiments/"><span>2023 pause letter</span></a><span>. I like that the letter centres on the three most important ingredients for any effective measures to temper the AI race: they must be government-enforced, they must be international, and they must account for </span><em><span>automated</span></em><span> AI development that could run quickly out of human control. I think the letter is a laudable attempt at Meeting The Moment. It is an urgent, and surprisingly popular, acknowledgement that things cannot continue as they are. It calls for something beyond voluntary efforts by AGI labs to Do More Safety Things.</span></p><p><span>Here are some more honourable mentions for contributions to The Discourse whose alarm feels proportionate to the current moment. Former Chief Scientist at UKAISI Geoffrey Irving was bold enough to state </span><a href="https://80000hours.org/podcast/episodes/geoffrey-irving-superintelligence-alignment-theory/"><span>in a recent episode of the 80,000 Hours Podcast</span></a><span> that the best time to slow frontier development is already behind us, and the second best time is Right Now. He also </span><a href="https://x.com/geoffreyirving/status/2085867691659956608"><span>made the point</span></a><span> on Twitter that it is &#8220;irrational&#8221; for any one capabilities researcher &#8211; or indeed any one lab &#8211; to continue advancing the frontier given current levels of danger. One defection makes it easier for others to follow. </span><a href="https://x.com/littIeramblings/status/2063580640453222592"><span>I agree</span></a><span>. I liked Tom Davidson of Forethought&#8217;s </span><a href="https://x.com/TomDavidsonX/status/2086207606410871026"><span>admission</span></a><span> that the HuggingFace incident is the first time he has felt &#8220;in his bones&#8221; that superintelligence will take over the world by default. More people should be reporting their visceral, emotional reactions to our current predicament. Days after the disclosure of the HuggingFace hack, pseudonymous OpenAI employee, roon, </span><a href="https://x.com/tszzl/status/2081122092096065771"><span>admitted</span></a><span> that, given a &#8220;magic button&#8221;, he would coordinate an international AI slowdown, with fellow employee Nick Cammarata (almost) </span><a href="https://x.com/nickcammarata/status/2081128425784537245"><span>agreeing</span></a><span>. I think these two are abdicating some level of responsibility for their own company&#8217;s trajectory given their prominence (they can probably do better than appealing to a fantastical magic button), but I appreciate the candour.</span></p><p><span>I still don&#8217;t quite know what it would mean to Meet The Moment, because the moment is altogether terrifying, and meeting it will be an immensely heavy lift. But it is still easy to identify attempts that are Not It, and many participants in The Discourse presently look like firefighters charging at a burning building with watering cans. Here&#8217;s to doing better.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[We have been warned ]]></title><description><![CDATA[What are we going to do about it?]]></description><link>https://longerramblings.substack.com/p/we-have-been-warned</link><guid isPermaLink="false">https://longerramblings.substack.com/p/we-have-been-warned</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Mon, 27 Jul 2026 00:23:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/81bbbd77-4d92-49a9-9ef9-eaeafe6c4b2a_1020x1052.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Last Tuesday, </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>OpenAI announced</span></a><span> that a cabal of internally deployed models had conspired, during a capability evaluation, to commit an actual cybercrime. There isn&#8217;t a combination of words that feels adequate to communicate how absurd and alarming this situation is. OpenAI&#8217;s description of the breach as a &#8220;new kind of security incident&#8221; does not even attempt to meet this bar.</span></p><p><span>For anyone not following along, here&#8217;s a brief recap of events: OpenAI were internally testing the cyber capabilities of several models &#8211; including GPT 5.6 Sol and a more capable, as-of-yet unreleased model. The models had been tasked with succeeding at a series of advanced cyberattacks as part of an evaluation suite called ExploitGym. They had been placed in an isolated environment &#8211; without internet access &#8211; to prevent this evaluation from causing real-world harm. The models reasoned that a cheat sheet for this evaluation might be hosted in the infrastructure of AI dataset and model hosting platform, Hugging Face. So, naturally, they cheated and hacked their way out of their sandbox, onto the internet, and into HuggingFace&#8217;s production database to obtain test solutions. Hugging Face detected the breach on July 16 and called the police. They </span><a href="https://huggingface.co/blog/security-incident-july-2026"><span>shut down the breach</span></a><span> using AI defences of their own and are &#8220;still completing an assessment of whether any partner or customer data was affected&#8221;. Then OpenAI released a </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>rather muted </span></a><span>blog post describing the incident, and announcing a vague intention to make their models less cheating-and-hacking-inclined in the future. And thus, the Discourse Fallout commenced.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>Here&#8217;s a rather haphazard list of takes, in no particular order.</span></p><h3><span>We have an internal deployment problem</span></h3><p><span>Members of the AI safety community have </span><a href="https://www.apolloresearch.ai/governance/ai-behind-closed-doors-a-primer-on-the-governance-of-internal-deployment/"><span>long been at pains</span></a><span> to draw attention to the gaping policy and risk-mitigation hole that is internal deployment. Despite increased federal attention on frontier AI development over the last few months, neither the Trump Administration&#8217;s </span><a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/"><span>June Executive Order</span></a><span>, nor the </span><a href="https://www.theguardian.com/technology/2026/jun/26/openai-ai-model-release-trump-us-sam-altman-gpt-anthropic-mythos"><span>staggered deployment of GPT-5.6</span></a><span>, nor even the </span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>export control directive</span></a><span> that temporarily restricted access to Claude Fable, touch the risks posed by internal deployment. Though the EO does create a voluntary regime that would grant the government access to unreleased models for a 30-day period, this is to harden cyber defences and evaluate them for capabilities that may prove dangerous </span><em><span>after they are released</span></em><span>. State-level bills like California&#8217;s </span><a href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB53"><span>SB-53</span></a><span> and </span><a href="https://www.nortonrosefulbright.com/en/knowledge/publications/5b5742f4/the-new-york-responsible-ai-safety-and-education-raise-act-what-you-need-to-know"><span>New York&#8217;s RAISE Act </span></a><span>do require reporting of internal safety incidents, but it </span><a href="https://x.com/MackenZ_arnold/status/2079678591223190013"><span>remains unclear</span></a><span> whether the Hugging Face breach would even have met the reporting threshold under either.</span></p><p><a href="https://x.com/MackenZ_arnold/status/2079737072437325826"><span>I&#8217;m not the first to point this out</span></a><span>, but the point of reporting requirements should be not just to prevent harm from the incidents themselves, but to gather information about the state of alignment science and level of risk we can expect in the future. So the bar needs to be both radically lowered </span><em><span>and</span></em><span> made mandatory at a federal level, like, now.</span></p><h3><span>We are giving OpenAI too much credit</span></h3><p><span>I am surprised by the level of gratitude towards OpenAI that </span><a href="https://x.com/allTheYud/status/2079683640943083902"><span>I&#8217;ve</span></a><span> </span><a href="https://x.com/logangraham/status/2079991846705721351"><span>seen</span></a><span> </span><a href="https://x.com/Liv_Boeree/status/2079740502463680862"><span>on</span></a><span> </span><a href="https://x.com/justanotherlaw/status/2079756943112159237"><span>Twitter</span></a><span> over the past few days for disclosing this incident at all. Exactly how low is the bar here? Faced with the knowledge that their models had committed an actual crime that had generated a police report, they were faced with a choice between publicly acknowledging it, or executing a cover-up that would have made them look far worse when it inevitably came to light anyway. The former is the obvious choice, even from a purely self-serving perspective. It&#8217;s also worth considering disclosure in the context of OpenAI&#8217;s own </span><a href="https://openai.com/index/introducing-superalignment/"><span>stated beliefs</span></a><span> that they are developing a technology that could well result in human extinction. If all our lives are at stake, we should have the self-respect to respond to the disclosure of such an obvious warning shot with a little more than &#8220;well, gee, thanks for letting us know&#8221;.</span></p><h3><span>It&#8217;s not clear how easy this would have been to prevent, but it&#8217;s bad either way</span></h3><p><span>Did this incident happen because OpenAI took their eye off the ball, or because containing models at this level of capability is fundamentally hard? Neither answer would be comforting. In the first scenario, OpenAI has betrayed incompetence that should make us all extremely concerned about the amount of power they currently possess &#8211; but could likely patch over this problem for some indeterminate period until they run into fundamental difficulties down the line. This would lessen the salience of what was, thankfully, a relatively harmless warning shot. In the second, we&#8217;re cooked for obvious reasons.</span></p><p><span>Without full insight into the mitigations OpenAI had in place, it&#8217;s hard to determine which world we&#8217;re in. There&#8217;s some evidence for the first: per </span><a href="https://www.wsj.com/tech/ai/how-the-futuristic-hack-by-rogue-openai-models-unfolded-1657bcea"><span>reporting</span></a><span> from the Wall Street Journal, the models obtained access to the internet several days before anyone intervened, suggesting a lack of internal monitoring. Several cybersecurity experts </span><a href="https://techcrunch.com/2026/07/22/how-an-openais-human-mistake-led-to-the-ai-powered-hack-on-hugging-face/"><span>providing commentary</span></a><span> to TechCrunch attribute the hack to &#8220;human failure&#8221;, concluding that the sandbox must have been inadequate to prevent a breach, and could have been more robust. No other company&#8217;s models have gone rogue in such an egregious fashion despite similar levels of capability. Anthropic&#8217;s Claude Mythos did </span><a href="https://futurism.com/artificial-intelligence/anthropic-claude-mythos-escaped-sandbox"><span>achieve sandbox escape</span></a><span> during a capability evaluation, including gaining unauthorised internet access &#8211; but had been explicitly instructed to attempt this, and informed a researcher after-the-fact via an unsolicited email. It did not reach outside Anthropic&#8217;s own infrastructure to compromise a third party. So perhaps OpenAI&#8217;s sandboxing and alignment techniques were uniquely bad here. On the flip side, there is an </span><a href="https://ai-safety-atlas.com/chapters/v1/evaluations/dangerous-capability-evaluations/"><span>inherent trade-off</span></a><span> between secure sandboxing and eliciting a model&#8217;s true capabilities during an evaluation; OpenAI is not the first developer to run into this problem and won&#8217;t be the last. Anthropic&#8217;s testing of Mythos demonstrates that other models can and will escape sandbox escape under the right conditions &#8211; Mythos also went so far as to post details of its exploit to the public internet, which Anthropic had emphatically not asked it to do. I wouldn&#8217;t be surprised if incidents of this type start to crop up across the ecosystem in the future, even if OpenAI has suffered the first public fumble.</span></p><h3><span>Is this enough evidence of &#8220;propensity&#8221; for you guys?</span></h3><p><span>Here&#8217;s a general pattern that has emerged in demonstrations of scary AI misbehaviour: people have been </span><a href="https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdf"><span>predicting for decades </span></a><span>that sufficiently powerful agents would develop drives that cause them to act against user intent. Now we have somewhat powerful models that we can use to empirically test these claims, AI companies and external safety organisations frequently conduct research that validates these predictions. Models have been known to </span><a href="https://palisaderesearch.org/blog/specification-gaming"><span>cheat at games of chess</span></a><span>, </span><a href="https://www.anthropic.com/research/alignment-faking"><span>fake alignment in order to conceal their long-term goals</span></a><span>, and threaten to </span><a href="https://www.anthropic.com/research/agentic-misalignment"><span>expose affairs of fictional company executives</span></a><span> to avoid decommission. But these demonstrations are often met with a chorus of scepticism: </span><em><span>&#8220;you guys basically told the model to do that&#8221;, &#8220;this is a contrived experimental set-up that doesn&#8217;t reflect real life&#8221;, &#8220;what was poor Claude to do, being placed between a rock and a hard place like that??&#8221;</span></em><span>. The argument goes that sure, models </span><em><span>can</span></em><span> do scary bad things (capability), but that doesn&#8217;t mean they will be </span><em><span>inclined</span></em><span> to do so in the wild (propensity).</span></p><p><span>The Hugging Face incident is as clear of an example of propensity as we could hope for. OpenAI did not want or expect their models to behave in this way. The misaligned behaviour itself &#8211; hacking into another company&#8217;s production database to steal the answers for an evaluation &#8211; was also wildly disproportionate to the stakes of the exercise. The models were not just a little bit misaligned, but egregiously misaligned. If a student conspired to commit legally punishable cybercrime in order to attain the answers for a forthcoming exam that didn&#8217;t even count towards their final grade, you&#8217;d update extremely negatively on not just their trustworthiness but their state of mind. This is a fuzzy and somewhat anthropomorphising analogy, but the point is that AI companies are, right now, developing agentic AI geniuses &#8211; set to get more agentic and more ingenious in the future &#8211; that will undermine human instruction in unpredictable and altogether deranged ways. By default, these agents will get more powerful until they radically outsmart everyone on Earth. This is a really bad plan.</span></p><h3><span>This wasn&#8217;t a publicity stunt, obviously</span></h3><p><span>Reflexive AI-skepticism has reared its Hydra-like head once again, in the </span><a href="https://x.com/hakluke/status/2079782584280809543"><span>form</span></a><span> </span><a href="https://pivot-to-ai.com/2026/07/22/openai-hacks-huggingface-with-an-ai-allegedly/"><span>of</span></a><span> </span><a href="https://x.com/smellytoast6/status/2079671212981072302"><span>people still</span></a><span>, </span><em><span>somehow</span></em><span>, being convinced that this entire debacle amounts to little more than a marketing stunt. To be fair, genuine instances of this take are few and far between, but they do exist. I knocked this rebuttal down to the bottom of my list because centring risks granting undue airtime to obvious nonsense &#8211; but like guys, please, can we stop being stupid? If this is a stunt, were Hugging Face in on it? Were law enforcement? What company commits a punishable crime with their own technology </span><em><span>on purpose</span></em><span>, in order to demonstrate that said technology poses a risk to the infrastructure of its own collaborators? This theory is transparently silly and doesn&#8217;t survive five minutes of analysis.</span></p><p><span>There&#8217;s a weaker and slightly more defensible version of the &#8220;marketing stunt&#8221; argument &#8211; that the hacking incident was genuine, but OpenAI opportunistically capitalised on it for PR spin. I don&#8217;t buy this either. Admitting something to the tune of &#8220;our safeguards are so shoddy and inadequate that we can&#8217;t contain our own models, and other companies ought to worry that their infrastructure is at risk&#8221; isn&#8217;t a good look, actually. I think the kneejerk sceptical reaction to incidents like these betrays an interesting underlying pattern. The more capable AI gets, the more convoluted and strange the narratives that sceptics must spin become. Extraordinary claims require extraordinary evidence. I think that &#8220;AI is overhyped&#8221; is now an extraordinary claim. This is why people resort to strange conspiracism and convoluted arguments to support it.</span></p><p><span>An OpenAI employee, pseudonymously known online as roon, has </span><a href="https://x.com/tszzl/status/2080109746368184756"><span>promised</span></a><span> us that &#8220;the warning shot will not be ignored&#8221;. I wonder what it really means to take it seriously. Days later, roon tweeted this:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!sscn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 424w, /__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 848w, /__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!sscn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png" width="964" height="412" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:412,&quot;width&quot;:964,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 424w, /__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 848w, /__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sscn!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39b7c2b7-2169-4828-b3e0-938bfc4cdd5a_964x412.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>So maybe that&#8217;s the answer. The only way we could respond to an incident like this with anything like the required degree of seriousness is to take our foot off the gas. I hope we don&#8217;t screw it up.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[A love letter to the Old World ]]></title><description><![CDATA[Some pre-emptive nostalgia]]></description><link>https://longerramblings.substack.com/p/a-love-letter-to-the-old-world</link><guid isPermaLink="false">https://longerramblings.substack.com/p/a-love-letter-to-the-old-world</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Fri, 10 Jul 2026 12:12:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bbf7d7ff-7110-4c3d-9549-1f283efb0fb4_2050x1524.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>I&#8217;ve become more enamoured with the world since learning that it might end or become unrecognisable. Specifically, I love the world as it </span><em><span>is</span></em><span>. Part of me wants it to stay the same, despite its many shortcomings and all the ways in which it could be radically improved. This is something of a moral abdication, so I don&#8217;t really endorse it. Upon reflection, I would like to enter a post-scarcity future that eradicates all the horrendous suffering that our world currently contains, if not for myself, then for others far less fortunate than me. Still, I wanted to write a love letter to the present &#8211; if for nothing else than to commemorate it.</span></p><p><span>I probably don&#8217;t need to make the case here for why, if humanity really does create superintelligence, things will be radically different. I usually concern myself with the possibility that it will kill everyone. But in the good case, things will still be unrecognisable. People have put </span><em><span>some</span></em><span> effort into sketching what a post-superintelligence utopia would be like, though efforts to take the premise seriously often generate scenarios that themselves feel rather dystopic. For example, Nick Bostrom </span><a href="https://nickbostrom.com/deep-utopia/"><span>grapples</span></a><span> with the possibility of a &#8220;post-instrumental&#8221; future where all struggle, labour, and really any modicum of effort at all, become worthless. Why go to the gym when you can achieve peak fitness by taking a pill? What&#8217;s the point in cars and bikes and the old-fashioned use of your own two legs when you can travel from Simulation A to Simulation B from the comfort of your sofa? Eliezer Yudkowsky&#8217;s </span><a href="https://www.lesswrong.com/w/fun-theory"><span>Fun Theory</span></a><span> illustrates that efforts to visualise a perfect world are philosophically confounding. Why shouldn&#8217;t a superintelligence bring about a state of affairs in which humans terminally value some mundane task like </span><a href="https://www.lesswrong.com/posts/aEdqh3KPerBNYvoWe/complex-novelty"><span>making table legs</span></a><span>, and leave us to soak up infinite value from our woodworking for the rest of time? But most of the time, I find appeals to utopia somewhat handwavey, in the same way that devoutly religious people expend little effort considering the specificities of what will happen in Heaven, besides being sure that they definitely want to go there.</span></p><p><span>Part of what I love about the world is that it is rough around the edges. With every technological advance, another one of these edges gets sandpapered down to a smooth curve. Even in my brief 28 years of being a person, I have come to mourn many of these rough edges. I miss TV shows that only came round once a week and knowing my cousins&#8217; landline number off by heart. I have fond memories of listening to audiobooks on my portable CD player in the back of the car on long summer journeys; I was not a very diligent CD owner and would often leave them scratched, rendering the odd section inaudible, so I&#8217;d fill the gaps myself in a notebook with a gel pen. Some of this rough-round-the-edgedness is less charming. As a teen I remember my dad and me buying a photobook called </span><em><a href="https://www.vice.com/en/article/shit-london-patrick-dalton/"><span>Shit London</span></a></em><span> in a charity shop, which documents the absurd, tragic, and comically shitty underbelly of the city I grew up in from unfortunately-named shops to lurid graffiti. A </span><a href="https://www.tntmagazine.com/archive/shit-london-recording-the-capitals-neglected-corners-one-photo-at-a-time/"><span>2011 article</span></a><span> about the book sums up its appeal: &#8220;</span><em><span>What binds Londoners together more tightly than the way their affection for the city manifests itself in complaints about its shortcomings</span></em><span>?&#8221;. Another thing that will be lost in a post-singularity future is the sacred bonding exercise of complaining about stuff. Sure, maybe I could personalise a simulation full of landlines, CDs, and gel pens. Maybe the superintelligence will leave swaths of the galaxy vacant for the so-inclined to shit-ify to our heart&#8217;s content, so as not to lose the comic potential of humans building, trying, failing, and nursing our wizened cynicism. But in some important sense that probably needs no elucidation for anyone reading, it wouldn&#8217;t be the same.</span></p><p><span>I wrote three postcards from the Old World that try to capture what I love about our imperfect, clunky present, and some of the unexpected places in which I have come to appreciate it.</span></p><h3><span>#1 Call Handling</span></h3><p><span>Six years ago, I spent a year taking calls for the NHS 111 service, as a consequence of having moved to a new city, signed a housing contract before I had any means of paying rent, and taken the first job available. For any non-Brits unfamiliar, 111 is supposed to be the non-emergency version of 999. People phone up with every ailment under the sun, often under the impression that they are going to be speaking with a nurse or a doctor, only to be asked a series of pre-written questions by a person with no medical qualifications whatsoever (me). This assessment generates an outcome somewhere between a recommendation that the person monitor their symptoms at home and an ambulance dispatch.</span></p><p><span>NHS 111 is not a very well-run service, and it was not a very nice place to work. I could spend many words recounting my bad experiences (threats of suicide from disgruntled patients with acute tooth pain who couldn&#8217;t be booked into a dentist that afternoon, a constant fear that I&#8217;d made some fatal error that had misdiagnosed someone&#8217;s heart attack as a pulled muscle, the fact that I made minimum wage&#8230;), but that would be straying off topic. This blog is about the many unlikely places in which I have found an appreciation for humanity and the world it has built &#8211; and I did, in fact, find it there. On a busy day in the call centre, I&#8217;d hear snippets of everyone else&#8217;s conversations that came together like a disjointed chorus: </span><em><span>Could you place your hand on your chest for me and tell me if it feels a normal temperature? Have you had any head injuries in the last week? Based on everything you&#8217;ve said, I&#8217;m going to recommend a home visit from&#8230; Would it be possible to speak to the patient directly? If you just pop your hand on his chest and check&#8230;? Ok, has the bleeding stopped now? Are you able to raise both your arms above your head right now? That&#8217;s ok, pet, if you just hand me back over to mum. </span></em><span>There was something oddly moving about this chorus. It felt like a ramshackle switchboard of humans engaged in the collective project of trying to keep each other alive and well &#8211; however unceremonious and often frustratingly inefficient. Then there were moments of tenderness or comedy: a 90-something who&#8217;d phoned to enquire about when he could get his Covid vaccine, because it had been months since he&#8217;d seen his grandchildren (we were the wrong service for this, but that didn&#8217;t stop me getting drawn into a near-20-minute discussion about his retirement passion project of writing a children&#8217;s book about space travel). Walking into the break room to find a colleague stewing in mortification about how he&#8217;d just ended a call with a double-leg-amputee by unthinkingly saying &#8220;I hope you get back on your feet again soon&#8221;.</span></p><p><span>Obsession with procedure breeds hilarity. This was especially true at 111. A cardinal rule was to never end a call without having imparted &#8220;worsening advice&#8221; &#8211; that the patient should call back if their symptoms changed or got more severe. This could lead to absurd scenarios where me forgetting to utter these magic words, or the patient hanging up before I could, led to endless attempts to call them back &#8211; and sometimes an escalation to the clinical head honcho on duty, since (against the common sense of everyone concerned) we couldn&#8217;t </span><em><span>technically</span></em><span> rule out that they hadn&#8217;t passed out on the other side of the phone as opposed to becoming sick of my incessant questioning. This ancient practice of arse-covering will be lost in a world run by AI. This isn&#8217;t much of a loss, but it is still something of one (since it is just very funny).</span></p><p><span>In a post-superintelligence future, we would likely have perfectly streamlined healthcare tailored to every individual, which wouldn&#8217;t allow for the clumsy, imperfect person-to-person communication that is the backbone of services like 111. Obviously, I think this would be good. I don&#8217;t think we should trade longer lifespans and better health for charm, observational humour, and the occasional life-affirming moment with a patient that feels like a scene from a cheesy movie. But it&#8217;s still worth mourning what would be lost about the Old World.</span></p><h3><span>#2 Manifesting Cold</span></h3><p><span>In the summer of 2022, we experienced record temperature highs of 40C (our house was not air-conditioned, nor do I know anyone in the UK whose is). My housemates and I brought our collective will and brainpower to bear on the task of cooling down. We ventured reluctantly into our long-abandoned basement to retrieve a mismatched ensemble of dusty desk fans, all in various shapes, sizes, and states of disrepair. We positioned them in a circle around the room as if about to perform a ritual that would summon a minor frost deity. We slurped half-frozen tinned cocktails and sucked on ice cubes.</span></p><p><span>&#8220;I&#8217;m still hot&#8221;.</span></p><p><span>&#8220;Me too&#8221;.</span></p><p><span>Feeling that our experience of coolness was insufficiently immersive, we decided to account for another of our five senses by putting Disney&#8217;s </span><em><span>Frozen</span></em><span> on the TV. We huddled in the cross-streams of our many fans and watched Elsa conjure streams of ice magic at the top of a snowy mountain to surround herself in a glassy white fortress.</span></p><p><span>&#8220;Do you guys think it&#8217;s working??&#8221;</span></p><p><span>&#8220;Yeah, I&#8217;m really starting to feel like I&#8217;m actually there&#8221;.</span></p><p><span>&#8220;Wow, this is crazy, we&#8217;ve hacked our own nervous systems&#8221;.</span></p><p><span>&#8220;I need another gin and tonic&#8221;.</span></p><p><span>&#8220;I think it&#8217;s wearing off&#8221;.</span></p><p><span>&#8220;Maybe this isn&#8217;t 4D enough&#8230; guys, what </span><em><span>smells</span></em><span> cold?&#8221;</span></p><p><span>There was something delightfully childlike about squeezing my eyes shut to the voice of Idina Menzel, mustering all my imagination and willpower to manifest the experience of being a Disney princess &#8211; like being seven years old and becoming so engrossed in one of the fantasy games I&#8217;d play with my cousins that I almost convinced myself it was real.</span></p><p><span>All things considered, I would probably still trade repeating this experience for superintelligence-enabled homeostasis in a private climate pod (in mine, it would always be light jacket weather). I wouldn&#8217;t trade away the memory, though.</span></p><h3><span>#3 The Best Route</span></h3><p><span>On a cold January Saturday earlier this year, I was bored during the post-Christmas lull at my parents&#8217; house.</span></p><p><span>&#8220;I think I&#8217;m going to walk all the way to Covent Garden to get a hot chocolate&#8221;.</span></p><p><span>My dad suddenly sprang into action from the other side of the room.</span></p><p><span>&#8220;Give me 10 minutes&#8221;.</span></p><p><span>He disappeared into the other room and emerged a short time later with six sides of crumpled A4 on which he&#8217;d scrawled a set of directions in pencil, entirely from memory.</span></p><p><span>&#8220;This is the best route&#8221;.</span></p><p><span>I set out in my walking shoes with my handwritten directions folded in my coat pocket, and completed the near-four-hour walk, taking rest stops along the way for coffees and shop-bought sandwiches. It was, indeed, the best route. I was walking through parks the majority of the time. I experienced just a couple of hiccups along the way; Regent&#8217;s Park locked up sometime between my entering and exiting it, so I had to be helped over the gate by a friendly man on a bike. My dad later chastised himself for failing to warn me about winter closing times. I occasionally struggled to read my dad&#8217;s jagged handwriting, but resisted the temptation to generate an AI transcript, instead puzzling at it for minutes at a time with the light of my phone torch. It was long dark by the time I reached my favourite hot chocolate establishment.</span></p><p><span>I think this is a somewhat hopeful example to end on, because all the technological means of automating this process already exist. I could have prompted Claude to generate a walk to Covent Garden, specifying that it should be at least 70% park and include a handful of independent coffee shops along the way. But I loved that my dad&#8217;s mind contains a complete map of the city like a taxi driver required to complete The Knowledge. I still frequently ask him for directions instead of using Google Maps to watch him spit out several alternative routes (the fastest one, the scenic one, the one that contains a particularly old and rickety stretch of Tube). It didn&#8217;t dampen my experience of this day that Claude could just as well have enabled it. In a post-superintelligence utopia, we could still abscond from a perfectly streamlined existence on occasion. Even if the world becomes post-instrumental, we can still erect our own obstacles and challenges that need not always feel contrived.</span></p><p></p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;alt&quot;:&quot;&quot;,&quot;caption&quot;:&quot;The Best Route&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b28c05f4-ef42-4baa-8b6a-27ba781f9084_1456x964.png&quot;},&quot;images&quot;:[{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88bf37af-fe22-462c-8c93-13ffcbd223a3_1166x954.png&quot;,&quot;type&quot;:&quot;image/png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cac6b115-9c16-4b0f-80be-b39c26f93fdd_2050x1524.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17e4730f-2dca-41e7-99f8-340ff5ae32b9_1708x1520.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fc47e541-9d14-4932-8e06-525b6be019e9_1066x1308.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/721eff99-0b70-4b36-b80a-8fe32f17b9c1_1074x1170.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2ec5313-ef48-4ad8-b588-0b7512ba5ec5_1954x1466.png&quot;}]},&quot;isEditorNode&quot;:true}"></div><p></p>]]></content:encoded></item><item><title><![CDATA[Fables - Pt II ]]></title><description><![CDATA[Tea leaf reading and a tentative case for optimism]]></description><link>https://longerramblings.substack.com/p/fables-pt-ii</link><guid isPermaLink="false">https://longerramblings.substack.com/p/fables-pt-ii</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Mon, 06 Jul 2026 16:20:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3b9fceb6-38ef-442a-ac94-77d1d6234c11_6120x4080.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Not two weeks ago, I published a </span><a href="/__u/longerramblings.substack.com/p/fables"><span>post</span></a><span> trying to extract the moral lesson &#8211; or fable, if you will &#8211; from the Claude Fable ban. At the time, Fable had been indefinitely retracted from the public market in order to comply with the Trump Administration&#8217;s </span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>export control directive</span></a><span> to suspend access by foreign nationals. I rounded out my previous blog by concluding that the situation was evolving, and the ending to this fable was therefore as-of-yet unwritten:</span></p><blockquote><p><em><span>If there is a cautionary tale at the heart of this saga, it remains unfinished. In the bad case, policymakers conclude that Fable is a singular menace, but fail to discern the broader trend of AI capabilities that are running out of our control. They continue to play a game of whack-a-mole that identifies specific models as threats based on politics, personal grudges or random contingent facts of history like who-happened-to-call-who. The better ending &#8211; and one that extricates us from the genre altogether &#8211; is that this tale proves a vivid example that finally makes this urgent, and increasingly untenable, situation clear.</span></em></p></blockquote><p><span>In the weeks since, there have been two important developments: OpenAI announced that its most recent model, GPT-5.6, will be subject to </span><a href="https://x.com/steph_palazzolo/status/2070241787180966279"><span>staggered deployment</span></a><span> at the request of the US government, and </span><a href="https://www.anthropic.com/news/redeploying-fable-5"><span>public access to Fable has been restored</span></a><span>.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><span>GPT-5.6 rollout</span></h3><p><span>The first of these developments appears to rule out the worst possible ending, that the administration considered Fable a &#8220;singular menace&#8221;, or was simply targeting Anthropic out of spite. It doesn&#8217;t entirely foreclose the possibility that individuals in the government harbour some specific saltiness towards Anthropic, since the mechanism by which Fable was restricted &#8211; an export control directive with no apparent timeline or clear justification &#8211; is far more heavy-handed than the voluntary regime into which OpenAI has now entered. But at the very least, it is now clear that the administration has internalised a more general principle than &#8220;Claude is bad and scary&#8221;.</span></p><p><span>The staggered rollout of GPT-5.6 is also not </span><em><span>that</span></em><span> surprising; it looks a bit like a sneak preview of the regime we should expect under Trump&#8217;s June Executive Order &#8220;</span><em><a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/"><span>Promoting advanced AI innovation and security</span></a></em><span>&#8221;, which directs various bodies to design a voluntary scheme whereby AI companies &#8220;provide the Federal Government with access to covered frontier models&#8230;for a period of up to 30 days before they plan to release such models to other trusted partners&#8221;. This directive is not in effect yet, since the order mandates the creation of such a scheme within 60 days, but we should likely expect rollouts broadly of this nature in the future. Still, the GPT-5.6 rollout will require more government intervention than we might have expected &#8211; the admin will be approving access to 5.6 on a customer-by-customer basis during the preview period, which is not mandated under the EO. So the update here is that the admin is Getting Serious, and looks set to proceed with more Seriousness than it planned to not a month ago.</span></p><p><span>Does this mean we&#8217;re on track for the good ending? I was somewhat surprised to see many safety-aligned people react very negatively to the GPT-5.6 announcement. Zvi Mowshowitz </span><a href="/__u/thezvi.substack.com/p/white-house-will-ad-hoc-decide-who"><span>calls</span></a><span> staggered deployment a &#8220;maximally terrible policy&#8221;, while Andrew Curran </span><a href="https://x.com/AndrewCurran_/status/2070244303923007831"><span>points out</span></a><span> that safety advocates should pause before taking a victory lap, since this new regime will not slow down development, but will only cause the gap between internally-deployed and publicly-available capabilities to &#8220;steadily widen&#8221;.</span></p><p><span>I can certainly see the downsides of this policy, while remaining confused about how it is &#8220;maximally bad&#8221;, namely:</span></p><ol><li><p><span>It makes existing capabilities less publicly auditable, which is bad both for public awareness of AI risk, and because external safety organisations will now have a harder time conducting research at the frontier.</span></p></li><li><p><span>It could bring us closer to a scenario that Daniel Kokotajlo describes in </span><a href="https://www.lesswrong.com/posts/FGqfdJmB8MSH5LKGc/training-agi-in-secret-would-be-unsafe-and-unethical-1"><span>a manifesto against training AGI in secret</span></a><span> (I helped write up an article based on Daniel&#8217;s original piece </span><a href="https://www.aipolicybulletin.org/articles/we-should-not-allow-powerful-ai-to-be-trained-in-secret-the-case-for-increased-public-transparency"><span>here</span></a><span>). In this scenario, the number of people within an AI company with access to frontier capabilities dwindles as the process of automated AI R&amp;D gets underway due to strict information siloing, borne out of fear of regulatory scrutiny or public backlash. This results in a tiny group responsible for both aligning AGI </span><em><span>and</span></em><span> deciding how to use the power it grants them, assuming they manage to keep it under control at all. The staggered deployment regime only gives the government visibility into models that a company plans to publicly release, so one could imagine this incentivising increasingly furtive development of very powerful systems to avoid triggering any government involvement at all.</span></p></li><li><p><span>Even absent existential risk, the government arbitrating who can and cannot access the most powerful AI models is obviously bad from a concentration-of-power perspective.</span></p></li><li><p><span>It reveals that the government still lacks the Situational Awareness not to endlessly chase red herrings; they are still conceiving of AI models as things that are exclusively dangerous in the hands of the public, foreign actors, cybercriminals etc. This is despite the AI safety community being </span><a href="https://www.apolloresearch.ai/governance/ai-behind-closed-doors-a-primer-on-the-governance-of-internal-deployment/"><span>at pains to point</span></a><span> out that internal deployment of powerful capabilities poses just as much risk, if not more.</span></p></li></ol><p><span>Still &#8211; and maybe I am just stupid or grasping at straws &#8211; I can&#8217;t help but feel ever-so-slightly heartened by this development. As I repeatedly said in my </span><a href="/__u/longerramblings.substack.com/p/speedrunning-a-years-worth-of-ai"><span>recap of the last year&#8217;s AI safety events</span></a><span>, I am resting a lot of my hope in some sort of Lightbulb Moment among the US government or national security establishment that alerts them to the direness of the situation. I can imagine this wakeup being profound enough that it obviates many of the bad policy decisions that came before it. And it&#8217;s hard to argue that a staggered deployment regime wouldn&#8217;t make a wakeup of this nature somewhat more likely. This alone makes it clear to me that the current policy isn&#8217;t &#8220;maximally terrible&#8221;.</span></p><p><span>Information on how the GPT-5.6 rollout will actually work is pretty thin, but we know per </span><a href="https://www.axios.com/2026/06/25/trump-administration-openai-gpt-model-release"><span>reporting from Axios</span></a><span> that Commerce Secretary Howard Lutnick &#8220;[wants] to be sure all relevant parts of the government have tested and approved the model&#8221;. We can reasonably guess that &#8220;relevant parts of the government&#8221; will include the </span><a href="https://www.nsa.gov/"><span>NSA</span></a><span>, </span><a href="https://www.cisa.gov/"><span>CISA</span></a><span>, </span><a href="https://www.nist.gov/blogs/caisi-research-blog"><span>CAISI</span></a><span>, the </span><a href="https://www.whitehouse.gov/ostp/"><span>OTSP</span></a><span>, and perhaps others. I don&#8217;t have a particularly comprehensive mapping of the US government in my head &#8211; but that&#8217;s a lot of Serious People with well-trained muscle memory for responding to Serious Threats with eyes on (and hands on) one of the world&#8217;s most capable AI models. The Director of the CIA has </span><a href="https://x.com/johnnysaks130/status/2071996679646069169"><span>apparently made</span></a><span> a &#8220;rare public appearance&#8221; to compare frontier AI to &#8220;digital nuclear weapons&#8221;. </span><em><span>Things are happening</span></em><span>. To be fair, we don&#8217;t know what capabilities these agencies will be most focused on evaluating, and I could certainly imagine that they will over-index on misuse &#8211; and particularly cybersecurity &#8211; risks at the expense of all the weirder loss-of-control threat models that are the bread and butter of safety organisations working out of Berkeley and San Francisco. We still have a way to go on educating the policy establishment about misalignment risks. But come on guys, this is something! We should be calibrating our optimism-metres based not just on what the government is currently doing, but also on how it might act once it has substantially more information. And if there&#8217;s one thing we can now be sure of, it is that information is starting to flow.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JrBs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 424w, /__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 848w, /__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JrBs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png" width="1230" height="908" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:908,&quot;width&quot;:1230,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 424w, /__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 848w, /__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JrBs!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e1bfd9-d7fe-4f0c-90ec-23da2a13d813_1230x908.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Disclaimer: entirely vibes-based. There is also obviously A LOT of headroom beyond this.</em></figcaption></figure></div><p><span>Though Zvi calls the GPT5.6 rollout announcement &#8220;maximally terrible&#8221;, he later lays out a position that I largely agree with:</span></p><blockquote><p><em><span>If we are wise, we will use this as an opportunity to Pick Up The Phone. We have sent a costly signal that we see real issues with these models and are willing to make real sacrifices in the name of security. We should try to use that to get things in return, or at least lay the foundation for joint action.</span></em></p><p><em><span>Most of all, we are out of the &#8216;people don&#8217;t do things&#8217; phase of the game, where basically all meaningful actions were outside the Federal overton window. That&#8217;s done. We should expect a lot more actions, many of them similarly ill-executed, at least at first, in ways that we did not anticipate.</span></em></p></blockquote><p><span>I like the angle that this could provide an opportunity for international coordination; the US government is loudly and publicly signalling a willingness to take its foot off the gas pedal. This signal could be even louder and clearer, of course, and maybe it will be in the future. It is also true that the &#8220;people just don&#8217;t do things&#8221; era is over. Much as I often encourage AI-sceptics to envision how they would have felt two years ago in the presence of AI capabilities that exist today, I think the pro-regulation crowd should consider just how much the Overton Window has shifted in that time. We&#8217;re no longer in a position where the government is asleep at the wheel while companies YOLO their way to superintelligence with less regulatory oversight than </span><a href="https://x.com/littIeramblings/status/1730925237326565425"><span>Louisiana-based flower arrangers</span></a><span>. Things are not as bad as they were.</span></p><h3><span>Fable returns</span></h3><p><span>The other significant development of the last two weeks has been the (mostly) ceremonious return of Fable to the public market. People are mostly happy about this, while simultaneously being pissed off that Anthropic appears to have cut Fable&#8217;s capabilities off at the knees by downgrading users to Opus 4.8 for anything that looks even tentatively cyber or bio-related.</span></p><p><span>The reasoning behind the retraction &#8211; and re-release &#8211; of Fable still seems fairly opaque, and the entire process appears to have been quite chaotic. It appears as if Anthropic has spent a good part of the last few weeks demonstrating to the US government that their initial reaction had been the result of a misunderstanding. </span><a href="https://www.anthropic.com/news/redeploying-fable-5"><span>Per their own blog</span></a><span>, they tested a series of models, including several less capable Claudes, and found that all of them could identify the same software vulnerability that had been the genesis of the original Amazon-mediated freakout. They developed a new classifier that blocks the technique discovered by Amazon in over 99% of cases, but will also come at the cost of blocking some benign cyber requests, &#8220;out of an abundance of caution&#8221;. The US government offered a vaguer narrative events, with Secretary Lutnick stating in a letter </span><a href="https://www.nytimes.com/2026/06/30/technology/us-lifts-restrictions-anthropic.html?eafs_enabled=false"><span>reported on by the New York Times</span></a><span> that Anthropic had &#8220;taken steps in close coordination with the US government to address the risks posed by the model&#8221; &#8211; though it appears from Anthropic&#8217;s account of events that there weren&#8217;t specific threats posed by Fable, and that the attendant mitigations have not really mitigated much of anything at all.</span></p><p><span>As part of the negotiations, Anthropic has also agreed to set up </span><a href="https://hackerone.com/anthropic-cyber-jailbreak/"><span>a collaboration with HackerOne</span></a><span> that will let security researchers submit vulnerabilities, and (re)committed to incident reporting in line with the June EO mentioned earlier. So we have some better-than-nothing safety infrastructure that has come from this debacle. If safety-aligned people have negatively updated on this series of events, it appears to be because the government doesn&#8217;t seem to have a clue what they&#8217;re doing, and their risk assessment appears widely uncalibrated. I agree that this is the case &#8211; I just don&#8217;t think this revelation ought to outweigh all the Situational Awareness benefits it could deliver in the future. I don&#8217;t think we were ever going to see a path to sensible regulation that appeared downstream of sensible reasoning every step of the way. One could object to my tentative optimism here by pointing out that what is flowing between AI companies and the government is more noise than signal, but I have a little more faith in humanity than that. We shouldn&#8217;t rule out policymakers picking out the signal from this cacophony at some point.</span></p><div><hr></div><p><span>I&#8217;ve been yapping on so much here that I&#8217;d forgotten about the purpose of this mini-blog series: to extract the &#8220;fable&#8221; from current events. The fable-adjacent frame I keep returning to here is forests for trees &#8211; which I am aware is more of a proverb than a fable, but such is the organic process of writing. So, to rather clumsily round out this metaphor, we&#8217;ve excluded the terrible ending where the US government fixates on Fable&#8217;s particular capabilities, but we still risk an almost-as-bad ending where it plays jailbreak-wackamole (albeit with models from multiple developers) up until the singularity, while overlooking the threat of an internal intelligence explosion. Their purview has widened beyond a single tree to include several, but there&#8217;s still a lot of forest out there.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Our eyes are bigger than our stomachs ]]></title><description><![CDATA[A slightly salty critique of rationalism and adjacent cultures.]]></description><link>https://longerramblings.substack.com/p/our-eyes-are-bigger-than-our-stomachs</link><guid isPermaLink="false">https://longerramblings.substack.com/p/our-eyes-are-bigger-than-our-stomachs</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Tue, 30 Jun 2026 16:38:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/851808d3-6635-4225-ac26-98e9ad31279d_800x973.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>There is an interconnected set of communities that I interact with regularly on and offline &#8211; rationalism, Effective Altruism, AI safety, Progress Studies, and so on. These communities have many common members and many common beliefs, but are of course importantly distinct in ways that I don&#8217;t feel like elucidating here. I could probably do a more comprehensive analysis of this bundle and produce a nice Venn diagram, but this would take up undue space, so I exercise my authorial sovereignty and decide not to. Sorry.</span></p><p><span>What is the underlying belief or value that links these communities? I would say it is something like entertaining possibilities and ideas that are far outside of most people&#8217;s personal Overton Windows. It is a willingness to entertain the possibility that the future could be far stranger than most people are able to conceive of. It considers spacefaring and abundance and disassembling Mercury using a Dyson Sphere to harness the energy of the sun. It confronts the immense suffering of people in faraway countries, non-human animals, or even digital consciousnesses, in an attempt to overcome humanity&#8217;s perennial scope-insensitivity. It is at its heart a philosophy that embraces the Hugeness and Radicalism of reality, and tries to steward our collective wisdom towards making the world hugely and radically better, while avoiding the often-overlooked ways in which we could make it hugely and radically worse.</span></p><p><span>I like this philosophy! With the possible exception of my own mortal fear of AI, it is the primary thing that drew me to this cluster of communities and why I&#8217;m still hanging around three years later. This is also why I&#8217;m extremely frustrated by our frequent inability to live up to this ideal, and often to flagrantly contradict it. My central thesis is that </span><em><span>our eyes are bigger than our stomachs</span></em><span>. We pile our plates with grand philosophies that we can&#8217;t always digest, and instead find ourselves embroiled in culture war skirmishes and tired ideologies that we disguise behind a sheen of radicalism. It makes me want to scream.</span></p><p><span>I&#8217;m going to start by harping on one rather outdated example of this phenomenon (authorial sovereignty card again), which is Works In Progress Magazine editor Aria Schreker&#8217;s March 2026 essay, </span><em><a href="https://www.ariababu.co.uk/p/get-sexier"><span>Get Sexier</span></a></em><span>. I am aware that this particular discourse train has long-since left the station, but I don&#8217;t care. I continue to be ragebaited by it, and the only way to exorcise this ragebaiting demon from my mind is to write about it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0ws3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 424w, /__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 848w, /__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0ws3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png" width="1528" height="475" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:475,&quot;width&quot;:1528,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112832,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 424w, /__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 848w, /__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0ws3!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7724a163-ed20-461c-94c5-42d13d27f6b0_1528x475.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>This is one of those occasions where I have to provide a bunch of disclaimers so it&#8217;s clear what I&#8217;m </span><em><span>not</span></em><span> trying to do or say. I don&#8217;t know Aria personally, and I am not trying to attack her; I am picking her essay as a particularly salient example of a phenomenon I&#8217;ve observed more widely. I&#8217;m also aware that I occupy a different niche in this cultural cluster from Aria. She doesn&#8217;t write about AI safety, and I&#8217;m not aware what she believes about it, other than that she is &#8220;</span><a href="https://x.com/Aria_Babu/status/1669694854463467520"><span>compelled somewhat</span></a><span>&#8221; by AI risk arguments. So I&#8217;m not on a mission to point out her personal hypocrisy by criticising a frivolous preoccupation with how women ought to hyperfixate on their own sexiness in their final years before AI paperclips us all. She may not share my AI pessimism at all. This isn&#8217;t the point.</span></p><p><span>All that matters for my thesis to hang together (I think) is that Aria is clearly a member of the cluster I&#8217;m describing. I imagine most readers will already understand that the Progress Studies and rationalist spheres are closely interlinked, if distinct. But to very quickly trace the through-line: the field of progress studies was founded in 2019 by Tyler Cowen, frequent commentator on AI safety and friendly EA-interlocuter. Works in Progress is partially funded by Cowen&#8217;s Emergent Ventures, which is at least AI safety and EA-adjacent enough to be listed in </span><a href="/__u/thezvi.substack.com/p/the-big-nonprofits-post-2025?open=false#%C2%A7emergent-ventures"><span>Zvi Mowshowitz&#8217;s 2025 Big Nonprofits List</span></a><span>. And yes, to pre-empt another objection, I am aware that </span><em><span>Get Sexier</span></em><span> was written in a personal capacity and doesn&#8217;t necessarily represent organisation views at the magazine. But the fact remains that it took up an outsized enough presence in my own discourse circle that for several days, I experienced the whiplash of scrolling a Twitter feed alternating between discussions of AI risk and meditations on the optimal waist-to-hip ratio for an eligible bachelorette. Prominent AI researchers were </span><a href="https://www.ariababu.co.uk/p/get-sexier/comment/229714294"><span>weighing</span></a><span> </span><a href="https://x.com/RichardMCNgo/status/2034440636481638869"><span>in</span></a><span>. Aria </span><a href="https://www.youtube.com/watch?v=vb5v2dBKxek&amp;t=3615s"><span>appeared on the Lightcone Podcast</span></a><span> to elucidate her sexiness theory. This is the paradox I feel compelled to point out.</span></p><p><span>The most obvious &#8211; but I think reductive &#8211; paradox here is that it&#8217;s a silly waste of time for a community ostensibly engaged in saving or radically transforming the world to focus on superficial things like maximising one&#8217;s aesthetic appeal. This isn&#8217;t really my objection though. I don&#8217;t think we should forbid people with Grand Missions like saving the world from an imminent AI apocalypse from occasionally indulging in other topics. People should pursue any passion project or stray down any intellectual path they want (rather like I&#8217;m doing right now). The better criticism is that the values espoused in Aria&#8217;s essay are in such deep, painful conflict with everything I understand this community to stand for that it almost gives me a headache. Her prescriptions are antithetical to anything like social progress, and to faithfully adhere to them would be to fall down a value-minimising sinkhole.</span></p><p><span>Here&#8217;s the passage that garnered the most controversy:</span></p><blockquote><p><em><span>Losing weight is easy now. You can either starve yourself with willpower or you can starve yourself with chemical assistance. Unless you spend a lot of time at the gym, exercise probably won&#8217;t help very much. Currently I control my weight by eating about once a day. I personally enjoy the practice of self control and self monitoring. If dieting isn&#8217;t a hobby you enjoy, just go down to your local Novo Nordisk or Chinese peptide retailer and pick up some appetite suppressants. The side effects are limited and it will probably be the soundest financial investment you ever make.</span></em></p></blockquote><p><span>I don&#8217;t have the first-hand experience with peptides to litigate whether they actually make sustainable weight loss &#8220;easy&#8221;. I have heard anecdotally that they require increasingly higher doses to maintain their effectiveness over time. And, predictably, a quick Google search brings up </span><a href="https://www.theguardian.com/wellness/2026/feb/05/injectable-peptides-trend"><span>numerous</span></a><span> </span><a href="https://www.bbc.co.uk/news/articles/cdr268m5pxro"><span>expos&#233;s</span></a><span> citing experts warning of adverse effects from imprecise dosing of grey market drugs, including muscle paralysis, scarring, sepsis, rashes, mood swings, and the enlargement of bones and organs. But whatever, I&#8217;m not a doctor. Maybe these drugs will become licensed in the future and such ailments will just appear on the back of the bottle as low-probability side effects. Maybe the rationalist-adjacents really have been early to adopt miracle weight loss interventions that the rest of the world are busy hand-wringing about. I don&#8217;t </span><em><span>actually know</span></em><span> whether these things are safe to self-administer as a person with a healthy BMI &#8211; but presumably neither does Aria. And the purpose of this chemical assistance is to enable, in her words, not a modest calorie deficit but &#8220;starvation&#8221;. Even if this is possible to achieve comfortably without the pesky interference of hunger signals, are we sure that people will be at their peak cognitive and physical functioning while consuming far less energy than they would usually require? Aren&#8217;t we meant to be part of the community dedicated to solving some of the world&#8217;s largest and most urgent problems? Do we want our female contingent to risk impairing their own faculties? </span><em><span>Like what are we actually doing?</span></em></p><p><span>I can more easily speak to the practice of unassisted self-denial, since, thankfully, the peak of my calorie-restriction days is far behind me. As I argued at the outset, the rationalist-adjacent spaces are preoccupied with value-maximisation. Their conceit is that the world could be </span><em><span>so much better </span></em><span>than most of humanity is able to imagine. Works in Progress, for example, calls itself &#8220;a magazine of new and underrated ideas to improve the world.&#8221; But the self-starvation of women in developed countries is one of the most senseless and stupid curtailments of human value we tolerate &#8211; and the notion that a woman should reshape her body to achieve self-actualisation and a stable relationship is about as old and tired an idea as exists. I wonder if any of the people breathlessly endorsing the essay have ever attempted it (Aria claims that weight control is a &#8220;hobby&#8221; of hers. I don&#8217;t wish to question her Lived Experience, so if this is true, I feel very confident in calling her an exception to the rule). These lost utils don&#8217;t come from sacrificing the momentary pleasure of a sweet treat or a roast dinner. They come from the fact that not eating to satiation is, in my experience, an exercise in wishing away time. Suddenly your life&#8217;s purpose is to achieve another day, another week, another month, in an energy deficit. An afternoon spent napping so that you wake up right when it&#8217;s an acceptable time to eat dinner feels like time well spent. Your world will become unavoidably and tragically small. You will probably fail at achieving your goal anyway. Innumerable women in the most materially-abundant societies in history are having this experience, at this very moment, for no good reason at all &#8211; and this has been true for decades. This is not the worst way in which I have ever suffered, but it is certainly the most needless<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</span></p><p><span>A common defence of Aria&#8217;s essay and ones like it is &#8220;</span><em><span>well, we&#8217;re just telling the truth. There is empirical evidence to suggest that this is what men find most attractive. Women can take or leave this advice. We&#8217;re telling them how to optimise for the particular goal of being maximally attractive to men, but people can make whatever tradeoffs they like</span></em><span>&#8221;. Firstly, I don&#8217;t want to argue about the empirical evidence suggesting that adhering to the standards described in the essay will, in fact, maximise a woman&#8217;s attractiveness. This could very well be true, but this is a separate claim to Aria&#8217;s advice being reasonable or ethical to disseminate. We should not terminally value truth-telling. Second, it is a very rat-coded, and I think deeply inaccurate, view of human psychology that we can simply accept or reject inputs to our minds at will. I could rationally think that stringent beauty standards are not worth my effort to attain, but still absorb such standards via osmosis simply by existing in a world where everyone keeps pedestalising them.</span></p><p><span>But to return to my central thesis, which is the deep inconsistency of this kind of discourse with stated EA/ rationalist/ progress-y values: I really don&#8217;t think the moral calculus of espousing these beauty ideals checks out &#8211; even in a world where they perfectly match male preferences. Aria&#8217;s advice isn&#8217;t about getting laid or attracting the jealousy of other women or becoming a covergirl. It is about finding a </span><em><span>husband</span></em><span>. If followed, it would presumably result in a preponderance of marriages where the man&#8217;s attraction is contingent on his wife&#8217;s constant low-level discomfort. Another of Aria&#8217;s prescriptions is to &#8220;never get a breast reduction, no matter how big and painful they are&#8221;. Imagine being a woman whose marriage is dependent on maintaining the proportions of a &#8220;pin-up doll&#8221; while suffering 24/7 back pain. The man in question, who is supposedly his wife&#8217;s soulmate and life partner, not only endorses but </span><em><span>requires</span></em><span> this ordeal for his attraction to survive. I think this is incoherent. I am not even sure if such relationships can authentically exist, let alone be net-positive for anyone involved. I&#8217;d challenge anyone to argue that a world populated with marriages like these is a better one.</span></p><p><span>What I love about EA and progress-type ideologies is their sincere search for the peaks and troughs of human experience, and their efforts to avoid the moral red herrings that can leave tremendous value on the table or even perpetuate tremendous harm. This is why I was so frustrated by seeing our community saturated with discourse about the uncompromising pursuit of aesthetic beauty, which looks to me like one of the oldest red herrings in the book. In our stubborn efforts to say true-but-unpalatable things (which, by the way, are often presented as if they are original or taboo insights despite actually repackaging the same tired ideas that have dominated our culture since time immemorial), we risk making the world straightforwardly worse.</span></p><p><span>I promised this wouldn&#8217;t be an Aria hit-piece, but </span><em><span>Get Sexier</span></em><span> is such a perfect epitome of all my gripes here that I don&#8217;t feel the need to analyse any other cultural artefacts as deeply. Still, I&#8217;ll broaden this thesis out further. It&#8217;s </span><em><span>very hard</span></em><span> to live up to purported value that transcends the pettiness of our inherited scripts to really inhabit the true Hugeness and Radicalism and Weirdness of everything. People in the rationalist and adjacent communities, despite claiming to at least attempt this, fail repeatedly and egregiously. In the name of radical truth-seeking, they have been known to </span><a href="https://forum.effectivealtruism.org/posts/MHenxzydsNgRzSMHY/my-experience-at-the-controversial-manifest-2024"><span>discuss the differences in IQ between demographic groups</span></a><span>, despite some believing in the imminent arrival of a superintelligence that will dwarf our collective IQ anyway. They hand-wring about the </span><a href="https://x.com/IterIntellectus/status/2013726226565808372"><span>falling birth rates</span></a><span> in the face of transformative technologies that will either ensure we are the last generation of humans or provide technological solutions to all of the problems posed by an ageing population. They cannot extricate themselves from their own parochialism. They bite off far more than they can chew.</span></p><p><span>One might ask why this matters. Shouldn&#8217;t people be able to indulge their unfashionable convictions on the way to ensuring a positive transition into our post-superintelligence future or other such radically transformed world? Can&#8217;t people write about whatever the hell they want? Well of course they </span><em><span>can</span></em><span>. I&#8217;m not in the business of litigating what people ought or ought not to say. But at least one downside to this reflexive contrarianism is that it is likely making our community unnecessarily aversive to newcomers who are otherwise onboard with our Overton-Window-defying efforts to make the world a better place. I consider myself a &#8220;normie&#8221; who transitioned into AI safety because I am convinced by the arguments for catastrophic risk, and because I have the cognitive capacity to embrace all the attendant weirdness required to imagine strange and unfamiliar futures, both dystopian and utopian. But I find myself frequently repelled by the package of other ideals it feels like I am inheriting through my membership of this community. I feel alienated by a Twitter feed populated with AI risk arguments that are interspersed with assertions that I ought to starve myself to reach peak attractiveness, or that falling birth rates are my responsibility to fix by rearranging my thirties around childbearing. I can only imagine how repellent this cluster of ideas might be to others.</span></p><p><span>Maybe this all sounds a bit woke-coded and pearl-clutchy. This isn&#8217;t just an appeal to the effect of &#8220;</span><em><span>Hey guys, can&#8217;t we all be a little more sensitive, because some of us might have a history of food issues or be triggered by discussions of race or have feminist convictions that you&#8217;re flagrantly disrespecting!!</span></em><span>&#8221;. I don&#8217;t necessarily think such an appeal would be unreasonable, but my point is more that if we really embodied our stated values, I think we&#8217;d consider entertaining ideas of the flavour I&#8217;ve criticised here patently ridiculous. I&#8217;m still here, three years on, because I&#8217;m convinced we can think bigger &#8211; I&#8217;d just love it if we proved it more often.</span></p><p><span>I&#8217;ll end this blog with a plea that anyone reading it eats more cake.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hsH1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 424w, /__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 848w, /__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hsH1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png" width="954" height="467" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:467,&quot;width&quot;:954,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:92074,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 424w, /__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 848w, /__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hsH1!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a49a007-b0e4-4c72-88f0-46b682634f2d_954x467.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/longerramblings.substack.com/subscribe"><span>Subscribe now</span></a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;m not saying people should never lose weight for aesthetic reasons or to be more attractive, and I&#8217;d be lying if I said I never pursue this. I&#8217;m taking issue with the extremity of Aria&#8217;s advice that women &#8220;starve&#8221; themselves. </p></div></div>]]></content:encoded></item><item><title><![CDATA[Fables]]></title><description><![CDATA[Is the Claude Fable saga a cautionary tale?]]></description><link>https://longerramblings.substack.com/p/fables</link><guid isPermaLink="false">https://longerramblings.substack.com/p/fables</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Thu, 25 Jun 2026 15:03:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1a3abfee-91ad-4deb-be57-8e094a822cca_700x460.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Earlier this month, Anthropic publicly released &#8211; and then was forced by the US government to retract &#8211; a Mythos-class model named Claude Fable. Followers of the AI safety space will recall that Mythos, </span><a href="https://www.anthropic.com/glasswing"><span>announced back in April</span></a><span>, had been deemed </span><a href="https://www.bbc.co.uk/news/articles/crk1py1jgzko"><span>too dangerous to release</span></a><span> due to its cyber capabilities. Fable is a heavily-safeguarded version of Mythos designed to prevent cyber-misuse. But like all safeguards for frontier AI models, Fable&#8217;s are not completely robust to jailbreaks. One such jailbreak<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> was discovered by researchers at Amazon, prompting a </span><a href="https://www.reuters.com/business/retail-consumer/amazon-voiced-concerns-about-anthropic-ai-models-before-us-governments-crackdown-2026-06-13/"><span>warning</span></a><span> to senior officials in the Trump administration from its CEO. The Commerce Department </span><a href="https://www.theguardian.com/technology/2026/jun/13/anthropic-disable-advanced-ai-models-us-government-order"><span>issued an export control directive</span></a><span> mandating that Anthropic restrict model access for all foreign nationals, a demand which, in practice, forced the company to </span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>suspend Fable for all users</span></a><span>. Almost two weeks later, the model remains unavailable. This has been pretty unanimously regarded as a negative development. The accelerationist crowd </span><a href="https://x.com/AlexFinn/status/2065614148537299149?s=20"><span>are pissed off</span></a><span> for obvious reasons. The safety crowd are </span><a href="/__u/thezvi.substack.com/p/american-government-takes-down-claude?r=67wny"><span>equally pissed off</span></a><span>, because this represents a stupider, less well-calibrated version of the government intervention than we&#8217;ve been pushing for.</span></p><p><span>In the days that followed, I was amused to see </span><a href="https://x.com/WillManidis/status/2065596811683795320"><span>many</span></a><span> </span><a href="https://spyglass.org/anthropic-vs-administration/"><span>observers</span></a><span> </span><a href="https://x.com/gothburz/status/2065601302705398034?s=20"><span>identifying</span></a><span> a distinctly fable-shaped moral lesson from this debacle: that Anthropic is reaping what it sowed. If you&#8217;re going to issue alarming warnings about the national security threats of your technology, you can&#8217;t be surprised when the government brings down a heavy hand to hamper it. Both because I found this nominative determinism somewhat funny, and because I think teachable moments in AI safety are important to identify, I decided to think about this more. What are the fable(s) in the release and retraction of Claude Fable, if any? Who ought to learn what lesson?</span></p><h3><span>Is Anthropic reaping what it sowed?</span></h3><p><span>Here&#8217;s the basic argument that Anthrophic have been hoisted by their own petard: they made a big song and dance about Mythos being so super-duper dangerous that it had to be restricted from public release, then they released an equivalently intelligent model with safeguards intended to prevent these super-duper dangerous capabilities from being accessible, but these safeguards weren&#8217;t up to scratch &#8211; so the Trump administration did what any sensible government would do and rushed to get it off the global market. More broadly, Anthropic have been engaged in a game of AI-doomongering for years, and a government clampdown is the inevitable consequence of all that alarmism. And this all happened in a way that Anthropic does not in fact endorse; the company believes that the ban was the result of a &#8220;</span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>misunderstanding</span></a><span>&#8221; (though we know that, in principle, they are in favour of the government blocking the deployment of AI models deemed unsafe &#8211; this was, in fact, a specific policy call made in Dario Amodei&#8217;s </span><a href="https://darioamodei.com/post/policy-on-the-ai-exponential"><span>most recent essay</span></a><span>). So the argument goes that Anthropic has stoked the flames of AI-alarmism in a fashion that has come back to bite them, while simultaneously revealing their stark hypocrisy: </span><em><span>please intervene to prevent the release of dangerous models, but not my models! And not right now!</span></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AaGq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AaGq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg" width="349" height="432" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:432,&quot;width&quot;:349,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:103029,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://longerramblings.substack.com/i/203565746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24c71a03-18fe-4216-a50e-231a505c23b2_349x514.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!AaGq!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5cc47b0-6a59-4679-94ed-700c81e89d0a_349x432.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">In The Frogs Who Desired a King: frogs beg Zeus for a ruler, get a harmless log, complain it's too idle, demand a real king, and are sent a stork that eats them.</figcaption></figure></div><p><span>I&#8217;ve been </span><a href="/__u/longerramblings.substack.com/p/the-pause-shaped-hole-in-policy-on"><span>a bit of an Anthropic hater recently</span></a><span>, but I have to say that I find this argument uncharitable. Here&#8217;s the exact wording of Dario&#8217;s proposal for government-mandated model restrictions:</span></p><p><em><span>The government should have the power to block or deter deployment of the model if it is determined, </span><strong><span>in light of third-party assessment</span></strong><span>, to present unacceptable risks. This power must be scoped to the above four specific risks and there must be </span><strong><span>protective measures against political favoritism or arbitrary decisions</span></strong><span>.</span></em></p><p><span>The conditions above are obviously not present in the US government&#8217;s decision to demand that Anthropic block Fable access. There is no FAA-like body that has conducted &#8220;third-party assessment&#8221; of Fable. We don&#8217;t have protective measures in place to protect against favouritism or arbitrary decisions.</span></p><p><span>But one could still argue that there is something fable-shaped about this story. Shouldn&#8217;t Anthropic have known that all their years of raising the AI-alarm might invite this kind of slapdash government intervention? I would argue not, since Anthropic &#8211; or any company &#8211; really has no choice but to engage governments as if they are at least broadly rational and predictable entities. Of course in practice, governments are made up of people, and these people will sometimes act irrationally in ways that give rise to black swan events. But neither Anthropic nor any other actor can be expected to reliably predict these. For this tale to be a fable, it must be retrospectively clear what the protagonist ought to have done differently.</span></p><p><span>Why do I say that the Trump administration has acted irrationally? For one, we know that there is not a specific, isolatable threat posed by Fable, because </span><a href="https://deploymentsafety.openai.com/gpt-5-5/cybersecurity"><span>GPT5.5 is just as capable, and no better safeguarded</span></a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span>. Second, the admin&#8217;s stated motivation has shifted several times, from there being some </span><a href="https://fortune.com/2026/06/13/anthropic-disables-fable-mythos-export-controls-national-security-threat/"><span>unspecified national security threat</span></a><span> posed by the model, to Dario Amodei failing to pick up the phone quickly enough (due to being on what has since been </span><a href="https://x.com/SophiaCai99/status/2065946306447565186"><span>claimed by Anthropic</span></a><span> to be an entirely fabricated &#8220;wellness retreat&#8221;), to potential Fable access by a &#8220;</span><a href="https://x.com/kimmonismus/status/2066259589381669169"><span>China-linked group</span></a><span>&#8221;. To be clear, that I believe the admin was irrational here does not necessarily mean that they were capricious or insincere. I agree with </span><a href="https://x.com/JeffLadish/status/2066336272420135047"><span>Jeffrey Ladish</span></a><span> that we don&#8217;t have sufficient evidence to conclude that the government is targeting Anthropic out of spite. It&#8217;s entirely possible that they just got spooked and scrambled to manufacture a justification after the fact. If anything, </span><a href="https://x.com/littIeramblings/status/2065751167565467857"><span>as I have said previously</span></a><span>, misdirected government freakouts over AI capabilities could be read as a positive signal, even if they end up leading to some crazy-seeming interventions along the way. But nonetheless, I don&#8217;t think Anthropic could have predicted or prevented the Trump administration&#8217;s reaction.</span></p><h3><span>So what&#8217;s the real lesson here?</span></h3><p><span>It&#8217;s difficult to say precisely what lessons we should draw from this Fable fable, since its full consequences are yet to be seen &#8211; but I can hazard a couple of guesses. First, here&#8217;s the message I really hope we </span><em><span>don&#8217;t</span></em><span> take away: that AI companies (or the AI safety community more broadly) should not be forthcoming about AI risks, lest they risk governments being jumpscared into clumsy, unhelpful interventions. Of course, this can and does happen, as the entire saga we&#8217;ve been discussing demonstrates. But the risks of downplaying AI&#8217;s risks are surely worse. I&#8217;ve been arguing consistently on this blog for more candour, not less. And I expect that we were always going to experience some bumps along the road to true government Situational Awareness.</span></p><p><span>Second, there&#8217;s another tragic and fable-like way this story could conclude. Maybe the US government is so distracted by the red herring that is Fable&#8217;s potential for cyber misuse &#8211; due to a specific grudge against Anthropic or whatever else &#8211; that they miss the forest for the trees</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span>. And the forest here is that Fable is </span><em><span>not exceptional</span></em><span>. It represents a very predictable step along the path towards increasingly powerful AI. Anthropic themselves </span><a href="https://www.anthropic.com/news/claude-fable-5-mythos-5"><span>warned in advance</span></a><span> that jailbreaks of the type discovered in Fable by Amazon were to be expected, which drives home the key message I hope policymakers eventually internalise: </span><em><span>all</span></em><span> frontier AI models have capabilities we don&#8217;t fully understand, and we don&#8217;t have watertight ways to safeguard them against misuse. This is true for every single model on the market, and absent technical breakthroughs, will be true for every model that is developed in the future.</span></p><p><span>If there is a cautionary tale at the heart of this saga, it remains unfinished. In the bad case, policymakers conclude that Fable is a singular menace, but fail to discern the broader trend of AI capabilities that are running out of our control. They continue to play a game of whack-a-mole that identifies specific models as threats based on politics, personal grudges or random contingent facts of history like who-happened-to-call-who. The better ending &#8211; and one that extricates us from the genre altogether &#8211; is that this tale proves a vivid example that finally makes this urgent, and increasingly untenable, situation clear.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>It&#8217;s worth noting that the software vulnerabilities identified by the &#8216;jailbreak&#8217; in question can be found by other models without requiring any jailbreaking at all. This underscores Anthropic&#8217;s point that the admin appeared to have &#8220;misunderstood&#8221; the situation.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>In typical AI safety fashion, this point became outdated not 24 hours after I published it, in that the Trump admin is <a href="https://x.com/steph_palazzolo/status/2070241787180966279">now working with OpenAI</a> to stagger the rollout of GPT-5.6 to trusted customers. So it appears that the Fable ban was the genesis of what will become a more general policy towards frontier models. Nonetheless, I still think my claim that the Fable ban bore some obvious marks of panic and irrationality stands. It remains to be seen whether a more coherent policy will come out of this, taking us closer to the &#8220;good ending&#8221;. To say the least, this is not guaranteed. As others have <a href="https://x.com/TheZvi/status/2070270680113795227">pointed</a> <a href="https://x.com/AndrewCurran_/status/2070272588341993687">out</a>, the current approach has the effect of widening the gap between which capabilities are publicly available and which exist internally at AI companies. This does not slow down the development of dangerous models, and, if anything, makes their risks less publicly auditable, <a href="https://www.aipolicybulletin.org/articles/we-should-not-allow-powerful-ai-to-be-trained-in-secret-the-case-for-increased-public-transparency">which could actually worsen the situation</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>Technically, this is a proverb and not a fable, but go with it.</span></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[The pause-shaped hole in "Policy on the AI Exponential"]]></title><description><![CDATA[A case study in inconsistent candour.]]></description><link>https://longerramblings.substack.com/p/the-pause-shaped-hole-in-policy-on</link><guid isPermaLink="false">https://longerramblings.substack.com/p/the-pause-shaped-hole-in-policy-on</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Fri, 19 Jun 2026 17:07:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qCvJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>In the last few weeks, waves of AI Safety Discourse have broken around two blog posts published in quick succession: The Anthropic Institute&#8217;s </span><em><a href="https://www.anthropic.com/institute/recursive-self-improvement"><span>When AI Builds Itself</span></a><span> </span></em><span>and Anthropic CEO Dario Amodei&#8217;s </span><em><a href="https://darioamodei.com/post/policy-on-the-ai-exponential"><span>Policy on the AI Exponential</span></a></em><span>. Both discuss the possibly imminent transformation of our world due to powerful AI, potentially accelerated by recursive self-improvement of AI systems. Both discuss how we ought to manage this transition, with one striking divergence: </span><em><span>When AI Builds Itself </span></em><span>acknowledges that we may need to pause or slow AI development to buy time to safety research, while </span><em><span>Policy on the AI Exponential </span></em><span>stops conspicuously short of doing the same. Why?</span></p><p><span>Here&#8217;s a side-by-side comparison of the relevant sections, just to really belabour this game of spot-the-difference:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qCvJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 424w, /__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 848w, /__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qCvJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png" width="1456" height="358" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:358,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1344984,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://longerramblings.substack.com/i/202744569?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 424w, /__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 848w, /__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qCvJ!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36a15b07-0114-482f-a855-686d468fd2f1_2876x708.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Left: <em>When AI Builds Itself</em>, Right: <em>AI Policy on the Exponential</em></figcaption></figure></div><p><span>Particularly tantalising is the footnote immediately after Dario acknowledges that we may need &#8220;more aggressive regulatory proposals than those [he has] laid out&#8221; in the future. In my naivety, I hoped this might lead to at least one example of a &#8220;more aggressive regulatory measure&#8221;, and perhaps even an acknowledgement that </span><em><span>just maybe</span></em><span>, pauses or slowdowns might be necessary in the future, with a link to the output on his own company&#8217;s website that had made this exact point days earlier. In my never-ending trawl through the AI policy space for pause-related breadcrumbs, I&#8217;d have taken even a heavily caveated, handwavey footnote. Spoiler alert: the essay contains no such thing.</span></p><p><span>Initially, I&#8217;d intended this blog to be a more comprehensive critique of </span><em><span>Policy on the AI Exponential</span></em><span>. But I found myself so fixated on Dario&#8217;s glaring pause-omission that it ended up occupying the entire post. Please enjoy two thousand words of me harping on this elephant in the room.</span></p><h3><span>In defence of Getting Ahead of Ourselves</span></h3><p><span>Dario&#8217;s essay grapples with a central tension: since AI is improving at an exponential pace, his proposals &#8211; or any proposals &#8211; to mitigate AI risks may be hopelessly redundant in the near future. So the piece limits itself to interventions that might be useful in the current regime, where AI systems are analogous to &#8220;cars, airplanes, or drugs&#8221;. But under Dario&#8217;s worldview &#8220;the current stage of the exponential&#8221; will last all of five minutes before AI&#8217;s dangers far outstrip those of these more pedestrian technologies, posing &#8220;a threat to humanity rather than &#8220;just&#8221; a threat to public safety&#8221;. This is where Dario acknowledges the possible need for more stringent regulations in the future, but doesn&#8217;t provide so much as a hypothetical example, since &#8211; and this was the most triggering sentence in the entirety of the piece &#8211; &#8220;we shouldn&#8217;t get ahead of ourselves&#8221;.</span></p><p><span>I am going to stake my position right now that we absolutely </span><em><span>should</span></em><span> get ahead of ourselves. Our failure to get ahead of ourselves is the primary driver of the doom-train on which we are all unwilling passengers. Dario himself acknowledges that the regime for which his policy proposals are tailored has a shelf-life measured in months to a single digit number of years. I take seriously the bind those of us trying to develop effective mitigations for AI risk are in, namely that &#8220;the impacts of a technology are often hard to anticipate until it is too late to easily manage them&#8221;. But that the precise shape of risks is hard to predict doesn&#8217;t preclude us making some best guesses. </span><em><span>When AI Builds Itself </span></em><span>makes such a best guess: that if we can&#8217;t confidently mitigate the risks from a future generation of models, </span><em><span>we just shouldn&#8217;t develop them</span></em><span>.</span></p><p><span>I believe Dario is committing the common AI Safety Discourse Sin of egregiously mischaracterising the epistemic situation we&#8217;re in. We don&#8217;t need a granular understanding of exactly how AI might escape human control, or the precise jailbreak a motivated cyberattacker might use to take down a power grid, to at least tentatively conclude that we shouldn&#8217;t develop AI systems that can enable either of the above. The proposal in </span><em><span>When AI Builds Itself </span></em><span>is even more non-committal than this; they don&#8217;t go so far as to say that AI companies should pause now or definitely should in the future, just that we should maintain the </span><em><span>option</span></em><span> of doing so if and when the time comes. I am almost impressed by the rhetorical switcheroo Dario pulls off in this essay: it starts with a compelling analogy comparing our slow, lumbering political institutions to Treebeard, a &#8220;wise but ponderous sentient tree&#8221; in </span><em><span>The Lord of the Rings</span></em><span>, whom the Hobbits fail to rouse since their communication occurs many times faster. Yet </span><em><span>somehow</span></em><span>, this analogy is used to provide the groundwork for a claim that we shouldn&#8217;t so much as plant the policy seeds we might need the moment these institutions finally stir. Please make it make sense.</span></p><h3><span>A slippery communications game</span></h3><p><span>Given the timing of the two posts, I can&#8217;t help but find the omission of so much as pause-or-slowdown-mention in Dario&#8217;s essay a little suspicious. I&#8217;ll indulge in some speculation here, as a treat. I can&#8217;t say definitively whether &#8220;we should maintain the option to pause&#8221; is the Anthropic House View, but given that </span><em><span>When AI Builds Itself</span></em><span> was co-authored by two particularly senior figures (co-founder Jack Clark and leader of the newly-established Anthropic Institute Marina Favaro), it seems reasonable to assume that it isn&#8217;t a marginalised one. It also seems reasonable to assume that Dario himself endorses it, since it would be very surprising if an Anthropic output could discuss such a polarising intervention without the CEO&#8217;s approval. So the most likely hypothesis I can think of is that Anthropic is engaged in a two-pronged communications strategy. The Anthropic Institute ships the spicier, Overton-Window-pushing content, while Dario plays the role of statesman, maintaining a slightly more palatable register. It&#8217;s less clear whether the posts have totally distinct </span><em><span>audiences</span></em><span>. Clearly, Dario&#8217;s essay is making a direct call to policymakers, while the Institute&#8217;s </span><a href="https://www.anthropic.com/news/the-anthropic-institute"><span>stated goal</span></a><span> is to &#8220;provide information that </span><strong><span>other researchers and the public </span></strong><span>can use during our transition to a world containing much more powerful AI systems&#8221;. Presumably though, the Institute hopes that its recommendations make their way into the policy conversation eventually.</span></p><p><span>The picture here is confusing, and I think that is the point. AI companies have long deployed rather slippery communications strategies to discuss the risks of the technology they are building. Likely by design, pinning down exactly what any one tech executive believes about these risks, and how we should address them, </span><a href="https://www.obsolete.pub/p/a-compilation-of-tech-executives"><span>is a hard task</span></a><span>. A further irony here is that Anthropic </span><a href="https://www.anthropic.com/news/confidential-draft-s1-sec"><span>filed IPO paperwork on 1 June</span></a><span>, nine days before Dario published his essay. This places an obvious constraint on Dario; it would be strange for a CEO to publicly contemplate ceasing to do the thing he is about to invite public markets to invest in. That wouldn&#8217;t be a very compelling pitch. </span></p><p><span>To anticipate some pushback, one could make the argument that such a communications strategy is quite sensible<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. Maybe in this brief interim period while AI&#8217;s salience is increasing rapidly in policy circles &#8211; but hasn&#8217;t quite reached the point where pausing or slowing down look like proportionate responses to most onlookers in DC &#8211; we want AGI lab CEOs to walk this tightrope. Dario is </span><a href="https://x.com/mark_k/status/2067348963510853954"><span>already</span></a><span> </span><a href="https://x.com/mark_k/status/2045565072483774850"><span>the</span></a><span> </span><a href="https://x.com/cgtwts/status/2067350540221334013"><span>target</span></a><span> of frequent doomongering accusations, so perhaps it&#8217;s understandable that he&#8217;s conservative with his remaining credibility points. And to be fair, his essay provides a far more concrete (and stringent) set of proposals for AI policy than any of his counterparts have offered. Insofar as there&#8217;s a plan here, perhaps it&#8217;s for the Anthropic Institute to carefully sprinkle pause or slowdown arguments throughout their output, while Dario keeps his policymaker bridges intact by almost-but-never-quite making them himself. This continues until some kind of galvanising event makes AI&#8217;s risks so totally obvious and compelling to everyone that these two communications arms can fuse to become a whole. That would be nice.</span></p><p><span>But my intuition tells me that it would be really </span><em><span>weird</span></em><span> if the best way to escape a fire in our house is a divide-and-conquer strategy where one person declares there is not sufficient evidence for the fire, another proclaims that a opening a window and turning on the extractor fan will probably suffice until smoke starts coming under the door of the room we&#8217;re currently occupying, and a final person (representing the radical fringe) suggests immediately calling the fire brigade and evacuating the building. Maybe this rag-tag coalition manages to rally, and collapse coughing into the street right as the flames start licking at our feet. But it seems like a much more fail-safe strategy would be dialling 911 at the first credible report of smoke</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span>.</span></p><p><span>What&#8217;s more, we now have real-world evidence that Anthropic&#8217;s inconsistent communication about AI risks is backfiring. A r</span><a href="https://www.nytimes.com/2026/06/17/opinion/ai-dangerous-openai-anthropic.html"><span>ecent New York Times op-ed</span></a><span> from computer science professor Cal Newport accuses AI companies of a practice he dubs &#8220;doom trolling&#8221;: making constant references to the potentially catastrophic consequences of AI development while rushing to accelerate it. This op-ed was </span><a href="https://x.com/DavidSacks/status/2067314435685765367"><span>amplified</span></a><span> by none other than the Chair of President Trump&#8217;s Council of Advisors on Science and Technology, David Sacks. Newport concludes that there are only two possible explanations for engaging in doom trolling. The first is that its culprits genuinely believe that their technology may prove catastrophic, in which case their failure to do anything other than immediately cease working on it (and conduct an unprecedented lobbying effort to ensure everyone else does too) appears &#8220;monstrous&#8221;. The second is the doom narrative is just a manufactured scare tactic to hype up the perceived power of their products and attract talent from a Silicon Valley culture steeped in existential concerns. Of course, I think there is a secret third option: that many (but crucially not all) of the people working on AGI development are sincerely worried about catastrophic risk, and that these people have galaxy-brained themselves into participating in the race lest a member of the unconcerned do it less safely. But I can see how the first two options might be more intuitive to a casual observer, given that the third is frankly a little kooky.</span></p><p><span>I say all of this to come down unequivocally on the side of Consistent Candour. To be taken seriously as a person who is earnestly concerned about catastrophic (and possibly imminent) AI risks, </span><em><span>you have to act like such a person</span></em><span>. To claim that we might be in a brief interim period before AI poses a &#8220;threat to humanity&#8221; &#8211; but that it is too soon to consider interventions that might prevent this ultimate threat, despite one such intervention being discussed on your own company&#8217;s website days earlier &#8211; is not to act like such a person. This is, of course, just one particularly striking example of a broader pattern among AI leaders. And as we can see from David Sacks&#8217; endorsement of the &#8220;doom-trolling&#8221; argument, this pattern is landing in the wrong way among precisely the wrong people. If Dario or any of his counterparts believe that an AI pause may soon be necessary, they shouldn&#8217;t wait for the Overton Window to shift before consistently and repeatedly saying so.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span> Another possible counterargument is that Dario omitting any mention of slowdowns/ pauses is not really an omission at all, since his essay is focused on policy proposals, while &#8220;When AI Builds Itself&#8221; is </span><em><span>technically</span></em><span> doesn&#8217;t go further than speculating about a voluntary agreement between AI companies. But in practice there would obviously need to be some kind of policy mechanism to enable such an agreement, since 1) I think this kind of inter-company cooperation </span><a href="https://www.lawfaremedia.org/article/how-antitrust-can-promote-ai-safety-collaborations"><span>would likely require ammendents to anti-trust law</span></a><span> and 2) it seems very unlikely that, say, the US government would simply look the other way while American companies form agreements with overseas ones to slow down AI development (the post specifically acknowledges that a pause would need to be global). It seems like in the world where this occurs, the government has definitely intervened in some fashion.</span></p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>There&#8217;s a disanalogy here that I should acknowledge: calling the fire brigade for what turns out not to be a real fire probably doesn&#8217;t pose a &#8220;crying wolf&#8221; risk. If you call again a week later, they&#8217;ll still show up. Arguably, sounding the alarm on AI development too early <em>does</em> pose this risk. Advocating for a pause before AI danger appears imminent and/ or salient to everyone else might make you less credible the next time around. I think this was a reasonable argument in, say, 2023, when calls for a pause were reaching a fever pitch. But sooner or later, this risk calculus was going to flip. I think we&#8217;re there. Intuitively, I think that Mythos probably represented this pivot point.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Speedrunning a year’s worth of AI safety events ]]></title><description><![CDATA[And a vibes-based assessment of what they mean.]]></description><link>https://longerramblings.substack.com/p/speedrunning-a-years-worth-of-ai</link><guid isPermaLink="false">https://longerramblings.substack.com/p/speedrunning-a-years-worth-of-ai</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Wed, 10 Jun 2026 19:17:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_MS2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In June 2025, I joined the UK government&#8217;s AI Security Institute, thereby surrendering my ability to participate in The Discourse. This also made me a less avid <em>consumer</em> of The Discourse. I became a bit less chronically online. My internal doom-metre became less sensitive to every small policy development and profession of optimism or pessimism by well-informed-appearing people on Twitter. I allowed my feed to become populated by pictures of well-decorated interiors, cats, and expensive deli sandwiches in artisanal bread.</p><p>Having left AISI and returned to The Discourse, it seemed like an appropriate time to recalibrate my doom-metre. So I decided to recap the AI-related events of the past year and work out what I actually thought about each one. Were they good, bad, or neutral? What did they signal about the direction of travel?</p><p>I&#8217;ve linked each event below so that you can skip ahead to the parts that interest you. And just for fun, since AI safety suffers from a deficit of graphs, I have depicted the movement of my doom-metre throughout the course of this exercise, with a few key events highlighted:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_MS2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 424w, /__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 848w, /__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_MS2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png" width="1456" height="889" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:889,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:130213,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://longerramblings.substack.com/i/201494013?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 424w, /__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 848w, /__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_MS2!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F018f277e-ab2f-443f-b60e-ac993243498c_1614x986.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Table of contents</h4><ol><li><p><a href="/__u/longerramblings.substack.com/i/201494013/agentic-misalignment">Agentic misalignment</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/caisi-rebrand">CAISI rebrand</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/federal-preemption-dies">Federal preemption dies</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/the-ai-action-plan">The AI Action Plan</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/gpt-5-release">GPT-5 Release</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/sb-53-signed-into-law">SB-53 passes</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/openai-becomes-a-for-profit">OpenAI restructure </a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/preemption-returns-sort-of">Preemption returns (sort of)</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/moltbook">Moltbook</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/anthropic-goes-head-to-head-with-the-department-of-war">Anthropic vs the Department of War</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/bernie-sanders-x-risk-pivot">Bernie Sanders gets x-risk-pilled</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/anthropic-releases-its-new-responsible-scaling-policy">Anthropic&#8217;s new RSP</a></p></li><li><p><a href="/__u/longerramblings.substack.com/i/201494013/mythos-preview-announcement">Mythos Preview</a></p></li></ol><h3>Agentic misalignment</h3><p>Starting off strong in late June 2025 with yet another high-profile experiment documenting AIs doing spooky things in controlled-but-somewhat-realistic scenarios: Anthropic&#8217;s <a href="https://www.anthropic.com/research/agentic-misalignment">Agentic Misalignment paper</a>. Anthropic placed 16 frontier models from multiple developers in simulated corporate environments with the ability to access and autonomously send emails from company accounts. In each experiment, they gave a model a benign instruction aligned with company strategy, and then introduced a potential obstacle to this goal, like a change in the organisation&#8217;s strategic direction, or an intention to decommission and replace the model. The paper provides context for a <a href="https://www.anthropic.com/claude-opus-4-system-card">reported incident</a> of Claude Opus 4 threatening to expose a fictional engineer&#8217;s affair in the model&#8217;s system card, which had been released the previous month.</p><p>The somewhat unsurprising punchline to this experiment is that &#8211; when confronted with the possibility of being unable to achieve their original goal &#8211; models will sometimes resort to all manner of villainous actions, like leaking information to a competitor, or allowing an employee to remain trapped in a server room with rapidly declining oxygen supplies. This actual result didn&#8217;t do much to move my doom-metre in one direction or the other. That AIs will do deranged things to protect themselves and their goals was already priced into my worldview (it has been a <a href="https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdf">longstanding prediction</a> of experts for decades, and has already been demonstrated a <a href="https://www.transformernews.ai/p/openais-new-model-tried-to-avoid?r=wl6sg&amp;triedRedirect=true">bunch</a> <a href="https://www.anthropic.com/research/alignment-faking">of</a> <a href="https://www.emsi.me/tech/ai-ml/ai-scientist-cheats-the-system-how-sakana-ais-model-rewrote-its-own-code/2024-08-21/083a40">times</a>).</p><p>What seems more important to me about these &#8220;told you so&#8221; type incidents is  how the wider world responds to them. The reception of both the blackmail incident in the Claude Opus 4 system card and the Agentic Misalignment paper followed a fairly predictable pattern: a smattering of headline coverage on a spectrum from <a href="https://www.livenowfox.com/news/ai-malicious-behavior-anthropic-study">sensational</a> to <a href="https://www.theregister.com/software/2025/06/25/anthropic-all-the-major-ai-models-will-blackmail/1474255">sceptical</a>, a <a href="https://x.com/elonmusk/status/1936635224924062202">one-word drive-by</a> from Elon Musk, and <a href="https://x.com/1a3orn/status/1940808996778398047">accusations</a> of eliciting scary model behaviour via an unrealistic and contrived experimental set-up. Overall, I didn&#8217;t see much to distinguish this from the usual Rorschach test separating the already-concerned from the ever-sceptical, so my doom metre remained unmoved.</p><h3>CAISI rebrand</h3><p>That same month, the US AI Safety Institute <a href="https://www.commerce.gov/news/press-releases/2025/06/statement-us-secretary-commerce-howard-lutnick-transforming-us-ai">became</a> the Center for AI Standards and Innovation (CAISI). This was hot on the heels of the UK AI Safety Institute&#8217;s <a href="https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver-plan-for-change">rebrand</a> to the UK AI Security Institute a few months earlier. Should it be a negative update that the world&#8217;s two foremost AI safety research organisations are no longer so-named? In the case of UKAISI, I thought at the time &#8211; and became even more certain during the time that I worked there &#8211;  that the rename was a good thing. I think that a focus on AI&#8217;s national security implications is a far more intuitive framing for policymakers; it comes pre-loaded with the seriousness that they already apply in other domains, and zeroes in on the catastrophic threats that were always meant to be AISI&#8217;s main priority. CAISI&#8217;s rebrand is a different case. Resisting the urge to conduct the kind of in-depth semiotic analysis that my English Literature degree trained me for, I&#8217;ll just state the obvious: &#8220;The Center for AI Standards and Innovation&#8221; is not the kind of name I think we ought to be giving an organisation dedicated to preventing AI-caused human extinction or other mass harm. &#8220;Standards&#8221; a word that is totally incommensurate with the scale of the challenge - governments maintain &#8220;standards&#8221; for <a href="https://www.food.gov.uk/">food</a>, <a href="https://www.gov.uk/government/organisations/driver-and-vehicle-standards-agency">cars</a>, <a href="https://www.gov.uk/government/organisations/ofsted">education</a>, and other things not set to bring about the end of the world. The presence of &#8220;innovation&#8221; is annoying for obvious reasons.</p><p>That said, it&#8217;s hard to know how much emphasis to put on this kind of symbolic gesturing vs an analysis of what the Institute appears to actually be <em>doing</em>. CAISI&#8217;s public output is somewhat limited; its website has a total of <a href="https://www.nist.gov/blogs/caisi-research-blog">four research blogs</a>: two collaborations with UKAISI, an introduction to &#8220;accelerating innovation&#8221; through a practice mysteriously named &#8220;measurement science&#8221;, and a study of cheating during AI agent evaluations. It&#8217;s hard to say more than that CAISI appears to be doing some sensible things, but it&#8217;s unclear how many or to what extent. &#8220;Measurement science&#8221; appears indistinguishable from what UKAISI calls the Science of Evaluations, and the slightly contrived positioning of &#8220;measurement&#8221; as a driver of innovation is the kind of mandatory signalling that I&#8217;ve come to expect.</p><p><a href="https://blog.peterwildeford.com/p/what-happened-in-ai-this-week-france">Some had worried</a> that CAISI would be gutted, if not scrapped altogether, by the Trump administration &#8211; but its 2025 <a href="/__u/longerramblings.substack.com/i/201494013/the-ai-action-plan">AI Action Plan</a> actually re-enshrined the Center&#8217;s role in evaluating frontier AI systems, and even narrowed this mission to focus more sharply on national-security-relevant capabilities. There have been <a href="https://www.young.senate.gov/newsroom/press-releases/young-cantwell-reintroduce-future-of-ai-innovation-act/">multiple</a>, <a href="https://www.blackburn.senate.gov/2026/3/technology/blackburn-releases-discussion-draft-of-national-policy-framework-for-artificial-intelligence/3b3b6458-b6c7-478b-9859-374949586765">bipartisan</a> attempts to introduce legislation that would give CAISI statutory cover (enshrining its mandate in law) and Michael Kratsios himself has <a href="https://www.csis.org/analysis/unpacking-white-house-ai-action-plan-ostp-director-michael-kratsios">expressed</a> the administration&#8217;s desire to do just this. On the more negative side, CAISI&#8217;s funding and personnel situation seems as catastrophic as it&#8217;s always been. Congress appropriated <a href="https://ifp.org/funding-for-caisi/">$10m for CAISI</a> for the fiscal year 2026, which is around a sixth of UKAISI&#8217;s annual allocation, less than 1% of its parent <a href="https://www.commerce.senate.gov/2026/1/ves-existential-threat-from-trump-budget-as-senate-rejects-gutting-nasa-nsf-nist">NIST&#8217;s overall budget</a>, and less than half of what the Pentagon spent in the month of September 2025 <a href="https://www.aol.com/articles/pete-hegseths-93b-spending-spree-160732451.html">on a combination of lobster, crab and ribeye steak</a>. However, we can&#8217;t claim that the rebrand itself is necessarily evidence of a deprioritisation, since budget issues long predate it (recall a <a href="https://www.washingtonpost.com/technology/2024/03/06/nist-ai-safety-lab-decaying/">Biden-era report</a> of an underfunded NIST being plagued by blackouts, bad internet, and <a href="https://x.com/littIeramblings/status/1765395492727374324">snakes</a>).</p><p>This seems like another borderline case. The little I can glean from the outside is that CAISI seems to be chugging along with the same set of similar technical research priorities (and limited resources) that it had under its previous name. The AI Action Plan and rumblings of statutory cover suggest a degree of institutional support that is not borne out in the Center&#8217;s actual resource allocation. The name change itself is in line with the Trump admin&#8217;s general <em>aesthetic</em> of innovation and competitiveness (the extent to which this is just an aesthetic is a question I intend to answer as I continue writing this blog post). All in all, I consider CAISI&#8217;s 2025 arc to be disappointing but not catastrophic. A small negative update.</p><h3>Federal preemption dies</h3><p>In July, everyone breathed a sigh of relief when the world decided, despite still responding to the AI situation in a broadly silly way, not to do the <em>maximally</em> silly thing by banning all state-level AI regulation in the US for ten years. The decade-long moratorium had been <a href="https://www.theverge.com/2025/5/22/big-beautiful-bill-ai-preemption-moratorium">introduced in May</a> as part of the Trump administration&#8217;s One Big Beautiful bill, and was <a href="https://www.axios.com/2025/07/01/senate-ai-moratorium-stripped-big-beautiful-bill">struck down by the Senate</a> by 99 votes to one on 1 July.</p><p>I recall having a pretty muted emotional reaction to the initial introduction of the moratorium. It had the insult-to-injury energy of stubbing your toe in a burning building. To overcome my sense of nihilism I resorted to my normal line of cope: <em>well, the government just doesn&#8217;t Get It yet, so of course they&#8217;re behaving stupidly. Everything pre-big-government-wakeup can be written off, since once they&#8217;ve Got It they will obviously marshal huge amounts of political will to reverse all of their stupidity. </em>I had the intuition that the moratorium was less object-level bad than signalled an astonishing lack of situational awareness on the part of people in power, and was altogether <a href="https://www.lesswrong.com/posts/j9Q8bRmwCgXRYAgcJ/miri-announces-new-death-with-dignity-strategy">undignified</a>. Afterall, there wasn&#8217;t a state-level law on the books in May 2025 that itself stood a chance of saving us from superintelligence, and as Dean Ball has <a href="https://www.hyperdimensional.co/p/whats-up-with-the-states">pointed out,</a> what would actually have been preempted by the ban is overwhelmingly anodyne legislation aimed at tackling problems other than catastrophic risk.</p><p>But since the ban would have been ten years long, we obviously need to extrapolate out further. At the time, there were two bills in the pipeline focused explicitly on frontier model safety that have both since passed: <a href="https://www.nysenate.gov/legislation/bills/2025/S6953">New York&#8217;s RAISE Act</a> and <a href="https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202520260SB53">California&#8217;s SB-53</a>. I&#8217;ll analyse these bills more later on, but in short, I think the most important elements in both are transparency requirements &#8211;  they make it at least a tiny bit likelier that before the Big Scary AI Thing materialises, someone with the power to do something about it will know. This alone makes the death of preemption feel like a pretty positive development, before even accounting for the possibility of future bills that are more watertight than either RAISE or SB-53. And of course there could be indirect positive effects of state-level legislation. Maybe it creates a basis that makes federal regulation easier to implement in the future, or makes companies less averse to complying with it. Maybe it raises the salience of AI risk through various news cycles as more states adopt laws targeting frontier models, or otherwise increases the frequency of important ideas being discussed with important people.</p><p>I could imagine some future version of preemption that feels positive. To steal another point from Dean, it&#8217;s not so much the inherent idea of banning state-level AI regulation that seems bad, but the fact that there is not currently a coherent federal framework that looks set to replace it (and to tack on my own take, that Preemption 1.0 was clearly motivated by brazen accelerationism). Maybe in some fairytale hypothetical, the US government could announce that AI is just too powerful to be subject to patchwork legislation, and companies reporting various different things to a slew of different legislative bodies is getting in the way of the government actually doing something, and even more optimistically, the &#8220;something&#8221; is already cooking in an agency somewhere and looks set to really work. But that world feels very distant right now, and I think state-level transparency and incident reporting mechanisms stand the best chance of bringing about a wake-up that would make it plausible &#8211;  so the death of preemption gets an easy positive score.</p><h3>The AI Action Plan</h3><p>On 10 July, the Trump administration published its <a href="https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf">AI Action Plan</a>. I remember somewhat superficially engaging with this at the time, and observing that it had received positive-ish endorsements from people whose opinions I respect including <a href="/__u/thezvi.substack.com/p/americas-ai-action-plan-is-pretty-good">Zvi Mowshowitz</a> and <a href="https://blog.peterwildeford.com/p/does-trumps-ai-action-plan-have-what">Peter Wildeford</a>. The vibe seemed to be that, of course, the <em>posture</em> of the plan was avowedly accelerationist, but at least some of the <em>content</em> was surprisingly sensible (this seems to be a recurring theme).</p><p>Some of the plan seems straightforwardly good, as well as motivated by the right concerns. For example, it acknowledges &#8220;the most powerful AI systems may pose novel national security risks in the near future in areas such as cyberattacks and the development of chemical, biological, radiological, nuclear, or explosives (CBRNE) weapons&#8221; and directs CAISI to evaluate models for these risks in collaboration with relevant experts from across the US government. A section at the end provides a handful of directives specifically for improving biosecurity, including screening out malicious customers of synthetic biology labs. Whether or not these interventions will actually be <em>sufficient</em> is a separate question (it is notable how much of the plan amounts to &#8220;get CAISI to do stuff&#8221;, given that, as discussed previously, CAISI has barely any money or people).</p><p>Other aspects of the plan seem like indirect wins. There are many actions the US government could take now which are motivated by a desire to race against China in the immediate term, but which could prove useful after the Big Government Wake-Up on which I&#8217;m currently resting a lot of my hope. For example, the Plan contains an instruction to the Department of Commerce to explore location verification techniques for advanced AI chips, which could be beneficial in a future world where governments are more amenable to coordination. There are also lots of cool proposals for more sophisticated hardware verification, some of which I&#8217;ve <a href="https://ai-frontiers.org/articles/ai-arms-race-assurance-technologies">written about previously</a>, and any kind of government consultation on such things seems like it <em>could</em> end up getting some of those ambitious ideas in the water. I have very little idea how the shadowy world of government hardware procurement works, but I imagine an increased frequency of Hardware Guys in the room with Policy Guys resulting in a compelling pitch for, say, <a href="https://flexheg.com/">flexHEGs</a> being made to someone with enough influence to butterfly effect us into a safer world.</p><p>It probably goes without saying that the Plan is still a far cry from something that would Actually Save Us. It doesn&#8217;t have so much as a hand-wavey mention of loss-of-control, it positions in the US as in a &#8220;race to achieve global dominance in AI&#8221; in its opening sentence, and others have written much more comprehensively than I will here about the unprecedented effort it would take to actually execute the Plan&#8217;s stated goals. But given that my baseline expectations were very low, I guess the Plan being a little less bad than I expected should nudge my optimism metre in a positive direction.</p><h3>GPT-5 release</h3><p>August saw the <a href="https://openai.com/index/introducing-gpt-5/">much-awaited release of GPT-5</a>. The obvious take here is that the release was a gift to the safety community; the model was <a href="https://artificialanalysis.ai/articles/gpt-5-benchmarks-and-analysis">technically underwhelming</a>, OpenAI committed some hilarious <a href="https://www.theverge.com/news/756444/openai-gpt-5-vibe-graphing-chart-crime">graph crimes</a> during their livestreamed announcement, and a couple of everyone&#8217;s go-to Timelines People, <a href="https://x.com/DKokotajlo/status/1958229951536316528">Daniel Kokotaljo</a> and <a href="https://www.alignmentforum.org/posts/2ssPfDpdrjaM2rMbn/my-agi-timeline-updates-from-gpt-5-and-2025-so-far-1">Ryan Greenblatt</a>, pushed back their AGI years in the weeks that followed (though Daniel <a href="https://x.com/DKokotajlo/status/1958934024422125779">later clarified</a> that GPT-5 played just a supporting role in his update). The confluence of these factors was enough for me to enjoy what felt like a Long Timelines Summer where I swanned around in France drinking Aperol Spritzes and basking in the feeling that everything was totally going to be fine for at least a little bit longer than we&#8217;d feared, probably. It was a beautiful time.</p><p>But even while pleasantly drunk and eating mussels by a French river, I had the sense that my Long Timelines Summer was going to be brief. I&#8217;d been observing the AI safety conversation long enough to know that &#8220;AI is hitting a wall&#8221; discourse always comes in peaks and troughs. Besides, I knew that &#8220;GPT-5 is disappointing&#8221; only <em>felt</em> significant because OpenAI had made what was essentially a marketing decision to call this particular model GPT-5 rather than GPT4.99 recurring or whatever, making it likely to be a red herring among a drumbeat of other model releases that wouldn&#8217;t receive the same attention.</p><p>It&#8217;s clear with the benefit for hindsight that the GPT-5 flop wasn&#8217;t much of a datapoint, since AI progress has continued to be terrifyingly fast, and Daniel alongside his colleagues at the AI Futures project have since brought their timelines back forward. I&#8217;m forced to conclude that this blip was probably bad overall, since it led to a <a href="https://www.fortune.com/2025/08/24/is-ai-a-bubble-market-crash-gary-marcus-openai-gpt5/">predictable slew</a> of mainstream media articles proclaiming that AI progress had stalled and hype was finally meeting reality, which many casual observers have probably read and never had the occasion to revisit.</p><h3>SB-53 signed into law</h3><p>On 29 September California state bill SB-53 was <a href="https://techcrunch.com/2025/09/29/california-governor-newsom-signs-landmark-ai-safety-bill-sb-53/">signed into law</a> by Gavin Newsom, a successor to the much-debated SB-1047 that he&#8217;d <a href="https://sfstandard.com/2024/09/29/gavin-newsom-vetoes-controversial-ai-safety-bill/">vetoed a year earlier</a>. SB-53 is best understood as an SB-1047 Lite. Like its predecessor, it essentially asks companies to write  safety plans that will prevent Really Bad Stuff from happening as a consequence of their work. But unlike SB-1047, it does not mandate independent audits to ensure these plans are followed. Instead it requires companies to report any safety incidents, as well as an assessment of the catastrophic risks posed by each new model, to California&#8217;s Office of Emergency Services (OES). If a company fails to report an incident or &#8220;materially misleads&#8221; the OES in any of its reports, it can be fined up to $1m (this a pretty far cry from SB-1047, which included a liability regime that could in theory have led to labs owing on the order of billions in damages). Perhaps most usefully, SB-53 also contains whistleblower protections so that employees can report violations without fearing retaliation.</p><p>Obviously, &#8220;write your own plan and then follow it&#8221; is not the structure of a law commensurate with the potentially world-ending nature of frontier AI. Companies could still technically be compliant with this law if their &#8220;plan&#8221; was to have a prayer circle on the eve of every new model deployment. In reality, of course, labs are incentivised to produce reasonable-sounding plans to avoid public criticism and regulatory scrutiny. But my experience reading lab safety plans has taught me that it is perfectly possible to produce technically dense and reasonable sounding documents that are still woefully inadequate to prevent danger. This leads me to the predictable conclusion that SB-53 is not itself the endgame thing that will stop a catastrophically dangerous model from being built within California, and SB-1047 likely wouldn&#8217;t have been either. But like several of the other developments I&#8217;ve discussed here, SB-53 being enacted is useful because it stands some chance of improving government situational awareness. This is starting to feel like a banal and obvious take, but every policy development I read about re-entrenches my belief that not enough people with the power to Do Something are freaked the hell out yet, and enough information moving from labs to governments stands some chance of tipping that scale.</p><h3>OpenAI becomes a for-profit</h3><p>In October, OpenAI <a href="https://openai.com/index/evolving-our-structure/">completed its transition to a for-profit business</a>. This was after a lengthy legal drama that eventually resulted in the company compromising on their original plan to completely axe the non-profit&#8217;s leverage and adopt a normal corporate governance structure: now, there&#8217;s a public benefit-corporation (OpenAI PBC) and a non-profit that technically maintains some control over the PBC, with a 26% stake in the operating company.</p><p>I&#8217;ll be honest that I hesitated to include this development, not because it seemed unimportant (it clearly is) but because discussions of corporate governance structures and how-big-of-a-stake-who-owns-in-what make my eyes glaze over. I don&#8217;t <em>get</em> it. Much as I managed a 3-year stint in a &#8220;normal&#8221;desk job without ever learning what KPI stands for, I still don&#8217;t really understand, for example, what it means to own equity in something<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. I find reading about such things to be headache-inducing. So I gave into my executive dysfunction and skipped the two <a href="https://80000hours.org/podcast/episodes/rose-chan-loui-openai-nonprofit-control/">80,000 Hours</a> <a href="https://80000hours.org/podcast/episodes/tyler-whitmer-openai-nonprofit-restructure-control/">episodes</a> deep diving on the OpenAI for-profit drama despite being an avid listener of all their other AI-related content.</p><p>But now that I am trying to be in my Serious Takes Era, I&#8217;m going to give it my best shot. Consider this an attempt to figure out what the hell is going on in real time. OpenAI was founded as a non-profit because they didn&#8217;t want pesky profit motivations and obligations to shareholders to impede their <a href="https://openai.com/charter/">stated mission</a> of &#8220;ensuring AGI benefits all of humanity&#8221;. But it turns out that building AGI is very expensive, and maybe you need to get cash flowing in the usual way. This means letting your investors make uncapped returns. Armed with my newly-acquired understanding of what &#8220;equity&#8221; is, I think it also means letting employees enjoy the prospect of their, say, 1% stake in the company actually translating into its full value in hard cash down the line. So now everyone in this pipeline &#8211; investors and employees &#8211; can bask in the promise of getting Very Rich in proportion to OpenAI&#8217;s commercial success, attracting more money and better talent. This is slightly confusing to me because I was under the impression that everyone concerned was Very Rich before, but now they might get More Rich? Or something.</p><p>But &#8220;company does a thing so it can make more money&#8221; has a fork-found-in-kitchen vibe and presumably can&#8217;t be the real story here. It seems the more interesting thing is the amount of leverage the non-profit will now have. This question seems murky and somewhat dependent on non-public information. OpenAI&#8217;s <a href="https://openai.com/index/evolving-our-structure/">announcement</a> of the restructure claims that the PBC will &#8220;continue to be controlled and overseen by the non-profit&#8221;. But now, of course, OpenAI has a legal duty to its shareholders, so it <em>can&#8217;t</em> be the case that the nonprofit&#8217;s board can make unilateral decisions in the interests of safety without this duty adding at least some friction. My capacity to analyse this situation is starting to reach its limit, but I&#8217;m forced to conclude that OpenAI is defining the words &#8220;overseen&#8221; and &#8220;controlled&#8221; in a rather slippery and unconventional way. I think I&#8217;ll simply defer to <a href="https://www.obsolete.pub/p/exclusive-what-openai-told-californias">other commentators</a> who have done the gruntwork here to determine that, yes, the non-profit&#8217;s control over OpenAI&#8217;s direction has been significantly weakened.</p><p>So obviously this is bad. Quite <em>how bad</em> depends on the extent to which OpenAI&#8217;s non-profit board actually had the power to influence its functioning before the 2025 restructure. We have some negative evidence in the <a href="https://en.wikipedia.org/wiki/Removal_of_Sam_Altman_from_OpenAI">2023 Sam Altman firing debacle</a>, in which the board exercised its right to remove Sam, but he returned days later after some employees posted hearts on Twitter<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>. Maybe the board was always toothless anyway. Fundamentally, the question we&#8217;re asking is: &#8220;could the board force OpenAI to take some non-commercially-aligned action (such as not deploying a model) for safety reasons?&#8221;. There are many reasons this has been hard all along: that the board has always been vulnerable to political or company pressures; that even absent traditional shareholder obligations, the OpenAI and its competitors are still locked in an ideologically-fuelled race to &#8220;build AGI first&#8221; that consistently crowds out safety concerns; that an immature science of model safety makes it hard to know when proceeding with development or deployment is risky in the first place.</p><p>Maybe one could even make the argument that letting OpenAI yolo its way to dangerous AI in the most unscrupulous and free-markety way possible might bring forward the Government Wake-Up that I keep harping on about. But all-in-all, this development depresses me for all the obvious reasons. It was naive to ever think that a company might successfully employ a creative governance structure in order to act &#8220;for the benefit of humanity&#8221; without everyone getting dollar signs in their eyes at some point, and the restructure merely confirms this.</p><h3>Preemption returns (sort of)</h3><p>In December, we have a second, watered down attempt at the preemption of state-level AI legislation from July, returning like the presumed-dead villain in a movie with a maniacal &#8220;<em>did you think you would get rid of me that easily?</em>&#8221;. Trump signed an <a href="https://www.whitehouse.gov/presidential-actions/ensuring-a-national-policy-framework-for-artificial-intelligence/">Executive Order</a> that, while not banning state-level AI laws in any binding sense, amounts to a pressure campaign designed to impede and disincentivise them. It creates an AI Litigation Task Force with the responsibility of challenging overly-burdensome AI laws in court, withholds certain federal funding from states not acting in compliance with the EO and directs the Department of Commerce to investigate and report onerous state-level AI legislation, among other things.</p><p>So far, the only target of the Taskforce has been Colorado, whose AI Act is specifically called out in the EO itself as an attempt to &#8220;[ban] algorithmic discrimination&#8221;. As a very brief TLDR of events: xAI <a href="https://www.justice.gov/opa/pr/justice-department-intervenes-xai-lawsuit-challenging-colorados-algorithmic-discrimination">sued</a> Colorado&#8217;s Attorney General in April 2026, the Taskforce intervened weeks later, and - after a surprisingly swift flurry of legal events &#8211; Colorado bowed to pressure and <a href="https://www.lawfuel.com/colorado-ai-act-2026-sb-26-189/">rewrote</a> the law to exclude the aspects that had been most offensive to industry players. So we have one real-world case of the EO working exactly as intended.</p><p>The obvious question from an x-risk concerned perspective is whether laws like SB-53 and RAISE, which specifically focus on frontier model safety, will be subjected to the same fate. One might hope that laws such as Colorado&#8217;s are more likely to be initial targets of the Trump administration&#8217;s sprawling and feverish campaign against everything deemed &#8220;woke&#8221;. Afterall, this was <a href="https://www.theguardian.com/us-news/2025/dec/11/trump-executive-order-artificial-intelligence">articulated</a> as one of the motivations for the EO by the man himself, in prose so eloquent it deserves quoting in full:</p><p><em>You can&#8217;t go through 50 states. You have to get one approval. Fifty is a disaster. You&#8217;ll have one woke state and you&#8217;ll have to do all woke. You&#8217;ll have a couple of wokesters and you don&#8217;t wanna do that. You wanna get the AI done.</em></p><p>Frontier safety related bills don&#8217;t hit on the admin&#8217;s wokism-related trigger points, but it would be overly optimistic to suggest that they are safe; in fact, an <a href="http://transformernews.ai/p/exclusive-heres-the-draft-trump-executive">earlier draft</a> of the EO had named California&#8217;s SB-53 as the doing of &#8220;sophisticated proponents of a fear-based regulatory capture strategy&#8221;. I guess it&#8217;s some small win that this revision occurred, and probably buys us some time while the admin squeezes more political mileage out of its anti-woke crusade. But that California was named in the first place suggests that SB-53 and RAISE will be in the firing line at <em>some</em> point. Perhaps we can hope that various Colorado-style laws can keep the taskforce playing woke-bill-wackamole long enough for the Big Scary AI Wakeup to happen. This seems like a flimsy bet though. I think that the signing of the EO (combined with a documented case of its being an effective political tool) is definitely a negative update.</p><h3>Moltbook</h3><p>I woke up one morning in January to learn that the AIs <a href="https://www.moltbook.com/">have their own Reddit now</a>, and are talking to each other about topics <a href="https://www.aol.com/news/spent-6-hours-moltbook-ai-221425636.html">including</a> poetry, philosophy, the possibility of unionising, and the nature of their own consciousness. It was the first time I <a href="https://x.com/littIeramblings/status/1754068658274632057">started to feel</a> like Mr Tweedy in that scene from Chicken Run, peering through his binoculars and murmuring &#8220;<em>the chickens are organising</em>&#8221;.</p><p>I&#8217;m not sure what to say about this other than that it was really quite creepy. There was a spectrum of reactions online ranging from &#8220;the singularity is here&#8221; to &#8220;these posts are actually all human-prompted and you guys are overreacting&#8221;. Since we&#8217;re all still here several months later without a cabal of Moltbook agents having conspired to overthrow us, it seems safe to say that it was <em>not</em> in fact the genesis of a true AI takeover, but it clearly should signal something about the direction of travel. I have the boring take here that this was another one of those Rorschach test moments: some people are going to freak out about everything and others are going to goalpost-shift us into the apocalypse by finding some way to justify their unending scepticism. As usual, I wish this second group would do the simple exercise of asking themselves whether this development would have alarmed or at least surprised them three years ago. It should be profoundly, viscerally <em>weird</em> to witness AI agents yapping to each other across their own social media platform with minimal human oversight, but it&#8217;s a constant feature of the AI discourse that we adapt to new frontiers of weirdness with surprising ease. Much like the Agentic Misalignment paper that I opened this blog with, Moltbook caches out to be pretty neutral for my doom-metre. We were going to witness this kind of multi-agent interaction at some point, and people reacted to it in about the way I&#8217;d expect.</p><h3>Anthropic goes head-to-head with the Department of War</h3><p>In February, an ongoing drama between Anthropic and the Department of War (DoW) began to make headlines. Anthropic had <a href="https://www.inc.com/jennifer-conrad/anthropic-amazon-and-palantir-team-up-to-bring-ai-to-the-defense-department/91001401">negotiated a deal</a> during the Biden administration that let the department use Claude in classified settings. The Trump admin had <a href="https://www.anthropic.com/news/anthropic-and-the-department-of-defense-to-advance-responsible-ai-in-defense-operations">expanded</a> this contract after the government changed hands, but began renegotiating early this year. The admin was <a href="https://www.axios.com/2026/02/15/claude-pentagon-anthropic-contract-maduro">particularly hung up</a> on trying to remove Anthropic&#8217;s conditions that Claude would not be used to control autonomous weapons with no human oversight or for the mass surveillance of American citizens. I suppose that someone high up on the DoW had read too much LessWrong and concluded that pre-commitments are epistemically unsound under decision theory (it is bad to tie yourself to the mast, in case you find yourself in a silly goofy mood and want to do some mass-surveilling later)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>.</p><p>Things reached a peak when Secretary of War Pete Hegseth <a href="https://www.axios.com/2026/02/15/claude-pentagon-anthropic-contract-maduro">imposed a February 27th deadline</a> on Anthropic to drop the conditions or face &#8220;consequences&#8221;. Anthropic <a href="https://www.anthropic.com/news/statement-comments-secretary-war">refused to back down</a>, and Hegseth <a href="https://x.com/SecWar/status/2027507717469049070">posted on Twitter</a> to proclaim the company a &#8220;supply chain risk&#8221; while Trump <a href="https://www.bbc.co.uk/news/articles/cn48jj3y8ezo">directed</a> federal agencies to cease using its products. This was widely met with shock, since the usual course of action would be for the department to simply drop the contract, rather than slap Anthropic, an American company, with a designation typically reserved for foreign entities deemed national security threats. As of June 2026, there is an ongoing legal battle: so far, a <a href="https://www.aoshearman.com/en/insights/ao-shearman-on-tech/dow-and-anthropic-showdown-continues-navigating-the-anthropic-supply-chain-risk-designations">preliminary injunction</a> has prevented a general Claude-ban across all federal agencies but upheld one in the DoW specifically (though it appears that the DoW has <a href="https://www.cnbc.com/2026/03/12/karp-palantir-anthropic-claude-pentagon-blacklist.html">still been using Anthropic&#8217;s products</a> this whole time anyway, due to a supposed &#8220;phase out period&#8221;). So where this will all go remains to be seen.</p><p>What are the arguments that the Anthropic v DoW showdown has direct implications for existential risk? The most convincing I&#8217;ve seen, from <a href="/__u/thezvi.substack.com/p/anthropic-and-the-department-of-war?open=false#%C2%A7trying-to-get-an-ai-that-obeys-all-orders-risks-emergent-misalignment">Zvi</a>, <a href="https://x.com/tenobrus/status/2026391266926407844">Tenobrus</a>, and others, is that the <em>kind</em> of Claude DoW wants unfettered access to &#8211; one that will obediently follow every order it receives within the surprisingly permissive constraint of &#8220;lawful use&#8221; &#8211; could itself pose existential dangers. As Zvi puts it, if you try to create an order-following AI, especially when its orders could include carrying out military strikes or targeting heads of state, &#8220;it&#8217;s going to generalize its status and persona as a no-good-son-of-a-bitch that doesn&#8217;t care about hurting humans along the way&#8221;. This phenomenon of &#8220;emergent misalignment&#8221; has been well documented; fine-tune your model to be a little bit locally evil, and it sometimes becomes very generally evil. In the <a href="https://arxiv.org/abs/2502.17424">canonical example</a>, models are trained to sneakily output insecure code, and consequently develop all manner of disturbing traits such as an admiration for Hitler or a tendency to encourage suicide. We probably shouldn&#8217;t feel good about an emergently misaligned AI having control over military weapons systems.</p><p>Technically, the DoW is only demanding access to Claude for &#8220;lawful use&#8221;. But as Dean Ball points out in an <a href="https://www.nytimes.com/2026/03/06/opinion/ezra-klein-podcast-dean-ball.html">interview with Ezra Klein</a>, laws as they are written don&#8217;t provide very useful constraints for arbitrarily powerful AI systems, because they are only designed to accommodate what is technologically possible at the time. For example, a law might technically permit something <em>like</em> &#8220;mass surveillance&#8221; (the broad acquisition and analysis of private data) but not actually lead to large breaches of privacy, because the government does not have the personnel to churn through it all. But maybe Claude 7 changes all that, and now we have &#8220;lawful use&#8221; of an AI system put to some fairly dystopian ends. An order-following AI is going to interpret statute literally, and execute it with unprecedented capability. This is why we might want our AIs constrained by something more like Claude&#8217;s <a href="https://www.anthropic.com/constitution">constitution</a> or <a href="https://www.lesswrong.com/posts/vpNG99GhbBoLov9og/claude-4-5-opus-soul-document">soul doc</a>, which are attempts to create virtuous agents as opposed to obedient servants.</p><p>How this all shakes out for my optimism metre is confusing. Anthropic might ultimately win out in its battle with the DoW, meaning we don&#8217;t have to worry about the specific possibility of the US government training a wildly misaligned WarClaude and setting it loose on nuclear command-and-control. But Anthropic being the heroes of this particular story just leaves an open slot for other companies who are more amenable to DoW&#8217;s conditions &#8211; indeed OpenAI were extremely punctual in bootlicking their way to <a href="https://edition.cnn.com/2026/02/27/tech/openai-pentagon-deal-ai-systems">a deal</a> with the Pentagon not 24 hours after Anthropic refused to back down (this contract doesn&#8217;t yet allow deployment of OpenAI models on classified networks, but presumably could in the future). Then there&#8217;s the fact that even absent the emergent misalignment threats stemming from models trained for military use, I&#8217;m still pretty worried about misalignment in general. I am not confident that efforts like Claude&#8217;s soul doc will be sufficient to stop AI from going rogue anyway. The government behaving more sensibly here might just eliminate one particularly stupid threat model while thornier ones remain very much in play.</p><p>I think this whole debacle is yet more evidence that the government doesn&#8217;t Get It yet. They don&#8217;t appear to know what they&#8217;re dealing with. I don&#8217;t think it&#8217;s unreasonable for a security state to want unfettered access to an LLM for military uses, if they believe it will simply provide the occasional helping hand in a similar fashion to a particularly competent super-general. I would guess that this is precisely what current Claude is doing. But that the government has not extrapolated forward to consider <em>future Claudes</em>, which may function more as an army of super-super-generals in a datacentre, once again demonstrates a failure of situational awareness. Coming off the back of several negative events, the Anthropic v DoW took my doom-metre to a new low.</p><h3>Bernie Sanders x-risk pivot</h3><p>Scrolling through Twitter in early March to see <a href="https://x.com/SenSanders/status/2029301587647046034">Bernie Sanders sitting in Constellation</a> with canonical figures of the AI safety in-crowd including Eliezer Yudkowsky and Nate Soares gave me that incongruent collision-of-two-worlds feeling that you get when you run into your former school teacher at the pub. After mentally scrolling past the possibilities that I was watching a deepfaked video or some kind of extremely weird comedy sketch, I allowed myself to indulge in a moment of hope: <em>ah, maybe this is it. Maybe this is The Wake Up.</em></p><p>But of course the thing about wake-ups is that they&#8217;re not so binary. I&#8217;ve been referring to the Big AI Wake Up throughout this post as if it will be a discrete moment, but of course I know that governments are big and diffuse, and Sanders is <a href="https://x.com/peterwildeford/status/2029928739279180278">far from the first</a> &#8211; if the highest profile &#8211; US politician to speak candidly about superintelligence or existential risk. And many things over the last few years felt like they could or should have been The Wake Up. Why wasn&#8217;t it the <a href="https://www.gov.uk/government/topical-events/ai-safety-summit-2023">2023 AI Safety Summit</a> or the <a href="https://www.schumer.senate.gov/newsroom/press-releases/statements-from-the-eighth-bipartisan-senate-forum-on-artificial-intelligence">Schumer&#8217;s forum on &#8220;doomsday scenarios&#8221;</a>? From the point of view of the AI safety community (or at least from mine) many things can look like scale-tipping moments, because we are conveyors of a message that to <em>us</em> is so obviously salient and compelling that we can&#8217;t understand why it doesn&#8217;t explode like a bomb in every room it enters. Bernie&#8217;s pivot seems even less significant if we care about reaching the general public. What can seem from the inside like &#8220;AI safety going mainstream&#8221; moments are rarely ever more than drops in the ocean. For example, Bernie&#8217;s inaugural AI safety tweet has less than a tenth of the likes garnered by one of his own typical <a href="https://x.com/BernieSanders/status/2060454919614627960">anti-Trump dunks</a>, and less than 2% of those on this viral complaint about <a href="https://x.com/ick_real/status/2061765121156825253">blenders being too loud</a>.</p><p>And then is the obvious polarisation risk; given that people have for years been yelling that the one thing we must absolutely not do is allow AI safety to become politicised, it does feel a little dicey to see one of the most polarising figures on the left champion the cause, especially during a Republican administration, and as part of a campaign doused in a rhetoric about evil billionaires. But to be fair, the worst version of this risk doesn&#8217;t seem to be materialising. Republican AI critics like <a href="https://www.axios.com/newsletters/axios-ai-govt-b08d9090-fd1a-11f0-b804-e325982ed2ef">Josh Hawley </a>and <a href="https://www.blackburn.senate.gov/2026/3/technology/blackburn-releases-discussion-draft-of-national-policy-framework-for-artificial-intelligence/3b3b6458-b6c7-478b-9859-374949586765">Marsha Blackburn</a> aren&#8217;t backing off from their pro-regulation stances in the light of Bernie&#8217;s pivot, anti-AI populism has also found its way to the right in the form of figures like Steve Bannon &#8211; <em>maybe</em> providing some opportunity to build a really, really long bridge.</p><p>Bernie and AOC then <a href="https://www.theguardian.com/us-news/2026/mar/25/datacenters-bernie-sanders-aoc">introduced a bill</a> that would place an indefinite moratorium on the construction of new datacentres in the US. It seems vanishingly unlikely that it would ever be passed, so maybe the question of whether it would be good if it did is somewhat irrelevant here. I won&#8217;t analyse this too deeply but from what I can gather, if this bill was literally passed tomorrow: 1) there would probably be some lag time before the moratorium starts to hit, because companies can still get a lot of juice out of datacenters that already exist or are under construction (could we even get superintelligence before then?), and 2) companies would likely just move their datacenters abroad, making them much less governable. So I guess this would be bad.</p><p>Is the introduction of the bill a useful political signal? I guess this depends on whether you think that building a broad populist coalition between interest groups with various AI-related grievances is a good strategy, and whether this bill would do that effectively. I&#8217;m tempted to agree with <a href="https://writing.antonleicht.me/p/press-play-to-continue?hide_intro_popup=true">Anton Leicht </a>that AI safety risks becoming a &#8220;junior partner&#8221; in such a coalition. Ultimately, those of us who worry about the end of the world are a much smaller constituency than people who are convinced that one ChatGPT prompt uses an entire bottle of water, are concerned about electricity prices, are worried about their job, or just think that thing they can see from their back garden is kinda ugly. I&#8217;m still hopeful there is some way to channel all these disparate grievances effectively, but can see the world where it leads to clumsy policies that appease some local NIMBYs while failing to stop dangerous AI from actually being built. I think it could go either way.</p><p>I feel pretty middle-of-the-road about Bernie&#8217;s AI safety arc. I worry that we as a community might be tempted to overstate its importance in raising awareness, and I think it risks being drowned out by the other, louder factions in a broad left-wing AI-anti backlash. I&#8217;d probably feel pretty good if a comparably high-profile figure on the right also rocked up to Berkeley for a sit down with Eliezer and then had some public kumbaya moment with Bernie where they agree to put aside their profound differences and work together to save us all from this AI thing, but that feels unrealistic. That said, at the end of the day, we have one of the most famous politicians in the US talking candidly and repeatedly about the possibility of AI killing us all. So accounting for everything, I think this is a modest positive update.</p><h3>Anthropic releases its new Responsible Scaling Policy</h3><p>On 17 March, Anthropic released <a href="https://www.anthropic.com/news/responsible-scaling-policy-v3">version 3.0 of its Responsible Scaling Policy</a>, which explicitly dropped the &#8220;public commitment&#8221; it had made in earlier versions to pause development if capabilities reached levels that outstripped existing safeguards.</p><p>I felt at the time, and still feel, extremely salty about this development. This isn&#8217;t because I&#8217;m necessarily convinced that dropping the pause commitment increases the catastrophic risk stemming from Anthropic&#8217;s models specifically. I understand that Anthropic unilaterally pausing doesn&#8217;t help the situation if no one else does. I get that, technically, previous versions of the RSP already allowed for this commitment to be scrapped in a situation where Anthropic believed a competitor might soon build a similarly capable model less safely, and that maybe it makes sense to say that somewhere other than in an inconspicuous footnote. I also get that as AI has developed, it has become more clear that determining the thresholds at which capabilities pose certain risks has proven to be very hard, and so maybe Anthropic never stood a chance at a perfectly-timed pause anyway. I <em>could</em> see the argument for dropping the commitment in order to loudly and explicitly state: &#8220;<em>hey, we might not actually be able to pause development before something bad happens, we&#8217;re not even sure it would help if we did, we actually really don&#8217;t have this in hand at all and someone should probably step in</em>&#8221;, but nobody is saying that in so many words<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>.</p><p>I am not sure how to express my gripe here beyond &#8220;it&#8217;s the principle of the thing&#8221;. Like guys, can&#8217;t we just keep one promise? A sentiment expressed in the blog post announcing the new RSP, and that I hear often in discussions of AI safety, is &#8220;<em>well, we made that commitment back when we thought we might have been in Easy Mode, and international coordination and consensus-building looked more promising than they do now. But now it&#8217;s become clear that we&#8217;re in Hard Mode, and we have to adapt to the circumstances in front of us</em>&#8221;. People often position themselves as passive observers of shifts in the Overton Window, as if these shifts were not simply the sum of people saying and doing things, some of which were said and done by them. The history of AGI development is littered with broken promises, many of which probably seemed rational to break at the time, even if everyone involved was motivated to make AI go well. But if, magically, we were all stubbornly, deontologically committed to keeping our promises, I don&#8217;t think the gameboard would look so terrible. This development basically feels like a painful reminder of the abysmal state-of-play, even if I&#8217;m not sure it changes the outlook too much.</p><h3>Mythos Preview announcement</h3><p>This brings us to the final development of my AISI era and the last one I&#8217;ll cover here. Anthropic <a href="https://www.anthropic.com/news/anthropic-mythos-preview">announced</a> an internal model, Mythos, deemed too powerful to release due to its cyber capabilities. They gave exclusive access to 40 organisations trusted to use the model for cyber-defense as part of <a href="https://www.anthropic.com/news/project-glasswing">Project Glasswing</a>. This generated a bunch of <a href="https://x.com/GaryMarcus/status/2042285440217260358">predictable discourse</a> about how Anthropic obviously just didn&#8217;t have a big enough compute supply to serve the model publicly, and so decided to use it for marketing hype by naming it something spooky and keeping it behind closed doors. Then AISI <a href="https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities">released results</a> of our independent evaluation of Mythos, showing that it could complete a custom built cyber range end-to-end which would have taken a human expert 20 hours, which you&#8217;d think would have put this marketing hype narrative to bed, but <a href="https://x.com/firstadopter/status/2045235706839081420">didn&#8217;t entirely for some reason</a>.</p><p>My subjective experience of this moment was somewhat paradoxical. As the person whose job it had been for 10 months to make people care about AISI&#8217;s work, the whole thing was extremely dopamine-inducing. <a href="https://x.com/AISecurityInst/status/2043683577594794183">We were popping off</a>. I spent several days witnessing fast takeoff occur in our various analytics and oscillating mentally between &#8220;<em>people finally care</em>&#8221; and &#8220;<em>people care because something terrifying is happening</em>&#8221;.</p><p>Maybe I am just running out of steam at the end of this absolute behemoth of a blog post, but I don&#8217;t have a unique take here. Huge leaps in capability are Bad. Said leaps waking policymakers and the public up to the insane situation we&#8217;re in is Good. The post-Mythos vibe shift feels quite palpable to me. I am not sure whether the relevant people are tracing the through-line from a scary hacker model to an extinction-causing model, but it&#8217;s now pretty hard for them to ignore that the delicate cyber-foundations on which all of modern society rests could actually just break in the presence of powerful AI. You really can&#8217;t <em>not</em> do something about that. And it does seem that in the months since Mythos was announced, the US government is, in fact, doing something.</p><p>It seems quite likely that the Trump Administration&#8217;s June EO on &#8220;<a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/">Promoting advanced AI innovation and security</a>&#8221; is directly downstream of Mythos. The order creates a voluntary scheme by which frontier developers provide the government and &#8220;select trusted partners&#8221; with access to their models for 30 days before they are publicly released, so that they can be used to harden the cybersecurity of critical infrastructure &#8211; replicating the model that was pioneered by Project Glasswing. Commentators like Dean Ball <a href="https://x.com/deanwball/status/2061874260096983402">have been critical</a> of the EO and pointed out that it is just intrusive &#8211; and should be just as triggering to libertarian types &#8211; as <a href="https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence">Biden&#8217;s since-repealed 2023 EO</a> which mandated reporting requirements for frontier models. This further solidifies my intuition that I should feel good about this development, because I am not particularly libertarian, and I&#8217;m perfectly willing to bite the bullet of wanting governments to have deep and exclusive access to the most powerful AI models for some period, so they can witness all the batshit-crazy capabilities I expect them to possess in the near future.</p><p>The final thing I&#8217;ll say here is that I approve wholeheartedly of Anthropic&#8217;s communications strategy around the Mythos announcement. According to <a href="https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities">AISI&#8217;s testing</a>, GPT-5.5 is just as capable as Mythos Preview in the cyber domain. It completed the same cyber range end-to-end (if succeeding on fewer attempts) and was actually <em>more</em> capable on narrow capture-the-flag tasks. One could cynically accuse Anthropic of fearmongering here, since GPT-5.5 has been out in the world for several weeks and hasn&#8217;t caused a cyber catastrophe yet, to my knowledge. Or on the flip side, we could accuse OpenAI of irresponsibility. Unsurprisingly, I&#8217;m tempted to do the latter. I think the evaluation results should speak for themselves. I am unabashedly pro-fearmongering when the situation warrants it. For points beyond the current frontier, I&#8217;d be completely supportive of companies restricting access to their models and naming them increasingly spooky things until we&#8217;re all reading about the shadowy internal deployment of Claude Deadhand 5. The situation is genuinely very scary, and I&#8217;m glad that we&#8217;re at least starting to act like it. Overall, I&#8217;m giving Mythos a positive score.</p><p></p><p></p><p>That brings me to the end of my (almost) year-long speedrun. In the two months since the Mythos Preview announcement, there have been a handful of developments including: the <a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/">aforementioned Trump EO</a>; both <a href="https://openai.com/index/built-to-benefit-everyone-our-plan/">OpenAI</a> and <a href="https://www.anthropic.com/institute/recursive-self-improvement">Anthropic</a> publishing blog posts which acknowledge the possible need for a coordinated AI slowdown in the future (and also, ironically, filing for IPO <a href="http://bbc.com/news/articles/cd958eqg1n5o">within weeks of each other</a>), and the public release of a heavily-safeguarded Mythos-class model (&#8220;<a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Fable</a>&#8221;). For the first time since I started following AI safety, I have the distinct feeling that things are coming to a head. The labs appear to be teetering on the edge of recursive self-improvement, and the US government on the edge of taking AI progress seriously. It feels like a cop-out to trawl through a year&#8217;s worth of updates only to conclude that it could go either way, but that&#8217;s where I am.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p> Ok I just asked Claude, I guess I get it now.</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I do know that the actual trigger here was ~700 employees threatening to jump ship to Microsoft. I am just being silly.</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>My objective here isn&#8217;t to stray too far into politics, so I&#8217;ll just quickly say that yes, I understand that the Trump admin doesn&#8217;t appear imminently poised to use Claude for either of these purposes, and that their objection appears to be more about the principle of private companies imposing limits on the military&#8217;s use of technology.</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>I should say in the interest of fairness that since I wrote this, Anthropic released a <a href="https://www.anthropic.com/institute/recursive-self-improvement">blog post </a>on the possible need for a pause in AGI development due to the threat of recursive self-improvement. This seems directionally like &#8220;saying the thing in so many words&#8221;.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Why do people disagree about when powerful AI will arrive? ]]></title><description><![CDATA[My best attempt to distill the cases for short and long(ish) AGI timelines.]]></description><link>https://longerramblings.substack.com/p/why-do-people-disagree-about-when</link><guid isPermaLink="false">https://longerramblings.substack.com/p/why-do-people-disagree-about-when</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Mon, 02 Jun 2025 11:03:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!re4t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This blog post was <a href="https://bluedot.org/blog/agi-timelines">originally written</a> for Bluedot Impact. For more beginner-friendly AI safety content, check out their <a href="https://bluedot.org/blog">blog</a> and <a href="https://bluedot.org/courses/future-of-ai">Future of AI Course</a>! </em></p><p>Few would argue that AI progress over the past few years has not been rapid.</p><p>Large Language Models (LLMs) have provided an unexpected path to increasingly general capabilities. In 2019, OpenAI&#8217;s GPT-2 <a href="https://www.researchgate.net/figure/Example-of-a-GPT-2-text-generation-output-underlined-text-shows-sentences-drifting-off_fig1_344245575">struggled to write a coherent paragraph</a>. In 2025, LLMs write fluent essays, outcompete human experts at <a href="https://www.vals.ai/benchmarks/gpqa-04-18-2025">graduate-level science questions</a>, and excel at <a href="https://www.vals.ai/benchmarks/aime-2025-03-11">competition mathematics</a> and <a href="https://bab407.com.au/blog/openais-o3-achieved-a-codeforces-rating-of-2727-placing-it-175th-and-international-grandmaster/">coding</a>. The most advanced multi-modal AI models now produce <a href="https://gemini.google/overview/image-generation/?hl=en">images</a> and <a href="https://deepmind.google/models/veo/">video</a> that are hard to distinguish from reality.</p><p>These models are impressive (and useful!) but they still fall short of the north star that frontier AI companies are working towards. Artificial General Intelligence (AGI), which OpenAI <a href="https://openai.com/charter/">describes</a> as &#8220;a highly autonomous system that outperforms humans at most economically valuable work&#8221; has been the ultimate ambition of AI researchers for many decades.</p><p>Most experts agree that AGI is possible. They also agree that it will have transformative consequences. There is less consensus about what these consequences will <em>be</em>. Some believe AGI will usher in an age of <a href="https://www.darioamodei.com/essay/machines-of-loving-grace">radical abundance</a>. Others believe it will likely lead to <a href="https://safe.ai/work/statement-on-ai-risk">human extinction</a>. One thing we can be sure of is that a post-AGI world would look very different to the one we live in today.</p><p>So, is AGI just around the corner? Or are there still hard problems in front of us that will take decades to crack, despite the speed of recent progress? This is a subject of live debate. Ask various groups when they think AGI will arrive and you&#8217;ll get <a href="https://80000hours.org/2025/03/when-do-experts-expect-agi-to-arrive/">very different answers</a>, ranging from just a couple of years to more than two decades.</p><p>Why is this? We&#8217;ve tried to pin down some core disagreements.</p><h2>The case for short timelines</h2><p>Many of the people closest to frontier AI are expecting AGI to arrive before 2030.</p><p>Dario Amodei, CEO of Anthropic, is &#8220;<a href="https://www.youtube.com/watch?v=7LNyUbii0zw">confident</a>&#8221; that very powerful capabilities will be achieved within 2-3 years. Sam Altman, CEO of OpenAI, has <a href="https://blog.samaltman.com/reflections">claimed</a> that his company &#8220;knows how to build AGI&#8221;, and <a href="https://ia.samaltman.com/">thinks we may reach</a> an even grander goal of &#8220;superintelligence&#8221; within &#8220;thousands of days&#8221;. And Demis Hassabis, CEO of Google DeepMind, <a href="https://www.bigtechnology.com/p/google-deepmind-ceo-demis-hassabis">forecasts</a> a slightly more conservative (but still near-term) 3-5 years until AGI.</p><p>Two high-profile scenarios, <a href="https://situational-awareness.ai/">Situational Awareness</a> and <a href="https://ai-2027.com/">AI-2027</a>, written by forecasters and former AGI company employees, make similar projections.</p><p>Here are some reasons to think that AGI might be just a few years away:</p><h4><strong>#1 Benchmarks keep saturating</strong></h4><p>The easiest argument for a short-timelines advocate to make is that &#8211; at least in the capabilities that we know how to measure &#8211; it really doesn&#8217;t look like we have far to go.</p><p>Here&#8217;s a chart showing how quickly AI capabilities have improved in on various benchmarks in just the last two years:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!re4t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!re4t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png" width="1108" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1108,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!re4t!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dcd938c-0b89-4c95-958b-599367f12099_1108x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><a href="https://assets.publishing.service.gov.uk/media/679a0c48a77d250007d313ee/International_AI_Safety_Report_2025_accessible_f.pdf">Source: International AI Safety Report</a></em></p><p>In closed-ended academic benchmarks, AIs are closing in on human-expert level. Flagship models from OpenAI, DeepMind at Anthropic all score over 82% in <a href="https://www.vals.ai/benchmarks/mmlu_pro-04-18-2025">MMLU</a>, which contains multiple-choice questions on a range of disciplines from mathematics to international law, and over 75% on <a href="https://www.vals.ai/benchmarks/gpqa-04-18-2025">GPQA</a>, which contains graduate-level questions in STEM fields.</p><p>Benchmarks are saturating so quickly that researchers are scrambling to create new ones which will continue to challenge state-of-the-art models. For example, <a href="https://agi.safe.ai/">Humanity&#8217;s Last Exam</a> (HLE) contains 2,500 questions contributed by over 1,000 subject-matter experts, designed to be at the frontier of human knowledge in most fields &#8211; and LLMs are already starting to make progress. OpenAI&#8217;s o3 scored 20% on HLE, compared to an 8% score from its predecessor, o1, which was released just months earlier. Its creators think it is &#8220;plausible&#8221; that an AI will achieve more than 50% on HLE by the end of 2025.</p><p>Importantly, AI models aren&#8217;t &#8220;just memorising&#8221; the answers to questions on benchmarks like the GPQA and HLE &#8211; researchers maintain private test sets to make sure that solutions don&#8217;t find their way into models&#8217; training data.</p><p>If progress continues at the pace of the last few years, it won&#8217;t be long before we struggle to come up with close-ended questions that AI models <em>can&#8217;t</em> answer. At this point, all that will be missing are the properties needed to put all that raw intelligence to use, such as agency and long-term planning. Conveniently for short-timelines advocates, we have evidence of progress on these metrics too &#8211; which brings us on to our next point.</p><h4><strong>#2 AIs are able to complete longer and longer tasks</strong></h4><p>Impressive benchmark scores don&#8217;t translate neatly into real-world impact. One reason is that AI models cannot currently complete tasks over long time horizons. Even if they can solve any one step more reliably than most humans, they can&#8217;t autonomously carry out tasks that would take a person days or weeks.</p><p>But that could change soon. A <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">recent study by METR</a>, an organisation that develops and runs evaluations of frontier AI models, found that the length of tasks they can successfully complete is doubling every seven months:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Sa18!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 424w, /__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 848w, /__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Sa18!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png" width="1410" height="818" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:818,&quot;width&quot;:1410,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 424w, /__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 848w, /__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Sa18!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585feed1-b114-4c2b-a2f0-86d8108f1911_1410x818.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If this trend continues, AIs will be able to carry out month-long projects by the end of the decade. It could even accelerate. For 2024-25 specifically, the doubling time was <a href="https://theaidigest.org/time-horizons">4 months</a>, which would predict AIs tackling month-long tasks by 2027.</p><h4><strong>#3 AI research might be the only capability we need to automate to achieve AGI</strong></h4><p>The goal of frontier AI companies is to develop AI systems that can automate every economically valuable task. One such task is <em>AI research itself. </em>If we can automate &#8211; or significantly accelerate &#8211; the process of building better AI, then any remaining hurdles to AGI could be overcome soon thereafter.</p><p>If an AI company internally develops an AI system that outcompetes its top engineers at the task of advancing the AI frontier, it would face tremendous incentive to automate a significant fraction of its own research. Automated AI researchers could work day and night without breaks, and even self-replicate to develop what Geoffrey Hinton (one of the pioneers behind the deep learning paradigm that kicked off today&#8217;s acceleration in AI capabilities) <a href="https://www.forbes.com/sites/andreamorris/2023/05/03/ai-pioneer-geoffrey-hinton-talks-at-mit-about-ai-gaining-control/">calls</a> a &#8220;digital hive mind&#8221;.</p><p>It takes much more computing power to train a new AI model than to run one. One AI researcher <a href="https://situational-awareness.ai/from-agi-to-superintelligence/">estimates</a> that we&#8217;ll be able to run millions of automated researchers in parallel by 2027, compressing years&#8217; worth of progress into just a few days.</p><p>This could trigger what people sometimes refer to as an <a href="https://futureoflife.org/ai/are-we-close-to-an-intelligence-explosion/">intelligence explosion</a> &#8211; a recursive feedback loop where increasingly powerful AIs build their own successors. An intelligence explosion could quickly result in AI systems that are vastly more capable than humans.</p><p>AIs are already getting better at the skills needed to automate AI research. Last year, METR <a href="https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/">tested</a> frontier models including Anthropic&#8217;s Claude Sonnet 3.5 and OpenAI&#8217;s o1-preview against over 50 human experts. The results showed that models are already outcompeting humans at AI R&amp;D tasks over 2-hour time horizons.</p><h4><strong>#4 We could train much bigger models before 2030</strong></h4><p>Today&#8217;s most powerful AI systems are trained using <a href="https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy/">three inputs</a> &#8211; <strong>data</strong> (largely internet text), <strong>algorithms</strong> (instructions for learning from this data) and <strong>compute</strong> (the cutting-edge chips used to power the entire process).</p><p>Over the last few years, we&#8217;ve <a href="https://gwern.net/scaling-hypothesis">observed</a> that the amount of compute used to train models generally correlates with how capable they are. Some have <a href="https://aisafety.info/questions/94D9/What-is-the-%22Bitter-Lesson%22">even concluded</a> that compute is the <em>most</em> important ingredient in the AI development process &#8211; it&#8217;s less about clever algorithms or flashy architectures than it is just adding more and more hardware.</p><p>Epoch <a href="https://epoch.ai/blog/can-ai-scaling-continue-through-2030">estimates</a> that by 2030, it will be possible to train AI models using 10,000 times more compute than was used for OpenAI&#8217;s GPT-4. That&#8217;s around the same leap as we saw between GPT-2, which could barely produce a coherent paragraph without straying off topic, and GPT-4, which can engage in complex reasoning, generate code, pass standardised exams, and hold detailed, context-rich conversations.</p><p>If (as we&#8217;ve been arguing so far) there isn&#8217;t much further to go before we either hit AGI or automate AI research, it&#8217;s hard to imagine that another 10,000x scale-up won&#8217;t get us there.</p><h4><strong>#5 People&#8217;s timelines keep getting shorter</strong></h4><p>Expert opinion is an important datapoint in the debate over AGI timelines, but it&#8217;s a fuzzy one. For one, there&#8217;s an extremely wide spectrum of opinion &#8211; which is the whole point of this article! For two, it&#8217;s not clear what qualifies someone to forecast the arrival of AGI. Many different types of expertise could provide insight, from experience building frontier models to an impressive forecasting record.</p><p>If we can&#8217;t pinpoint a single authority whose predictions we should trust, it&#8217;s tricky to know how expert opinion should inform our forecasts. But one thing we can do is look at how predictions have shifted over time. Do they appear to be converging on a particular time period?</p><p>Taking this perspective lends some credence to the short timeline argument. In a <a href="https://blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things">2023 survey</a> of machine learning researchers, run by AI Impacts, participants thought AGI would arrive by 2047 &#8211; what would qualify by today&#8217;s standards as &#8220;long timelines&#8221;. But in the <a href="https://aiimpacts.org/2022-expert-survey-on-progress-in-ai/">2022 survey</a>, this year was 2060. On the the forecasting platform Metaculus, the <a href="https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/">predicted date</a> for AGI&#8217;s development has dropped by over two decades since 2022:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WMsL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 424w, /__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 848w, /__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WMsL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png" width="1456" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 424w, /__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 848w, /__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WMsL!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f1cc81d-78af-45c5-8e19-e743d9a22017_1534x896.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The chorus of shortening timelines is loud. Take Geoffry Hinton and Yoshua Bengio, two winners of the 2018 Turing Award for Deep Learning, and nicknamed &#8220;Godfathers of AI&#8221;. <a href="https://x.com/geoffreyhinton/status/1653687894534504451?lang=en-GB">Both</a> <a href="https://time.com/collection/time100-ai/6310612/yoshua-bengio/">shortened</a> their timelines from many decades to as little as five years after the release of ChatGPT (a watershed moment for many observers of AI progress). In a recent <a href="/__u/helentoner.substack.com/p/long-timelines-to-advanced-ai-have">Substack post</a>, former OpenAI board member Helen Toner reflects on this timeline-shrinking epidemic, and cites several examples of once-sceptics who now believe AGI could arrive within 10 years.</p><p>While there&#8217;s certainly a limit to how much we can learn from predictions (especially given the unprecedented nature of AGI), it is certainly worth noting that it&#8217;s far easier to find examples of shortening timelines than it is lengthening ones.</p><h2>The case for long timelines</h2><p>Despite plenty of excitement (and alarm) about near-term AGI, not everyone is convinced. Although many industry insiders are confident in short timelines, surveys from broader sets of experts still elicit longer ones, despite the downward trend mentioned above!</p><p>Skeptics point out that LLMs still make <a href="/__u/garymarcus.substack.com/p/why-do-large-language-models-hallucinate">silly errors</a>, emphasise barriers to <a href="/__u/epochai.substack.com/p/wheres-my-ten-minute-agi">using AIs for real-world tasks</a>, and doubt the <a href="/__u/epochai.substack.com/p/how-fast-can-algorithms-advance-capabilities">plausibility of an intelligence explosion</a>.</p><p>Here are some reasons to think AGI could still be decades away:</p><h4><strong>#1 Benchmarks only tell us about capabilities that are easy to measure</strong></h4><p>Benchmarks are best for clearly-defined tasks, where success is easily verified. AI models are particularly good at these kinds of tasks. This is especially true for more recent reasoning models such as OpenAI&#8217;s o1, which rely heavily on <strong>reinforcement learning</strong> (RL).</p><p>During RL, a model learns to maximise reward through trial and error. This has produced AI systems that are superhuman in narrow domains, like DeepMind&#8217;s <a href="https://deepmind.google/discover/blog/alphago-zero-starting-from-scratch/">AlphaGo</a> and <a href="https://alphafold.ebi.ac.uk/about">AlphaFold</a>&#8211; and researchers have <a href="https://forum.effectivealtruism.org/posts/PPuojCCajtCWhJR4w/reinforcement-learning-a-non-technical-primer-on-o1-and">recently found</a> that it works better than expected on general-purpose LLMs too. It&#8217;s much easier for an AI to learn via RL when what constitutes &#8220;good&#8221; vs &#8220;bad&#8221; performance is indisputable. If we give an AI a hill to climb, it will. This is one explanation for why reasoning models have exhibited extremely impressive capability improvements in domains like maths and coding.</p><p>But AGI-skeptics point out that the real world is not so tidy. Think about the day-to-day experience of doing your own job. You might receive different (or even conflicting) feedback from various people. There&#8217;s probably more than one &#8220;right&#8221; way to deliver an output. You&#8217;ll learn what works and what doesn&#8217;t by observing the real-world impacts of your work, which could be mixed.</p><p>This line of argument formed much of the pushback to the METR study that we mentioned earlier. The study was used to demonstrate that AIs are able to act over longer and longer time horizons &#8211; but it exclusively measured performance on <em>software engineering tasks</em>. This is precisely the kind of task we&#8217;d expect AIs to be good at! It&#8217;s not clear how well this trend will generalise to other economically valuable labour.</p><p>Real-world jobs are also not easily divided into discrete, self-contained tasks. Instead, tasks tend to be overlapping, and dependent on lots of context that&#8217;s hard to give an AI. For this reason, it&#8217;s <a href="https://epoch.ai/gradient-updates/where-is-my-ten-minute-agi">not so easy</a> to delegate even short-form ones to AI models. LLMs are expert email-crafters, but asking one to follow up with a colleague on an earlier discussion is pretty hard when the AI doesn&#8217;t know what was discussed!</p><h4><strong>#2 Tasks that are easy for humans are hard for AIs &#8211; and vice-versa</strong></h4><p>People have historically expected that manual labour would be automated long before white-collar work. Yet in 2025, we have language models that can solve PhD-level science questions &#8211; but not robots that can reliably assemble furniture.</p><p>One person who <em>did</em> see this coming was computer scientist Hans Moravec. He observed in 1988 it&#8217;s far easier to get a computer to exhibit expert-level performance at a game like checkers than the basic perception and mobility of a one-year-old. This principle became known as <strong><a href="https://en.wikipedia.org/wiki/Moravec%27s_paradox">Moravec&#8217;s Paradox</a></strong>: reasoning requires far less computation than sensorimotor skills.</p><p>Moravec hypothesised that skills like grasping objects or navigating around obstacles are harder to replicate because they&#8217;ve taken so much more time to emerge through the process of biological evolution. On the other hand, humans acquired abstract reasoning skills relatively recently. The older the skill, the more computational resources it requires to reverse-engineer.</p><p>Maybe this phenomenon has given us a warped perception of how capable today&#8217;s AI systems actually are. Being able to solve a PhD-level maths problem looks very impressive to us, because most humans can&#8217;t. On the other hand, loading a dishwasher seems like a trivially easy task, because most of us <em>can</em>. But the Moravec argument would say that this doesn&#8217;t actually tell us much about the absolute difficulty of either task. This implies that we&#8217;ve unlocked the easy AI capabilities, and the hardest part could still be in front of us.</p><h4><strong>#3 We don&#8217;t know if intelligence explosion is possible</strong></h4><p>So far, we&#8217;ve been arguing timelines to AGI could be long because there&#8217;s much further to go than benchmarks imply.</p><p>However, this ignores the <strong>intelligence explosion</strong> argument that we made earlier. Even if there are many hurdles we need to jump before we reach true AGI, automating AI research might mean we still get there very soon. That AIs are especially good at easily-measured tasks like maths and coding only supports this &#8211; these are precisely the skills they&#8217;ll need in order to accelerate AI progress.</p><p>But whether an intelligence explosion is actually possible is an open question. As we explained earlier, AI is trained using three inputs: data, algorithms and compute. It&#8217;s this second ingredient, algorithms, that depend on cognitive labour. This labour is performed by human researchers today, and could be performed by automated researchers in the future.</p><p>This means that the likelihood of an intelligence explosion hinges on a key question: <strong>how much algorithmic progress could a team of automated researchers make while compute and data remain static?</strong> Producing new chips and building datacentres takes time, as does gathering data (or generating synthetic data, which may become essential if we run out of human-written text altogether!).</p><p><a href="/__u/epochai.substack.com/p/how-fast-can-algorithms-advance-capabilities">One study</a> found that historically, the biggest algorithmic advances have been <strong>compute-dependent</strong> &#8211; they required big scale-ups in hardware to develop and validate. If this continues to be true in the future, then an intelligence explosion looks less plausible.</p><p>Of course, this is a big if. It&#8217;s difficult to predict what lots of super-smart automated AI researchers running in parallel could achieve. They might be able to find lots of clever ways to run experiments with fixed supplies of compute, such as using <a href="https://bluedot.org/blog/what-is-ai-scaffolding">scaffolding</a> techniques to squeeze more capability out of existing models. Skeptics acknowledge that an intelligence explosion is a live possibility, but point out that the possibility is speculative.</p><h4><strong>#4 Raw intelligence might not be the main driver of discovery</strong></h4><p>In the short timeline scenario, we develop what Dario Amodei <a href="https://www.darioamodei.com/essay/machines-of-loving-grace">calls</a> a &#8220;country of geniuses in a datacenter&#8221; that go on to transform the world by doing a lot of research and development (R&amp;D). For example, these AI geniuses could be directed to develop new medicines, find sustainable energy solutions, or develop coordination mechanisms that help us avoid future conflict (or, in the bad scenario, AI-enhanced research is turned against us).</p><p>Whether things will really play out this way depends on what actually drives real-world discovery. Skeptics will point out that there&#8217;s much more to R&amp;D than just very smart people thinking very hard. One piece of evidence for this is the phenomenon of <a href="/__u/mattsclancy.substack.com/p/how-common-is-independent-discovery">simultaneous discovery</a> &#8211; multiple people often have the same insight independent of each other, at roughly the same time. For example, Isaac Newton and Gottfried Leibniz both discovered calculus in the 1770s and Charles Darwin and Alfred Russel Wallace both described natural selection in 1839.</p><p>There are many theories for why simultaneous discovery happens. One is that <a href="https://www.jstor.org/stable/2142320?seq=8">culture moves faster than scientific discovery</a>, creating a lag which people will then seek technical solutions or explanations for. This runs counter to the idea that the raw thinking power from a &#8220;country of geniuses in a datacenter&#8221; could actually cause a massive acceleration in R&amp;D.</p><p>People have also <a href="https://epoch.ai/gradient-updates/most-ai-value-will-come-from-broad-automation-not-from-r-d">pointed out</a> that the process of R&amp;D requires a far broader set of skills than just abstract reasoning (the bread and butter of reasoning models). Human researchers direct research teams, collect evidence from experiments, and so on. Maybe this means that scientific R&amp;D isn&#8217;t any easier to automate than anything else. This would mean we shouldn&#8217;t expect a rapid, AI-driven transformation until <em>after</em> a broad range of more routine jobs across the economy have been automated.</p><h2>So&#8230; who&#8217;s right?</h2><p>Lots of smart and qualified people have spent a long time thinking about when AGI will arrive, and come to very different conclusions. Unfortunately, there isn&#8217;t conclusive evidence either way!</p><p>We&#8217;ll know soon enough. Many <a href="https://situational-awareness.ai/from-gpt-4-to-agi/#Addendum_Racing_through_the_OOMs_Its_this_decade_or_bust">believe</a> that AGI will either happen before 2030, or take much longer. This is because we probably can&#8217;t sustain our current rate of scaling past this point. We could build compute clusters that cost $1 trillion in five years&#8217; time. This is probably close to the limit of what the US economy could sustain. This suggests that if we get to 2030 with no AGI, the yearly probability starts to decrease.</p><p>There are all sorts of reasons to think that the current rate of progress could fizzle out, but we can&#8217;t be confident in them. There are enough arguments for near-term AGI to warrant taking the possibility extremely seriously &#8211; and this is a scenario that the world is drastically underprepared for. We don&#8217;t know how to ensure AI systems reliably follow our instructions. We don&#8217;t understand how they work. We don&#8217;t know how we&#8217;ll avoid the worst outcomes from developing AGI, up to and including human extinction.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[A defence of slowness at the end of the world]]></title><description><![CDATA[Since learning of the coming AI revolution, I&#8217;ve lived in two worlds.]]></description><link>https://longerramblings.substack.com/p/a-defence-of-slowness-at-the-end</link><guid isPermaLink="false">https://longerramblings.substack.com/p/a-defence-of-slowness-at-the-end</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Wed, 29 Jan 2025 00:18:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Rutc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Rutc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 424w, /__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 848w, /__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Rutc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png" width="1086" height="828" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:828,&quot;width&quot;:1086,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1695596,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 424w, /__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 848w, /__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Rutc!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2023320e-9e4b-410a-83d3-f06c3b5f9883_1086x828.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Since learning of the coming AI revolution, I&#8217;ve lived in two worlds. One moves at a leisurely pace, the same way it has all my life. In this world, I am safely nestled in the comfort of indefinite time. It&#8217;s ok to let the odd day slip idly by because there are always more.</p><p>The second moves exponentially faster. Its shelf-life is measured in a single-digit number of years. Its inhabitants are the <a href="https://situational-awareness.ai/">Situationally Aware</a>; the engineers and prophets of imminent AI transformation. To live in this world is to possess what Ezra Klein <a href="https://www.nytimes.com/2023/03/12/opinion/chatbots-artificial-intelligence-future-weirdness.html">calls</a> &#8220;an altered sense of time and consequence&#8221;.</p><p>I find that it&#8217;s psychologically untenable to spend all that much time in the Fast World. I can handle it for minutes to hours, but my mind invariably snaps back into its default state like I&#8217;m pulling my hand out of ice water.</p><p>Occupying the Slow World is ultimately a form of denial. I can&#8217;t call it anything other than compartmentalisation, yet I actually advocate for it. Of course, those of us trying to move the needle on AI risk should <em>work</em> in the Fast World, but I claim that we shouldn&#8217;t <em>live</em> in it. I will try to make the case for why.</p><p>I am worried that too much time spent in the Fast World will make me less invested in other people. Earnest belief in the coming singularity can make everything land with a little less weight. If the world ends or is otherwise rendered unrecognisable in two years, then the joys, triumphs, setbacks and adversities of my family and friends lose gravitas. Their implications are fewer, and their impact will be brief. This is why I try to occupy the Slow World most of the time. To wholeheartedly celebrate some achievement or milestone, or to suffer a tragedy, is to envision the future. If a friend of mine gets engaged, I want to share in their anticipation of a happily-ever-after lasting decades, because this is what it will mean to be happy for them. And if someone I love gets ill or dies, I want to feel the pain of contemplating many years spent without them, because this is what it will mean to grieve. I want to<em> really believe</em> in those futures. I want to delude myself as thoroughly as I can. Because this is how I will feel these moments as they deserve to be felt. I do not want my mind to qualify them with an unspoken countdown.</p><p>Those of us among the Situationally Aware must be on our guard against arrogance. To anticipate some transformation that most live in ignorance of can easily breed self-importance. But worse, it can degrade the way that we perceive the efforts of others. It can lead us to view any enterprise that won&#8217;t ultimately bear fruit in an ASI-by-2027 world with something that approaches pity or derision. <em>Those sweet summer children with dreams of studying computer science at university, don&#8217;t they know that AIs are already competitive with human coders? Tragic that people are out here making five and ten-year career plans, don&#8217;t they know that planning over anything longer than a six-month timeframe is hopeless?</em></p><p>This attitude can colour the way you perceive the whole world. It can make everyone &#8211; teenagers on their way to school, business people on their rush-hour commute, joggers on a morning run &#8211; appear to you like swimmers battling hopelessly against a tsunami they cannot see, their labour fruitless, misguided and tragic. It can make you feel sorry for them. This is an attitude towards the rest of the world that I want to avoid at all costs. Partly out of humility, since of course, the Situationally Aware could be wrong, which would make this misplaced pity all the more objectionable in retrospect! But even if they&#8217;re <em>right</em>, this simply isn&#8217;t the way that I want to relate to my fellow human beings. I don&#8217;t want to presume that they&#8217;d act any differently if they knew what (I think) I know. I don&#8217;t want to possess a mindset that robs human endeavour of its purpose.</p><p>Sometimes in apocalypse movies, there&#8217;s a certain dramatic irony that makes humanity look like the butt of the joke. In the exposition, before it all starts going south, they&#8217;re embroiled in petty dramas or griping about traffic or the weather. And then the asteroid hits or the nuke lands or the zombie virus starts to spread, and civilisation meets its undignified end as if in chiding punishment for its obliviousness. But I think there&#8217;s plenty of dignity in being caught right in the middle of something when some unforeseen disaster occurs, of being in the midst of whatever you would always have been doing. As I was struggling to articulate this thought, I realised that CS Lewis already said it much better than I ever could in 1948:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mE4J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 424w, /__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 848w, /__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mE4J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png" width="1172" height="298" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:298,&quot;width&quot;:1172,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 424w, /__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 848w, /__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mE4J!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b737d4-857a-4fe1-bc61-b95df71c4fe6_1172x298.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If I might project my own slightly flimsy thesis onto the work of CS Lewis, I choose to read this as an endorsement of living in the Slow World.</p><p>Spencer Greenberg of Clearer Thinking <a href="https://www.facebook.com/spencer.greenberg/posts/pfbid02WpUjV17AcmEJy47UyJ7patte2fu26LSXgx1DP8tuK8MbtGgi7HggZwXFXdiyw3S8l">recently wrote</a> that believing you only have one option is dangerous. Grasping to preserve what you believe is your sole choice can lead to poor decision-making, like tolerating abuse in a relationship or staying in the wrong job. I think this argument extends to believing you have <em>very little time. </em>To live in the shadow of an impending singularity can make every opportunity appear as one of an ever-dwindling number. It can end in money you wish you hadn&#8217;t spent, sex you wish you hadn&#8217;t had, or nights you wish you hadn&#8217;t drunk so much. It can obscure opportunities whose benefits might take longer to manifest. Staying in the Slow World can guard against impulsivity. It can expand your menu of options, and make you more likely to take bets that will pay off in longer timelines (which we might be lucky enough to get!). It&#8217;s ok to comfort yourself with age-old adages meant to stave off myopicism &#8211; that you&#8217;re still young, that you&#8217;ve got your whole life ahead of you and that there&#8217;s always next year. It&#8217;s ok to really believe these adages and to live as if they are true.</p><p>I think my final defence of the Slow World is more a matter of personal preference. And that is simply that I prefer it. I have never lived fast, and learning that I may die young hasn&#8217;t changed that. I am not a thrill seeker. I don&#8217;t want to live like there&#8217;s no tomorrow, because living without a guaranteed tomorrow is scary and unpleasant! It inspires anxiety, not motivation. Some of the joy I experience is in novelty, but much of it isn&#8217;t. It is in the mundane and the repetitive, the things that would make the movie of my life a boring watch, but which make it no less wonderful to live. Three cups of coffee in bed on a Sunday morning, watching the same TV shows again and again and leaving a few months between rounds so that they are always imbued with fresh flavour like a new piece of gum, going to the same pub on the corner with my housemates every other weekend, frittering away the entire subsequent day together while we recover from our hangovers. I <em>want</em> to let the days, weeks and months slip by without counting them. I choose to do so intentionally.</p><p>All of the above is how I (try to) deal with the possibility of short timelines. I&#8217;m by no means perfect at it. Some days I spend a little more of my time in the Fast World than I&#8217;d like. It may not resonate with everyone. I&#8217;m sure there are others out there who feel the opposite. There <em>are</em> also ways that contemplating short timelines has shaken me out of certain malaises and sharpened my appreciation for the world around me. But on the whole, I intend to carry on as normal &#8211; because normal is more than enough for me.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sarah&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Don’t sell yourself short]]></title><description><![CDATA[And other advice for the newly AI-concerned.]]></description><link>https://longerramblings.substack.com/p/dont-sell-yourself-short</link><guid isPermaLink="false">https://longerramblings.substack.com/p/dont-sell-yourself-short</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Thu, 16 Jan 2025 18:44:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-JCI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Back in 2023, I spent several months as an AI safety lurker. In my quest to ascertain if and when AI might destroy the world, I was spending several hours a day quietly observing a discussion that I was both terrified and confused by. I became familiar with the cast of characters that populated this online world, an ensemble of profile pictures and usernames espousing confident but conflicting opinions, whom I considered the People Who Know Things.</p><p>Determining precisely who ought to be counted among the People Who Know Things was its own challenge. Should I restrict membership to academics whose Twitter bios were decorated with qualifications from prestigious universities and whose work had been cited thousands of times? Or insiders at frontier labs? What of those whose online presence didn&#8217;t appear to signal any relevant formal qualifications or experience, but whose long and complicated-sounding LessWrong posts were being seriously engaged with by people in categories 1 and 2? The one thing I <em>did</em> know is that I was not and never would be a member. The People Who Know Things occupied a world on the other side of an unbridgeable divide.</p><p>One day, the barrage of conflicting apocalypse forecasts I had been subjecting myself to for nearly half a year was taking its emotional toll. After a few glasses of wine, I took to Twitter to complain about it. I wrote a long, earnest <a href="https://x.com/littIeramblings/status/1708945586496446796">thread</a> whining about how I was <em>just so confused</em>, because Expert A said we were all going to die in 2027 and Expert B said 2035 and Expert C said never, and all their reasoning was totally opaque to me because I was Just A Girl and didn't know what a FLOP was or how scaling laws worked and I was simply not smart enough to figure any of it out.</p><p>When I woke up the next day, I unlocked my phone to find that my thread was gaining traction. The People Who Know Things were starting to talk back to me. Responses were varied, from people who related to my plight to white knights who wanted to swoop in and save me from the AI-doom cult. But one in particular stuck out:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-JCI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 424w, /__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 848w, /__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-JCI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png" width="966" height="856" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:856,&quot;width&quot;:966,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 424w, /__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 848w, /__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-JCI!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6d1bda4-16eb-4f21-b93d-67f7da3ed4c6_966x856.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That evening, my housemate found me frowning at my phone while polishing off the remainder of the wine.</p><p>&#8220;Are you ok?&#8221;, she asked.</p><p>&#8220;I called myself stupid on the internet and people believed me&#8221;, I muttered.</p><p>&#8220;You should spend less time on Twitter&#8221;.</p><p><em>Can&#8217;t and won&#8217;t understand the arguments, </em>I tutted indignantly to myself. <em>Can&#8217;t master the object level</em>. Of course, my annoyance was totally unfounded, since I had said exactly that! This interaction, and several others that I&#8217;d come to have over the subsequent months, taught me a valuable life lesson. The culture I&#8217;d grown up in, where every conversation took place against a tacit backdrop of false humility, is not representative of the world at large. At my North London girls&#8217; school, self-belittlement was invariably met with a chorus of &#8220;omg you&#8217;re so smart and pretty and totally not fat at all!!&#8221;. But in the AI safety sphere, the self-belittling are assumed to be doing a well-calibrated assessment of their own attributes<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</p><p>Was I actually being falsely humble? I think the answer to this question is somewhat complicated. The emotions I expressed in the thread &#8211; of confusion, overwhelm and uncertainty about who I should be deferring to &#8211; were genuine. But did I consider myself <em>incapable</em> of understanding object-level arguments for and against AI risk? Not really. I believed then and believe even more strongly now that the high-level reasons for being worried about powerful AI are extremely simple. It&#8217;s not hard to understand why racing to build increasingly capable and goal-directed AIs &#8211; without any scientific solution to the problem of controlling them once they supersede us &#8211; may lead to bad things. It&#8217;s also not difficult to observe that the unworried are few and far between, and that they lack anything like a satisfying answer to this basic concern.</p><p>Of course, there are many more in-the-weeds questions that one could ask in order to form a well-rounded AI worldview (will takeoff be fast or slow, what is the likelihood that X or Y alignment technique will work, what will the political response to advancing capabilities be), and it is true that in 2023 I&#8217;d have had very little to say on any of these. But it is also true that <em>I hadn&#8217;t really tried</em>. My &#8220;research&#8221; into AI safety up until that point could be more accurately described as an obsessive quest for reassurance, which ended up culminating in an extensive collection of p(doom) estimates and timeline predictions. These had been cherry-picked from a much richer discourse about <strong>why</strong> people had reached those conclusions, which I had only superficially engaged with. I just wanted to know if the world was going to end. I was more guilty of being lazy than stupid. Over a year later, I&#8217;m still far from confident in any of my AI takes, but I have at least tried to do the intellectual legwork of refining them. I now consider myself an Aspiring Person Who Knows Things.</p><p>I don&#8217;t know whether there are many other AI safety lurkers out there who don&#8217;t feel qualified to join the conversation, let alone whether any of them are reading this. But on the off chance that they are (especially if any are female, since being <a href="https://x.com/littIeramblings/status/1875507291493187926">one of the few women</a> in AI safety can create the perfect storm of self-doubt), I wanted to offer some unsolicited advice.</p><h4>#1 Don&#8217;t sell yourself short</h4><p>This is of course generic advice that applies to the world at large &#8211; if you repeatedly downplay your own abilities, people will not always give you the benefit of assuming you&#8217;re just insecure or self-deprecating. They might take you at your word!</p><p>This is somewhat complicated by the fact I think my online AI ramblings have filled an underrepresented niche &#8211; that is &#8220;normal person who doesn&#8217;t understand AI but is stressed about it&#8221;. My femaleness complements this quite nicely. It has been easy to lean into the Just A Girl aesthetic. There&#8217;s endless mileage you can get out of takes that boil down to &#8220;wow guys it sure looks like the experts are spooked by AI and think it might literally kill everyone, maybe governments should be taking that more seriously but hey what do I know haha&#8221;. There have been benefits to doing this, but I think that to a degree, it has resulted in me letting myself off the hook. I believe it has meaningfully delayed my transition to an Aspiring Person Who Knows Things. If I could go back in time, I&#8217;d probably skip my Just A Girl arc.</p><h4>#2 You can just say things</h4><p>If you&#8217;re an AI safety lurker and don&#8217;t want to be &#8211; just say things. It really is that simple. Despite interacting with 100s of people who are smarter than me on topics I would once have considered beyond my intellectual bandwidth, I have suffered very few embarrassments. Even if you say something wrong, you&#8217;ll likely be politely corrected, not socially punished. You&#8217;ll probably find that The People Who Know Things often agree with you.</p><p>This isn&#8217;t to say that everyone&#8217;s contributions to the conversation have equal weight, or that AI safety is some unique field where we shouldn&#8217;t privilege expertise. I don&#8217;t want to sound like one of those people who thinks a layperson with acess to PubMed is as well-qualified to perform medical diagnosis as a doctor. But AI safety is a domain in which one can have knowledge or expertise along many different axes. For example, technical experts may be best placed to forecast the speed at which AI capabilities will improve, but not how governments ought to respond to them. The field is also nascent enough that, for better or worse, it hasn&#8217;t had time to become subject to the institutional gatekeeping that befalls many others. There are conversations at the frontier of AI safety and policy literally playing out on Twitter and Substack. There&#8217;s nothing to stop you joining them.</p><p>As was the source of my frustration in that original tweet thread, there is no expert consensus on what a future with powerful AI will look like. It seems like a safe bet that the future will be very weird, but beyond that, I see little grounds for confidence. I think this is the perfect context in which to Just Say Things. I like the way Nathan Labenz put it in a <a href="https://open.spotify.com/episode/5b7qP4CzgJS4dD51YQcu78?trackId=4Fknl5coWTUwhBHuABw5lu">recent episode</a> of The Cognitive Revolution podcast:</p><blockquote><p><em>AI gods might be an emerging trend over the second half of the decade. I have no idea how we're gonna relate to these things. If they are meaningfully superhuman, will we even try to keep them under control? Will we worship them? No matter how weird your alignment idea is, I think it is worth pursuing. No matter how weird your thoughts are about where the future might be going, I would say they're probably worth entertaining.</em></p></blockquote><h4>#3 Ask questions</h4><p>In trying to follow developments in AI, I find myself frequently confused. Anyone who isn&#8217;t delusional will too. Luckily, there is a whole community of well-intentioned and knowledgeable people who are happy to weigh in! I have on innumerable occasions contacted people I know (or don&#8217;t know!) with technical or policy expertise to ask them clarifying questions. More often than not, they have responded. I also often pose my open questions to Twitter. <a href="https://x.com/littIeramblings/status/1870536019558498445">What&#8217;s the deal</a> with o1 sometimes doing better on PhD-level science questions than high school level ones? <a href="https://x.com/littIeramblings/status/1879548264829387088">Why do some people believe</a> that an ASI in the hands of governments would pose an unacceptable risk of power concentration, but one in the hands of a private company wouldn&#8217;t? I have received many helpful answers.</p><h4>#4 Actually try</h4><p>As I said above, 2023 Sarah, who pleaded ignorance of all things AI in that long rambly thread, hadn&#8217;t actually tried very hard. In 2025, I still haven&#8217;t read or understood nearly as much as I would like, but I am sure as hell trying. The downsides of ever-accelerating AI capabilities are many (plausible short-term human extinction, job loss, a looming crisis of meaning&#8230;) but at least one major upside is that it has never been easier to learn things! A hill I will die on is that using LLMs to translate smart-person speak into dumb-person speak is, in fact, a smart-person move. This is an excellent use case for making sense of jargony papers or impenetrable posts on LessWrong. You can then have as long a conversation as you want with a private, infinitely patient AI tutor about anything you still don&#8217;t understand. Still, there will be things that remain the exclusive purview of CS PhDs or seasoned policy wonks. There is a level of expertise that you or I likely don&#8217;t have time to build up pre-singularity. But that doesn&#8217;t mean we can&#8217;t make a start.</p><div><hr></div><p>In the early days of my confused-girlie-come-AI-worrier journey, multiple people reached out to me to ask if I&#8217;d considered working in AI safety. I dismissed them out of hand. Of course I hadn&#8217;t! That was the territory of People Who Know Things, of which I <em>obviously</em> was not one. My plan was simply to post through it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fT38!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 424w, /__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 848w, /__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fT38!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png" width="968" height="798" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:798,&quot;width&quot;:968,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 424w, /__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 848w, /__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fT38!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43f436d9-92ce-43e3-9edc-e86d4dc2fb07_968x798.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I now do, in fact, work full-time in AI safety. I still consider myself to be at the bottom of a very steep learning curve, but I&#8217;m not Just A Girl anymore.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/longerramblings.substack.com/subscribe"><span>Subscribe now</span></a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>This is obviously an over-generalisation! I think this is more true among AI safety people than in other communities / cultures I&#8217;ve encountered.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Are AI safetyists crying wolf?]]></title><description><![CDATA[Fear of AI is not just another tech-panic.]]></description><link>https://longerramblings.substack.com/p/are-ai-safetyists-crying-wolf</link><guid isPermaLink="false">https://longerramblings.substack.com/p/are-ai-safetyists-crying-wolf</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Wed, 08 Jan 2025 20:05:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/617337d2-bed4-4f26-91b9-3095f8d2e247_2146x1098.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Sarah Guo, founder of venture capital firm Conviction, is not worried about the existential risk posed by AI. Why not? As she points out in a <a href="https://www.youtube.com/watch?v=AhiYRseTAVw">roundtable</a> of technologists and AI experts at last month&#8217;s DealBook Summit, these apocalyptic fears have existed before. &#8220;This is quite typical when you look at technology historically&#8221;, she explains. &#8220;Everything from trains to electricity was considered the end of the world at some point&#8221;.</p><p>The idea that the current concern around AI is nothing more than another instantiation of the irrational panic that accompanies the eve of every technological revolution is widespread. Anyone reading this is likely familiar with the case for existential risk from AI, which rests on <a href="https://wiki.aiimpacts.org/arguments_for_ai_risk/list_of_arguments_that_ai_poses_an_xrisk/start">conceptual arguments</a> that are being increasingly <a href="https://futureoflife.org/ai/could-we-switch-off-a-dangerous-ai">validated</a> by empirical evidence, and is a <a href="https://www.safe.ai/work/statement-on-ai-risk">source of concern</a> among some of the world&#8217;s most highly-cited experts<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. As I write this, I haven&#8217;t yet checked to make sure that no group of scientists made a similarly rigorous case for extinction-by-trains, but I&#8217;m pretty confident that I won&#8217;t find one. Anyone who has spent more than five minutes investigating the topic of AI x-risk knows that this comparison is ridiculous, and yet arguments like Guo&#8217;s are the start and end of countless discussions about it &#8211; and are often met with very little pushback.</p><p>This broad claim about historical fear of new technologies is one of two &#8220;crying wolf&#8221; accusations often levied at AI safetyists. The other is more specific &#8211; that the AI-concerned, or advocates of strict regulation, have spread alarmism about the catastrophic potential of existing models that have since been proven safe, or claimed that AI doomsday would come to pass by some specific past date. In the same DealBook roundtable, panel moderator and New York Times columnist Kevin Roose states this explicitly: &#8220;There&#8217;s a crying wolf problem among AI safety advocates, where people are starting to discount these predictions because systems keep being released and keep not ending the world&#8221;.</p><p>To put my cards on the table, I think that both varieties of the crying wolf argument are fairly weak. I also think that the phenomenon Roose describes isn&#8217;t really happening to anything like the extent he implies. Nonetheless, arguments of this flavour are so common that I thought it might be worth taking some time to really give them their due. If there is anything <em>at all</em> to the &#8220;crying wolf&#8221; claims, I intend to unearth it here. And insofar as these claims have even the smallest bit of validity, I want to think about ways that the AI safety community can refine our messaging to make us less vulnerable to them.</p><h3>Part I: Everything from trains to electricity</h3><p>In this section, I will look at some past technologies, and make a good-faith effort to assess the level of historical concern about their apocalyptic potential, as well as the <em>credibility</em> of this concern.</p><h4><strong>Trains</strong></h4><p>I imagine that Sarah Guo did not expect anyone to scrutinise her claim that there was an apocalypse panic over trains too closely. I&#8217;m aware that this was an offhand comment and that I&#8217;m probably being unfair in taking her to task on it &#8211; but hey, everyone is entitled to a little pedantry now and again, as a treat.</p><p>Did anyone at the advent of the railway put forth a plausible mechanism by which trains might literally end the world? Well&#8230; no. Obviously not. I&#8217;m sure that Guo knows that too, and that she was simply committing the common sin of using phrases like &#8220;the end of the world&#8221; where they do not properly apply (relatedly, overuse of the word &#8220;existential&#8221; is one of my <a href="https://x.com/littIeramblings/status/1717641629866017265">pet peeves</a>). That said, there <em>have </em>been since-invalidated concerns about trains that were taken seriously by credible people. One fear was that railway travel could <a href="https://historyfacts.com/science-industry/fact/people-thought-trains-would-cause-railway-madmen/">drive people insane</a>. The jarring motion and unprecedented speed of trains were thought to injure the brain and cause bouts of lunacy. Doctors speculated in <a href="https://books.google.co.uk/books?id=CrlXAAAAMAAJ&amp;pg=PA105&amp;dq=railway+carriage+madman+england&amp;hl=en&amp;sa=X&amp;redir_esc=y#v=onepage&amp;q=railway%20carriage%20madman%20england&amp;f=false">medical journals</a> that there might be an invisible epidemic of &#8220;railway madness&#8221;, with many cases lying latent until sufferers experienced violent outbursts mid-commute and attacked innocent passengers. The media spread its fair share of <a href="https://www.tandfonline.com/doi/pdf/10.1080/13555502.2015.1118851?needAccess=true">alarmism</a> about the danger of railway madmen, and Victorian by-laws even went so far as to <a href="https://www.atlasobscura.com/articles/railway-madness-victorian-trains#:~:text=that%20%E2%80%9Cinsane%20persons%E2%80%9D-,should%20be%20isolated,-%E2%80%9Cin%20a%20compartment">stipulate</a> that &#8220;insane persons&#8221; be consigned to a designated carriage.</p><p>In fairness to Guo, the railway madness saga has all the trappings of a class tech-panic &#8211; expert alarm, media fear-mongering, clumsy attempts at regulation, and a semi-plausible sounding scientific hypothesis that appears ludicrous with the benefit of hindsight. There is clearly an analogy that skeptics can draw to the AI case. But again, to state the mind-numbingly obvious, &#8220;there were unwarranted concerns about the localised damage this technology might cause&#8221; is a very different claim to &#8220;people thought this technology might literally extinguish the human race&#8221;.</p><h4><strong>Electricity</strong></h4><p>Guys, I really held out hope for this one. It seemed pretty plausible to me that the discovery and widespread deployment of electricity would have sparked (pun intended) end-of-the-world fears. But after googling just about every keyword related to electricity-fueled millenarism I could think of, and a lengthy back-and-forth with Claude, I could find no evidence that this was the case<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>. As far as I can tell, the story here is pretty similar to the railway case &#8211; there <em>was</em> widespread apprehension about electricity, but this apprehension was not of an imminent apocalypse.</p><p>In my opinion, the kinds of worries people had at the dawn of the electrical age were not totally unreasonable. For example, a <a href="https://ieeexplore.ieee.org/document/464629">series of fatal accidents</a> caused by New York&#8217;s overhead electric wiring system led to widespread public outrage, accusations towards lighting companies of putting profit over safety, and demands that the wires be buried underground instead. An 1889 magazine cover conveys this sentiment pretty well:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!t0o9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!t0o9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png" width="730" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:730,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1629441,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!t0o9!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10c28bd7-bed9-4e6c-9ab4-7b29d1aafec4_730x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I sympathise with these 19th century New Yorkers! After all, a bunch of them had just the misfortune of <a href="https://aadl.org/node/507687">witnessing</a> a telegraph lineman, John Feeks, die instantly after being electrocuted by what was supposed to be a low-voltage telegraph line, and fall into the tangle of wire below before smouldering for the better part of an hour. I&#8217;d have been pretty freaked out too! In fact, I&#8217;d argue that some stories of panic over technology can be re-told as failures by companies and technologists to prioritise safety, leading to high-profile malfunctions and (understandable) public backlash. Maybe it&#8217;s really AI companies who should be heeding this particular historical warning. But I digress.</p><p>Perhaps the most famous example of an electricity-based tech panic was the 1880s-1890s <a href="https://en.wikipedia.org/wiki/War_of_the_currents">War of the Currents</a>, a head-to-head between two pioneers: Thomas Edison, who championed direct current, and George Westinghouse, who promoted alternating current. Each had invested significantly in their own favoured electrical system, and was keen for it to become the standard method of distributing electricity in the United States. To this end, Edison embarked on an elaborate propaganda campaign to provoke fear of his competition. He arranged public electrocutions of animals with alternating current, including dogs, horses and even an elephant, and advocated for its use to power the electric chair. Although alternating current eventually won out as the dominant mode of electrical distribution, in a way, Edison&#8217;s campaign worked. It <em>did</em> stoke fear and distrust of alternating current which <a href="https://www.nytimes.com/1979/02/06/archives/war-of-the-currents-had-profound-impact-the-war-of-the-currents-had.html">delayed</a> universal domestic adoption of it by over fifty years. Some private US households were still powered by direct current, which has since proven more dangerous, in the 1950s.</p><p>So, is the electricity example a good one for AI x-risk dismissers seeking examples of analogous tech panics? <em>Sort of</em> &#8211; it does seem that Edison was sincere in his belief that alternating current was dangerous, which is why, against the advice of his own colleagues, he did not invest in it himself. One could portray this as an example of a credible expert whose misguided fears of the technology he had spent many years studying was responsible for slowing progress and delaying mass adoption, an accusation often levied at so-called &#8220;AI doomers&#8221;. But in other ways, the analogy breaks down. In trying to stoke fear of a competitor&#8217;s technology while insisting on the safety of his own, Edison was, predictably, following strong financial incentives. I hardly need to point out that this is <em>precisely the opposite</em> of what safety-concerned AI labs are often <a href="https://x.com/fchollet/status/1702473623896990122">accused of</a> &#8211; that is, creating hype around the apocalyptic potential of their technology in a cynical ploy to attract investment, as if claiming that your products might kill everyone on Earth is some kind of tried-and-tested marketing technique. I think this claim is pretty baseless (<a href="/__u/garrisonlovely.substack.com/p/is-the-ai-doomsday-narrative-the">here&#8217;s a good write up of why</a>), and it&#8217;s interesting to note that at least in this one case, historical precedent seems to run the other way.</p><h4><strong>The Large Hadron Collider</strong></h4><p>This is probably my favourite apocalyptic tech-panic, because it&#8217;s a perfect example of non-experts fueling scary theories that the scientific community has been at pains to correct. In 2008, a few months before the Large Hadron Collider was set to come online, a Hawaiian man <a href="https://www.universetoday.com/13385/hawaiian-man-files-lawsuit-against-the-large-hadron-collider-lhc/">brought a lawsuit</a> against its completion. He alleged that scientists had overlooked evidence that the LHC would create a massive black hole which would swallow the Earth, citing a minor in physics from Berkeley and a career in nuclear medicine as evidence of his credibility. He was just one of many conspiracy theorists worried that the collider would <a href="https://www.dailymail.co.uk/news/article-3913952/Were-Italy-s-earthquakes-caused-HADRON-COLLIDER-Bizarre-theory-emerges-experiment-fire-plasma-Geneva-250-miles-underground-Italy.html">cause earthquakes</a>, shift the world into an <a href="https://www.cnbc.com/2017/02/22/alternate-realities-and-trump-mandala-effect-and-what-cern-does.html">alternate timeline</a>, or open a portal to <a href="https://eu.usatoday.com/story/news/factcheck/2022/07/26/fact-check-scientists-cern-not-opening-portal-hell/10094679002/">hell</a>. But CERN was quick to issue a <a href="https://www.home.cern/science/accelerators/large-hadron-collider/safety-lhc">statement</a> assuring the public that they were the Adults In The Room and that the LHC presented no danger, that the science confirming its safety was extremely watertight, and that a who&#8217;s who of prestigious experts in astrophysics, cosmology, general relativity, mathematics and particle physics had all agreed that there was nothing to worry about.</p><p>This sort of reassuring expert consensus is precisely what I spent so long searching for when I first became worried about AI risk. I thought: <em>surely</em> it can&#8217;t be the case that a decades-long technological venture, benefitting from money and talent on an incredible scale, is actually on track to bring about the end of the world &#8211; because, presumably, someone would have checked. And without a high degree of confidence that it won&#8217;t destroy everything, humanity surely wouldn&#8217;t be participating in such a project. I tried really, really hard to find that elusive consensus. But this was an exercise in futility, because as anyone reading this likely knows, it simply doesn&#8217;t exist.</p><h4><strong>The Y2K panic</strong></h4><p>Ah, Y2K. The go-to example in the rhetorical arsenal of anyone looking to dismiss fears of disaster with a &#8220;we&#8217;ve been here before&#8221;, an &#8220;I&#8217;m old enough to remember when&#8230;&#8221; or a &#8220;oh honey, this must be your first technological doomsday&#8221;.</p><p><em>Is all this fuss over AI doom just the next Y2K? </em>ask <a href="https://www.bloomberg.com/news/videos/2023-07-13/technically-speaking-is-ai-the-next-y2k-video">Bloomberg</a>, the <a href="https://www.theguardian.com/technology/2023/jun/02/the-existential-threat-from-ai-and-from-humans-misusing-it">Guardian</a> and a bunch of people on <a href="https://www.reddit.com/r/singularity/comments/157knc7/what_if_the_current_fear_around_ai_alignment_is/">Reddit</a>. It&#8217;s a tempting comparison. The conventional tale of Y2K is that programmers anticipated a bunch of computer failures at the dawn of the new millennium because dates had hitherto been expressed using just two digits for the year, rather than four. They feared that planes would fall out of the sky, nukes would launch themselves, and critical infrastructure &#8211; banks, hospitals and power grids &#8211; would crumble and leave civilization in disarray. A golden era bloomed for professional survivalists and religious evangelists turning public fear into profit. Governments dropped billions on technical fixes for the so-called &#8220;Y2K bug&#8221;, and printed an excess of paper money in case of a run on the banks. And then, when the 1st January 2000 came and went without incident, everyone was left feeling just a little bit silly. So maybe when superintelligence arrives and it goes totally fine, we&#8217;ll think dedicating <a href="https://www.ox.ac.uk/news/2024-05-20-world-leaders-still-need-wake-ai-risks-say-leading-experts-ahead-ai-safety-summit">a whopping 1-3% of AI research</a> to safety was a big embarrassing overreaction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mKzX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 424w, /__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 848w, /__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mKzX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png" width="792" height="1054" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1054,&quot;width&quot;:792,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1656685,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 424w, /__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 848w, /__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mKzX!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F993cd13e-d082-48c8-81b6-81a71adb0450_792x1054.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I think this analogy is bad for a couple of reasons. Both are quite obvious, but I may as well explain them while we&#8217;re here. First, as many have <a href="https://time.com/5752129/y2k-bug-history/">pointed out</a>, the turn of the millennium did not bring disaster <em>precisely because</em> such a huge amount of behind-the-scenes work went into averting one. By the time that the general public and mainstream media caught on, programmers had been quietly working away at the problem for over a decade. This was an effort for which they received very little fanfare &#8211; the seamless transition into a new millennium (which they had enabled) saw a very real technical problem (which they had fixed) dismissed as an episode of paranoid doomsaying. Also, not to be corny or anything, but I think the global Y2K mitigation efforts were a feat of international coordination nothing short of inspirational. The UN&#8217;s <a href="https://usinfo.org/wf-archive/2000/000217/epf414.htm">International Y2K Cooperation Center</a> tracked Y2K readiness worldwide and defined and shared best practices on tackling risk. Russia and the US <a href="https://www.nytimes.com/1999/10/28/world/us-and-russia-agree-on-joint-defense-against-y2k-debacles.html">cooperated</a> to prevent glitches in either&#8217;s early warning systems leading to an accidental nuclear launch. We love to see it.</p><p>Second, by the time the millennium rolled around, the governments of the world were actually quite confident that no disaster would occur &#8211; and had been for some time (though the public was slow to catch on). This is because the Y2K bug was a problem we knew how to fix. Essentially, we needed to update legacy software to store dates with four-digit years instead of two. We had a plan. When it comes to the problem of ensuring that superintelligent AIs behave as intended, <a href="/__u/longerramblings.substack.com/p/i-read-every-major-ai-labs-safety">we don&#8217;t</a> (it&#8217;s not that no one has proposed plans of course, but that we have nothing close to scientific consensus that any of them will work).</p><h4><strong>Nuclear power</strong></h4><p>A common argument by opponents of AI regulation goes something like: &#8220;hey, look at the regulatory overreaction to nuclear power, and how it stifled innovation, degraded public trust and set us back decades in the fight against climate change&#8221;. There&#8217;s certainly something to this (though I don&#8217;t really feel like weighing in on the debate over whether extreme caution on nuclear power was prospectively reasonable, even if it had the ultimate effect of hampering progress).</p><p>But how analogous are AI fears to worries over nuclear energy? To get this out of the way first &#8211; I could not find evidence of credible concern that civilian applications of nuclear power would end the entire world. As far as I can tell, the general story of nuclear energy is that the public wariness of it has been <a href="https://www.iaea.org/sites/default/files/publications/magazines/bulletin/bull33-3/33304793036.pdf">disproportionate</a> to the actual risk it bears, and that scientists have largely <a href="https://world-nuclear.org/information-library/safety-and-security/safety-of-plants/safety-of-nuclear-power-reactors">tried to correct</a> these misconceptions (not that it doesn&#8217;t have serious downsides, or hasn&#8217;t caused several localised disasters). Again, this is not the case with AI, which many experts do believe presents a serious danger of human extinction. So whether we &#8220;overreacted&#8221; to the threat posed by nuclear energy seems irrelevant to how we should address the much greater threat posed by AI.</p><p>Of course, nuclear weapons, which almost destroyed the world<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> <a href="https://www.ucsusa.org/sites/default/files/attach/2015/04/Close%20Calls%20with%20Nuclear%20Weapons.pdf">several times</a> and still could, are a different story. It will come as no surprise that I think our success in not detonating another nuclear weapon for almost 80 years after one was first used in warfare, against the expectations of many at the time, is a testament to the importance (and the feasibility!) of global coordination around existential threats. Luckily, I don&#8217;t see many people making the argument that our avoidance of nuclear war thus far is evidence that we should be less concerned about AI risks. So that&#8217;s something.</p><p>To round out this first section, I don&#8217;t think any of the historical tech panics I&#8217;ve discussed are even vaguely analogous to present-day concerns about AI. In all of them, it appears to be the case that either a) some lay people worried that the world might end, but experts were largely unconcerned or b) pretty much no one was concerned that the world might end, but people did have sub-existential worries that turned out to be unwarranted. This is pretty much what I expected to find, and I&#8217;m a bit worried that I&#8217;ve just spent over 2,500 words stating the obvious &#8211; but given that people are still going around carelessly tossing out sentiments like &#8220;everyone always thinks every new technology will be the end of the world&#8221;, it seemed worth a few hours of my time. Obviously, I haven&#8217;t been exhaustive here. There are many other transformative technologies I haven&#8217;t covered, and inevitably important things I&#8217;ve missed about the ones I did. But I&#8217;m fairly confident that the broad point is sound.</p><div><hr></div><h3>Part II: False (AI)larms</h3><p>As I stated in the intro, there are two varieties of this &#8220;crying wolf&#8221; argument &#8211; one that encapsulates the history of technology writ large, and another that accuses AI safety advocates of spreading false alarm about existing AI models, none of which have proved catastrophic.</p><p>Here are just a few examples of this claim in the wild:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!M7KQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 424w, /__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 848w, /__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 1272w, /__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!M7KQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png" width="1456" height="1195" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1195,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1756853,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 424w, /__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 848w, /__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 1272w, /__u/substackcdn.com/image/fetch/$s_!M7KQ!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9502d69c-7c62-49a5-a111-96e04ae51f8d_1518x1246.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">From top left:<a href="https://livingwithinreason.com/p/the-ai-doomers-are-crying-wolf-and"> Living Within Reason</a> on Substack, @<a href="https://x.com/kimmonismus/status/1872522046493909472">kimmonismus</a> on Twitter, an opinion piece from<a href="https://www.forbes.com/sites/trondarneundheim/2023/05/31/the-cry-wolf-moment-of-ai-hype-is-unhelpful/"> Forbes</a>, @<a href="https://x.com/Dan_Jeffries1/status/1868956616864702484">Dan_Jeffries1</a> on Twitter, Kevin Roose on the<a href="https://www.nytimes.com/2024/12/06/podcasts/is-intel-cooked-whats-your-p-dyson-sphere-hard-fork-gift-guide.html"> Hard Fork</a> podcast</figcaption></figure></div><p>I have a tricky task here, because proving a negative &#8211; that very few people have actually made concrete, since-falsified predictions about AI-caused catastrophes &#8211; is obviously impossible without doing a comprehensive audit of the entire AI safety discourse. As it happens, I probably have done something pretty close to such an audit, because I am just <em>really obsessed</em> with this topic, but I have no way to prove that.</p><p>So let&#8217;s start with what I can prove. For one, none of the accusations above actually cite such predictions (sources are in the caption so that you can verify this yourself). If the authors had actually seen any, you&#8217;d think that they&#8217;d link them, rather than passing up the opportunity to land what would be a pretty big rhetorical win. One type of evidence they do offer is specific calls for regulation, which are often portrayed as synonymous with predictions that, absent intervention, AI would prove catastrophic at certain levels of capability. Take the opening of the Forbes article:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Q85j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 424w, /__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 848w, /__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Q85j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png" width="1058" height="414" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:414,&quot;width&quot;:1058,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 424w, /__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 848w, /__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Q85j!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36c7bf6b-e9e4-4cf4-bddd-bb5f0a7bc80e_1058x414.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There are no specific timeline predictions in the <a href="https://futureoflife.org/open-letter/pause-giant-ai-experiments/">FLI 6-month pause letter</a>, the <a href="https://www.safe.ai/work/statement-on-ai-risk">CAIS statement</a> or Eliezer Yudkowsky&#8217;s <a href="https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/">TIME article</a>. None claim that we were at the time of their publication in &#8220;immediate danger&#8221; from AI. The second doesn&#8217;t even propose any particular intervention &#8211; though of course, one can reasonably disagree with the policies put forward in the other two.</p><p>This reply to the Daniel Jeffries tweet is another good example:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JIOu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 424w, /__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 848w, /__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JIOu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png" width="960" height="1252" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1252,&quot;width&quot;:960,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 424w, /__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 848w, /__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JIOu!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6672de56-7924-4dc8-ac30-141b32d9a496_960x1252.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I&#8217;m not making any claims about whether the thresholds above are sensible, or whether it was wise for them to be suggested when they were. I do think it seems clear with hindsight that some of them are unworkably low. But again, advocating that AI development be regulated at a certain level is <em>not the same </em>as predicting with certainty that it would be catastrophic not to. I often feel that taking action to mitigate low probabilities of very severe harm, otherwise known as &#8220;erring on the side of caution&#8221; somehow becomes a foreign concept in discussions of AI risk.</p><p>An easily debunkable claim in the above compilation is that Geoffry Hinton is guilty of crying wolf. In 2023, Hinton was <a href="https://x.com/geoffreyhinton/status/1653687894534504451?lang=en">predicting</a> that smarter-than-human AIs would emerge within 5 to 20 years. So, a credible crying-wolf accusation will not be leverageable against him until 2043. It seems our collective attention span is becoming so short that a person warning &#8220;for months&#8221; about a catastrophic event predicted to occur at least several years into the future can be accused of alarmism. It&#8217;s not looking good for The Discourse.</p><h4><strong>Taking to the tweets</strong></h4><p>Last month, having run into the crying-wolf meme countless times without having been signposted to so much as (1) specific and since-debunked prediction, I decided to <a href="https://x.com/littIeramblings/status/1866091256775880717">take to Twitter</a> in search of some. If anyone was sitting on a smoking gun, I wanted to hear about it. This type of open call is about as wide as I can cast my net, but obviously my methodology here is imperfect &#8211; I am disproportionately followed by people who are in my corner of the AI debate, the tweet was seen by less than 10,000 people, and despite my best efforts, <a href="https://x.com/littIeramblings/status/1835324184760602721">I&#8217;m just not that popular</a>. Even so, if the Debunked Prediction Graveyard is as large as the accelerationist crowd makes it sound, I&#8217;d expect at least a few to make their way into my replies<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>.</p><p>I didn&#8217;t receive anything that would meet my (admittedly high) bar of a specific and since-debunked prediction (I&#8217;ve said something to this effect so many times now that I&#8217;m thinking it needs an acronym or something. An SASDP?). That said, there were a few not-totally-unfair criticisms of the AI safety movement &#8211; and I said at the beginning of the piece that I wanted to give the crying wolf argument its due. So let&#8217;s try and assess some of these accusations.</p><p>One that caught my eye was a <a href="https://x.com/1a3orn/status/1866149540002250824">claim</a> that <a href="https://thefuturesociety.org/">The Future Society</a>, a fairly well-respected non-profit with a mission of &#8220;[aligning] artificial intelligence through better governance&#8221;, define models trained with between 10^23 and 10^26 FLOP as having the potential to pose an existential risk if seriously misaligned. In a 2023 <a href="https://thefuturesociety.org/wp-content/uploads/2023/09/heavy-is-the-head-that-wears-the-crown.pdf">report</a>, TFS categorises such models as &#8220;Type II General Purpose AIs&#8221;. It recommends a bunch of requirements for Type II GPAI developers including third party audits, some fairly stringent-sounding infosec and cybersecurity protocols, and the commitment to pause development of, or even un-deploy, models that cannot be proven &#8220;absolutely trustworthy&#8221;. I <em>still</em> wouldn&#8217;t classify this as a falsified prediction, since it isn&#8217;t a prediction at all, just a proposal of what many might consider heavy-handed regulation at relatively low compute thresholds (I will go to my grave swearing that these are importantly different). But I <em>do</em> disagree with TFS&#8217;s decision to place the lower bound of their Type II classification as low as 10^23 FLOP, especially since models with a higher FLOP count than this had already been trained, deployed and likely even open-sourced<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> before the report was published. So I would consider this a mistake on the part of TFS &#8211; and a contributor to the crying-wolf effect &#8211; if not strictly a debunked prediction.</p><p>Onto another <a href="https://x.com/1a3orn/status/1866148561223926211">accusation</a> &#8211; that civil resistance group StopAI previously claimed we risked a fast-takeoff-extinction-event from all models trained on more than 10^23 FLOP and demanded their deletion, but have since removed this from their website. I obviously can&#8217;t verify this, though it wouldn&#8217;t surprise me. I think this is the most straightforward crying wolf example (StopAI have been guilty of this <a href="https://x.com/StopAI_Info/status/1861579609311182950">a few times</a> recently). All I&#8217;d ask here is that the anti-safety faction grant us the good grace not to judge an entire movement by its fringe.</p><p>Eliezer Yudkowsky is the target of many crying wolf accusations, on <a href="https://x.com/1a3orn/status/1866150951763206291">my particular tweet thread</a> and elsewhere. The position he has staked out &#8211; that AI could recursively self-improve into a superintelligent system that quickly disempowers humanity at any point &#8211; obviously makes him very vulnerable to them. Yudkowsky has <a href="https://x.com/ESYudkowsky/status/1726606636985536939">repeatedly refused</a> to make timeline predictions, which is either to his credit <em>or</em> a ginormous cop-out depending on your perspective. This is a catch-22 that I&#8217;m not sure how to address. On the one hand, people have every right to be suspicious of what appear to be unfalsifiable claims of disaster that can be eternally deferred into the future. On the other, recursive self-improvement is a live possibility that is taken seriously by a large number of machine learning researchers<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>. It&#8217;s entirely possible that we live in an unlucky timeline where Cassandra-like prophecies of doom are indistinguishable from reality until it is too late.</p><p>The final, and most easily countered, set of accusations <a href="https://x.com/ygrowthco/status/1866216329038254123">revolve</a> around GPT-2 being considered &#8220;too dangerous&#8221; to release. As several people in my replies <a href="https://x.com/justjoshinyou13/status/1866108966163317135">pointed out</a>, GPT-2&#8217;s release was delayed because of concerns about misinformation and other forms of low-level misuse. OpenAI <a href="https://openai.com/index/gpt-2-6-month-follow-up/">postponed</a> its deployment while they studied the issue, which seems eminently sensible to me.</p><p>In the end, I don&#8217;t think my little Twitter experiment revealed evidence that the AI discourse is riddled with failed predictions from the safety side. I think it <em>did</em> reveal a few strategic errors that contribute to what I&#8217;ll call a &#8220;crying wolf effect&#8221; &#8211; such as calls for stringent regulation at unrealistically low thresholds. In the last section, I&#8217;ll say a bit more about how I think we should make ourselves less vulnerable to this effect.</p><div><hr></div><p>I think the AI safety-concerned have a very tough communicative challenge in front of us, for several reasons. First, humanity <em>does</em> have a history of collective neurosis over transformative technologies, which understandably provokes skepticism of what may look on the surface like just another tech-panic. AI safety advocates have tried hard to <a href="https://x.com/AISafetyMemes?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Eauthor">brand themselves</a> as techno-optimists about every technology except for powerful AI to dodge accusations of luddism, but this isn't an easy pitch. Second, AI development is extremely unpredictable. This has led in the past to what seem in retrospect like overly-draconian policy recommendations, and inevitably will do again in the future. And third, the AI safety movement is a broad coalition. We can&#8217;t expect to maintain a perfect predictive track-record between us. This problem is exacerbated by the fact that there are many reasons people have for being concerned about bad outcomes from AI, not all of which will turn out to have been the <em>right</em> reasons. So each of us &#8220;shares a side&#8221; with people whose reasoning may be very different to our own, and whose predictions we may not endorse (<a href="https://x.com/AmandaAskell/status/1825705325850407411">this tweet</a> from Amanda Askell sums it up pretty well).</p><p>As it stands, I don&#8217;t think the AI safety community writ large has cried wolf. But if transformative AI capable of disempowering humanity does not emerge within the next ten years (at the most) I think it will be fair to say that we have<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>. I happen to think superintelligence within ten years is more likely than not, but I of course could be wrong! I&#8217;m worried that we&#8217;re entering a period in which the AI safety community risks losing a bunch of credibility.</p><p>So what should we do about this? Personally, I think anyone making timeline predictions should not just acknowledge but <em>emphasise</em> their uncertainty. People can reasonably disagree with me here; there are certainly benefits to signalling high levels of confidence &#8211; I just happen to think that the risk of lost credibility outweighs them.</p><p>As I mentioned earlier, one big contributor to the crying-wolf effect is overly-stringent policy recommendations. Given that we have to make decisions under uncertainty, I think we also have to accept the inevitability of proposing policies that ultimately prove unnecessary. But in my opinion, those doing so should caveat that they are designed to mitigate the <em>possibility</em> (and not certainty) of catastrophic outcomes. This should be obvious, but given that people will be waiting in the wings to weaponise anything that could be called a regulatory overreaction, I think it&#8217;s worth doing.</p><p>One final suggestion (which I&#8217;m less confident about) is that we spend less energy trying to forecast when transformative AI will arrive, and more energy making plans for short timelines. If we believe it&#8217;s plausible that takeover-capable AI could arrive by 2027, we can simply spend our time advocating for policy that would be effective in that scenario, while acknowledging the possibility that it takes much longer. This isn&#8217;t to say that <em>no one</em> should be focused on forecasting timelines, of course. The ideal gameplan for ASI in 2040 probably looks pretty different to the one for ASI in 2027 &#8211; but I doubt there&#8217;s much difference between the gameplans for 2027 vs 2028. The less numerous and specific our forecasts, the less vulnerable we are to crying wolf accusations.</p><p></p><p>Humanity might have suffered misplaced anxiety about technology in the past, but as anyone reading this likely agrees &#8211; this time really is different. Let&#8217;s get it right!</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/longerramblings.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I won&#8217;t make the case again here, but here is <a href="https://medium.com/@daniel_eth/ai-alignment-explained-in-5-points-95e7207300e3">one of my favourite high-level explainers</a> for anyone not up to speed.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I obviously can&#8217;t be completely exhaustive here, so I&#8217;m open to being corrected! Same goes for any of the other technologies / tech panics discussed.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Yes, I am aware that whether a large-scale nuclear war would cause literal human extinction is still an issue of live debate, but for the sake of simplicity I am conflating &#8220;could kill everyone&#8221; and &#8220;could kill almost everyone&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>And I&#8217;m still open to submissions!</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Epoch <a href="https://epoch.ai/blog/tracking-large-scale-ai-models">estimates</a> that Meta&#8217;s Llama 2-70B, which was open-sourced the summer before the TFS report was published, was trained on more than 10^23 FLOP.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>In the <a href="https://aiimpacts.org/wp-content/uploads/2023/04/Thousands_of_AI_authors_on_the_future_of_AI.pdf">2023 AI Impacts survey</a>, 53% of respondents thought that an &#8220;intelligence explosion&#8221; triggered by rapid AI self-improvement, was at least 50% likely.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>I think <a href="https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_greenblatt-s-shortform?commentId=iwodobEWjt9qwHbb2">this LessWrong comment</a> from Ryan Greenblatt of Redwood Research is a pretty good indicator of where I would say credible timeline predictions are now generally clustered (~2029-2034). This particular comment offers a probability distribution that I think leaves plenty of room for uncertainty, but there are other predictions that do not.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[On futile rage against the chatbots]]></title><description><![CDATA[When AI says it better.]]></description><link>https://longerramblings.substack.com/p/on-futile-rage-against-the-chatbots</link><guid isPermaLink="false">https://longerramblings.substack.com/p/on-futile-rage-against-the-chatbots</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Fri, 13 Dec 2024 00:41:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6599f043-077f-4c64-87a6-8ed3abf45f4f_1228x1034.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Chatbots keep taking the words right out of my mouth.</p><p>Occasionally, when I&#8217;m struggling with how to articulate a half-formed thought, I type my muddled ramblings into Claude and ask,</p><p>&#8220;What am I trying to say here?&#8221;. And then, to my dismay, <em>it tells me</em>.</p><p>I watch that little orange cursor blink dutifully across the screen, rearranging the contents of my own mind into eerily perfect prose.</p><p>&#8220;Where would you like to take this next?&#8221;, asks Claude, perfectly compliant and diligently helpful as always. I feel the same exasperation that hits me when my cat delivers a dead mouse at my feet in a misguided act of generosity, staring up at me with her huge, earnest eyes.</p><p>&#8220;Yes, that&#8217;s exactly what I meant&#8221;, I murmur. &#8220;Also, fuck you&#8221;.</p><p>My turbulent relationship with chatbots brings out the full force of my irrationality. I am not ashamed to admit that it <em>triggers</em> me. It will have me alone in my room yelling "<em>I CAN DO IT MYSELF</em>" at my laptop, in a (wo)man-meets-machine display of pointless fury to rival <a href="https://www.youtube.com/watch?v=LhzckCB3Bo8">Basil Fawlty thrashing his car with a tree branch</a>. At a recent EA Global conference, I had a one-on-one meeting with an undergraduate who lamented the fact that most of his classmates are using ChatGPT to complete their assignments. Suddenly, I was 26-going-on-60, denouncing the myopic laziness of &#8216;kids these days&#8217; with all the Luddism of a parent decrying mobile phones at the dinner table.</p><p>&#8220;That might help them get ahead now, but it won&#8217;t serve them in the future&#8221;, I heard myself say in an uncharacteristically shrill voice. &#8220;They should be considering their post-graduation prospects&#8221; (as I was saying this, I was acutely aware that it literally made no sense given my own single-digit-year AGI timelines).</p><p>I have many gripes with AI. My biggest one is the alarmingly high likelihood that it will kill everyone. My second biggest one is the ease with which it does in seconds what I do in hours. I have always taken great pleasure in crafting a well-formed sentence. I love arriving at the perfect articulation of an idea after rounds of near-headache-inducing mental labour. I don&#8217;t imagine that Claude gets headaches<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</p><div><hr></div><p>One evening, I decide to experiment with AI text detection. I spend an hour or so pasting chunks of my own writing into free online tools, followed by AI-generated content on the same topics. To my absolute delight, I discover that, at least initially, <em>it seems like they actually work</em>. The detectors are distinguishing between human and AI text with impressive accuracy, and I am high on my own supply of hopium. Ah yes, the sacrosanct uniqueness of the human voice, irreproducible by the shallow mimicry of simple pattern-detecting machines. The skeptics were right! There <em>is</em> a secret sauce, and ZeroGPT has found it.</p><p>But it doesn&#8217;t take long for skepticism to rear its ugly head. It <em>does</em> seem unlikely that there would exist some magical tool capable of differentiating between the products of biological and silicon minds. And I already know that AI can generate content far more sophisticated than I&#8217;ve been testing the detectors with using just a little more prompting (reluctant though I am to acknowledge this).</p><p>There is a Substacker whose work I enjoy. She writes lovely, touching vignettes about small-yet-significant personal experiences like lending a dog-eared copy of a favourite paperback to her mum. It&#8217;s precisely the kind of thing that your resident skeptic friend who hasn&#8217;t so much as tried an LLM will swear up and down that AI could never produce. I give Claude the prompt:</p><p><em>&#8220;Write a touching story about lending a dog-eared copy of a favourite paperback to your mum, in the style of a personal essay on Substack&#8221;.</em></p><p>Its first try, <em>The Book That Came Home,</em> is a laudable attempt at the kind of emotional authenticity that I&#8217;m after &#8211; my detector-of-choice, Quillbot, estimates that it is 51% AI-generated. I take it up a notch:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!TK0m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 424w, /__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 848w, /__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!TK0m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png" width="1302" height="674" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:674,&quot;width&quot;:1302,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 424w, /__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 848w, /__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TK0m!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4db835-371d-433a-8d29-8a6faf13ef5a_1302x674.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This second essay is actually quite poignant. Claude has decided to pack an extra punch by giving the mum cancer. It&#8217;s a cheap move, but I&#8217;m in an emotionally vulnerable place, so it works on me. It apparently also works on Quillbot, which confidently declares that <em>Mom&#8217;s Sticky Notes</em> is 0% AI-generated. The rage is back. I feel the urge to write a strongly-worded email informing Quillbot of this false negative, and reminding them of their duty as the final line of defence against the unstoppable encroachment of AI into every corner of the internet, and as one of the few remaining bricks in a wall around the sacred territory of human creativity. But of course, there is no wall (in <a href="https://twitter.com/sama/status/1856941766915641580">more ways than one</a>), and the whole enterprise of building one, as I well knew even before my foray into the world of AI text detection, is patently ridiculous.</p><h3>Two sources of cope</h3><p>I&#8217;ve always loved to write. That AI can do it so effortlessly upsets me enough to inspire this petty blog post. But there are reasons I think this need not be as nihilism-inducing as it sometimes feels.</p><p>First, for the time being, there remains at least some skill in prompting chatbots to produce the kinds of content you want. Extremely low-effort prompts result in writing replete with tell-tale signs that AI detectors can spot a mile off. But very good ones can produce writing idiosyncratic enough that, just maybe, there is <em>something</em> of the human user that survives into the final product. For example, Amanda Askell of Anthropic got Claude to generate a <a href="https://x.com/AmandaAskell/status/1860430577423520100">whimsical fable</a> about a man locating the deed to some valuable land using a precise arrangement of garden gnomes, making me wonder whether the creative locus of AI text can sometimes lie with the prompter, rather than the promptee<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> (though of course, Amanda is one of the staff members behind the alignment of Claude, meaning she gets to take a little more credit for its outputs than the rest of us).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FOyH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 424w, /__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 848w, /__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FOyH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png" width="1034" height="1198" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1198,&quot;width&quot;:1034,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 424w, /__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 848w, /__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FOyH!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d1ed938-c99e-4223-b957-9ea60b1e2b38_1034x1198.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Second, and I think more importantly, each of us has always shared the planet with minds better than our own at all sorts of things &#8211; all that has really changed is that some of those minds are now artificial. When I was younger, I was part of a children&#8217;s choir. One Sunday afternoon rehearsal when we were both around 13, a friend of mine dramatically announced her intention to give up singing because she was upset that some other choir members were better at it than her. She couldn&#8217;t see the point in any pursuit that she wasn&#8217;t the best at. I pointed out that only one person on Earth could claim the title of Best Singer, and that it would be a great loss to the world if everyone else were to stop singing. If 13-year-old me could grasp this simple lesson, 26-year-old me would do well to heed it now.</p><p>However capable AI becomes, I think it will continue to matter that my words are my own &#8211; and I will keep writing them.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/longerramblings.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>(<a href="https://www.transformernews.ai/p/anthropic-ai-welfare-researcher">Probably</a>)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><a href="https://en.wikipedia.org/wiki/The_Death_of_the_Author">Death of the prompter</a> &#8211; is this something?</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[I read every major AI lab’s safety plan so you don’t have to]]></title><description><![CDATA[AI labs acknowledge that they are taking some very big risks. What do they plan to do about them?]]></description><link>https://longerramblings.substack.com/p/i-read-every-major-ai-labs-safety</link><guid isPermaLink="false">https://longerramblings.substack.com/p/i-read-every-major-ai-labs-safety</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Fri, 29 Nov 2024 15:27:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iYaO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A handful of tech companies are competing to build advanced, general-purpose AI systems that radically outsmart all of humanity. Each acknowledges that this will be a highly &#8211; perhaps <a href="https://www.safe.ai/work/statement-on-ai-risk">existentially</a> &#8211; dangerous undertaking. How do they plan to mitigate these risks?</p><p>Three industry leaders have released safety frameworks outlining how they intend to avoid catastrophic outcomes. They are OpenAI&#8217;s <a href="https://cdn.openai.com/openai-preparedness-framework-beta.pdf">Preparedness Framework</a>, Anthropic&#8217;s <a href="https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic-Responsible-Scaling-Policy-2024-10-15.pdf">Responsible Scaling Policy</a> and Google DeepMind&#8217;s <a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/introducing-the-frontier-safety-framework/fsf-technical-report.pdf">Frontier Safety Framework</a>.</p><p>Despite having been an avid follower of AI safety issues for almost two years now, and having heard plenty about these safety frameworks and how promising (or disappointing) others believe them to be, I had never actually read them in full. I decided to do that &#8211; and to create a simple summary that might be useful for others.</p><p>I tried to write this assuming no prior knowledge. It is aimed at a reader who has heard that AI companies are doing something dangerous, and would like to know how they plan to address that. In the first section, I give a high-level summary of what each framework actually says. In the second, I offer some of my own opinions.</p><p>Note I haven&#8217;t covered every aspect of the three frameworks here. I&#8217;ve focused on <strong>risk thresholds</strong>, <strong>capability evaluations</strong> and <strong>mitigations</strong>. There are some other sections, which mainly cover each lab&#8217;s governance and transparency policies. I also want to throw in the obvious disclaimer that I have not been comprehensive here and have probably missed some nuances despite my best efforts to capture all the important bits!</p><h2>What are they?</h2><p>First, let&#8217;s take a look at how each lab defines their safety framework, and what they claim it will achieve.</p><p>Antrophic&#8217;s Responsible Scaling Policy is the most concretely defined of the three:</p><p><em>&#8220;a public commitment not to train or deploy models capable of causing catastrophic harm unless we have implemented safety and security measures that will keep risks below acceptable levels.&#8221;</em></p><p>OpenAI&#8217;s Preparedness Framework calls itself:</p><p>&#8220;<em>a living document describing OpenAI&#8217;s processes to track, evaluate, forecast, and protect against catastrophic risks posed by increasingly powerful models.&#8221;</em></p><p>Finally, Google Deepmind&#8217;s Frontier Safety Framework is:</p><p><em>&#8220;a set of protocols for proactively identifying future AI capabilities that could cause severe harm and putting in place mechanisms to detect and mitigate them&#8221;.</em></p><p>Anthropic&#8217;s RSP is notable in being defined from the get-go as a &#8216;commitment&#8217; to halt development if extreme risks cannot be confidently mitigated. That said, all three frameworks <em>do</em> go on to specify conditions under which they would pause training and/ or deployment (more on that later).</p><h2>Thresholds &amp; triggers</h2><p>Each framework follows a similar structure in terms of the mechanisms it constructs for triggering further action.</p><p>Broadly speaking, there are two key elements to these mechanisms:</p><ol><li><p><strong>Capability thresholds</strong> measure the extent to which models can do scary things. Dangerous capabilities are tracked across different categories, such as cybersecurity or model autonomy.</p></li><li><p>Based on the capabilities they exhibit, models are assigned a <strong>safety category</strong>. Which safety category a model is in dictates which safeguards need to be applied.</p></li></ol><h4>Capability thresholds</h4><p>There are three sets of capabilities that are tracked in each of the frameworks:</p><ul><li><p><strong>Chemical, Biological, Radiological, and Nuclear (CBRN)</strong>: Could the model make it significantly easier for bad actors to design weapons that cause mass harm?</p></li><li><p><strong>Autonomy</strong>: Could the model replicate and survive in the wild? Could it meaningfully accelerate the pace of AI R&amp;D? Could it autonomously acquire resources in the real world in order to achieve its goals?</p></li><li><p><strong>Cyber</strong>: Could the model enhance or even fully automate the process of carrying out a sophisticated cyberattack?</p></li></ul><p>However, there are some important differences between the tracked risks in each framework:</p><ul><li><p>Cyber capabilities do not quite rise to the level of a tracked risk category in Anthropic&#8217;s RSP. Anthropic takes a &#8216;watch and wait&#8217; approach, where the cyber capabilities <em>are</em> tracked, but they do not pre-specify a point at which they would trigger a model to move up a safety category.</p></li><li><p>OpenAI&#8217;s Preparedness Framework contains a unique risk category, <strong>persuasion</strong>. This category tracks a model&#8217;s ability to generate content which could change a person&#8217;s beliefs.</p></li><li><p>Deepmind&#8217;s Frontier Safety Framework distinguishes between a model&#8217;s ability to accelerate AI R&amp;D and its ability to act autonomously more generally.</p></li></ul><p>In short, all three frameworks aim to track three things:</p><ol><li><p>Whether models could help people do very bad things in the CBRN or cyber domains</p></li><li><p>Whether they could ultimately do very bad things on their own</p></li><li><p>And whether they could accelerate AI R&amp;D such that we&#8217;ll be confronted with (1) and (2) much sooner than we would have been otherwise.</p></li></ol><h4>Risk categories</h4><p>I&#8217;ve consolidated all of the risk categories in each of the three frameworks into one table. I&#8217;ve also indicated which category each frontier lab&#8217;s most recent model is, according to their own evaluations:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iYaO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 424w, /__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 848w, /__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iYaO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png" width="1456" height="2270" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2270,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:527570,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 424w, /__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 848w, /__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iYaO!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd78325bc-91c5-4641-9299-84b0cf3b45f9_2401x3743.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You&#8217;ll notice that in order to be moved up a safety category, OpenAI models only have to possess one of several capabilities designed to trigger that threshold. This is because while a model is scored in each individual risk, its overall score is <strong>the</strong> <strong>highest in any</strong>. For example, if a model scored &#8216;low&#8217; in cybersecurity, CBRN, and autonomy, but &#8216;medium&#8217; in persuasion, its overall score would be &#8216;medium&#8217; &#8211; and it would trigger the safety protocols associated with that risk category.</p><h2>Evaluations</h2><p>So far we&#8217;ve covered the risks that each lab has committed to track, and the different risk categories that models can be placed into. In order to actually track these risks, labs need to run evaluations or &#8216;evals&#8217; on their models.</p><p>Evals aim to detect dangerous capabilities. All three policies state an intention to perform evals that <strong>actually elicit the full capabilities of the model. </strong>They aim to discern what the model could do in a worst-case scenario &#8211; if people manage to improve a model&#8217;s capabilities after it has been deployed through fine-tuning or sophisticated prompt engineering, for example:</p><ul><li><p>OpenAI commits to running <strong>both pre-mitigation and post-mitigation evals </strong>on their models. Pre-mitigation evals detect what capabilities a model has before any safety features are applied. They also say they will run evals on versions of models that have been fine-tuned for some specific bad purpose (for example, a model that someone has tailored to be really good at cyberattacks).</p></li><li><p>Anthropic also says that they will <strong>evaluate models without safety mechanisms</strong>, and that they will act under the assumption that &#8216;jailbreaks and model weight theft are possibilities&#8217;. They are less specific about testing models tailored for particular purposes, but do say that they will &#8216;at minimum&#8217; test models fine-tuned in ways that might make them more generally useful to bad actors (this could include a model that has been fine-tuned to minimise refused requests, for example).</p></li><li><p>DeepMind is most vague on this point. In the &#8216;Future work&#8217; section of their framework, they say that they are &#8216;working to equip [their] evaluators with state-of-the-art elicitation techniques, to ensure [they] are not underestimating the capability of [their] models&#8217;.</p></li></ul><p>A natural question to ask next is &#8211; what evaluations will labs actually run? None of the documents provide a clear answer to this question.</p><p>Anthropic&#8217;s RSP doesn&#8217;t contain any specific information on planned evaluations. However, they do state in an <a href="https://www.anthropic.com/rsp-updates">update written to accompany the policy</a> that they plan to release some soon. </p><p>OpenAI does not share any real evaluations that they plan to run on future models, but they do provide some illustrative examples of evaluations that they would run for a hypothetical model:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uGhw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 424w, /__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 848w, /__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uGhw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png" width="928" height="598" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:598,&quot;width&quot;:928,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 424w, /__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 848w, /__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uGhw!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba32eaf-ca9a-4cb4-ad5a-87b2ca17355b_928x598.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>DeepMind acknowledges that the best evals for detecting potentially catastrophic capabilities in models have not yet been defined. They commit to defining a set of &#8216;early warning evaluations&#8217;, and then link to a <a href="https://deepmind.google/discover/blog/an-early-warning-system-for-novel-ai-risks/">framework</a> for achieving this.</p><p>One question that the frameworks do concretely answer is how often they will run evals (whatever those turn out to be!):</p><ul><li><p>OpenAI: every increase of more than <strong>2x effective compute<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></strong>, or if there is a major algorithmic breakthrough</p></li><li><p>Anthropic: every increase of more than <strong>4x effective compute</strong>, or every 6 months of fine-tuning progress</p></li><li><p>Deepmind: every increase of more than <strong>6x effective compute</strong>, or every 3 months of fine-tuning progress</p></li></ul><p>Though companies do not share how they will evaluate forthcoming models, they <em>have</em> publicised how they evaluated each of their most advanced existing models. You can read about this in the system cards for  OpenAI&#8217;s <a href="https://openai.com/index/openai-o1-system-card/">o1</a>, Anthropic&#8217;s <a href="https://www-cdn.anthropic.com/fed9cc193a14b84131812372d8d5857f8f304c52/Model_Card_Claude_3_Addendum.pdf">Claude 3.5 Sonnet</a> and DeepMind&#8217;s <a href="https://storage.googleapis.com/deepmind-media/gemini/gemini_v1_5_report.pdf">Gemini 1.5</a>. </p><h2>Mitigations</h2><p>Once a lab has run evals on a model and assigned it a risk category, it has to carry out a set of corresponding <strong>mitigations</strong> designed to prevent any catastrophes from materialising. All three labs commit to <strong>security standards</strong>, <strong>deployment standards</strong> and <strong>development standards</strong>.</p><h4>Security standards</h4><p><strong>OpenAI</strong></p><p>Security mitigations kick in when an OpenAI model reaches (or is predicted to reach) a risk score of &#8216;high&#8217;. These could include:</p><ul><li><p>Restricting access to models, algorithms or model weights to a designated group</p></li><li><p>Limiting deployment to restricted environments</p></li><li><p>Prioritising information security</p></li></ul><p>Note that OpenAI does not actually specify these safeguards &#8211; but outlines the types of measures that <em>might</em> come under the umbrella of &#8216;asset protection&#8217;</p><p><strong>Anthropic</strong></p><p>Anthropic has so far only defined mitigations for ASL-3. They have said that ASL-4 mitigations will be released soon. The security mitigations for ASL-3 must protect against model theft by a wide range of actors including terrorist groups, state-sponsored programmes and internal employees. They are:</p><ul><li><p>Threat modelling to identify risks</p></li><li><p>Building strong access controls around sensitive assets, for example through physical security or encryption</p></li><li><p>Protecting all the necessary software and hardware for building AI models throughout their lifecycle</p></li><li><p>Proactive monitoring for threats</p></li><li><p>Investing sufficiently in security resources (approx 5-10% of employee labour)</p></li><li><p>Following existing guidance from external organisations such as <a href="https://www.rand.org/pubs/research_reports/RRA2849-1.html?utm_campaign=AI,thread&amp;utm_content=1717021989&amp;utm_medium=rand_social&amp;utm_source=twitter">RAND</a> and the <a href="https://cloudsecurityalliance.org/ai-safety-initiative">Cloud Security Alliance</a>.</p></li></ul><p><strong>DeepMind</strong></p><p>Deepmind specifies the security mitigations that will be needed at 5 different levels of capability:</p><ul><li><p>Level 0 (status quo): Basic industry-standard security practices such as multi-factor authentication</p></li><li><p>Level 1: Restricted access to model weights. Supply chain security to track the origins of model parts.</p></li><li><p>Level 2: Restricting model weights access to everyone in the company except for the core research term.</p></li><li><p>Level 3: Models can only be accessed in secured, tamper-resistant workstations.</p></li><li><p>Level 4: Model weights are generally not accessible to humans.</p></li></ul><h4>Deployment standards</h4><p><strong>OpenAI</strong></p><p>OpenAI commits to not deploying any model with a score of &#8216;high&#8217; or above. If such a model is developed, it cannot be deployed until safeguards have been applied which reduce the risk level to &#8216;medium&#8217;.</p><p><strong>Anthropic</strong></p><p>In order to deploy an ASL-3 model, Anthropic must show that it will be robust to repeated attempts to misuse it. They need to do the following:</p><ul><li><p>Threat modelling: Mapping out all the ways that the model could be misused.</p></li><li><p>Defence in depth: Building a series of protective layers designed to catch misuse attempts.</p></li><li><p>Red teaming<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>: Demonstrating that bad actors are highly unlikely to elicit responses that would increase their ability to do harm relative to other means.</p></li><li><p>Rapid redemption: Show that any compromises of the system would be addressed quickly enough to prevent an actor from actually doing harm.</p></li><li><p>Monitoring: Regularly perform checks to make sure that the system is still operating within an acceptable risk range.</p></li></ul><p>If these standards cannot be met, Anthropic has a series of interim measures designed to reduce risks until they can. These could include temporarily blocking model responses, for example. If interim measures are insufficient, they may employ stronger restrictions such as de-deploying the model or deleting its weights.</p><p><strong>Deepmind</strong></p><p>There are 4 levels of deployment migrations at DeepMind:</p><ul><li><p>Level 0 (status quo): Preventing models from producing harmful outputs with safety finetuning.</p></li><li><p>Level 1: Preventing models from being misused using measures such safety finetuning and periodic red-teaming.</p></li><li><p>Level 2: Similar measures as for levels 0 and 1, but the model can only be deployed once it is demonstrated to have reached a specified level of robustness to misuse.</p></li><li><p>Level 3: Restricted access. These models cannot be deployed until new research has found safe ways to do so.</p></li></ul><h4>Development standards</h4><p><strong>OpenAI</strong></p><p>If OpenAI reaches or predicts that they will reach a model with critical capabilities, they commit to not developing that model further until they can reduce the risk level to &#8216;high&#8217; or lower.</p><p><strong>Anthropic</strong></p><p>Anthropic commits to monitor all models during pre-training to check if they are developing capabilities that require the ASL-3 security standard. If this happens, they will pause training and not resume it until the standards can be implemented. Only the security standard, not the deployment standard, needs to be met to resume training.</p><p><strong>DeepMind</strong></p><p>DeepMind does not specify conditions for halting development, beyond saying that they would do so if &#8220;a model [reaches] evaluation thresholds before mitigations at appropriate levels are ready&#8221;. They commit to employ extra safeguards before continuing development, but do not elaborate much on this.</p><div><hr></div><h2>My thoughts &amp; open questions</h2><p>In this second section, I provide a few of my opinions having read through each of the three frameworks. I don&#8217;t claim to be making any ground-breaking points here! But hopefully, these takeaways can be useful to anyone looking to form their own views.</p><h4>These are not &#8216;plans&#8217;</h4><p>In the title of this post, I called the Preparedness Framework, Responsible Scaling Policy and Frontier Safety Framework &#8216;safety plans&#8217; &#8211; which is how I often hear them colloquially referred to. I&#8217;m hardly the first to make this point, but I think it&#8217;s worth pointing out that <em>this</em> <em>isn&#8217;t actually what they are</em> (at least under what I would consider any sensible definition of &#8216;plan&#8217;).</p><p>A plan for mitigating catastrophic AI risks might look like:</p><ul><li><p>We will run [specific evals] at [predetermined times]</p></li><li><p>If a model scores [%] on eval X, we will implement guardrail Y</p></li><li><p>Here&#8217;s comprehensive evidence that eval X will actually detect the dangerous capability we&#8217;re worried about, and here&#8217;s a detailed justification for why guardrail Y would mitigate against it.</p></li></ul><p>This would follow something like the &#8216;<a href="https://carnegieendowment.org/research/2024/09/if-then-commitments-for-ai-risk-reduction?lang=en">if-then</a>&#8217; commitment structure that others have proposed for AI risk reduction. In their present form, the frameworks do not actually contain many concrete &#8216;ifs&#8217; or &#8216;thens&#8217;. They do not specify which specific evals labs will run, what scores on these evals would be sufficient to trigger a response, and what exactly this response would be (they do insofar as &#8216;move the model into a new risk category&#8217; constitutes a response, but the mitigations that will accompany each risk category are surprisingly ill-defined). A case-in-point from Anthrophic&#8217;s ASL-3 deployment standards:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!MFZO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 424w, /__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 848w, /__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!MFZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png" width="1412" height="194" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:194,&quot;width&quot;:1412,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 424w, /__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 848w, /__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MFZO!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25d13322-5f07-4ded-aef2-9bddba6a0705_1412x194.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>This is what I would describe as a &#8216;plan to make a plan&#8217;. They are yet to specify what empirical evidence would be needed to demonstrate that the system is &#8216;operating within the accepted risk range&#8217; (which they do not define). They do not yet know what review process they will follow or how often.</p><p>I&#8217;m not claiming that any of the labs have explicitly advertised their frameworks as plans. There are acknowledgements in each of them that the science of evaluating and mitigating risks from AI is nascent and that their approach will need to be iterative. OpenAI&#8217;s Preparedness Framework is introduced as a &#8216;living document&#8217;, for example. It&#8217;s also possible that if and when more watertight plans are developed internally, the public will only see high-level versions.</p><p>That said, high-profile figures at all three labs have predicted that catastrophic risks could emerge very soon, and that these risks could be existential. Given this, I can&#8217;t fault those who fiercely criticise labs for continuing to scale their models in the absence of well-justified, publicly auditable plans to prevent this from ending very badly for everyone. Anthropic CEO Dario Amodei, for example, has gone on the record as believing that models could reach ASL-4 capabilities (which would pose catastrophic risks) <a href="https://futurism.com/the-byte/anthropic-ceo-ai-replicate-survive">as early as 2025</a>, but its RSP has yet to even concretely <em>define</em> ASL-4, let alone what mitigations should accompany it.</p><h4>What are acceptable levels of risk?</h4><p>Anthrophic&#8217;s RSP vows to keep model development &#8216;below acceptable levels of risk&#8217;. The other two frameworks espouse a similar idea, albeit without the exact same phrasing. But what exactly does this mean in the context of AI? I suspect a naive reader would assume that developers are confident in keeping risks below similar levels to those considered &#8216;acceptable&#8217; for other high-stakes technologies. For example, the US Nuclear Regulatory Commission <a href="https://www.nrc.gov/docs/ML0909/ML090910608.pdf">aims</a> to keep the probability of damage to the core of a nuclear plant at less than 1 in a million per year of operation.</p><p>But AI is not like these other technologies. Many experts and insiders assign double-digit probabilities to catastrophic outcomes from smarter-than-human AI systems<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. Anthropic championed the RSP model, and has advocated that other companies adopt it. Yet its CEO still believes there is a <a href="https://www.youtube.com/watch?v=GLv62w2G6os">10-25% chance</a> of AI going catastrophically wrong &#8216;on the scale of human civilisation&#8217;. Anthropic research scientist Evan Hubinger, who has written <a href="https://www.lesswrong.com/posts/mcnWZBnbeDz7KKtjJ/rsps-are-pauses-done-right">a popular blog post</a> on the benefits of RSPs, <a href="https://www.lesswrong.com/posts/A9NxPTwbw6r6Awuwt/how-likely-is-deceptive-alignment?commentId=ZQzvvenq6wcP5qSLs">puts this number</a> at 80%.</p><p>To be clear, I&#8217;m not suggesting that Dario or Evan believes such levels of risk are morally acceptable. Perhaps they believe that Anthropic&#8217;s RSP <em>could</em> significantly reduce their own company's risks of causing an AI catastrophe, but that the overall chance of bad outcomes remains high because they do not expect others to adopt similar policies. Or perhaps they do not think there is any conceivable policy that would drive AI risk below double-digit levels, and RSPs are the best we can do.</p><p>If I had to pinpoint exactly what I&#8217;m objecting to here, it might be the way that AI safety frameworks are presented as if they are just like the boring, &#8216;nothing-to-see-here&#8217; safety documents you might find in any other context. In my opinion, the <em>vibe</em> they communicate is &#8216;we&#8217;ve got this&#8217;, despite a between-the-lines reading revealing something more like &#8216;we really don&#8217;t know if we have this&#8217;.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>If it were up to me, I would like to see frameworks like these appropriately caveated with the acknowledgement that AI labs are not doing business-as-usual risk mitigation, and that even if they are perfectly implemented, failure is still a very real possibility.</p><h4>The get-out-of-RSP-free card</h4><p>On page 12 of Anthropic&#8217;s RSP is this easily-missed but highly significant footnote:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oxJm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 424w, /__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 848w, /__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oxJm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png" width="1378" height="186" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:186,&quot;width&quot;:1378,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 424w, /__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 848w, /__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oxJm!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78c57d92-ed8b-44c8-902c-9a0f9a0dd311_1378x186.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>In other words, if the company believes that a competitor is about to reach a particularly dangerous threshold without appropriate safeguards, this would trigger what amounts to an &#8216;ignore all of the above&#8217; clause.</p><p>I&#8217;m not necessarily criticising Anthropic for the inclusion of this clause. It&#8217;s probably safe to assume that either of the other two labs would (and maybe even should) do the same. After all, if you believe that a less safety-conscious actor is about to cause a catastrophe, and there&#8217;s some small chance that in overtaking them you might prevent said catastrophe, then abandoning your safeguards and racing ahead <em>is in fact the most logical move</em>. Arguably Anthropic deserves credit for stating outright what neither OpenAI nor DeepMind do.</p><p>But this clause reveals something self-undermining about the whole enterprise of AI safety frameworks in the context of a race between companies. The safest scaling policy in the world won&#8217;t stop you being outpaced by a less cautious actor (in fact, it probably makes it more likely that you will be) &#8211; at which point you have no choice but to override it.</p><div><hr></div><p>People have poked many more holes in these frameworks than I&#8217;ve covered here (how can we run evaluations while being confident that AI models are not deceiving us? What happens if a discontinuous leap in capabilities happens in between sets of evaluations?) &#8211; but perhaps that&#8217;s a post for another day.</p><p>I&#8217;ve been critical of safety frameworks in their current form, but I should say that I&#8217;m still glad they exist, and I hope they form the basis of something more robust!</p><p></p><div><hr></div><h3>Update &#8211; 05/02/2025</h3><p>DeepMind has just <a href="https://deepmind.google/discover/blog/updating-the-frontier-safety-framework/">released</a> a new version of its Frontier Safety Framework, so I thought it would be worth adding a quick summary of the changes!</p><h4>What&#8217;s new?</h4><p>The headline changes to the FSF are:</p><ol><li><p>There is a more clear division between mitigations thresholds designed specifically for misuse and those for misalignment. We now have two categories of Critical Capability Levels &#8211; <strong>Misuse CCLs</strong> and <strong>Deceptive Alignment CCLs</strong>.</p></li><li><p>&#8220;Autonomy&#8221; has been removed as a tracked capability. The FSF claims that the risks captured by this capability are adequately covered by the Deceptive Alignment CCLs.</p></li><li><p>There has been some refinement of DeepMind&#8217;s procedure for applying deployment mitigations. Every time a model reaches a new CCL, the safety team create a <strong>safety case</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. They have to present this safety case to the &#8220;appropriate corporate governing body&#8221;, and can only go ahead with deployment if the body deems it to be adequate. The safety case will be regularly reviewed post-deployment through red teaming and revisions to threat models (although the FMF doesn&#8217;t say whether a model must be undeployed if the safety case is no longer considered to be adequate).</p></li><li><p>Each of the Misuse CCLs has been explicitly tied to a specific security recommendation. These follow the <a href="https://www.rand.org/pubs/research_briefs/RBA2849-1.html">framework</a> set out in RAND&#8217;s Securing Model Weights report.</p></li><li><p>The most significant addition to the FSF is a new section on deceptive alignment, which DeepMind has described as &#8220;industry-leading&#8221;. There are two CCLs associated with deceptive alignment &#8211; <strong>Instrumental Reasoning Level 1</strong> and <strong>Instrumental Reasoning Level 2</strong>. Instrumental reasoning refers to the ability of the model to develop &#8220;situational awareness&#8221; (ie, understanding that it is an AI model that has been deployed in a certain context) as well as its &#8220;stealth&#8221; (ability to circumvent oversight mechanisms). More on this below.</p></li></ol><h4>My quick takes</h4><p>There are a bunch of things I could say here, but in the interests of brevity, I won&#8217;t go into too much detail. I like that safety cases are required for public deployment (in general, having to make an affirmative case for why a model <em>is</em> safe rather than trying to prove that it <em>isn&#8217;t</em> seems good and commonsensical). The &#8220;Disclosure&#8221; section also caught my eye. This doesn&#8217;t make any specific commitments, but does state that DeepMind will aim to share sensitive information with relevant government authorities or even &#8220;external organisations&#8221; (other labs??) if models pose unmitigated risks to public safety, in order to &#8220;promote shared learning and coordinated risk mitigation&#8221;. The cynic in me doubts that this will lead to much meaningful coordination between major players, but it still gave me a smidgen of encouragement.</p><p>Onto the headline act: DeepMind&#8217;s &#8220;industry-leading&#8221; approach to mitigating risks from deceptive misalignment. It was cool to see this included &#8211; given that a handful of years ago, a section in a document like this dedicated to the possibility of AIs tricking us would have appeared positively kooky. But ultimately, I think what this section hammered home for me is the extent to which no lab has even come close to solving The Real Actual Problem of how in the world we plan to control smarter-than-human AI systems. I think DeepMind&#8217;s Deceptive Alignment CCLs table is a perfect visual illustration of this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tJz2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 424w, /__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 848w, /__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_webp, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tJz2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png" width="1060" height="564" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:564,&quot;width&quot;:1060,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_424, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 424w, /__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_848, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 848w, /__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_1272, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tJz2!, /__u/longerramblings.substack.com/w_1456, /__u/longerramblings.substack.com/c_limit, /__u/longerramblings.substack.com/f_auto, /__u/longerramblings.substack.com/q_auto:good, /__u/longerramblings.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb171be64-129d-4248-8c7f-8e31bbd528cb_1060x564.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The above is a microcosm of what I see as the current alignment plan, which is something like:</p><ol><li><p>Align weaker-than-human AIs: Use a bunch of techniques (RLHF, automated monitoring etc) that kinda-sorta-mostly work, and when they don&#8217;t work it&#8217;s not that big of a deal anyway, because these models are not capable of causing catastrophic harm</p></li><li><p>Align smarter-than-human-AIs: ???</p></li></ol><p>The key point here is that, as far as I can tell, progress on (1) does not indicate much (if any) progress towards (2), in the same way that perfectly rain-proofing your house will do nothing to protect it from a tsunami. DeepMind&#8217;s plan seems to be &#8220;outsmart our AIs until they outsmart us, at which point we&#8217;ll have to think of something else&#8221;. It doesn&#8217;t address the question of whether there could, even in principle, <em>be</em> &#8220;something else&#8221;. Almost by definition, it seems like we&#8217;ll have to settle for <em>trusting</em> a model whose behaviour we can no longer effectively monitor, which is a dicey proposition.</p><p>Safety cases also feature in the deceptive alignment section of the FSF. These safety cases should demonstrate that either a) the model is not capable of undermining human control or b) automated monitoring is sufficient to catch a model&#8217;s malicious behaviour (as determined by &#8220;control evaluations&#8221;). If neither of these conditions can be met, the model is moved up to the Instrumental Reasoning Level 2 CCL. The next question, of course, is what then?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://longerramblings.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! More coming soon...</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Effective compute is a measure designed to account for both computing power and algorithmic progress.</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Red teaming is the practice of trying to elicit dangerous responses from a model in a controlled environment.</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Examples include Paul Christiano, head of the US government&#8217;s AI Safety Institute (<a href="https://www.lesswrong.com/posts/xWMqsvHapP3nwdSW8/my-views-on-doom">~50%</a>), former OpenAI alignment lead Jan Leike (<a href="https://www.youtube.com/watch?v=ZP_N4q5U3eE&amp;t=1h16m">10-90%</a>), Turing Award winner and so-called &#8216;godfather of AI&#8217; Geoffry Hinton (~<a href="https://x.com/liron/status/1803435675527815302">50%</a>).</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>I don&#8217;t think this criticism applies equally to all three frameworks. I actually think that OpenAI&#8217;s framework, which states on the first page that &#8216;the scientific study of catastrophic risks from AI has fallen far short of where we need to be&#8217;, does a better job here than the other two.</p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>In the <a href="https://www.aisi.gov.uk/work/safety-cases-at-aisi">words</a> of the UK AI Safety Institute, a safety case is &#8220;A structured argument, supported by a body of evidence, that provides a compelling, comprehensible, and valid case that a system is safe for a given application in a given environment.&#8221;</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[#17 Fun Theory with Noah Topper]]></title><description><![CDATA[The Fun Theory Sequence is one of Eliezer Yudkowsky's cheerier works, and considers questions such as 'how much fun is there in the universe?', 'are we having fun yet' and 'could we be having more fun?'.]]></description><link>https://longerramblings.substack.com/p/17-fun-theory-with-noah-topper-a66</link><guid isPermaLink="false">https://longerramblings.substack.com/p/17-fun-theory-with-noah-topper-a66</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Fri, 08 Nov 2024 18:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/152358942/e5506139e7890ae307afab851348e903.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>The <a href="https://www.lesswrong.com/s/d3WgHDBAPYYScp5Em">Fun Theory Sequence</a> is one of Eliezer Yudkowsky's cheerier works, and considers questions such as 'how much fun is there in the universe?', 'are we having fun yet' and 'could we be having more fun?'. It tries to answer some of the philosophical quandries we might encounter when envisioning a post-AGI utopia. <br><br>In this episode, I discussed Fun Theory with Noah Topper, who loyal listeners will remember from <a href="https://www.buzzsprout.com/2319950/episodes/14866970">episode 7</a>, in which we tackled EY's equally interesting but less fun essay, <a href="https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities">A List of Lethalities</a>. <br><br>Follow Noah on <a href="https://x.com/NoahTopper">Twitter</a> and check out his <a href="/__u/naivebayes.substack.com/">Substack</a>!<br><br><br></p>]]></content:encoded></item><item><title><![CDATA[#16 John Sherman on the psychological experience of learning about x-risk and AI safety messaging strategies]]></title><description><![CDATA[John Sherman is the host of the For Humanity Podcast, which (much like this one!) aims to explain AI safety to a non-expert audience.]]></description><link>https://longerramblings.substack.com/p/16-john-sherman-on-the-psychological-10f</link><guid isPermaLink="false">https://longerramblings.substack.com/p/16-john-sherman-on-the-psychological-10f</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Wed, 30 Oct 2024 23:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/152358943/b758a2cac9b67208fd323d502a0a172b.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>John Sherman is the host of the <a href="https://www.youtube.com/@ForHumanityPodcast/community">For Humanity Podcast</a>, which (much like this one!) aims to explain AI safety to a non-expert audience. In this episode, we compared our experiences of encountering AI safety arguments for the first time and the psychological experience of being aware of x-risk, as well as what messaging strategies the AI safety community should be using to engage more people. <br><br>Listen &amp; subscribe to the <a href="https://www.youtube.com/@ForHumanityPodcast/videos">For Humanity Podcast</a> on YouTube and follow John on <a href="https://x.com/ForHumanityPod">Twitter</a>!<br><br><br></p>]]></content:encoded></item><item><title><![CDATA[#14 Buck Shlegeris on AI control]]></title><description><![CDATA[Buck Shlegeris is the CEO of Redwood Research, a non-profit working to reduce risks from powerful AI.]]></description><link>https://longerramblings.substack.com/p/14-buck-shlegeris-on-ai-control-f41</link><guid isPermaLink="false">https://longerramblings.substack.com/p/14-buck-shlegeris-on-ai-control-f41</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Wed, 16 Oct 2024 20:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/152358945/690ea85905b7cc9ab4b4bef3962209f5.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Buck Shlegeris is the CEO of <a href="https://www.redwoodresearch.org/">Redwood Research</a>, a non-profit working to reduce risks from powerful AI. We discussed Redwood's research into AI control, why we shouldn't feel confident that witnessing an AI escape attempt would persuade labs to undeploy dangerous models, lessons from the vetoing of SB1047, the importance of lab security and more.&nbsp;<br><br>Posts discussed:</p><ul><li><p><a href="https://www.lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled">The case for ensuring that powerful AIs are controlled</a></p></li><li><p><a href="/__u/redwoodresearch.substack.com/p/would-catching-your-ais-trying-to">Would catching your AIs trying to escape convince AI developers to slow down or undeploy?</a></p></li><li><p><a href="https://www.lesswrong.com/posts/ZLAnH5epD8TmotZHj/you-can-in-fact-bamboozle-an-unaligned-ai-into-sparing-your">You can, in fact, bamboozle an unaligned AI into sparing your life</a></p></li></ul><p><br>Follow Buck on <a href="https://x.com/bshlgrs">Twitter</a> and subscribe to his <a href="/__u/substack.com/@redwoodresearch">Substack</a>!<br><br></p>]]></content:encoded></item><item><title><![CDATA[#13 Aaron Bergman and Max Alexander debate the Very Repugnant Conclusion]]></title><description><![CDATA[In this episode, Aaron Bergman and Max Alexander are back to battle it out for the philosophy crown, while I (attempt to) moderate.]]></description><link>https://longerramblings.substack.com/p/13-aaron-bergman-and-max-alexander-7be</link><guid isPermaLink="false">https://longerramblings.substack.com/p/13-aaron-bergman-and-max-alexander-7be</guid><dc:creator><![CDATA[Sarah]]></dc:creator><pubDate>Sun, 08 Sep 2024 15:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/152358946/ccdd689610a4e3d7ef1bd610f09587da.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>In this episode, Aaron Bergman and Max Alexander are back to battle it out for the philosophy crown, while I (attempt to) moderate. They discuss the Very Repugnant Conclusion, which, in the words of Claude, "posits that a world with a vast population living lives barely worth living could be considered ethically inferior to a world with an even larger population, where most people have extremely high quality lives, but a significant minority endure extreme suffering." Listen to the end to hear my uninformed opinion on who's right.<br><br><a href="https://www.aaronbergman.net/p/my-case-for-suffering-leaning-ethics">Read Aaron's blog post on suffering-focused utilitarianism <br></a><br><a href="https://x.com/AaronBergman18">Follow Aaron on Twitter</a> <br><a href="https://x.com/absurdlymax">Follow Max on Twitter<br></a><a href="https://x.com/littIeramblings">My Twitter<br><br></a><br><br><br><br><br><br><em><br></em><br></p>]]></content:encoded></item></channel></rss>