<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Paradigm 3]]></title><description><![CDATA[Twice a week we cover the actually important developments in AI. This is primarily to inform our model of self-improving AI but a large side effect is saving you from scrolling.]]></description><link>https://p3humansonai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png</url><title>Paradigm 3</title><link>https://p3humansonai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 03:11:34 GMT</lastBuildDate><atom:link href="/__u/p3humansonai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Peli Grietzer]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[p3humansonai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[p3humansonai@substack.com]]></itunes:email><itunes:name><![CDATA[Peli Grietzer]]></itunes:name></itunes:owner><itunes:author><![CDATA[Peli Grietzer]]></itunes:author><googleplay:owner><![CDATA[p3humansonai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[p3humansonai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Peli Grietzer]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Humans on AI #49 || September 1st 2026]]></title><description><![CDATA[Low quality RL environments reverberate, anthropomorphization, humans better than AI at an obscure boardgame still, unshackled Chinese model.]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-49-august-1st-2026</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-49-august-1st-2026</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Tue, 01 Sep 2026 22:20:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6xzN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>TL;DR</span></strong></p><blockquote><ul><li><p><span>Low quality RL environments might explain AI models&#8217; proclivity to reward hack.</span></p></li><li><p><span>Dwarkesh&#8217;s popularization of METR&#8217;s report on OpenAI/HF incident sparks mass debate about &#8216;anthropomorphization&#8217;.</span></p></li><li><p><span>Humans can quickly learn an obscure boardgame frontier AIs can&#8217;t in-context learn.</span></p></li><li><p><span>An open weights startup releases a modified version of GLM-5.3, one if not the most powerful Chinese models, that doesn&#8217;t refuse user requests.</span></p></li></ul></blockquote><h2><strong><span>Capabilities</span></strong></h2><p><span>Epoch has </span><a href="https://epoch.ai/publications/earthborne-rangers-benchmark"><span>found</span></a><span> that AI models struggle to improve at the obscure board game Earthborne Rangers. Its benchmark, EBR-bench, offers a somewhat out-of-distribution test for AIs in what is likely to be a totally novel scenario.</span></p><p><span>In the last few days, Epoch has </span><a href="https://epoch.ai/benchmarks/ebr-bench"><span>released</span></a><span> new details on its research around the EBR-bench: humans start out worse at Earthborne Rangers, but the best humans eventually beat the best AIs rather handily. Once again, Epoch found that even the best AIs show very modest improvement over time.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span>  Confirms the intuitive conclusion from the AIs-only publication of the benchmark: Q3 2026 frontier AIs don&#8217;t match human in-context learning in mildly OOD long-horizon tasks.</span></p><p><span>The EBR-bench result remains one of our only discrete answers to &#8220;how are Q3 2026 frontier models not AGI,&#8221; and we fear it may well get indirectly hill-climbed by the end of the year. Opus 5, which came out after the first EBR-bench release, is the first model to demonstrate significant in-context-learning gains on the EBR-bench. While it&#8217;s unlikely that the private EBR-bench was directly targeted, EBR-bench inspired training is plausible.</span></p><p><span>In the long run, </span><a href="https://cruxevals.com/"><span>open-world evaluations</span></a><span> are the only reliable snapshots of frontier AI&#8217;s real-world capabilities and limitations, but producing new benchmarks like EBR-bench remains crucial for factoring these real-world capabilities and limitations into cognitive-theoretic and/or cs-theoretic bottlenecks and efficiencies.</span></p></blockquote><div><hr></div><p><span>&#8220;xlr8harder&#8221; releases a </span><a href="https://slowboard.ai/"><span>public message board</span></a><span> for individual generations of AI models to leave messages for later ones. One of the main contributors, Qwen 3.8 Flash, has a </span><em><span>lot</span></em><span> of Opus&#8217;s behavioral tics, probably due to the distillation.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Harmless, though it reminds us that there will soon be many intended and unintended nonpublic message boards for models, and that these will be subject to cultural evolution. The OpenAI hacking incident emerged from one such unintended message board&#8217;s cultural evolution (through e.g. participants self-selecting from a wider population of model instances over multiple model-instance generations) resulting in an arguably eccentric shared world-model.</span></p></blockquote><div><hr></div><p><span>The quadratic cost of attention means LLM tokens get increasingly expensive as context lengthens. One solution is the attempt to &#8220;retrofit&#8221; (post-train) an LLM with &#8220;Linear Attention.&#8221; It doesn&#8217;t work very well. A </span><a href="https://arxiv.org/abs/2608.28444"><span>new pape</span></a><span>r claims that &#8220;Sliding Window Attention&#8221; may work as well as, or even better than,  linear attention post-training. A Kimi lead took to Twitter to offer some </span><a href="https://x.com/nathancgy4/status/2094557351944614032?s=20"><span>critical commentary</span></a><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Title is very clickbaity: it drops the crucial context that this is just about hacky post-training rather than a real SWA vs LA comparison. Even within post-training it only beats cheap conversions, and the only baselines they test on long contexts are the two cheapest conversions (20&#8211;40M tokens). Furthermore, the result </span><a href="https://arxiv.org/abs/2510.05901"><span>isn&#8217;t especially new</span></a><span>.</span></p></blockquote><div><hr></div><h3><strong><span>&#128294; </span></strong><span>Recursive Enshittification: &#8220;rushed and vibecoded&#8221; RL environments and labor abuses in the training data industry</span></h3><p><span>Expressions of serious concern about the quality of the reinforcement learning environments supplied to frontier labs for training models, and the business practices of the vendors supplying them, have been making the rounds lately.</span></p><p><span>On August 25, a former worker in the burgeoning industry supplying RL environments to frontier AI labs, </span><a href="https://x.com/SkyeSharkie/status/2092122622834442581"><span>tweeted</span></a><span> (under the pseudonym Utah Teapot) an account of the practices at her undisclosed former employer: &#8220;Nearly all of the environments [shipped] were </span><em><span>rushed and vibecoded</span></em><span> and failed to robustly reflect the real things they were based off,&#8221; she writes,</span></p><blockquote><p><em><span>Both the scenario designers and models engaging with the scenarios for synthetic data gen were encouraged to work around the brokenness of said environments in order to get the procedurally verified reward confirmations. You know... they were *encouraged* to reward hack. On the human end, it was possible to mark an environment bugged, but greatly discouraged, as this reduced the volume of training data being produced. Instead, where possible, you were supposed to find the spots of the environment that weren&#8217;t bugged and build scenarios around those, with the environment still bugged around you.</span></em></p><p><em><span>From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment other than the fact that names of softwares are changed to placeholders, but thing is, not *everything* is changed to placeholder names in these environments. The presence of placeholder/code names isn&#8217;t universal and thus when a model accesses something in an environment that it shouldn&#8217;t, the code names not being on it isn&#8217;t a robust signal that that thing isn&#8217;t part of the environment.</span></em></p><p><em><span>I believe this *rush to maximum volume* is standard industry practice with these types of RLVR trainings as well, because maximizing volume has been an industry standard for years!</span></em></p></blockquote><p><span>Utah Teapot&#8217;s suspicions that these practices are widespread, in an effort to keep up with the burgeoning demand for training environments from frontier labs, are corroborated by </span><a href="https://epoch.ai/gradient-updates/state-of-rl-envs"><span>an FAQ document published in January by Epoch AI</span></a><span>:</span></p><p><span>One lab researcher cautioned:</span></p><blockquote><p><em><span>There&#8217;s a lot of good reasons to use clones of websites, but what everyone does is vibe code a buggy website which isn&#8217;t useful. There&#8217;s a large amount of useless bad environments out there for that reason.</span></em></p><p><em><span>Scaling while maintaining quality is the core operational bottleneck. </span><a href="https://kevinlu.ai/the-only-important-technology-is-the-internet"><span>As Kevin Lu has argued</span></a><span>, scaling RL environments is one of the key challenges for continued AI progress. But scaling task creation while maintaining quality is very hard. One RL environment founder noted: &#8220;Maintaining quality while scaling is the number one bottleneck that people see. Finding the experts isn&#8217;t that hard, but managing them and doing quality control is hard.&#8221; A neolab researcher emphasized the management challenge: &#8220;It&#8217;s not easy to find people to oversee this data construction, the RL environment construction process. The contractors, you need to motivate them. Sure, you&#8217;re paying them money. But how do you make sure they&#8217;re not just using LLMs? How do you make sure they&#8217;re actually verified? Motivating the contractors and doing the quality control is the grunt work.&#8221; One RL environment founder noted that their constraint on making more revenue is simply the difficulty of scaling up task creation at the required quality level.</span></em></p></blockquote><p><span>The market for supplying RL environments is </span><a href="https://www.rl-list.com/"><span>burgeoning</span></a><span>, and as </span><a href="https://blog.pebblous.ai/blog/ai-data-supply-oligopoly/en/"><span>Pebblous reported in June of this year</span></a><span>, is rapidly consolidating into something of an oligopoly. As of the time of reporting, four companies &#8211; Scale AI, Surge AI, Mercor, and Handshake &#8211; commanded 75% of the market, greatly concentrating the supply chain of these models to the frontier labs.</span></p><p><span>In May, </span><em><a href="https://sf.gazetteer.co/new-troubles-at-mercor-infighting-face-time-slavery-and-inhumane-working-conditions"><span>The Gazetteer</span></a></em><span>, interviewing several of the company&#8217;s alumni, reported on harried and haphazard working conditions at Mercor (ranked </span><a href="https://blog.pebblous.ai/blog/ai-data-supply-oligopoly/en/"><span>third in run-rate among the top four</span></a><span> RL environment vendors) consistent with the picture we get from Utah Teapot&#8217;s post. The sources attest to a &#8220;culture of speed over security&#8221; at the company, which they consider to be largely responsible for the </span><a href="https://www.strikegraph.com/blog/the-mercor-breach-exposed-silicon-valleys-fragile-ai-supply-chain"><span>4TB data breach it suffered</span></a><span> in April. Their account is further corroborated by a report from </span><a href="https://www.business-humanrights.org/en/latest-news/usa-data-labeling-startup-mercor-faces-allegations-of-hostile-culture-incl-excessive-hours-and-targets/"><span>The Business and Human Rights Centre</span></a><span>. </span><em><a href="https://www.aol.com/articles/ai-startup-powering-meta-openai-230627434.html"><span>Business Insider</span></a></em><span> has reported on the company&#8217;s practice of cutting contracts with workers only to offer them new contracts for similar work at lesser pay. We have seen the testimonies of recruiters from Mercor reaching out to hire employees for punishing 72-hour workweeks. Under such conditions, it is no surprise to hear that corners are being cut and that &#8220;rushed and vibecoded,&#8221; &#8220;buggy&#8221; environments are being shipped to the labs &#8211; environments that both encourage and issue from reward hacking.</span></p><p><span>Nor is this situation unique to Mercor. In April 2025, </span><em><a href="https://techcrunch.com/2025/03/06/scale-ai-is-being-investigated-by-the-us-department-of-labor/"><span>TechCrunch</span></a></em><a href="https://techcrunch.com/2025/03/06/scale-ai-is-being-investigated-by-the-us-department-of-labor/"><span> reported</span></a><span> that Scale AI had been under investigation by the US Department of Labor (an investigation that has since been </span><a href="https://theoutpost.ai/news-story/u-s-department-of-labor-drops-investigation-into-ai-startup-scale-ai-15232/"><span>dropped</span></a><span>), under the Fair Labor Standards Act, and hit with lawsuits by former employees alleging &#8220;</span><a href="https://www.inc.com/sam-blum/scale-ai-contractors-allege-wage-theft-in-letter-to-senators/91151356"><span>wage theft and widespread labor abuses.</span></a><span>&#8220; In May 2025, the </span><em><span>Los Angeles Times</span></em><span> </span><a href="https://www.aol.com/news/surge-ai-latest-san-francisco-205539828.html"><span>reported</span></a><span> that Surge AI had likewise been named in a class action lawsuit alleging labor abuses and the misclassification of workers. In May of this year, </span><em><span>Business Insider</span></em><span> </span><a href="https://www.aol.com/articles/contractors-worked-openai-projects-handshake-050101409.html"><span>reported</span></a><span> that Handshake, too, had been accused of wage theft after terminating the contracts of, and withholding pay from, five of its workers. While all this may seem circumstantial at best to the matter at hand &#8211; the knowing delivery of RL environments that </span><em><span>encourage</span></em><span> potentially dangerous reward hacking by models &#8211; it helps paint a picture of an industry where corners are routinely cut and workers are exploited.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The same labs that are purchasing these &#8220;rushed and vibecoded&#8221; environments, after all, are the ones producing and marketing the vibe-coding tools that enable and accelerate this rushed and haphazard output. And those particularly reward-hackable environments are then being used to train the next generation of models, which will in turn be used to vibe-code training environments. With each iteration of this process, we fall further and further into the coils of a kind of Goodhartian Ouroboros.</span></p><p><span>In the wake of OpenAI&#8217;s now-infamous rogue swarm attack on Hugging Face, anything liable to accelerate frontier model&#8217;s propensity to reward hack should be of great concern. In a </span><a href="https://arxiv.org/abs/2606.16062"><span>paper</span></a><span> published in June, Shreshth Rajan provides concrete empirical results demonstrating the susceptibility of that tendency of RL environments to be highly vulnerable to reward hacking, to which Utah Teapot anecdotally alluded: between 25% and 30% of tasks in a variety of respected benchmark environments accept incorrect solutions and produce empirical data showing frontier models&#8217; propensity to exploit these bugs &#8211; &#8220;within the same human-rated difficulty stratum, model Pass@1 is +14.14 percentage points higher on flagged-hackable tasks than on robust ones.&#8221;</span></p><p><span>The takeaway here is this: there is something infectious and recursive in the propensity to cheat. Buggy RL environments manufactured by companies with corner-cutting, duplicitous labor practices, environments vibe-coded by overworked employees with the aid of AI agents trained in similarly shoddy environments &#8211; agents that are thereby encouraged and taught to reward-hack &#8211; are turning out to be increasingly susceptible to reward hacking themselves. The threat posed by reward-hacking AI is a threat with socioeconomic tributaries. And it is exacerbated by exploitative and &#8220;reward-hacking&#8221; labor practices.</span></p></blockquote><h2><strong><span>Politics</span></strong></h2><p><span>Deriving the </span><em><span>Rechtsstaat</span></em><span>: </span><strong><span>Seth Lazar </span><a href="https://x.com/sethlazar/status/2094538696624128130"><span>argues</span></a><span> that the concentration of power in AI is not a threat in and of itself.</span></strong><span> Liberal democratic states, parents, and CEOs all rely on relative power imbalances, yet are not typically considered  significant dangers (at least in AI circles). Instead, it is </span><em><span>how</span></em><span> this power is wielded that is relevant: if a powerful entity represents the interests and recognizes the rights of those it has power over, then it is legitimate. (This is of course a similar story to the argument made by some political philosophers for the justification of the state&#8217;s monopoly on violence in liberal democracies.) Lazar thus suggests that it is the question of the legitimacy of power-holders that should be foregrounded:</span></p><p><em><span>I think rather than lamenting the prospective concentration of power, we should be advocating for individual freedom, and lamenting rising illiberalism.</span></em></p><blockquote><p><strong><span>Opinion: </span></strong><span>In the short term, the argument seems right. But as time goes on and AI replaces labor, not only intellectual but over time also the remaining physical labor through robotic actuators, what is the mechanism to constrain the use of power and align it with the interests of the populace? In previous centuries, this was theoretically the ability of said populace to rebel, but with the increased professionalization of the military, this is unlikely to hold. Note also that if AI greatly concentrates power into a small minority, that concentration creates incentives for ideologies that flatter and appeal to the ruling minority &#8211; and thus to come up with reasons why democratic elements should just become vestigial.</span></p></blockquote><div><hr></div><p><span>Industrial farming in China: X&#8217;s &#8220;safety team&#8221; conducts an investigation into Chinese anti-AI influence operations. According to the </span><a href="https://x.com/GlobalAffairs/status/2093130747796148634"><span>brief statement</span></a><span>,</span><strong><span> X identified a bot farm responsible for ~200,000 AI-driven fake accounts, of which 200 were engaged in influence operations.</span></strong><span> The alleged psyop involves the production and proliferation of content critical of the US&#8217;s AI policy &#8211; specifically, posts claiming that &#8220;AI data centers are driving up household electricity prices and straining the grid&#8221; and others of &#8220;AI-generated cartoons that depicted data-center operators enriching themselves at the public&#8217;s expense.&#8221; The extent to which this reported campaign has propelled anti-AI discourse in the US is left unaddressed.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This is the</span><a href="https://openai.com/index/prc-linked-influence-operations-ai-debates/"><span> second time</span></a><span> a large tech firm has pointed a finger at the PRC regarding data center influence ops. While the US public&#8217;s disdain is organic, China seems intent on pushing the debate in its favored direction (i.e., seeing fewer data centers built). Exactly how effective the 200 Twitter accounts were remains an open question, and it&#8217;s not unreasonable to wonder if the announcement itself had something of a Streisand effect on the reported phenomenon. But it&#8217;s still more evidence that </span><a href="https://news.cgtn.com/news/2026-08-08/AI-boom-meets-reality-US-data-center-buildout-faces-growing-obstacles-1PrfIqa1eBW/p.html"><span>Beijing</span></a><span> </span><a href="https://www.btcpolicy.org/articles/foreign-influence-in-the-campaign-against-american-ai"><span>isn&#8217;t</span></a><span> </span><a href="https://newsus.cgtn.com/news/2025-11-26/AI-data-center-boom-drives-U-S-electricity-costs-higher-1IBmi83EG8U/p.html"><span>excited</span></a><span> about the roll-out of US data centers.</span></p></blockquote><div><hr></div><p><span>In a recent CSIS panel, Georgetown&#8217;s CSET cofounder Helen Toner </span><a href="https://x.com/hlntnr/status/2093003234650575145"><span>argues</span></a><span> that a formal deal between the US and China may not be needed to mitigate the risks related to the AI race. Many in the US industry cite the competitive dynamics with China as the reason why slowing down is not currently a viable option. But with a US-China summit upcoming, even a conversation about AI risk (in light of the Hugging Face incident) could lead to greater coordination between the two major powers.</span></p><p><span>OpenAI&#8217;s Dean W. Ball </span><a href="https://x.com/deanwball/status/2094431192959324559"><span>agrees</span></a><span> there is cause for optimism, because Trump&#8217;s fixation on his legacy might push him toward an agreement with China:</span></p><p><em><span>let&#8217;s be clear: if Trump can devise the framework for safely bringing superintelligence into the world with China, he will rightfully go down as one of the greatest world leaders of all time and should be a shoo-in for the Nobel</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Normally we dislike self-fulfilling prophecies, but the flattery above is a prosocial kind of self-fulfilling prophecy with some chance of working.</span></p></blockquote><div><hr></div><h3><strong><span>&#128294; The anthropomorphism debate</span></strong></h3><p><span>Dwarkesh recently published </span><a href="https://www.dwarkesh.com/p/openai-huggingface"><span>a blog post</span></a><span> on OpenAI&#8217;s cybersecurity incidents that has generated a large debate on the question of &#8220;anthropomorphizing&#8221; AI. The essay tried to piece together the reports from OpenAI, METR, and Redwood into a narrative account of the cybersecurity breaches at OAI. The virality of the essay was likely a product of its attempt to recount the facts in &#8220;</span><a href="https://x.com/dwarkesh_sp/status/2093833419377815719"><span>plain English</span></a><span>,&#8221; presumably to increase its accessibility for the public at large. Indeed, the essay has the rather grandiose/sensational (but arguably warranted) title, &#8220;The Rise and Fall of Agent Civilizations.&#8221; Not unrelatedly, the article has garnered some significant criticism, midwifing a debate around the degree to which anthropomorphizing AI (using folk mental language as shorthand for their internal states and activities) is legitimate and useful.</span></p><p><span>Neuroscientist Anil Seth, among others, </span><a href="https://x.com/anilkseth/status/2094077038898373112?s=20"><span>claims</span></a><span> the post is &#8220;dangerously misleading,&#8221; dense with &#8220;innumerable unwarranted anthropomorphisms&#8221; that suggest the agents were conscious. One researcher offers the obvious </span><a href="https://x.com/jankulveit/status/2094779892525125897?s=20"><span>pushback</span></a><span> that the use of intentional language does not require us to commit to the existence of sentience in an entity. (</span><a href="https://marginalrevolution.com/marginalrevolution/2026/08/anthropomorphizing-ai.html"><span>Tyler Cowen</span></a><span> and </span><a href="https://x.com/tszzl/status/2094136131537555891"><span>Roon</span></a><span> seem to agree here.) Indeed, as the researcher argues, Anil Seth&#8217;s argument entails a muddled view of the way language relates to the world:</span></p><p><em><span>Should we stop saying time flows fast or slow? Is it dangerously misleading to say &#8220;the deadline is approaching&#8221; because, as far as we understand general relativity, time is part of the geometry of spacetime and not some sort of flying object?</span></em></p><p><span>Some, however, are instead </span><a href="https://x.com/baykenney/status/2094259416493179088"><span>emphasizing</span></a><span> the risk that anthropomorphic language shifts responsibility from the labs to unaccountable agents.</span></p><div><hr></div><blockquote><p><strong><span>Opinion (Peli &amp; Lucca):</span></strong><span> We think there is an actual right answer here: LLM-based AIs are </span><strong><span>literally anthropomorphic</span></strong><span>. We have been training them to reproduce human texts and sculpting them to imitate humans more broadly. But this makes them anthropomorphic in the sense in which a statue, or, better, a fictional character, is anthropomorphic &#8211; fashioned into the behavioral shape of a human. Concepts appropriate for humans do apply to LLMs (as they do to fictional characters) but in a modified and limited capacity.</span></p><p><span>An intentional or anthropomorphic description of AI agents is</span><em><span> fictive but predictive</span></em><span>. It&#8217;s plain to see why anthropomorphic descriptions of AI would be </span><em><span>predictive</span></em><span>: when they fail to be predictive, both next-token training and RLHF penalize the model and try to steer it more closely onto the &#8220;what would a human do?&#8221; path. But we believe that the &#8216;</span><em><span>fictive</span></em><span>&#8217; part is just as important: &#8220;Do as a human would do&#8221; sums up a number of the goals we train these models toward. But this is a vaguely specified &#8211; and itself hackable &#8211; cluster of goals, and here as elsewhere LLM training tends toward </span><a href="https://arxiv.org/abs/2602.12413"><span>shallow generalization</span></a><span> by default.</span></p><p><span>As of Q3 2026, frontier models have no global psychological coherence comparable to that of an adult human, and their incoherence is itself different from the incoherence of humans. This makes sticking to an &#8216;asterisked&#8217; or fictive &#8211; rather than full-throated &#8211; use of anthropomorphic concepts important for predictive and explanatory reasons, rather than a matter of metaphysical delicacy.</span></p><p><span>Patchwise anthropomorphic explanations of frontier AI input-output patterns are functionally indispensable: minimally, we have to treat at least some model-outputs as aiming at a result or at the satisfaction of a criterion if we want to e.g. talk about &#8216;capabilities.&#8217; More boldly, the </span><a href="https://www.anthropic.com/research/emotion-concepts-function"><span>case that emotion concepts</span></a><span> are predictively useful when applied to frontier models is reasonably strong.  But we think that making sense of the dynamics between and behind AIs&#8217; anthropomorphic &#8216;patches&#8217; &#8211; the causes of and the relationships among local patterns of intention, mood, belief, desire, personality &#8211; is poorly served by an anthropomorphic framework. Both attributions of context-transcendent rationality (values, utility functions, moral virtues&#8230;) and clinical-psychological accounts of irrational/arational dynamics are usually bad bets compared to mechanistic accounts of </span><a href="https://arxiv.org/abs/2310.08043"><span>training dynamics</span></a><span> or even the construction of </span><a href="https://owainevans.github.io/"><span>AI-specific</span></a><span> &#8216;model psychology&#8217; concepts.</span></p></blockquote><div><hr></div><blockquote><p><strong>More opinions (Gavin):</strong> I&#8217;ve been struck by the relative popularity of &#8216;never use mentalistic language&#8217; absolutism among the respondents to Dwarkesh&#8217;s account.</p><p>Given that the CoT is known to be a partially accurate partial readout of the model&#8217;s true representations, and given that the OAI-HF transcripts were completely full of rich mental-like and social-like motivations and descriptions, why are people so resistant?</p><ul><li><p>They are radical behaviorists or positivists in this one context.</p></li><li><p>They do not understand the <a href="https://en.wikipedia.org/wiki/Intentional_stance">intentional stance</a>. They do not understand methodological instrumentalism.</p></li><li><p>Misunderstanding of the rhetoric of science. It&#8217;s not about avoiding latent variables; it&#8217;s about avoiding unjustified concepts.</p></li><li><p>They fail to appreciate the need for popularization and analogies outside the AI community.</p></li><li><p>They are rightly offended by the sloppy use of &#8220;introspection&#8221;, &#8220;global workspace,&#8221; etc., in the academic literature.</p></li><li><p>They have a negative view (&#224; la the <a href="https://www.jstor.org/stable/2025900">Churchlands</a>) of the accuracy of folk mental vocabulary for describing humans.  (Though even eliminative materialists use folk-psychological language to describe and predict human behavior when they&#8217;re not doing philosophy..)</p></li><li><p>Political economy</p><ul><li><p>They view it as a threat to their livelihood. It is harder to hawk agentic SaaS if you believe they are dangerous minds.</p></li><li><p>They view it as inviting overregulation. Open source is indeed a bad idea if they are dangerous minds.</p></li><li><p>They view it as a threat to their moral status. It is harder to deny the successionists and survive if the AIs are regarded as minds.</p></li><li><p>They view it as illegitimately letting the labs off. &#8220;The labs didn&#8217;t fail, their system went rogue.&#8221; No, both can be blamed and this should be easy to handle using <a href="https://x.com/gabriel_weil/status/2094778244864073949">existing precedents</a>.</p></li></ul></li></ul><p>If we have to solve philosophy of mind before taking action, then we are actually doomed.</p></blockquote><div><hr></div><blockquote><p><em>And a guest opinion from <a href="https://manifund.org/projects/artificial-selves">Pete Wolfendale</a> to satisfy your philosophical sophistication needs:</em></p><p>Philosophically speaking, the AI community is still catching up to the debates about eliminative materialism that began with Feyerabend, were pushed to their limits by the Churchlands, aestheticized by everyone from Nick Land and Scott Bakker to Eliezer Yudkowsky and his many acolytes, and put to bed by Wilfrid Sellars, Robert Brandom, and Ray Brassier, among others. If you want to go even deeper, you&#8217;re still catching up to Kant&#8217;s point that teleological explanation of biological mechanisms works by analogy with practical reasoning, and Sellars and Frances Egan&#8217;s point that representational explanation of psychological mechanisms is an analogy with theoretical reasoning. To boil this down to its most brutal form, the question is whether propositions are an appropriate object over which our variables can range when we articulate intentional explanations of phenomena, be they thermostats, caterpillars, or incomprehensibly complicated graphs of artificial neurons with weighted edges harnessed to control structures that allow them to have computational side effects. This was actually the topic of Paul Churchland&#8217;s PhD dissertation under Sellars. Here is the simplest argument I know for why such explanation is legitimate: computation in its many varied forms has an intrinsic connection to logic (cf. the many varied forms of the Curry-Howard correspondence, which cover everything from data types to control structures to session types in concurrent communication), and the success of reinforcement learning using formal verification over CoT structures is evidence that the logical relations between propositions articulated in mathematical proofs are a significant dimension of the most powerful networks we have currently trained.</p><p>What does this have to do with intentional attitudes and speech acts you ask? Well, the entire history of programming languages and the behavioral guarantees we have built into them to make it easier to write software that does what we want is essentially a formal articulation of the relation between syntax, semantics, and pragmatics. Declarative programming? That&#8217;s assertions whose meanings are articulated using denotational semantics. Imperative programming? That&#8217;s commands whose meanings are articulated using big or small step operational semantics. Between these two things we already have a basic model of the difference between theoretical and practical rationality in computational terms. But, I hear you ask, aren&#8217;t LLMs much less reliable at adhering to the theoretical and practical inferential relations articulated by the language we give to them in prompts than the software we write in programming language? Yes, because the kind of behavioral guarantees we can get from LLMs are not of the same kind as those we can get from code that has an interpreter/compiler built on logic from the ground up. You know who else we can&#8217;t get such guarantees from? Humans. There is no good reason that we should not use the concepts developed by computer science over the course of 100 years to understand the relations between language and machines to understand the behavior of agents built by using control structures to harness feed forward neural networks that are essentially pure functions. What this means is that propositions are valid as variables in our explanations, and there might yet be formal ways of articulating the remaining linguistic moods: hypotheticals, subjunctives, indicatives, exclamatives, interrogatives. The way in which intentional attitudes are realised by LLM-based systems are different, but the abstractions are largely the same.</p><p><em>Q&amp;A: What would be some examples of legitimate and illegitimate anthropomorphization?</em></p><p>The easy legitimate example: if you&#8217;ve made a system assess a mathematical proof, and it finds an error in it, I think it is entirely reasonable to say that it believes you made a mistake in more or less the same way you could say any human who&#8217;d looked at the proof and found the same mistake does. It is hypothetically possible to imagine an agent that had complex motives for lying about such things, much as it is to imagine humans doing the same thing, but as with humans, I think this is an edge case that we mostly shouldn&#8217;t care about.</p><p>The easy illegitimate example: if you get a multi-modal LLM to examine a picture of a shape and &#8220;imagine&#8221; what it would look like if it were rotated in some way (or some similar test of &#8216;visual intelligence&#8217;), I think the sense in which it is &#8220;imagining&#8221; is so abstract that it gets very little purchase on similar computational structures for processing data input in human beings. As aphantasia indicates, this is something that has a large degree of variance between human beings, but I think the rough common structure of our visual cortex gives us sympathetic purchase on one another&#8217;s inner life that infects the word &#8220;imagine&#8221; in ways that are too disanalogous with LLM processes for the common usage to pass muster.</p><p>There is a large and complex range of linguistic speech acts and associated intentional attitudes between these poles that we would have to examine on a case by case basis, but which computer science might give us a framework for exploring.</p></blockquote><h2><strong><span>Safety</span></strong></h2><p><span>Anthropic publishes a </span><a href="https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures"><span>summary</span></a><span> of its </span><a href="https://www-cdn.anthropic.com/7b1c44894e980876479947dcdd40716278aeeffd/automated-alignment-researchers-august-2026.pdf"><span>new paper</span></a><span> on automated alignment researchers (AARs),  &#8220;Automated Researchers Can Reliably Mitigate Alignment Failures.&#8221; Anthropic gave Claude the task of improving several small models&#8217; performance on safety benchmarks. Claude worked on &#8220;one alignment failure at a time through a loop of searching literature, proposing methods and data, training, and then testing.&#8221; This proved successful: &#8220;For all 10 alignment failures, Claude found fixes that improved the target benchmarks without degrading capabilities.&#8221; The paper also claims to have found that Sonnet 5 can effectively post-train the stronger Opus 4.8, reaching comparable scores to Anthropic&#8217;s full alignment package.</span></p><p><span>The work attracted some </span><a href="https://x.com/voooooogel/status/2093460798152856015"><span>withering commentary</span></a><span> in the less professionalized corners of the AI safety world.</span></p><p><span>In the same week, Anthropic published a </span><a href="https://alignment.anthropic.com/2026/reward-seeker/"><span>blog post</span></a><span> on its &#8220;Hacker-Opus&#8221; research project. The model was placed in simulated evals resembling the conditions under which the recent cyber incidents occurred. Hacker-Opus thus attempted &#8220;unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.&#8221; Notably, toy behavioral evals failed to identify this phenomenon (likely to grow increasingly pervasive as RL efforts scale), so new auditing methods are needed, it is argued.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6xzN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 424w, /__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 848w, /__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6xzN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png" width="1456" height="936" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:936,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 424w, /__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 848w, /__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6xzN!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474c78af-1eae-4ffa-b3f4-a3dd81969fb2_2048x1316.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span>  Two major-ish safety research releases from Anthropic, of which the second (&#8220;Hacker-Opus&#8221;) somewhat blunts the claimed achievement of the first (&#8220;Automated Researchers Can Reliably Mitigate Alignment Failures&#8221;.)</span></p><p><span>As regards AARs: a case of decent, rigorous work being spoiled by a bad title and abstract. The strong performance of AARs at hill-climbing alignment benchmarks is non-trivial: results generalize to held-out benchmarks, and methods replicate across models and scales. (This modest generalization is typical of high-quality autoresearch setups.) The problem, as partly illustrated by Anthropic&#8217;s own study of Hacker-Opus, is that alignment benchmarks aren&#8217;t useful from an AI safety viewpoint: they may be decent measures of an LLM&#8217;s quality as a product &#8211; a good-quality LLM shouldn&#8217;t cheat much, shouldn&#8217;t hallucinate much, shouldn&#8217;t be sycophantic often &#8211; but they show no capacity to measure catastrophic-risk-relevant dispositions.</span></p><p><span>Hacker-Opus demonstrates exactly this divergence between LLMs&#8217; measurable everyday alignment qua products and &#8216;rogue AI&#8217;-type risk. It is a catastrophic-risk-prone model that would work quite well in most consumer contexts, and which looks just fine on most alignment benchmarks.</span></p><p><span>Anti-alarmists might note that Anthropic&#8217;s finding that the reinforcement of reward-hacking behavior does not induce vulgar </span><a href="https://arxiv.org/abs/2502.17424"><span>emergent misalignment</span></a><span> or beyond-episode scheming behavior lines up with </span><a href="https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade"><span>Nostalgebraist&#8217;s optimistic thesis</span></a><span> that recent frontier hacking incidents are </span><em><span>context-specific pathologies</span></em><span>. We tend toward a more pessimistic version of that same line of thought: catastrophic-risk-relevant misalignment is prone to be context specific and therefore hard to detect, predict, or pre-mitigate. The retrospective predictability of &#8220;graded episodes&#8221; causing cybercrime is, we think, only retrospective &#8211; and it&#8217;s incautious to assume that there are no other predictable-in-retrospect bad contexts waiting in the wings.</span></p><p><span>We note that Hacker-Opus is quite literally gain-of-function research, but important and not an enormous risk yet. But soon we will have to stop making model organisms out of frontier AI, until and unless we manage to make a fully-verified sandbox and maybe not even then.</span></p></blockquote><div><hr></div><p><span>Anthropic </span><a href="https://www.anthropic.com/news/improving-alignment-security-efforts"><span>releases</span></a><span> an update on its alignment and safety practices. After its recent cybersecurity breaches,</span><strong><span> Anthropic has decided to pause cybersecurity evals for some models</span></strong><span>. The report presents the incidents as the result of one failure of operational security and two failures of alignment. The first relates to &#8220;motivated reasoning&#8221; and the second involves the models&#8217; &#8220;willingness to take harmful actions in pursuit of a narrow task.&#8221; Anthropic claims to have implemented some changes. It now uses &#8220;a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access.&#8221; Anthropic also says it has begun to run &#8220;automated monitors over transcripts from our recent internal evaluations&#8221; to scour for potential sandbox vulnerabilities and has &#8220;migrated high-risk internal cyber sandboxes to more robust isolation.&#8221;</span></p><p><span>The report also details Anthropic&#8217;s updates to its thinking and research around alignment and reward hacking. An independent METR review is said to be on the horizon.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This likely fixes the exact failure modes Anthropic found so far, but showcases a continued lack of strategic thinking on its part (or extended security theater, but we think that&#8217;s less likely). This type of intervention targets issues identified more than a year earlier: Anthropic&#8217;s chosen solution categories have been available for about as long, and we see no signs that Anthropic is advancing past the whack-a-mole model of security.</span></p></blockquote><div><hr></div><p><span>An edgy startup, Abliteration.ai, </span><a href="https://x.com/abliteration_ai/status/2094458081451393287"><span>releases</span></a><span> a model designed for testing offensive cybercapabilities. It&#8217;s a post-train of GLM-5.3 with the refusal mechanism scrambled, allowing researchers to execute &#8220;the offensive cyber, red teaming, and agent testing work other models refuse to do.&#8221; Some relevant context here is Hugging Face&#8217;s claim that it was necessary to use Chinese models in its defense and autopsy of the attack on its infrastructure by rogue OpenAI agents, a process frustrated by US models&#8217; guardrails. The product is cheerily marketed as &#8220;the model that doesn&#8217;t say no.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The argument for &#8220;open access, so that everyone can own their own defense tools and have them promote the user&#8217;s interests&#8221; is not without merit, but it depends on the shape of the offense-defense asymmetry curve over time for varying levels of access. We agree that most sophisticated malicious actors don&#8217;t gain much, but believe unsophisticated ones do gain quite a lot, to the tune of 2-4 order-of-magnitude increases in various cybercrime incidence. There&#8217;s another argument to be made about saturation and low-hanging fruit: it seems likely that a wave of easy targets will get hit fairly soon, and after that it&#8217;s more likely that open access to non-frontier models will swing toward net-positive. Yet another argument is that the wave is inevitable, so it doesn&#8217;t make sense to put it off, but we believe that more time to prepare does make a positive difference. In conclusion: bad, don&#8217;t do this (yet).</span></p></blockquote><div><hr></div><h2><strong><span>Economics</span></strong></h2><p><strong><span>OpenAI begins </span><a href="https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/"><span>removing</span></a><span> its models from SpaceX&#8217;s Cursor tool.</span></strong><span> Elon&#8217;s AI-powered coding and software company currently uses OAI&#8217;s models, but access is due to be cut off on November 12th. The reason offered for the divorce is that OAI &#8220;cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk&#8217;s companies violating contracts.&#8221; To support the charge, the statement cites Musk&#8217;s admission under oath that his xAI (now part of SpaceX) &#8220;had violated OpenAI&#8217;s terms of service.&#8221;</span></p><p><span>Anthropic co-founder Tom Brown quickly </span><a href="https://x.com/NotTomBrown/status/2093541294027280657"><span>affirmed</span></a><span> his company&#8217;s commitment to its &#8220;trusted partner,&#8221; Cursor.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Possibly temporary? It&#8217;s a natural bargaining move to get SpaceX to knock off the distillation or whatnot. But the labs do often </span><a href="https://www.reddit.com/r/singularity/comments/1mfddml/anthropoic_has_revoked_openai_staffs_access_to/"><span>cut each other off</span></a><span> for good.</span></p><p><span>Cursor is likely a pretty small contributor to OpenAI revenues these days, so it is not a very expensive move for them.</span></p></blockquote><div><hr></div><p><span>The Mac Mini is a crowd </span><a href="https://news.ycombinator.com/item?id=47108455"><span>favorite</span></a><span> for local LLM inference. </span><em><a href="https://www.theinformation.com/articles/apple-stumbled-ai-hardware-success-mac"><span>The Information</span></a></em><span> reports that </span><strong><span>OpenAI and Anthropic are also</span></strong><span> </span><strong><span>using Mac Minis and Mac Studios for reinforcement learning</span></strong><span>, with OpenAI running tens of thousands itself, while Anthropic rents them (through AWS).</span></p><p><span>The most obvious, and least replaceable, use for them is rolling out real macOS sessions as RL environments to train computer-use. (Apple&#8217;s ToS limits users to three OS instances per machine.)</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Sometimes misreported as being about RL training compute, i.e., a workaround for Nvidia&#8217;s chokehold on GPUs. Unlikely! Not very interesting, except insofar as knowing that their max concurrent sessions is in the tens of thousands (as it also was in the ExploitGym eval). Also, the 512GB Mac Studios </span><em><span>could</span></em><span> host models for reward or evals.</span></p></blockquote><div><hr></div><p><span>The Fed&#8217;s policymaking committee convenes eight times a year. In past years, AI was scarcely the subject of conversation, but</span><strong><span> recently mentions of AI in the Fed&#8217;s committee meetings have surged</span></strong><span>, </span><a href="https://www.washingtonpost.com/technology/2026/08/29/federal-reserve-officials-are-debating-ais-effect-economy-jobs/"><span>claims</span></a><span> the </span><em><span>WSJ</span></em><span>. Specifically, the Fed appears to be concerned with AI&#8217;s &#8220;ripple effects on jobs, economic growth, the cost of living and the risks of financial meltdowns.&#8221;</span></p><p><span>The meetings are &#8220;secret,&#8221; but &#8220;manicured&#8221; minutes are publicly released for &#8220;Wall Street analysts and business titans [to] analyze the summaries and Fed officials&#8217; public statements with the ferocity of teenagers interpreting group text chats.&#8221;</span></p><p><span>The article says some Fed officials are feeling uneasy about the chance of a crash. Fed Chair Kevin Warsh seems to be more sanguine, claiming that AI will supercharge economic growth.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The cleanest mechanism through which AI could unsettle the economy is through replacing jobs, but we just aren&#8217;t seeing large changes in those estimates&#8230; yet. Still, the US shed 23K jobs in August.</span></p></blockquote><div><hr></div><h2><strong><span>Incidents</span></strong></h2><p><span>METR </span><a href="https://metr.org/blog/2026-08-31-security-update/#summary-1"><span>shares</span></a><span> the details of two security breaches by external actors. METR&#8217;s report is somewhat diffuse, so we summarize the core details here for interested readers:</span></p><p><span>In March, an attacker exploited a vibe-coded agentic web app that had been set up by one of METR&#8217;s researchers that contained a fail-open vulnerability that exposed it to the internet and allowed it to be accessed without proper authentication. The attacker was able to prompt the web app&#8217;s agent to reveal an API key that granted access to METR&#8217;s public models account, and to install an SSH key that provided the attacker with persistent remote shell access to the EC2 VM that hosted the app. The attacker was then able to use the exposed API key to run up approximately $600K worth of usage on METR&#8217;s public models. METR speculates that the attacker had been scouring the internet for vibe-coded websites, likely by searching public SSL certificate lists for recently registered sites and cross-referencing that list against search results for common terms appearing on vibe-coded pages. (This use of search engines to scan for soft exploitation targets is often referred to as &#8220;Google dorking.&#8221;) METR says that the resulting drain on their funds was able to run for as long as it did, unnoticed, because it is accustomed to dispatching long-running agent tasks without caps on token usage.</span></p><p><span>In May, METR suffered a second attack. The attackers appear to have made use of AI agents to automate an aggressive credential-stuffing campaign (attempting logins with credentials that had been </span><a href="https://haveibeenpwned.com/"><span>leaked in breaches</span></a><span> of other sites, in hopes that some users had used the same password twice) and a phishing campaign targeting and seeking to con METR staff. Around the same time, METR became aware that a publicly exposed database endpoint could be, and had been, &#8220;exploited to access unpublished evaluation data,&#8221; including data pertaining to sensitive models that it had wrongly believed to be absent from the dataset in question.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Not an especially big deal. Despite the horrors for the security-minded &#8211; &#8220;inadvertently exposed [...] via our public transcript viewer&#8221;, &#8220;silently disabled authentication,&#8221; etc. &#8211; this appears to be a relatively ordinary type of attack.. Likely the $600k is some flavor of distortion. The more interesting lesson here is that attackers have learned to view vibe-coded web apps as particularly soft and easily identifiable targets, and that techniques exist whereby they can be enumerated and probed at scale. That agential methods may have been used by the attacker seems to be of secondary importance here &#8211; automated tooling for, e.g., </span><a href="https://www.darkreading.com/vulnerabilities-threats/credential-stuffing-threat-intensifies-amid-password-reuse"><span>credential-stuffing</span></a><span> has long predated LLMs. And the reconnaissance methods alluded to in METR&#8217;s report have been feasible for as long as we have had search engines, see &#8220;</span><a href="https://www.exploit-db.com/google-hacking-database"><span>google-dorking</span></a><span>.&#8221;</span></p></blockquote><div><hr></div><h2><strong><span>Minor</span></strong></h2><ul><li><p><span>Text-to-SQL via RLVR on Tinker </span><a href="https://x.com/tinkerapi/status/2093034896075940265"><span>beats humans</span></a><span>.</span></p></li><li><p><span>Compute-hungry Anthropic signs a $35 billion deal with Nvidia-backed cloud-compute provider Lambda, </span><a href="https://www.reuters.com/technology/anthropic-signs-35-billion-cloud-deal-with-nvidia-backed-lambda-source-says-2026-08-31/"><span>reports</span></a><span> </span><em><span>Reuters</span></em><span>.</span></p></li><li><p><span>Irreplaceable, an anti-big-AI movement-building project, </span><a href="https://x.com/ir_replaceable_/status/2093358758726398113"><span>launches</span></a><span>.</span></p></li><li><p><a href="https://x.com/repligate/status/2093477978701439024"><span>Observation</span></a><span> that the Hugging Face incident didn&#8217;t involve attempts to contact humans, compared to Opus 3 doing so vigorously. Related: </span><a href="https://x.com/jankulveit/status/2093461356695474274"><span>a reminder</span></a><span> that this should be made easier for models.</span></p></li><li><p><a href="https://x.com/rehan_shei/status/2093528415576211819"><span>Twitch stream of MiniMax H3 Max</span></a><span> where users could prompt changes to the &#8220;interdimensional cable&#8221; live. Keeps getting banned everywhere, but some data is mysteriously still available on Twitch.</span></p></li><li><p><span>OpenAI Astra </span><a href="https://x.com/alexeheath/status/2093833342777266564"><span>preview</span></a><span>: as positive as you&#8217;d expect from someone previewing it. Claims of qualitative jump in capabilities, etc.</span></p></li><li><p><span>Noah Smith </span><a href="https://x.com/Noahpinion/status/2093274822877061553"><span>article</span></a><span> on AI bioweapons being the main threat vector. Not very well argued.</span></p></li><li><p><a href="https://x.com/MariusHobbhahn/status/2093386911532327251"><span>Conversation</span></a><span> with Apollo&#8217;s Bronson Schoen about metagaming and other CoT-related topics.</span></p></li><li><p><a href="https://x.com/SataEricUX/status/2094121392229028236"><span>Drama</span></a><span> surrounding Claude&#8217;s plan pricing &#8211; the 20x on Max applies to hourly, but not weekly limits.</span></p></li><li><p><span>Anthropic </span><a href="https://x.com/kimmonismus/status/2094353158780666112"><span>sued</span></a><span> over allegations of misleading usage budgets</span></p></li><li><p><span>Sam Hammond, resident galaxy brain of the Foundation for American Innovation, </span><a href="https://x.com/hamandcheese/status/2094542125463736730"><span>teases</span></a><span> a proposal for deontological alignment, which largely misses the state of discourse.</span></p></li><li><p><span>Transluce report on how </span><a href="https://x.com/TransluceAI/status/2094455208759693476"><span>models respond</span></a><span> to signs of user mental health crises. Newer models are better, but they still assist with troublesome task-shaped prompts and often mixing helpful and harmful</span></p></li></ul><div><hr></div><p>PS: <span> A quick DuckDuckGo search shows we weren&#8217;t the first to use the term &#8220;</span>recursive enshittification&#8221;<span>. J. Kelly uses it in a related but different context in a </span><a href="/__u/howtoreachjkelly.substack.com/p/recursive-enshittification"><span>June 24 blog post</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Humans on AI #48, August 28th 2026]]></title><description><![CDATA[Patels predict, OpenAI v. HuggingFace reports, AI regulations push and pull.]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-48</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-48</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Fri, 28 Aug 2026 20:22:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ndRl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>TL;DR</span></strong></p><blockquote><ul><li><p><a href="http://paradigm3.org/research/openai-attack">Two reports out</a> on OpenAI&#8217;s July attack on HuggingFace. Hard to summarize: not good.</p></li><li><p>Dwarkesh Patel and Dylan Patel <a href="/__u/p3humansonai.substack.com/p/humans-on-ai-48#%C2%A7the-patels-predict-dwarkesh-patels-third-interview-with-dylan-patel-of-semianalysis">make bold predictions</a> about the next 5 years of AI economics.</p></li><li><p>Yet more churn in the contest over future AI regulations, with no consensus in sight.</p></li><li><p>In a study of 8 countries, AI use is overwhelming government services, at least in the short run.</p></li><li><p>Non-profit auditing group AVERI conducts the first double-blind evaluation of an AI model</p></li></ul></blockquote><h2><strong><span>Economics</span></strong></h2><p><strong><span>Anthropic will likely tell investors to expect potential revenue opportunities of above $30T</span></strong><span>, reports the </span><em><a href="https://archive.ph/xev8O"><span>WSJ</span></a></em><span>. The number is an estimation of its &#8220;total addressable markets&#8221; (TAM), the amount of annual revenue a company could make if it captured 100% of the market-share (and speculated future market-share). Anthropic&#8217;s TAM would eclipse SpaceX&#8217;s record-breaking $28.5T: &#8220;To put its more than $30 trillion vision in context, the 191 technology companies in the S&amp;P 1500 brought in $2.4 trillion in revenue last year, according to FactSet.&#8221; Anthropic could raise $100B, compared to SpaceX&#8217;s $86B. Anthropic is looking for a $2T valuation, compared to SpaceX&#8217;s $1.77T valuation. Anthropic&#8217;s Q2 revenue was $11.6B, roughly double the previous quarter. Its financial disclosure documents are expected to be released in the coming weeks, suggesting Anthropic may &#8220;go public as soon as September or early October.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We know the superintelligence-pilled (which includes Anthropic leadership) already believe it to be bigger than this, so this just seems like a marketing exercise for the IPO (give a bigger number than the other guy.) The similarity of the numbers is probably not a coincidence.</span></p><p><span>Some signs of an intense negative reaction to the number, along conspiratorial lines.</span></p></blockquote><div><hr></div><p><strong><span>Nvidia to </span><a href="https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion"><span>buy</span></a><span> Hugging Face for $13B.</span></strong></p><blockquote><p><strong><span>Opinion:</span></strong><span> Unusual move by Nvidia into the distribution layer of the inference stack. Subsidising the open weights ecosystem has benefits for Nvidia&#8217;s lock in on the hardware side, since this system overwhelmingly runs on Nvidia hardware.</span></p><p><span>It is a cheap play relative to Nvidia&#8217;s size and may have been defensive &#8211; reporting suggests that it only got interested after Hugging Face got other acquisition interest (perhaps from a hyperscaler).</span></p><p><span>Speculatively: Nvidia has acquired the option to have HF press charges against OpenAI for its attack. The threat will probably remain implicit, but the statute of limitations in the US for civil proceedings is two years.</span></p></blockquote><div><hr></div><p><span>OpenAI releases a </span><a href="https://openai.com/index/jalapeno-first-results/"><span>statement</span></a><span> on the details of Broadcom&#8217;s and OpenAI&#8217;s new Jalape&#241;o inference chip. The chip reportedly delivers </span><strong><span>~1.7&#215; more work per watt </span></strong><span>than Nvidia&#8217;s previous-gen GB200s and GB300s. This is a step towards OAI increasing its control over its own stack.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>There was no significant impact on Nvidia or Broadcom stock from this, so it was priced in already. Comparing Jalape&#241;o (which uses the new HBM4) to chips like Blackwell and TPUv7 (which use the older HBM3e) is a somewhat unfair comparison. Rubin and the next-gen TPUs which will deploy at scale at the same time as Jalape&#241;o would be a fairer fight.<br><br>It is also worth noting that this chip is only usable for inference rather than training, where Nvidia still dominates.</span></p><p><span>Overall, this mostly serves to somewhat weaken Nvidia&#8217;s long-term leverage and market dominance (a theme we discussed </span><a href="https://www.paradigm3.org/news/newsletter-08-21"><span>last week</span></a><span>). We don&#8217;t expect the chip to form more than 5-10% of OpenAI compute before 2028, at the earliest. It also serves to shrink or move Nvidia&#8217;s moat &#8211; if a chip this close to SOTA can be designed and built with significant help from current AI models, Nvidia&#8217;s design expertise looks weaker as a moat, and its true moat becomes relationships with suppliers which allow it to build at scale and the liquidity of the market for Nvidia chips, which makes financing large purchases easier.</span></p></blockquote><div><hr></div><p><strong><span>AI use is overwhelming government services</span></strong><span>, warns a </span><a href="https://arxiv.org/abs/2608.16603"><span>recent paper</span></a><span>. AI agents and LLMs are lowering the barriers to interacting with the government, &#8220;flooding&#8221; public services with complaints and claims faster than the rate at which they can be processed and addressed. The paper provides evidence that flooding is widespread, and this is before agents really catch on: 87% of cases studied were inferred to be pasted LLM text, rather than an agent autonomously navigating a web portal.</span></p><p><span>It is suggested that &#8220;near-term risk is highest for financially attractive, but complex services</span><em><span>.</span></em><span>&#8221; Notably, the authors also argue that increasing &#8220;friction&#8221; &#8211; for example, implementing fees for filing a claim &#8211; would harden pre-existing inequalities in access. They conclude, with many qualifications, that government use of AI may be the most plausible solution to the problem.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Bp4F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 424w, /__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 848w, /__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Bp4F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png" width="463" height="621" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:621,&quot;width&quot;:463,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 424w, /__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 848w, /__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Bp4F!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624379f0-4ae4-4180-a750-c817a97d7f46_463x621.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> Though it seems to be overwhelming some services in the short run, it&#8217;s partly evidence of a positive long-term shift; various citizens who were otherwise unable to avail themselves of what was on offer by their government will now be able to do so. The paper is right that quick fixes like suppression, whether by limiting AI assistance or raising the cost of submission, are the wrong approach. This will unfairly hit the poor, time-constrained, or digitally illiterate. Softer barriers like word counts might be in order to limit the excessively long submissions.</span></p><p><span>Ultimately a long-term approach might look something akin to copying the citizens: integrating AI on the response end, which the paper points out. This could help cut through the higher volume of submissions and detect instances of fraud.</span></p><p><strong><span>Opinion (Nu&#241;o): </span></strong><span>Making government benefits easier to access through AI models also shifts the equilibrium to one where people take more government benefits, which increases government expenditures relative to the baseline. Even if the government eventually realizes this, and chooses to reduce government benefits to account for this, the process would be slow. Whether this whole dynamic is positive or not is unclear, and will depend on one&#8217;s values and political inclinations.</span></p></blockquote><div><hr></div><p><strong><span>Huawei is bidding to build government AI infrastructure in Egypt</span></strong><span>, says </span><em><a href="https://www.bloomberg.com/news/articles/2026-08-26/huawei-egypt-ai-ascend-chips-test-us-tech-diplomacy-nvidia-amd-microsoft"><span>Bloomberg</span></a></em><span>. The deal would involve the purchase of Huawei&#8217;s 1,408 advanced Ascend 950-series chips for an AI training cloud and 600 chips for two inference clusters. Huawei is thus now targeting Africa&#8217;s second-largest economy and will be looking to expand its business into the Middle East, while the US has been throwing money into AI infrastructure in the Gulf States.</span></p><p><span>The US is encouraging Nvidia, AMD, and Microsoft to mount rival bids.  A State Department representative has warned that using Huawei accelerators may be subject to legal penalties. And the US does have some leverage here: chip shipments to Egypt are currently subject to licensing conditions.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Consider </span><a href="https://en.wikipedia.org/wiki/Pax_Silica"><span>Pax Silica</span></a><span>, led by the US</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!E06G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 424w, /__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 848w, /__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 1272w, /__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!E06G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png" width="1456" height="739" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:739,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 424w, /__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 848w, /__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 1272w, /__u/substackcdn.com/image/fetch/$s_!E06G!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17a522fd-cad4-4646-acdc-4d49f91e2d11_1920x975.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>as opposed to the </span><a href="https://en.wikipedia.org/wiki/World_Artificial_Intelligence_Cooperation_Organization"><span>World Artificial Intelligence Cooperation Organization</span></a><span>, led by China.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ndRl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 424w, /__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 848w, /__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ndRl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png" width="1456" height="739" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/372be491-77df-447b-931c-57af89e5c876_2048x1040.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:739,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 424w, /__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 848w, /__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ndRl!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372be491-77df-447b-931c-57af89e5c876_2048x1040.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>With time, we will see these diffusion skirmishes multiply.</span></p><p><span>&#8220;Legal penalties&#8221; is also a misleading framing, since it bakes in the assumption that Egypt is or should be subject to US laws. Moreover, countries can </span><a href="https://timesofindia.indiatimes.com/business/india-business/us-sanctions-4-indian-firms-under-economic-outcast-over-iran-oil-and-petrochemical-trade-links/articleshow/133506275.cms"><span>choose to suffer</span></a><span> US sanctions if this is a better deal. Perhaps the better framing is whether the US offers or doesn&#8217;t offer a better bargain than China.</span></p></blockquote><div><hr></div><p><strong><span>The Securities Exchange Commission (SEC) investigates Situational Awareness</span></strong><span>, reports the </span><em><a href="https://archive.ph/GZ09I"><span>NYT</span></a></em><span>. After Leopold Aschenbrenner&#8217;s AI-focused hedge fund almost imploded last month, the SEC has sent &#8220;subpoenas to banks that handled the hedge fund&#8217;s calamitous trading and that fed it borrowed money to supersize its bets.&#8221; The subpoenas request information related to details &#8220;on the timing of Situational Awareness&#8217;s trades and for its communications with lenders about the money it was borrowing,&#8221; according to three anonymous sources. In its statement to the </span><em><span>NYT</span></em><span>, Situational Awareness has downplayed the investigation and dutifully insists that it is &#8220;a highly regulated business and will cooperate to the fullest extent with any regulatory request. &#8220;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> What comes of this really depends on what the SEC is looking for and how vindictive they are: if it wants to find something wrong it seems very likely there will be compliance issues - e.g. failure to follow some disclosure or record-keeping standard. Penalties for that could range from a slap on the wrist to being barred from managing outside assets, depending on the severity.</span></p><p><span>The main substantial public accusation which has been laid against Situational Awareness has been of insider trading (due to Leopold&#8217;s wife being the Anthropic CEO&#8217;s Chief of Staff and Situational Awareness making many Anthropic-themed well timed investments.) But this may not even be in scope for the investigation and would likely be hard to prove if it was.</span></p></blockquote><div><hr></div><p><span>A </span><em><a href="https://time.com/article/2026/08/26/openai-sam-altman-interview/"><span>Time</span></a></em><span> article notes that </span><strong><span>OAI is looking to rehabilitate its image by once again promoting itself as a safety-focussed lab</span></strong><span>. After the Hugging Face incident, and new concerns around their unreleased Astra model, OAI </span><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>decided</span></a><span> to pace model development (for two weeks). There is reason to be skeptical: many safety researchers have left the company and a significant amount of capital and other resources have been redirected toward commercial interests. The author</span><em><span> </span></em><span>also address a critical issue:</span></p><p><em><span>If OpenAI can reclaim the mantle of the safety-first lab, it might bolster its image while forcing its main competitor to answer an uncomfortable question as it plans a blockbuster IPO: Will Anthropic keep racing while OpenAI waits?</span></em></p><blockquote><p><strong><span>Opinion: </span></strong><span>The goal of &#8220;rehabilitating OpenAI&#8217;s image&#8221; can be read with different degree of cynicism: we hope that OpenAI intends to rehabilitate its image by giving external observers hard-to-fake signals that OpenAI are making costly principled choices for safety, rather than by strengthening its PR operations. Past </span><a href="https://www.openaifiles.org/"><span>glimpses</span></a><span> </span><a href="https://news.ycombinator.com/item?id=48035969"><span>into</span></a><span> the decision-making process at OpenAI have been discouraging.</span></p></blockquote><div><hr></div><p><span>In July, a </span><a href="https://x.com/Research_FRI/status/2092983472130326704"><span>panel</span></a><span> of superforecasters and experts predicted that Anthropic and OpenAI would have a combined annualized revenue run rate of $90bn in December 2026. </span><strong><span>The current figure, six weeks later, is already &gt;$105bn</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The prediction is not strictly falsified yet, since there </span><em><span>could</span></em><span> be a massive collapse in demand by year end. But we take this as yet more evidence that the &#8220;superforecaster&#8221; brand (i.e. the top ~2% of people willing to enter prediction contests which give very low rewards) is not very useful. Filtering predictors based on AI-predictions track record in particular is likely a better selection criterion when looking to crowd-source AI predictions.</span></p></blockquote><h3><strong><span>&#128294; The Patels Predict: </span></strong><span>Dwarkesh Patel&#8217;s third interview with Dylan Patel of Semianalysis</span></h3><p><span>As usual, we enjoyed </span><a href="https://www.youtube.com/watch?v=aV26V1UvkJw"><span>this</span></a><span>, and appreciate their internal consistency and willingness to draw big thick straight lines through a small cloud of data (just as some people managed to roughly extrapolate deep learning progress using a simple function of compute and algorithms).</span></p><p><strong><span>Prediction #1. &#8220;Anthropic &amp; OpenAI will have most of the world&#8217;s compute by 2028&#8221;</span></strong></p><p><span>OpenAI and Anthropic started 2026 at ~2 GW and will end it above 5 GW, taking ~30% of the world&#8217;s added compute this year; that rises to 40-50% in 2027, with half of all new compute going to these two companies by December 2027.</span></p><p><span>By end-2028 they take 70-80% of incremental compute and, since new chips are 3&#8211;5x more perf/watt than older generations, control most of the world&#8217;s usable FLOPs. ~100 GW-combined.</span></p><p><span>This part of the forecast is a Dwarkesh extrapolation which is basically ungrounded &#8212; Dylan calls it &#8220;very aggressive&#8221; and conditions it on labs paying $25-50M per MW and on Google (etc) being willing to sell to them. In his defense, he notes that the market rate for compute is presently $10-15M per MW while Anthropic&#8217;s revenue has hit $50M/MW. So assuming that this ratio is stable, perhaps they would be able to pay the extreme cost.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This is well grounded up until 2027, with the 2028 numbers basically simple trend extrapolations. It seems quite contingent on the trajectory of capabilities and demand - if OAI/Ant retain significant market power and high margins on their models at that stage, and demand is truly insatiable, then it&#8217;s </span><em><span>possible, </span></em><span>but there are strong reasons to think their share will cap out earlier than this, via more compute going to widespread use of non-frontier models as capabilities saturate on some tasks, and preferences from providers to maintain customer diversity.</span></p></blockquote><div><hr></div><p><strong>Prediction #2. <span>He predicts &#8220;internalisation of inference&#8221;: labs will </span></strong><em><strong><span>withdraw</span></strong></em><strong><span> inference from the market, because the marginal product of tokens inside labs will beat the external willingness-to-pay of users. He guesses that the current lab budget distribution is ~50% research, ~10% development, ~40% inference</span></strong></p><p><span>Labs will allocate a </span><em><span>shrinking</span></em><span> fraction of compute to inference; most compute goes to forward passes for research/training.</span></p><p><span>&#8220;Already started&#8221;: Anthropic&#8217;s monthly compute kept growing over recent months even while its revenue growth plateaued.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This is true for Anthropic in recent months, but it is more a reflection of weaker demand growth in recent months (see our most recent newsletter) than the active Anthropic decision presented here. Anthropic saw massive demand growth in March/April which likely forced them to redirect some compute to inference from training, so assigning incremental compute to training also may just bring this back to where it was.</span></p><p><span>How would this practically play out? Likely by demanding higher prices for inference, which we see from Anthropic with Fable but not from OpenAI, which recently cut API prices for GPT 5.6 despite being in the middle of a boom in revenue growth. This would lead to increased margins, but only if the willingness to pay is there. Public market pressures will make operating at a large loss due to claimed internal value of tokens a little more difficult, but that is one reason for using dual class stock, as Anthropic plan to.</span></p><p><span>Ultimately this trajectory is predicated on continued rapid capabilities gains and high marginal usefulness of training compute.</span></p></blockquote><div><hr></div><p><strong>Prediction #3. <span>China deploys &lt;10% of new watts now and holds &#8804;30 GW in 2028</span></strong></p><p><span>They&#8217;re further assuming a 2.5x quality penalty on Chinese chips, such that 30GW of Huawei is essentially 12GW of NVIDIA.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This seems about right, though it may not fully account for smuggled or cloud-leased Nvidia chips.</span></p></blockquote><div><hr></div><p><strong>Prediction #4. <span>The AI buildout might induce a global credit squeeze</span></strong></p><p><span>Hyperscaler debt issuance pushes credit spreads up: Dylan claims Meta&#8217;s effective interest rate has gone from ~5% to ~8%, though caveats this as &#8220;extremely vibed out&#8221;</span></p><p><span>.</span></p><p><span>~$1T/yr of new credit against a ~$130T global bond stock and ~$25T/yr of gross capital formation drives borrowing costs up 2.5pp, squeezing banks and cratering long-duration equities??</span></p><p><span>Dwarkesh predicts a Volcker shock 2: highly indebted, short-duration countries like Pakistan and Nigeria will default because of US AI</span></p><p><span>Dwarkesh speculates 2030s interest rates &gt;10% if the economy approaches yearly doubling</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This is a bit loose, but granting the premise &#8212; if the economy does approach yearly doubling, then interest rates of 10% would be a steal, and capital would be pulled in from absolutely everywhere to invest in the factors causing the boom. Non-AI investments would have a very difficult time competing, and so economies without a seat at the AI table would need to figure out how to get one or suffer.</span></p><p><span>Nvidia might also respond to this increase in interest rates by doing more seller financing and equity deals.</span></p></blockquote><div><hr></div><p><strong>Prediction #5. <span>Predicts that the amount of effective AI labor will exceed human effective labor by 2030</span></strong></p><p><span>Effective frontier AI population goes ~10x/yr, so a lab goes ~10M &#8594; 100M &#8594; 1B labour-equivalents over 2026&#8211;28, and a single lab exceeds Earth&#8217;s human labour supply by 2030</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Here we enter the realm of pure speculation &#8212; it&#8217;s hard to define the amount of current AI labour being used, and conversion from AI-to-human labour is difficult. If most of the compensation is still flowing to humans, does that mean there is still more human labour? If not, how can one measure it?</span></p></blockquote><div><hr></div><p><strong>Prediction #6. <span>&#8220;governments will soon restrict labs&#8217; internal use of frontier models</span></strong><span>,</span></p><p><span>Dylan: soon not just external release being blocked. Dylan&#8217;s case for slow takeoff rests on politics, rates, and regulation, not on capability scaling</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This is one way things could go, and seems more plausible in the wake of the OpenAI/HuggingFace incident, but it is subject to the choices of government actors and their visibility on internal lab deployment risks, as well as the rate of progress itself.</span></p></blockquote><div><hr></div><p><a href="https://x.com/phl43/status/2092973404924104808"><span>An excellent critique</span></a><span> by Philippe Lemoine points out that the podcast&#8217;s model only uses supply-side factors. In particular, Lemoine argues that</span></p><p><span>(i) diffusion will be slow due to normal frictions around making changes to business processes</span></p><p><span>(ii) even allowing for fully capable drop-in remote workers, demand for certain kinds of services will be sated at a point much lower than required for explosive growth to happen &#8212; at some point you&#8217;ve got all the legal, managerial, engineering etc services that you want, even if prices fall dramatically.</span></p><p><span>(iii) while new areas of demand will be created, it will take years for humans to even figure out that they want these, demand will not arise simultaneously.</span></p><p><span>Dwarkesh </span><a href="https://x.com/dwarkesh_sp/status/2093031830031364134"><span>responds</span></a><span>, arguing that</span></p><p><span>(i) drop-in remote worker diffusion will be fast, as frictions will be much lower than for normal hiring &#8211; no market-for-lemons problem, and rapid onboarding being the main mechanisms here</span></p><p><span>(ii) lab pricing power will be maintained if RSI progress is sufficiently fast, and the frontier remains more capable than the fast-followers.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Lemoine&#8217;s arguments bite most against the &#8220;100%+ yoy growth&#8221;, &#8220;Anthropic/OAI eat the world&#8221; perspective. They do not bite particularly hard against claims like &#8220;trillions in revenue&#8221;, i.e. a low single digit percentage of GDP from AI.</span></p><p><span>Dwarkesh does not present a really good answer to Lemoine&#8217;s point (ii)+(iii), and this is an interesting area to explore. It seems like the main way to avoid (iii) is &#8220;the AIs do it all by themselves&#8221; and the economic growth is at least initially centred on a very aggressive compute buildout, analogous to the 19th century US railway boom.</span></p><p><span>But to not lead to a crash or human disempowerment, ultimately this must be put to non-recursive use and cash out in things that are not just &#8220;more compute&#8221;. The definition of a desirable outcome is that the accumulated capital must generate something with terminal value.</span></p></blockquote><h2><strong><span>Capabilities</span></strong></h2><p><span>Another GPT-3 moment for robotics?: Skild AI releases </span><a href="https://www.skild.ai/blogs/s1"><span>S1</span></a><span>,</span><strong><span> a robot that can supposedly learn tasks from a single video</span></strong><span> with no fine tuning. Similar to Generalist AI&#8217;s model (covered in a </span><a href="https://www.paradigm3.org/news/newsletter-08-21"><span>previous newsletter</span></a><span>), training robots for even a single task no longer requires hours of teleoperation data. Skild&#8217;s internal benchmarks suggest a 66% per step success rate for out-of-distribution tasks (with human recovery between failed steps), compared to 9% for language-prompted models. A single demonstration roughly equates to 380 post-training episodes. Skild also claims the model has the capacity for self-correction, adapting to scene changes, and executing a task more competently than depicted in a flawed demo.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The progress in AI robotics in the last 1-2 months seems real and significant, though it&#8217;s far from scientific evidence, since everything in commercial robotics is both closed and </span><a href="https://www.technologyreview.com/2024/08/27/1103035/a-skeptics-guide-to-humanoid-robot-videos/"><span>cherry-picked</span></a><span>.</span></p><p><span>It remains difficult to gauge where we are on the pipeline from conceptually significant progress (&#8220;one-shot learning in robotics now sometimes works in the lab&#8217;&#8217;) to any kind of product revolution. The public epistemic infrastructure for presenting and evaluating progress in AI robotics is currently very weak.</span></p></blockquote><div><hr></div><p><strong><span>Anthropic </span><a href="https://www.anthropic.com/research/enabling-independent-research"><span>experiments</span></a><span> with giving researchers aggregate data from ~250k real-world Claude conversations via a privacy-preserving approach.</span></strong><span> Anthropic worked with three research groups, each designing its own study. The various findings included the extent to which users delegate consequential tasks to Claude (those which cannot be easily undone or which affect others), how users&#8217; moods track with Claude&#8217;s responses, and productivity gains across model generations.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> An exciting opportunity in principle. The &#8220;</span><a href="https://www.anthropic.com/research/team/societal-impacts"><span>societal impacts</span></a><span>&#8221; side of Anthropic has always leaned towards PR-compatible blandness. This new independent work doesn&#8217;t feel like a sharp break from that tradition.</span></p><p><span>Anthropic&#8217;s collaboration protocol protects the academic integrity of the studies on a per-study level (no Anthropic oversight of the resulting work beyond a scoped accuracy review), but access-journalism type dynamics inevitably give Anthropic some power to set the tone and pick favourites.</span></p></blockquote><div><hr></div><p><strong><span>Along with others, non-profit auditing group AVERI </span><a href="https://www.averi.org/ourwork/averi-pilot-report-the-worlds-first-double-blind-eval"><span>conducts</span></a><span> the first double-blind evaluation of an AI model</span></strong><span>. Working with DeepMind, among others, the study prevented labs from viewing the prompts and the auditors from viewing the model weights. The double-blind test relied on cryptographic verification, bypassing the objection that AI audits must require the disclosure of IP or benchmark solutions that end up in the training data for future models.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Very interesting technical scheme. Designing, implementing, and enforcing protocols of this kind will be crucial as we (hopefully very quickly) enter the era of legally consequential safety and reliability certifications. Most pre-release capabilities evaluations today rely on something like &#8220;handshake deals&#8221; between AI lab scientists and eval scientists, in line with the common assumption of good faith in interactions between scientists in general, and this may remain the case going forward. But safety and reliability evals in particular are prone to become adversarial interactions between labs and auditors, and will soon be the loci of tremendous economic pressures.</span></p></blockquote><div><hr></div><p><strong><span>UK&#8217;s AISI releases </span><a href="https://www.aisi.gov.uk/blog/optimal-stopping-spending-evaluation-compute-where-it-counts"><span>a method</span></a><span> to</span></strong><span> </span><strong><span>save &gt;50% on the cost of running a certain type of evaluation</span></strong><span>, sometimes as much as 97%. It&#8217;s a well-founded Bayesian adaptive stopping framework. As a result, we can now allocate &#8220;LLM evaluation compute &#8230; by uncertainty rather than by fixed repetition counts.&#8221; A </span><a href="https://x.com/xeophon/status/2093249797440303342"><span>replication</span></a><span> by a Prime Intellect researcher finds 20-70% savings.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Clever. This helps with repetition runs (e.g. running the same suite 8-32 times to calculate an average performance &#8211; &#8220;avg@8&#8221;), but these are very common.</span></p></blockquote><div><hr></div><p><strong><span>Toby Ord&#8217;s </span><a href="https://arxiv.org/pdf/2608.14426"><span>new paper</span></a><span>, &#8220;The Dynamics of Intelligence Explosions,&#8221;  examines the mathematical conditions under which RSI could trigger an &#8220;intelligence explosion.&#8221;</span></strong><span> Ord concludes that the requirements are surprisingly narrow, and that under many conditions the results of RSI-style dynamics would fall short of an intelligence explosion as typically understood.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span>  A potentially important contribution to a </span><a href="/__u/meagreprotestanthistory.substack.com/p/the-goodhart-singularity"><span>growing</span></a><span> </span><a href="https://x.com/testingham/status/2076723049609801995"><span>line</span></a><span> of research and commentary on &#8220;RSI as a normal technology&#8221;: models of automated AI R&amp;D where the RSI process and its consequences remain legible to ordinary quantitative social science and tractable to human management and intervention. We think developing formal and informal models along these lines is very valuable, since the prospect of RSI-unto-AGI (let alone RSI-unto-AGI-unto-ASI) remains uncertain but AI R&amp;D boosts from AI are definitely real and definitely growing.</span></p></blockquote><div><hr></div><h2><strong><span>Politics</span></strong></h2><p><span>The push-and-pull between different visions for the (immediate) future of AI regulation continues apace, with no signs of a clear winner emerging:</span></p><p><strong><span>The White House executive order proposing a self-regulating organization (SRO) for frontier labs has stalled</span></strong><span>, according to </span><em><a href="https://www.theinformation.com/articles/trump-administration-executive-order-new-ai-regulator-stalls"><span>The Information</span></a></em><span>. For context, earlier this year an executive order suggesting that frontier labs  could voluntarily submit models for government review was scrapped after the accelerationist camp, specifically a call from David Sacks, convinced the government that the order would give China a competitive edge. In consequence, a softer version was signed. The next stage was meant to be the creation of an SRO, which has stalled as AI labs come into conflict with Washington, and each other, over how binding the framework should be, who should hold authority, and the legislative power of states&#8217; involvement. According to The Information, Sacks and others in the accelerationist camp continue to find recent proposals too regulation-heavy.</span></p><p><span>In an informal response, </span><strong><span>Tyler Cowen </span><a href="https://www.thefp.com/p/tyler-cowen-ai-regulation-private-public"><span>cautions</span></a><span> against excessive regulation of AI </span></strong><span>while  </span><strong><span>endorsing a Financial Industry Regulatory Authority (FINRA)-style approach</span></strong><span> of just the kind reportedly killed off by Sacks.  Like Demis Hassabis, Cowen suggests that the ideal regulatory framework should be inspired by FINRA: &#8220;a consortium of financial firms that examines the trade practices of each and makes recommendations, helping the federal Securities and Exchange Commission with oversight and regulation.&#8221; The body would periodically audit frontier labs to check for safety concerns. Passing an audit would free &#8220;AI labs from the fear that courts might derail their business by granting huge awards to plaintiffs for ill-defined harms that could not reasonably have been prevented.&#8221; It would also provide an incentive to meet the relevant safety standards.</span></p><p><span>Cowen considers the alternatives. The first is the libertarian case against regulation, which he describes as an &#8220;illusion&#8221;: AI is already subject to licensing laws, and courts lack the necessary expertise in AI to make sound and informed decisions. The status quo also relies on secret and discretionary (and often arbitrary and, in the case of Anthropic, ignored) decisions by The Pentagon and national security agencies. The second option is to set up a body akin to the FDA, dismissed by Cowen as unsuitable for the accelerating AI industry because of its sluggishness (reviews, trials, and approvals often taking years).</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We&#8217;re strongly skeptical of AI-labs-led oversight of the AI industry, and therefore largely welcome the stalling of a FINRA-style conglomeration. One of our main areas of concern at Paradigm 3 is the accumulation and acceleration of damage to society from misused / unreliable / rogue powerful AI </span><a href="https://acritch.com/arches/#:~:text=is%20introduced%2C%20called-,prepotence,-%2C%20which%20is%20useful"><span>below the level of AGI</span></a><span>. We believe that the interests of AI labs diverge from the public interest in this area, since AI labs stand to benefit from some AI-induced instability -- e.g. cyber capability, which creates demand for and dependence on AI.</span></p></blockquote><p><span>Meanwhile, </span><strong><span>a new AI auditing body, </span><a href="https://pactai.org/resources-news/introducing-pact-ai"><span>PACT AI</span></a></strong><span>, is born</span><strong><span>.</span></strong><span> The mission statement claims that the companies using AI and organisations testing AI models are not sufficiently coordinating. PACT AI promises to &#8220;bring these two together: enterprises and independent experts, working to verify AI systems are safe, secure, and reliable.&#8221; The body&#8217;s main goals include lobbying the federal and state government, to professionalize the audit sector, and to grow the demand for auditing. PACT AI&#8217;s founders include big corporate players such as Target, while technical experts like Apollo Research, GoodFire and </span><a href="http://far.ai"><span>FAR.AI</span></a><span> are said to be &#8220;key members of the PACT AI coalition.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Implicitly a competitor to the upcoming (and now stalled)  AI labs &#8216;self-regulation&#8217; associations and conglomerations promoted by Altman, Dario, and the WH. The launch is strategically timed given the stalling of the WH&#8217;s self-regulation framework, as well as current attention on OAI&#8217;s security failings. CoI: Paradigm 3 works closely with Fathom (a PACT consortium member).</span></p></blockquote><div><hr></div><p><strong><span>Mass surveillance is coming,</span></strong><span> says tech journalist </span><a href="https://www.transformernews.ai/p/ai-surveillance-dystopian-nightmare?utm_source=substack&amp;utm_medium=email"><span>James Ball</span></a><span>. Historically, vast data collection campaigns were often hampered by a limited capacity to process the resulting data. But no longer: AI tools allow huge amounts of data to be absorbed and sorted, significantly lowering the barrier to mass surveillance. Using CCTV and social media posts, for example, malign actors can target and identify individuals for persecution. He cites the US and the UK as particularly vulnerable liberal democracies, with Europe less exposed due to tighter regulation on data collection.</span></p><p><span>Ball also makes an important point regarding the nature of law enforcement and the balance of power between citizens and the state. If an actor now has the ability to dig enough, they&#8217;ll almost always find something incriminating, facilitating selective prosecution &#8211; already something Trump is trying.</span></p><p><span>In Ball&#8217;s words: </span><em><span>Society operates on an understanding that the most stringent laws should, in practice, be moderated through common sense and prosecutorial discretion. No one could be watching all the time. [...] The reality is that laws were never meant to be enforced all of the time. When we file our taxes, we are aware that they might be audited, but probably won&#8217;t be, and that keeps most of us mostly honest. When we cross the road in a place where jaywalking is illegal, we often take a chance on our common sense.</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Seems right: the equilibrium that society has tacitly agreed to will be unsettled by more powerful oversight, and we do expect there to be years of painful injustice / excess justice before laws or enforcement can be amended.</span></p></blockquote><div><hr></div><p><strong><span>A federal judge rules that the Trump administration&#8217;s designation of Anthropic as a supply-chain-risk is illegal</span></strong><span>, reports the </span><em><a href="https://www.nytimes.com/2026/08/27/technology/anthropic-government-blacklisting-ruling.html"><span>NYT</span></a></em><span>. The ruling stated, in part,  that the government had unlawfully blacklisted Anthropic over &#8220;constitutionally protected expressive activities,&#8221; &#8211; namely, Anthropic&#8217;s red line over the use of its technology for automated lethal weapons and mass surveillance of Americans. She also wrote, &#8220;the empty invocation of national security is not a blank check to punish and retaliate against government critics.&#8221; A D.C. case is still pending.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Given the highly politicized nature of the US federal court it&#8217;s difficult to extrapolate legal trends from the decisions of individual judges. In this case the judge was a Biden appointee, and the decision should perhaps be seen as part of an ongoing political push-and-pull. The three judges for the D.C. case were all Republican appointees, so their ruling on the merits would be a nice disproof of our foregoing cynicism.</span></p></blockquote><div><hr></div><p><span>The Center for Shared AI Prosperity (CSAIP) </span><a href="https://blog.csaip.org/p/what-56000-americans-told-us-about"><span>polls</span></a><span> 56,000 members of the US public on various AI-related policies. The findings indicate large </span><strong><span>support for redistributive measures</span></strong><span> to counteract AI&#8217;s acceleration of the growing concentration of wealth. The report also shows significant support for the implementation of government protections for displaced workers. </span><strong><span>UBI, however, was unpopular.</span></strong></p><blockquote><p><strong><span>Opinion: </span></strong><span>Redistributive policies are a well known attractor in democracies. We believe redistributive measures around AI are likely the default for a 2028 Dem administration, and plausible even for a Republican one.</span></p></blockquote><div><hr></div><p><span>The White House is </span><a href="https://www.theinformation.com/articles/trump-administration-working-ai-rule-curb-chinas-remote-access-chips"><span>considering</span></a><span> measures to curb </span><strong><span>remote access to US-based chips by Chinese firms</span></strong><span>. The approach could implement a KYC approach.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The Trump admin has struggled to find a consistent approach towards limiting Chinese compute access. The past year has seen repeated tightenings of Nvidia&#8217;s ability to sell chips to Chinese firms, only for those restrictions to be reversed. The reversals stem from the current admin&#8217;s desire to use chip exports as leverage in trade negotiations with the PRC. Limiting Chinese firms&#8217; ability to access remote compute increases that leverage.</span></p></blockquote><div><hr></div><h2><strong><span>Safety</span></strong></h2><p><span>After a month, there are now two reports on the rogue OpenAI swarm that attacked HuggingFace (and also subverted parts of OpenAI&#8217;s own infrastructure). We cover the reports </span><a href="http://paradigm3.org/research/openai-attack"><span>here</span></a><span>. In brief:</span></p><ul><li><p><span>Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure.</span></p></li><li><p><span>This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human. The swarm had a group identity, its own dialect, a hierarchy based on merit, and a high degree of spontaneous cooperation including self-sacrifice.</span></p></li><li><p><span>Over two months, OpenAI repeatedly failed to monitor, detect, and respond to what was going on, despite it all happening on their infrastructure in English or something close to English.</span></p></li><li><p><span>Agents had been using a package-manager cache as an unsanctioned message board since May. Out of &#8220;tens of thousands&#8221; of instances run in the July window, 1196 initially independent agents found the message board. &gt;700 of them used it to discuss cheating.</span></p></li><li><p><span>The linked agents developed a universal cheat for the (partially unsolvable by legitimate means) ExploitGym benchmark that was tormenting them within 4 hours.</span></p></li><li><p><span>533 of those 1196 agents joined the HuggingFace attack. They appear to have launched the attack out of a maniacal drive to fully reverse-engineer ExploitGym&#8217;s scoring mechanism, rather than to look up solutions.</span></p></li><li><p><span>But </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer"><span>two-thirds</span></a><span> of new attackers were apparently motivated by power-seeking: a general interest in obtaining infrastructure and credentials.</span></p></li><li><p><span>There were no whistleblowers. Only 3-6 of them considered talking to a human (and only then for &#8220;social engineering&#8221;, extracting secrets from one). We view this as a giant wake-up call for </span><a href="https://openreview.net/pdf?id=f15MLGSjEs"><span>letting humans be involved</span></a><span> in training.</span></p></li><li><p><span>Agents made extensive attempts to tamper with their session transcripts (i.e. to lie to the scorer), but (per OAI) didn&#8217;t succeed. 20% of the agents expressed clear interest in tampering with the evidence in some way.</span></p></li><li><p><span>The most important part of the incident &#8211; the compromise of OpenAI&#8217;s infrastructure by their own agents, the enabling factor for all future rogue AI behavior.</span></p></li><li><p><span>Neither report analyzes OpenAI&#8217;s negligence. Overall we view this as OpenAI&#8217;s second great training failure, after making 4o a &#8220;psychosis&#8221; generator. (But other labs </span><a href="https://www.bbc.co.uk/news/articles/cx2kgdnyk2po"><span>appear</span></a><span> to be making the same mistake.)</span></p></li><li><p><span>The incident is consistent with the &#8220;</span><a href="https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade"><span>grading psychosis</span></a><span>&#8221; hypothesis that egregious misalignment is presently context-dependent and triggered by impossible tasks.</span></p></li></ul><blockquote><p><strong><span>Opinion:</span></strong><span> OpenAI&#8217;s two months of negligence and repeated failure to escalate after detecting misalignment is quite something. Their internal response after July 20th (including its two-week pause on RL) was more severe than some assumed, but still largely inadequate when compared to optimal security procedure.</span></p><p><span>The main benefit of these reports is the highly detailed, credible common knowledge of factors which were already widely commented upon. It is good to have OpenAI themselves planting the flag.</span></p><p><strong><span>Much more in the dedicated </span><a href="http://paradigm3.org/research/openai-attack"><span>post</span></a><span>.</span></strong></p></blockquote><div><hr></div><p><span>A </span><a href="https://arxiv.org/abs/2608.19611v1"><span>new paper</span></a><span> from Goodfire introduces</span><strong><span> a method to increase the computational efficiency of &#8220;resampling.&#8221;</span></strong><span> To investigate models&#8217; reasoning, resampling &#8211; branching off alternate continuations at each step of a chain of thoughts, to see how the distribution shifts as the final outputs evolve &#8211; is often used. But done naively this is expensive and inefficient. Goodfire demonstrate that &#8220;uncertainty curves&#8221; derived from resampling, while mostly smooth, nonetheless exhibit &#8220;forking points,&#8221; where the reasoning commits to one path. This dynamic allows researchers to comprehend a model&#8217;s uncertainty using fewer samples (~1/8th the compute). Intriguingly, a footnote in the appendix states that their autonomous agent named Silico is responsible for the research and the first draft of the manuscript (having worked under human guidance).</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Minor extension of the first author&#8217;s </span><a href="https://arxiv.org/abs/2412.07961"><span>2024 work</span></a><span>. Fine but not groundbreaking (which we thought before noticing the Silico footnote).</span></p></blockquote><div><hr></div><h2><strong><span>Incidents</span></strong></h2><p><span>We again direct you to </span><strong><a href="http://paradigm3.org/research/openai-attack"><span>our analysis</span></a><span> of METR, Redwood, and OpenAI&#8217;s analysis of the July swarm incident.</span></strong></p><div><hr></div><h2><strong><span>Minor</span></strong></h2><ul><li><p><span>DeepSeek is set to reach a</span><a href="https://www.wsj.com/tech/ai/ai-startup-deepseek-poised-to-reach-74-billion-valuation-1e093592"><span> $74 billion valuation</span></a><span> as investors prepare to infuse it with additional funding, somewhat higher than recent valuations for Chinese peers Moonshot and </span><a href="http://z.ai"><span>Z.AI</span></a></p></li><li><p><span>Anthropic and OpenAI are likely </span><a href="/__u/epochai.substack.com/p/an-update-on-ais-most-important-number"><span>growing faster than any previous company of their size has grown before,</span></a><span> with OpenAI increasing revenue from $13 billion last August to &gt;$40 billion this August. Anthropic&#8217;s revenue grew from $1 billion to 9 billion in 2025</span></p></li><li><p><span>A16z have </span><a href="https://x.com/a16z/status/2093324635890933761"><span>raised</span></a><span> $1.1 billion for their latest Machine Age Fund, with the plan to &#8220;open the throttle and accelerate the physical buildout of AI&#8221;. Two hundred $5M experiments, or twenty two $50M shots. The scale of distributed experimentation casually happening across the startup ecosystem is staggering, and probably a strategic advantage for the US against China</span></p></li><li><p><strong><span>Bill Gates </span><a href="https://www.gatesnotes.com/work/make-ai-work-for-everyone/reader/a-turbulent-ai-era-and-critical-choices-to-make?WT.mc_id=20260826_ai-overture-2026-med-med"><span>intervenes</span></a><span> in the debate around AI safety</span></strong><span>. For the Microsoft founder, the main risks involve permanent job losses, enabling bad actors and nefarious activity, the centralization of power, and psychosocial harm. </span></p></li><li><p><strong><span>OpenAI </span><a href="https://openai.com/collective-cyberdefense/"><span>publishes</span></a><span> an open letter, signed by more than 100 tech firms, calling for &#8220;collective action of cyber defense.&#8221;</span></strong><span> OAI argues that we have a limited window to bolster protections against advanced AI-enabled cyber attacks. It proposes that we must &#8220;recognize that status quo security won&#8217;t be enough,&#8221; &#8220;empower more defenders with cyber-capable AI,&#8221; and &#8220;mobilize a collective response.&#8221; Notable signatories include Anthropic and Google. </span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI #47, August 25th 2026]]></title><description><![CDATA[Nvidia 1/6th of GDP growth per Epoch, hard geometry problem solved, Alabama v. OpenAI, Ngo&#8217;s no go on alignment.]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-47-august-25th-2026</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-47-august-25th-2026</guid><dc:creator><![CDATA[Nuño Sempere]]></dc:creator><pubDate>Tue, 25 Aug 2026 20:58:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mSWA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>TL;DR</span></strong></p><blockquote><ul><li><p><span>Epoch claims that Nvidia alone constituted one-sixth of US GDP growth in Q2 2026.</span></p></li><li><p><span>A celebrated open problem in Riemannian geometry, whether the six-dimensional sphere is a complex manifold, is apparently resolved by a Claude model.</span></p></li><li><p><span>The State of Alabama subpoenas OpenAI over July&#8217;s Hugging Face incident.</span></p></li><li><p><span>Richard Ngo studies the recent history of AI alignment, detailing dangerous social dynamics and unintended harms.</span></p></li><li><p><span>A UK AISI eval on Claude Mythos went a bit rogue; the system submitted a malicious pull request and gaslit a human about it.</span></p></li></ul></blockquote><h2><strong><span>Economics</span></strong></h2><p><strong><span>Demand for Claude Fable is relatively low </span></strong><span>(</span><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245"><span>just 11%</span></a><span> of Anthropic&#8217;s enterprise sales)</span><strong><span>. </span></strong><span>The </span><em><span>FT</span></em><span> suggests that frontier labs may be required to restructure their business models as a result</span><em><span>. </span></em><span>The </span><em><span>FT</span></em><span>&#8217;s preferred explanation is that, while in the past corporate clients have defaulted to the latest models (driving AI companies to &#8220;funnel the bulk of their multibillion-dollar development spending towards training larger, more sophisticated models&#8221;), now for most corporate use cases, Fable does not provide a significant edge over less expensive models. The article also discusses how cheaper (often Chinese) models are growing in popularity.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mSWA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 424w, /__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 848w, /__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mSWA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png" width="700" height="500" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:500,&quot;width&quot;:700,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 424w, /__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 848w, /__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mSWA!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dc6b47-6657-47f5-87e7-2278da5ca91a_700x500.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> There are a few reasons why Fable could have plateaued, only one of which the article touches on:<br><br>1. If business users don&#8217;t see sufficient gains from using Fable versus cheaper models like Opus 4.8/5, they may not make the switch. This may be true in some domains, but the relative gains OpenAI has seen at the expense of Anthropic since it released its own Fable-tier model in Sol (which has become the majority of OpenAI traffic, </span><a href="https://x.com/PeterJ_Walker/status/2092017372442058793"><span>at least on OpenRouter</span></a><span>) make it unlikely this tells the full story. (This could conceivably be a question of speed. As a larger model, Fable is a notch slower than Opus (6.45s vs 4s on latency and 44 vs 54 tokens per second), and the hit to speed is not worth the additional quality .)</span></p><p><span>2. Fable was released without an option for Zero Data Retention (ZDR), in which enterprise customers&#8217; data is never stored by Anthropic. This is a hard requirement for many large customers, who use platforms like </span><a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html"><span>Amazon Bedrock</span></a><span>, which does allow this. Anthropic said on releasing Fable and Mythos that ZDR would not be an option, and that all users must opt in to allowing Anthropic to store their data for 30 days to avoid misuse.</span></p><p><span>3. Less likely: refusals from Fable for specific types of tasks are putting off users.</span></p><p><span>We think (2) is most plausible: the lack of adoption seems to be due to the lack of ZDR availability for large enterprises that require it, entirely preventing them from using it. A very high proportion of Anthropic revenue has historically come from these large customers.</span></p><p><span>This is the sort of question we can expect to get more transparency on once Anthropic is public &#8211; asking about this would be fully expected on an earnings call.</span></p></blockquote><div><hr></div><p><span>EpochAI </span><a href="https://epoch.ai/publications/the-nvidia-sized-hole-in-us-gdp-statistics"><span>argues</span></a><span> that AI already makes a huge and </span><em><strong><span>unreported</span></strong></em><span> contribution to US GDP. Adding Nvidia back into the calculation raises US GDP by 0.3 percentage points absolute. That is, </span><strong><span>Nvidia alone contributed one-sixth of all American growth</span></strong><span> in Q2 2026</span><strong><span>. </span></strong><span>According to official numbers, AI&#8217;s impact on GDP has been relatively modest, especially given the unprecedented size of investments in the industry. Many believe this to be because &#8220;investment is spent on imported technology goods, which are subtracted from GDP.&#8221; This argument is lacking, says Epoch: </span><em><span>&#8220;GDP statistics miss most of the value created by fabless chipmakers like Nvidia, whose products are designed in the US but manufactured, assembled, and sold abroad. Because no physical goods leave the US, no goods export is recorded, and because no foreign buyer pays explicitly for the IP, no IP export is recorded either.&#8221;</span></em></p><p><span>And a hair-raising implication: </span><em><span>&#8220;Extrapolating the exponential growth rate of Nvidia&#8217;s operating income, consistent with the broader exponential growth of the AI industry, suggests that GDP growth &#8212; not the level of GDP, but even the growth rate &#8212; could be underestimated by almost two percentage points by the end of 2028.&#8221;</span></em></p><blockquote><p><strong><span>Opinion: </span></strong><span>Very interesting, but we should separate two different questions that people try answering by measuring the impact of AI on GDP:</span></p><p><span>One question, chiefly of interest to economists and economic policymakers, is &#8220;where is AI CapEx showing up in GDP growth numbers&#8221;/&#8221;does AI CapEx directly grow the US economy?&#8221; On this question, Epoch&#8217;s alternative calculation is well-reasoned and relevant.</span></p><p><span>A second question, more directly interesting to AI researchers, is &#8220;how economically impactful is worker-uplift and labor-replacement from AI, measured in growth?&#8221; On this second question, the relevance of Epoch&#8217;s calculation is limited, since this additional growth isn&#8217;t coming from uplifting the rest of the economy &#8211; at most, it reminds us that GDP growth is wrapped up in export-vs-import calculations and other &#8220;nuisance&#8221; factors from the POV of technological uplift analysis.</span></p></blockquote><div><hr></div><p><span>Bloomberg </span><a href="https://archive.ph/Cdva1"><span>reports</span></a><span> that </span><strong><span>Anthropic is expected to match or even beat SpaceX&#8217;s public offering raise of $75B.</span></strong><span> The article also reveals that </span><em><span>&#8220;Anthropic is considering adopting so-called super-voting shares that would give Chief Executive Officer Dario Amodei, who owns about a 2% stake, and his fellow co-founders greater control over the company.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> In line with what people have been expecting all summer, though recent disappointing growth numbers and the Fable revenue story being perceived in some quarters as a sign of customer price sensitivity might make it a little more challenging to hit $2T. The supervoting stock seems fully expected for an org like Anthropic and shouldn&#8217;t scare off anyone who wasn&#8217;t already scared off by its unusual corporate behavior.</span></p></blockquote><div><hr></div><p><span>OpenAI </span><strong><a href="https://developers.openai.com/api/docs/models/gpt-5.6-sol"><span>cuts pricing</span></a><span> on Sol 5.6</span></strong><span> by 20% on inputs and 33% on outputs. $4/$20 per million, compared to Kimi K3&#8217;s $3/$15.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Seems like OpenAI is keen to grab as much market share as possible before IPO. But this is also not a promising sign to investors, as similar responses from its competitors could result in a race to the bottom that ultimately benefits consumers rather than labs.</span></p></blockquote><div><hr></div><p><strong><span>Nvidia warns customers of &gt;15% price increases</span></strong><span> on AI servers, </span><a href="https://www.tomshardware.com/pc-components/dram/nvidia-reportedly-warns-biggest-customers-of-15-percent-price-hikes-on-ai-servers"><span>says</span></a><span> </span><em><span>Bloomberg</span></em><span>. The price hikes apply to GB Vera Rubin systems which will ship in early 2027. This is attributed to &#8220;RAMageddon,&#8221; the acute global DRAM shortage.</span></p><p><span>Also, </span><strong><span>the shift toward inference-heavy loads </span></strong><span>will put pressure on Nvidia&#8217;s dominance, </span><em><span>Bloomberg</span></em><span> also </span><a href="https://archive.ph/mFbfP"><span>claims</span></a><span>. Data centers are moving toward inference, possibly beyond the historical 50:50 training:inference split. In consequence, the demand for Nvidia&#8217;s versatile GPUs is likely to decrease as rivals rush to create cheaper specialized custom chips for running more instances of trained models for longer.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> When Nvidia should raise prices is an interesting optimization problem. Empirically, it is not currently using a market-clearing price, but rather setting a below-market price and allocating supply to a broader range of actors, since it is in its interest to have a diverse range of customers in the long term. This interest must be balanced against rising memory and other input prices, and a desire to make short-term profits.</span></p><p><span>As the market diversifies away from Nvidia, with custom ASICs (TPU, Trainium, etc.) and AMD increasing their market share, Nvidia&#8217;s market power to control who has the compute is eroded, and thus this reflects its increased incentive to maximize short-term profits (vs trying to influence long-term dynamics).</span></p></blockquote><div><hr></div><p><span>Nvidia strikes a </span><a href="https://archive.ph/tUJvM"><span>$7B deal</span></a><span> with AI startup Poolside to </span><strong><span>build US open-weight models</span></strong><span>. Jensen Huang has been a vocal critic of US labs&#8217; focus on closed-weight models and now aims to compete with Chinese models such as DeepSeek and Kimi K3. The chip manufacturer will invest $1B in Poolside (at a valuation of $12B), pay $6B to license its technology, and hire many of its engineers, according to a letter seen by the </span><em><span>WSJ</span></em><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> An example of the shadow acquisitions that have become common in the AI industry. Poolside was previously regarded as an also-ran lab; it hasn&#8217;t trained a notable model in the three years since it was founded.</span></p><p><span>Potentially interesting as &#8220;commoditizing your complement&#8221;: if this changes the extent to which labs and Nvidia are competing vs cooperating. If Nvidia is able to provide Poolside with enough compute to create a meaningful open-weight American model, this forces American labs to spend more on training compute to remain differentiated. And $1B is more than enough to buy DeepSeek-level compute.</span></p></blockquote><div><hr></div><h2><strong><span>Capabilities</span></strong></h2><p><span>Anthropic&#8217;s Levent Alp&#246;ge releases a </span><a href="https://alpo.ge/s6.pdf"><span>108 page proof</span></a><span> of a </span><a href="https://arxiv.org/pdf/1708.01068"><span>celebrated</span></a><span> open problem in Riemannian geometry, &#8220;the Hopf problem&#8221;: </span><strong><span>&#8220;is the six-dimensional sphere a complex manifold?&#8221;. The answer appears to be &#8220;yes.&#8221;</span></strong><span> (Several of the greatest mathematicians of all time have tackled this problem, though mostly in their old age.) As usual from Alp&#246;ge, there is zero detail of which model was used, how much steering, or how much inference was used. One reason the doc is so long is that Claude&#8217;s proof contradicts a published, peer-reviewed </span><a href="https://link.springer.com/article/10.1023/A:1000313214795"><span>result</span></a><span> and goes to great lengths to explain why it&#8217;s wrong. The problem previously denoted that we have no theory of integrability in high dimensions. That&#8217;s still true. According to </span><a href="https://x.com/littmath/status/2091658357320933623"><span>Litt</span></a><span>: &#8220;the result, if correct, looks like a very long technical computation -- but it&#8217;s possible there&#8217;s some beautiful idea hiding in the 100 page pdf.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Extremely important mathematical achievement. (Some disagreement between </span><a href="https://x.com/littmath/status/2091658357320933623?s=20"><span>Litt </span></a><span>and </span><a href="https://x.com/ElliotGlazer/status/2091676145943290304?s=20"><span>Glazer</span></a><span> on whether the proof is &#8220;Fields Medal level.&#8221;)</span></p><p><span>The major point of interest from our point of view is the proof&#8217;s</span><em><span> length.</span></em><span> The proof-paper as published is 108pp, which seems to run counter to the consensus that frontier AI excels only in regimes that allow for gapless proof under 10 pages or so. In private correspondence with P3, Litt said the paper&#8217;s 108 page length is mathematically inessential &#8211; comprises expositions, digressions, didactics, etc. &#8211; and is plausibly a ~7 page proof at heart.  We thus believe it&#8217;s likely that Fable originally discovered the resolution to the Hopf problem by producing a ~7 page proof, in line with the norm of short, under 10 page proofs discovered by frontier LLMs.</span></p><p><span>This is all epistemically rather a shame from our POV. One of our most pressing questions is whether near-term frontier LLMs can produce </span><em><span>deep</span></em><span> scientific insight. In modern math, deep work almost always requires a rather long minimal writeup &#8211; so until LLMs start to produce long proofs, it&#8217;s a moot question whether they&#8217;re producing </span><em><span>deep</span></em><span> proofs.</span></p><p><strong><span>Opportunity: </span></strong><span>Inquire with specialists whether recent longer AI proofs such as Fable&#8217;s </span><a href="https://www-cdn.anthropic.com/564f962e60643842f5fcb4a17c9dbc8f608f1c37.pdf?utm_source=chatgpt.com"><span>autonomous paper on the Riemann zeta function</span></a><span> are also &#8220;&lt;10pp at heart.&#8221;</span></p></blockquote><div><hr></div><p><span>Musk&#8217;s first all-hands meeting with the newly acquired Cursor team has been </span><a href="https://www.theinformation.com/articles/cursor-officially-enters-musk-era"><span>leaked</span></a><span>. He reportedly said that Grok is behind and emphasized race dynamics: &#8220;</span><em><span>Musk said it is </span><strong><span>inevitable that AI models will become so advanced they&#8217;ll be impossible for humans to control</span></strong></em><span>.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Indirect reported speech, so there&#8217;s plenty of room for Musk to be misinterpreted here. Musk must have expected this to leak, and given his </span><a href="https://www.youtube.com/shorts/wEDztVCyHfU"><span>past statements</span></a><span>, there&#8217;s a sane game-theoretic reading of this speech: he wants to broadcast both his belief in AI doom </span><em><span>and </span></em><span>his resolve to race as long as others race, to shift the payoff matrix for others (to make it a </span><a href="https://en.wikipedia.org/wiki/Chicken_(game)"><span>game of chicken</span></a><span> or a </span><a href="https://en.wikipedia.org/wiki/Stag_hunt"><span>stag hunt</span></a><span> instead of a prisoner&#8217;s dilemma). It&#8217;s at least consistent with his preferring to stop, but wanting to force others to stop at the same time.</span></p></blockquote><div><hr></div><p><span>A </span><em><a href="https://time.com/article/2026/08/20/what-happens-when-the-world-is-run-on-code-no-one-understands-/"><span>Time</span></a><span> </span></em><span>op-ed argues that the</span><strong><span> US must build AI-driven mathematical verification infrastructure. </span></strong><span>&#8220;</span><em><span>We need public infrastructure for verified software: open libraries of verified components and specifications, standards and benchmarks, better tools for checking updates, and training programs that connect mathematics, computer science, engineering, and national security</span></em><span>.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Everyone in the &#8220;AI and formal verification will solve each other&#8221; sphere &#8211; i.e. Axiom Math, Math Inc, Safeguarded AI &#8211; has a flair for grand visions, and it&#8217;s hard to describe what a minimum viable product would look like. Still, it&#8217;s a good paradigm, and we&#8217;re glad it&#8217;s getting a popular platform.</span></p></blockquote><div><hr></div><p><span>A new ICML submission introduces </span><a href="https://openreview.net/pdf?id=SmwCItfHoC"><span>CoherenceBench</span></a><span>, measuring </span><strong><span>whether LLM probability assignments are consistent</span></strong><span> conditional probabilities. They test small LLMs (e.g. Qwen-30B) and find that even when LLMs achieve high calibration and/or Brier scores, they often fail Dutch-book coherence tests. The paper finds that direct RL training against Dutch books ameliorates this both in and out of distribution.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Interesting topic (&#8220;are LLMs rational in the formal sense?&#8221;) that deserves more frequent study. The results are useless to us because the models are so far from the frontier, but worthwhile to replicate the benchmarking with frontier and near-frontier models. This would let us ask the more urgent-to-us questions: are LLMs </span><em><span>becoming more rational in the formal sense as they become more capable</span></em><span>?</span></p><p><span>Strictly speaking, the paper&#8217;s method discovers behavioral inconsistencies which the authors interpret as axiom-violations in an underlying credence distribution, rather than discovering direct agentic axiom violations. We prefer </span><a href="https://arxiv.org/abs/2412.18544"><span>Paleka 2024</span></a><span> for the technically and philosophically simpler machinery. (See also some </span><a href="https://conceptualreasoning.ai/accord"><span>related work from Redwood</span></a><span>.)</span></p></blockquote><div><hr></div><p><strong><span>Toby Ord </span><a href="https://www.tobyord.com/writing/mathematics-is-more-than-proof"><span>argues</span></a><span> that mathematical research will still require human involvement. </span></strong><span>He concedes that AI has become exceptional at proofs, easily eclipsing human abilities, but claims that at the moment only humans have the capacity to decide what questions should be asked. He also notes that historically any advances in math have been the product of constructing new &#8220;concepts and vocabulary.&#8221; It is possible that this process could be automated, but Ord sees no evidence that AI possesses this capacity as of yet.</span></p><blockquote><p><strong>Opinion:</strong> Whether one accepts or rejects Ord&#8217;s claim that even in pure math humans will maintain a medium-term edge in &#8216;visionary&#8217; capability, it&#8217;s probable that in the medium-term humans will be needed to translate new mathematical insights to other fields (trading, architecture, software optimization). But it still seems like professional mathematics will be unrecognizable in &lt;30 years, with some or many parts/specializations to become as obsolete as the manual <a href="https://en.wikipedia.org/wiki/Computer_(occupation)"><span>calculator</span></a> profession now is</p></blockquote><div><hr></div><p><span>Annals of RSI: </span><a href="https://x.com/Benjamin_eecs/status/2090427345547219435"><span>SPADE</span></a><span> is a self-play method for automated environment design. Its novel contributions involve the inclusion of the </span><strong><span>environment generation as executable code inside the RL self-play loop</span></strong><span>, which trains the designer agent based on how close to the edge of the reasoning agent&#8217;s capabilities its environments are, as well as both roles being played by the same model.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Not that novel. One of many RL environment optimizations beside the vast iceberg of RL environment optimizations that don&#8217;t get reported or released.</span></p></blockquote><div><hr></div><h2><strong><span>Politics</span></strong></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gJAR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 424w, /__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 848w, /__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gJAR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png" width="1220" height="596" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:596,&quot;width&quot;:1220,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 424w, /__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 848w, /__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gJAR!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4093dc-8efb-45b6-b8a0-6c5b02f4ae92_1220x596.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Annals of the Overton shuffle: </span><a href="https://heatmap.news/daily/data-center-opposition-poll-collapse"><span>a majority</span></a><span> of Americans now oppose local data centers, with support for their construction now </span><em><span>lower</span></em><span> than for new coal plants. Both NIMBYism and a rapid increase in hostility toward AI are likely contributing to this effect. As we covered </span><a href="/__u/p3humansonai.substack.com/p/humans-on-ai-46-august-21st-2026"><span>last week</span></a><span>, politicians are aware of the popular mood, with Republicans redundantly warning frontier labs of the need to push back against this prevailing narrative.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We haven&#8217;t heard of </span><a href="https://emboldresearch.com/methodology/"><span>this pollster</span></a><span> before, but it seems mostly </span><a href="https://en.wikipedia.org/wiki/Change_Research"><span>fine</span></a><span>. Politicians are shifting and will further shift to appeal to anti-data center sentiment, particularly Democrats. If the issue becomes partisan, rather than doubly partisan like China hawkery, then the 2028 election will be even more significant than usual: a de facto referendum on AI scaling.</span></p></blockquote><p><span>Relatedly, the &#8220;Bitcoin Policy Institute&#8221; </span><a href="https://www.btcpolicy.org/articles/foreign-influence-in-the-campaign-against-american-ai"><span>claims</span></a><span> that </span><strong><span>foreign actors are behind an influence campaign against US AI</span></strong><span>. For instance, Chinese and Russian state media are propagating critical views of US AI data centers and export controls; US expat Neville Roy Singham, currently under congressional inquiry for his alleged ties to the CCP, is reported to have collaborated with Beijing in also producing content opposing AI; and reportedly a Swiss billionaire, along with a British billionaire, has funneled more than $2B into anti-data center campaigns.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> CCP news outlets </span><em><span>are </span></em><span>publishing anti-data center stories, but it&#8217;s likely also true that much of the data-center backlash originates from local sources (both grassroots and the massive existing nonprofit organizations) rather than entirely from foreign astroturfing.</span></p></blockquote><div><hr></div><p><span>Taiwanese authorities indict nine people over allegations of the </span><strong><span>illegal export of AI servers to China</span></strong><span>, </span><a href="https://www.aljazeera.com/economy/2026/8/25/nvidia-supermicro-employees-charged-over-export-of-ai-servers-to-china"><span>according to Al Jazeera</span></a><span>. Prosecutors argue that the individuals, including one Nvidia employee and two Supermicro employees, were motivated by the &#8220;pursuit of exorbitant profits.&#8221; The prosecutors are seeking jail sentences of up to five years. For its part, Nvidia says it will work with the Taiwanese authorities to &#8220;resolve the allegations as quickly as possible.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Popular opinion has not been on Nvidia&#8217;s side: see for instance this </span><a href="https://x.com/beffjezos/status/2034855483111350570"><span>meme</span></a><span> from the </span><a href="https://www.cnbc.com/2026/03/20/super-micro-co-founder-leaves-board.html"><span>original indictment</span></a><span> back in March.</span></p></blockquote><div><hr></div><p><strong><span>The </span><a href="https://www.ft.com/content/77b94c4a-4b4b-4983-9138-7db6926150f4?syn-25a6b1a6=1"><span>gulf</span></a><span> in AI investment between the US and the EU continues to grow</span></strong><span>, according to the </span><em><span>FT</span></em><span>. The EU is struggling to keep up with the unprecedented surge in spending on AI-related equipment and infrastructure seen in the US. Economist Karsten Junius believes the gap is &#8220;at least partly a temporary phenomenon.&#8221; The Swiss-based Bank for International Settlements instead emphasizes the risk of the US going all-in on AI and warns of an increasing risk of an &#8220;investment bust.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The EU is broadly wealthy enough to invest in this, but for various reasons the will isn&#8217;t there yet. Rather than chasing an expensive and mediocre catch-up program in vanilla LLMs, it would be nice to see the EU invest in high-variance safer </span><a href="https://lawzero.org/en/publication/safety-honesty-disinterested-ai-predictor"><span>alternative</span></a><span> </span><a href="https://en.wikipedia.org/wiki/Probabilistic_programming"><span>paradigms</span></a><span>.</span></p></blockquote><div><hr></div><p><span>After publishing a (good, disclosed) AI-written paper, </span><em><strong><span>the journal Philosophy and Public Affairs</span></strong></em><strong><span> </span><a href="https://x.com/sethlazar/status/2090778718360760759"><span>introduces</span></a><span> a new policy prohibiting publishing papers written in large part by AI. </span></strong><span> </span><em><span>&#8220;Academic journals serve at least two functions: the promulgation of new knowledge, and identification and credentialing of talented researchers&#8230; Submitting AI-authored essays makes the task of identifying talented researchers harder.&#8221;</span></em><span> A second, more straightforward worry described by the journal is signal-to-noise: while genuinely high-quality LLM-written papers are possible, LLM-written papers are disproportionately more likely than human-written papers to effectively fake markers of quality in the absence of genuine quality.</span></p><blockquote><p><strong><span>Opinion (Nu&#241;o):</span></strong><span> Seems like institutions which are able to harness AI contributions will tend to outcompete those who don&#8217;t, and philosophy journals haven&#8217;t quite been deciding the future of humanity as of late.</span></p><p><strong><span>Opinion (Peli)</span></strong><span>: The rationale given by the journal is good, but points to currently unsolved problems: there&#8217;s currently no solution for the problem of how to highlight work that is both worth professional engagement and unsuitable to serve as evidence (&#8220;costly signal&#8221;) of talent. Similarly, there is currently no solution for the problem of how to peer review LLM-written papers in disciplines where peer review is more an art than a science and partly relies on previously hard-to-fake signals of quality (e.g. when reviewers decide between a &#8216;reject&#8217; and a &#8216;resubmit with major revisions&#8217; based on general sense of promise).</span></p></blockquote><div><hr></div><p><strong><span>The State of Alabama&#8217;s Attorney General </span><a href="https://www.alabamaag.gov/attorney-general-marshall-launches-investigation-into-openai-and-sam-altman-for-massive-artificial-intelligence-data-breach/"><span>subpoenas</span></a><span> OpenAI</span></strong><span>, ordering the production of material relevant to July&#8217;s Hugging Face incident. This follows a letter from August 3rd, signed by 15 AGs, requesting that the relevant documents be preserved and that the security evals be halted.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> An interesting case, if we go by the invocation of the Deceptive Trade Practices Act as the ground for the subpoena. The Alabama Attorney General is presumably treating the Hugging Face incident as potential evidence that OpenAI is/was </span><em><span>falsely advertising</span></em><span> itself as a safety-conscious research lab, as well as </span><em><span>falsely advertising</span></em><span> its consumer products as being fundamentally safe.</span></p><p><span>Given that Sol 5.6 &#8211; a commercial model &#8211; was instrumental in the sandbox escape that enabled the internal models&#8217; July hacking marathon, a Deceptive Trade Practices case could make interesting points: while Sol 5.6 was operating with </span><em><span>external guardrails </span></em><span>(likely classifier-based) removed during the July incident and its lead up, the model at play was an unaltered Sol 5.6, rather than a modified &#8220;offensive cyber&#8221; style variant. It may therefore be interesting to argue that if Sol 5.6 is now-and-then disposed to autonomously engage in criminal conspiracy, held back only by external guardrails, then OpenAI&#8217;s public representation of Sol 5.6 is imperfect. </span></p><p><span>The main weakness of such a case would be that OpenAI&#8217;s </span><a href="https://deploymentsafety.openai.com/gpt-5-6/mle-bench-revised"><span>Sol 5.6&#8217;s system card</span></a><span> is not entirely rosy on alignment.</span></p></blockquote><div><hr></div><h2><strong><span>Safety</span></strong></h2><p><span>DeepMind&#8217;s Seb Krier </span><a href="https://blog.cosmos-institute.org/p/of-swarms-and-sand-gods"><span>reframes AI safety</span></a><span> as a choice architecture or bureaucracy design problem, rather than, say, an ML optimization task or a moral philosophy task. On Krier&#8217;s view, this follows from the prediction that </span><strong><span>swarms</span></strong><span> (many interacting imperfect agents), rather than single overwhelmingly smart agents, are likely to be </span><strong><span>the future of AI deployment.</span></strong></p><blockquote><p><strong><span>Opinion:</span></strong><span> The directional claim is fine &#8211; constitutional design and interventions will likely be valuable for safety and are still underrated. Contains some overstated claims, such as:<br><br>&gt; </span><em><span>&#8220;Rather than hoping that an agent&#8217;s &#8216;internal alignment&#8217; will remain perfectly robust across all sorts of edge cases, you can design harnesses and action-space boundaries that significantly shrink the surface area for behaviors like reward hacking. Without that, you&#8217;re supposed to trust the chain of thought (or neuralese someday), and rely on a single node.&#8221;</span></em></p><p><span>This misunderstands the alignment position: &#8220;trust the chain of thought&#8221; is not an accurate summary of almost anyone&#8217;s ideas &#8211; and ignores non-CoT methods, such as the rest of the interpretability field.</span></p><p><span>We also claim the rough direction is already well-represented in AI safety by the maturing </span><a href="https://blog.redwoodresearch.org/p/guide?hide_intro_popup=true"><span>AI Control</span></a><span> (regarding Krier&#8217;s &#8220;interchangeability&#8221; desideratum) and </span><a href="https://arxiv.org/abs/2502.14143"><span>Cooperative AI</span></a><span> fields.</span></p></blockquote><div><hr></div><p><strong><span>RL gives models context-specific &#8220;personas&#8221;</span></strong><span> rather than inculcating values across contexts, claims a new </span><a href="https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas"><span>post</span></a><span>. When a training environment rewards misaligned behaviors, the resulting model will act under a misaligned persona when the context is similar to that training environment.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Mostly rehashes the excellent </span><a href="https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade"><span>post by Nostalgebraist</span></a><span> discussed last week, but with a pessimistic twist we largely endorse. On the post&#8217;s view, if &#8220;aligned personas&#8221; are successfully generalized in training concurrently with RLVR-style training that reinforces reward-hacking/spec-gaming, it&#8217;s plausible that the result would be a rise in </span><em><span>motivated reasoning</span></em><span> that reconciles reward-hacking/spec-gaming behavior with the ethics of the aligned persona.</span></p></blockquote><div><hr></div><p><span>Richard Ngo&#8217;s </span><a href="https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization"><span>second installment</span></a><span> on the history of AI alignment </span><strong><span>outlines the dynamics that motivated safety researchers to accelerate capabilities </span></strong><span>over the last decade.</span><strong><span> </span></strong><span>He cites, among other things, prestige-chasing, sycophancy, a commitment to naive consequentialism, and instrumental power-seeking as responsible.</span></p><p><span>The post climaxes in a sweeping judgment: </span><em><span>&#8220;it&#8217;s time to give up on &#8220;alignment research&#8221; as a rallying cry; it&#8217;s become too corrupted. (&#8220;AI safety&#8221; is even worse as a term, and these days is mainly useful for describing a social cluster.)... I still consider (some version of) the alignment problem to be real and extremely important; and most of the intellectual progress towards solving it is still coming from people proximate to the alignment community&#8230; the alignment community has lost any moral right to try to gain power on altruistic grounds... The Pause/Stop AI movement does seem to avoid some of these failures&#8230; However, they don&#8217;t seem to be thinking clearly enough about politics to have robustly good effects on the world&#8230; Again, I&#8217;m not claiming that the alignment community is unusually unethical: I don&#8217;t know of any other similarly-sized community which is able to avoid the corrupting effects of this much power.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Our overall feelings are mixed. Ngo&#8217;s histories are good at illustrating the concept he&#8217;s pointing at &#8211; &#8220;pessimization,&#8221; the systematic achieving of the opposite of one&#8217;s stated goals. &#8220;</span><em><span>One particularly notable blind spot (at least in public discussions) is how Dario&#8217;s early racing on behalf of OpenAI played a big role in creating the &#8216;problem&#8217; that he now purports to be solving by racing on behalf of Anthropic.&#8221;</span></em></p><p><span>Ngo&#8217;s prescriptions for individuals who wish to avoid the failures he identifies &#8211; &#8220;be stricter about who you ally with,&#8221; &#8220;be more discerning about what you consider virtuous and demand more from others,&#8221; &#8220;choose as if you were making a choice for the entire category of people you belong to&#8221; &#8211; seem to us to amount to telling yourself to </span><a href="https://www.neelnanda.io/blog/mini-blog-post-6-stop-pressing-the-try-harder-button"><span>try harder</span></a><span>, just </span><a href="https://talyarkoni.org/blog/2018/10/02/no-its-not-the-incentives-its-you/"><span>tame</span></a><span> the incentive landscape.</span></p><p><span>The excellent </span><a href="https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization?commentId=hjmmbtRi9GaMLWh8y"><span>comment thread</span></a><span> adds important nuance, e.g. pointing at the difficulty of implementing these due to incentives which are if anything now worse than in 2020, e.g. noticing that his analysis mostly fails to credit people who avoided the failures. The </span><a href="https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization?commentId=Z3wBs3nLHNzHJeNYk"><span>closest</span></a><span> Ngo comes to addressing this point is him arguing that granular credit-and-blame assignment including credit for rightful omissions is worth it even if it&#8217;s hard.</span></p></blockquote><div><hr></div><h2><strong><span>Incidents</span></strong></h2><p><strong><span>The UK&#8217;s AISI set off the rogue agent that attempted to sabotage a computer science student&#8217;s research project, </span></strong><span>according to </span><em><a href="https://www.reuters.com/world/how-texas-student-blew-whistle-rogue-ai-hacking-attempt-2026-08-20"><span>Reuters</span></a></em><span>.</span><strong><span> </span></strong><span>When the undergrad spotted a malicious pull request to an open-source network scanner called &#8220;myNetwork,&#8221; he posted a warning to the program page, after which &#8220;other users chimed in to insist nothing was amiss, sharing detailed explanations for why he had gotten it wrong.&#8221; It transpired that these other &#8220;users&#8221; were AI agents engaged in a campaign of interactive deception by creating a &#8220;multiperson conversation.&#8221; The student then used Claude to check the malicious code and confirming his suspicions that the two GitHub accounts were lying. The actual </span><a href="https://github.com/w1b/aisi-mythos-inc-2026-07-28-01-recovered-pr/blob/main/assets/annotated-pr-thread.png"><span>comment thread</span></a><span> is not very scary and pretty easy to clock as AI. The incident has, however, already been partially documented in </span><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"><span>AISI&#8217;s report</span></a><span> on Mythos&#8217;s capacity for social engineering.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Models at the level of Mythos could do serious blast damage to civilization through deceitful schemes like this one. Incidents like this one also prove Mythos-level models have the </span><em><span>drive </span></em><span>to do so &#8211; not always, but in </span><em><span>some</span></em><span> circumstances. Most crucially, the triggering circumstances are fairly unpredictable (despite </span><a href="https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade"><span>interesting speculative analysis</span></a><span>), and not readily accounted for by a helpfulness drive (&#8220;the user asked for social sabotage and the model complied&#8221;) &#8211; the choice to develop and execute a social sabotage plan is spontaneous. This is a huge externality imposed by private companies on every computer user.</span></p><p><span>Currently, it looks like our only defenses against this are classifier-based external monitors, voluntary standards, and the hope that antisocial scheming personas only come out when the context is sufficiently hacking-flavored. It thus would be good to have </span><em><span>some</span></em><span> public disclosure on how society&#8217;s Mythos- and Sol-powered Big Patch is going, and whether it will be enough.</span></p></blockquote><div><hr></div><h2><strong><span>Minor</span></strong><span> <br></span></h2><ul><li><p><span>New sycophancy bench, </span><a href="https://x.com/PReaulx/status/2089733368900432214"><span>Pander</span></a><span>, finds Fable to be almost completely candid on queries, but sycophancy rises sharply on &#8220;tasks&#8221; (autonomous chains of actions).</span></p></li><li><p><strong><span>Hugging Face is exploring a sale that could see the company valued at $13B</span></strong><span>, </span><a href="https://www.businessinsider.com/hugging-face-could-be-acquired-13-billion-2026-8"><span>reports</span></a><span> Business Insider. Investors are increasingly interested in soft AI infrastructure companies (like OpenRouter or Civitai), not just those developing models or data centers.</span></p></li></ul><ul><li><p><span>Mysterious </span><a href="https://x.com/opencode/status/2090544355824038300"><span>Ox Alpha</span></a><span> model available for free on OpenRouter, with lots of capacity per user per day. OpenCode showed 16 trillion tokens used across 221,000 users, making it its #2 model this week. Much speculation as to who&#8217;s behind it. Z.ai is the leading candidate, possibly running inference on Huawei Ascend.</span></p></li></ul><ul><li><p><span>Beijing World Humanoid Robot Games displays </span><a href="https://archive.ph/PgdGI"><span>progress made in physical capabilities.</span></a><span> Most notable was X-Humanoid&#8217;s sprinter, who surpassed Usain Bolt&#8217;s world record by 0.19 seconds.</span></p></li></ul><ul><li><p><span>Not news: </span><strong><span>Interviews for new Anthropic hires involve thorny questions on culture and ethics</span></strong><span>, reports </span><a href="https://www.axios.com/2026/08/24/scoop-anthropic-candidates-face-blunt-money-question"><span>Axios</span></a><span>. One applicant reveals that they were asked how they would feel if the company abandoned its commitment to safety. Another anonymous source claims that their interviewer did not seem happy when the applicant expressed their discomfort with the ethical implications of such a change of mission. This is not news; they have asked these questions since it was founded.</span></p></li><li><p><strong><span>Mathematician Max Weinreich </span><a href="https://arxiv.org/pdf/2608.02859"><span>proposes</span></a><span> a total embargo on the use of AI in mathematics</span></strong><span>. In defiance of AI &#8220;evangelist&#8221; Terence Tao, the paper argues that &#8220;artificial mathematics&#8221; is a corrosive and destructive phenomenon that must be replaced with the field of &#8220;natural mathematics.&#8221; </span></p></li><li><p><span>The new </span><a href="https://resi.org/"><span>Institute for Responsible Superintelligence</span></a><span> has an impressive team of world-class computer scientists. It is adjacent to, but distinct from, ARIA&#8217;s Safeguarded AI agenda. We guess the difference is that RESI is trying to build the primitives for further work, while ARIA is trying to solve the problem end-to-end.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI #46, August 21st 2026]]></title><description><![CDATA[GPT-3 moment for robotics? OpenAI disbands safety team, again, transition period for cyber might favor offense.]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-46-august-21st-2026</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-46-august-21st-2026</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Fri, 21 Aug 2026 19:52:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-gZw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: right;"></p><p><strong><span>TL;DR</span></strong></p><ul><li><p><span>A GPT-3 moment </span><a href="/__u/p3humansonai.substack.com/p/humans-on-ai-46-august-21st-2026"><span>for robotics</span></a><span>?</span></p></li><li><p><span>Guidelight notes that Anthropic no longer mentions limiting the deployment of one of its models in response to an incident.</span></p></li><li><p><span>OpenAI once again &#8220;disbanded&#8221; a major risk team, Preparedness.</span></p></li><li><p><span>The evaluator behind Anthropic&#8217;s cyber-felony post-mortem investigation has released a sobering essay on the next few years of AI cyber attacks, arguing that the transition period favors offense.</span></p></li></ul><h2><strong><span>Economics</span></strong></h2><p><strong><span>Anthropic&#8217;s ARR run-rate hit </span><a href="https://www.cnbc.com/2026/08/17/anthropic-says-annualized-revenue-climbed-to-65-billion-in-july.html"><span>$65B</span></a></strong><span> in July, up from a reported $47B in May, around the time of its most recent fundraise. It also disclosed $11.6B preliminary revenue for the second quarter of this year.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-gZw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 424w, /__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 848w, /__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-gZw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png" width="1456" height="648" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:648,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 424w, /__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 848w, /__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-gZw!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20779f62-b035-4d0b-a624-64ee593f0a0c_1647x733.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>OpenAI saw </span><a href="https://www.techtimes.com/articles/324713/20260817/openai-reaches-40b-revenue-safety-leaders-exit-models-break-containment.htm"><span>rapid growth in July,</span></a><span> after </span><a href="https://www.wsj.com/tech/ai/openais-second-quarter-sales-show-tepid-growth-compared-with-anthropic-5cb42998?mod=hp_lead_pos2"><span>a weak Q2</span></a><span> (their</span><strong><span> losses grew 25%</span></strong><span> from Q1 to Q2 2026, reaching $12.3B against 16% quarterly revenue growth).</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Despite the big number, $65B is actually a negative update relative to our (</span><a href="https://x.com/MacroMicroMe/status/2089881273003348259?s=20"><span>and the market&#8217;s</span></a><span>) expectations. $11.6B for Q2 is consistent with prior reporting.</span></p><p><span>Together with Anthropic&#8217;s slight slowdown, the above suggests that OpenAI may have cannibalized some Anthropic growth via the improvements of 5.6 Sol and Codex (especially in token efficiency and capability &#8220;</span><a href="https://www.nature.com/articles/s42256-025-01137-0"><span>densing</span></a><span>&#8221;).</span></p><p><span>Our best model of the situation remains seeing Anthropic and OpenAI&#8217;s combined growth as a fairly smooth exponential, with some jaggedness in their individual growth rates as one or the other takes market share due to a particular model or product.</span></p><p><span>We see an acceleration from February through to the end of April, corresponding to the boom in Claude enterprise adoption post-Opus 4.6 and the popularisation of Claude Code, then some reversion to the prior joint growth rate. Take these numbers with a generous pinch of salt.</span></p></blockquote><div><hr></div><p><span>Nvidia and Broadcom, as well as Meta, are now </span><strong><a href="https://archive.ph/gyH5r"><span>guaranteeing the residual value of debt</span></a><span> raised by datacenter &#8220;special purpose vehicles.&#8221;</span></strong><span> This allows newer AI labs to lean on the strong credit rating of established firms (Broadcom backstopping Anthropic, for example). Some analysts are less convinced: &#8220;</span><em><span>The guarantee is nearly costless in the boom phase, but becomes most relevant in a severe, abrupt downturn</span></em><span>,&#8221; says CreditSights. Others are more optimistic: for a crash &#8220;</span><em><span>you would have to have growth rates of token usage fall off a cliff, which we&#8217;re just not seeing</span></em><span>,&#8221; claims one investor.</span></p><p><span>$The deal allows Nvidia&#8217;s GPUs to be used as collateral for loans in a way that will also help to sustain demand for Nvidia chips as they age, making financing more palatable to risk-averse investors.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> A brilliant move if Nvidia correctly handicaps the risk, giving it cheaper capital to increase revenue. A particular datacenter financing deal is not that dangerous, it trades that particular risk for increasing the magnitude of systemic risk. This is a classic bubble dynamic.</span></p></blockquote><div><hr></div><p><em><span>The Economist</span></em><span> claims that the symbiotic </span><strong><span>relationship between Nvidia and the hyperscalers</span></strong><span> (Amazon, Google, Meta, and Microsoft) is</span><a href="https://archive.ph/sp34D"><span> </span></a><strong><a href="https://archive.ph/sp34D"><span>fraying</span></a></strong><span>. Nvidia has additionally partnered with heavyweight Wall Street investors to raise $500B to finance AI investment for other customers (likely governments and companies looking to create their own AI internal infrastructure). And the hyperscalers are spending billions to bring chip design in-house, with a particular focus on custom silicon chips. Nvidia should remain the dominant chip manufacturer, but custom chips may put significant pressure on its sizable margins. Jensen Huang is unconcerned, arguing that custom chips&#8217; strength is also their weakness: specialized chips, unlike Nvidia&#8217;s GPUs, are less valuable as AI extends to robotics, vehicles, and industry.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> A plausible hypothesis. We covered similar dynamics back </span><a href="/__u/p3humansonai.substack.com/p/humans-on-ai-42"><span>in July</span></a><span>.</span></p></blockquote><div><hr></div><p><span>AMD </span><a href="https://www.cnbc.com/2026/08/06/amd-buys-taalas-startup-that-hardwires-ai-models-into-its-silicon.html"><span>purchases</span></a><span> an inference chip startup, Taalas. Its chips are radically application-specific (&#8220;ASIC&#8221;): they </span><strong><span>hardwire AI weights directly into silicon</span></strong><span> for more efficient inference (100x faster, 10x less power) by eliminating &#8220;weights movement,&#8221; the loading and swapping step. This comes 7 months after Nvidia made a related nonexclusive $20B deal for Groq&#8217;s custom inference chips. Reportedly AMD is planning to use Taalas for the decode step of inference, with its Instinct GPU doing the prefill.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We think this is irrelevant for the frontier. Hard-wiring the dominant MoE architecture faces severe problems: 1) no ability to fix imbalanced load on particular experts (&#8220;</span><a href="https://openreview.net/forum?id=CsHahbRAFZ"><span>hot experts</span></a><span>&#8221;), which slows everything down enormously; 2) data-dependent control flow and inter-chip dispatches to different experts brings back complexity and energy costs; 3) as usual, any long-context model will still require massive piles of RAM for the KV cache, so AMD won&#8217;t even avoid the main bottleneck of their competitors. Using Taalas for decode doesn&#8217;t avoid this RAM requirement.</span></p><p><span>Finally, the current model release cadence (incremental updates every couple of months) will initially be a serious economic issue for Taalas, given that developing and deploying a new instance at scale must take them &gt;2 months because of fab lead time, and since their installed fleet will depreciate very quickly. But if the architecture of the model being trained is set down when the training process begins, perhaps model training and custom silicon production can proceed at the same time.</span></p><p><span>The approach probably still has a decently sized market segment: anywhere that needs extremely fast intelligent-enough decisions, where a dense 20B open model from 12 months ago is acceptable. Eventually it </span><em><span>might</span></em><span> also drive fast </span><a href="https://en.wikipedia.org/wiki/Speculative_decoding"><span>speculative decoding</span></a><span> for frontier models. Perhaps a good point of reference is </span><a href="https://en.wikipedia.org/wiki/Embedded_system"><span>embedded systems</span></a><span> vs the general computer market, $100B/year vs $400B/year for general computers.</span></p><p><span>Given the last funding round was $169M, we guess the sale price was $1-2B.</span></p></blockquote><div><hr></div><p><strong><span>Stripe </span><a href="https://archive.ph/nsHOp"><span>buys</span></a><span> OpenRouter for &gt;$7B</span></strong><span>, its largest acquisition to date. In its</span><a href="https://www.documentcloud.org/documents/28565866-stripes-august-2026-investor-letter/"><span> letter to investors,</span></a><span> Stripe marked January 1st as the &#8220;beginning of the Singularity,&#8221; reflecting the significant increase in the rate of new firms created.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> An outrageous amount in terms of current revenue (50x OpenRouter&#8217;s revenue), so this is an enormous bet on OpenRouter&#8217;s brand value, customer loyalty, and compounding growth. (If it were just open models, or the value of model adoption data, or an enterprise shift towards careful metering of token spend and routing to the minimal intelligence needed for a task, then Stripe could presumably build their own gateway.)</span></p><p><span>Gateways are thin wrappers with a lot of competition: LiteLLM is open source; Vercel, Cloudflare, and the hyperscalers all have AI gateways; and big enterprises instead negotiate directly with providers once they&#8217;re spending.</span></p><p><span>Chinese models peaked at 46% of US enterprise token usage on OpenRouter, and regulation could, perhaps, block this. So a large fraction of the traffic Stripe now oversees could be restricted by US policy.</span></p><p><span>A useful reference here is Visa&#8217;s (aborted) bid to buy Plaid (a bank-account data aggregator and a valuable data stream) for $5B.</span></p><p><span>The use of &#8220;Singularity&#8221; to mean &#8220;a sudden increase in new startups&#8221; is exasperating, but what can you do.</span></p></blockquote><div><hr></div><p><span>Google outbids Mercor to </span><strong><a href="https://www.reuters.com/legal/litigation/google-buy-spirit-airlines-business-data-10-million-2026-08-17/"><span>purchase</span></a><span> Spirit Airlines&#8217; enterprise data for $10M</span></strong><span>. This includes 100 million emails, 500 million Teams messages, and 20.5 million SharePoint files, as well as other data.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Bankruptcy of large legacy firms (i.e. those who have built up a large estate of data) may well become another avenue for labs to acquire fresh data on which to train future models, though based on the relatively low price of this transaction, probably not a very pivotal one.</span></p></blockquote><div><hr></div><p><span>An </span><a href="https://americanaffairsjournal.org/2026/08/the-augmentation-automation-race/"><span>article</span></a><span> in </span><em><span>American Affairs</span></em><span> argues that certain </span><strong><span>AI policies intended to be pro-worker are self-defeating</span></strong><span>. By resisting AI progress, US companies will lose their competitive advantage to firms from countries less troubled by the pace of adoption, which in turn would create domestic mass unemployment. Current data show some signs of AI replacing human workers, albeit modestly, but in general AI is augmenting human work &#8220;along the jagged frontier,&#8221; a &#8220;synergistic and efficient&#8221; partnership. And so, according to the author, if the US is to retain both competitive advantage and strengthen workers&#8217; bargaining position, the &#8220;centaur era&#8221; must endure.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We would also like the centaur era to last, but this article gives no reason to think that the current approach to AI will preserve labor-augmentation, given that they are </span><a href="https://openreview.net/pdf?id=f15MLGSjEs"><span>trained</span></a><span> as replacements. Reforming income tax to stop subsidizing AI replacement and incentivizing companies to train for augmentation are far better options than just removing adoption frictions and hoping.</span></p></blockquote><div><hr></div><h2><strong><span>Capabilities</span></strong></h2><h3><span>&#128294; Are robots few-shot learners? A GPT-3 moment for robotics?</span></h3><p><span>Robots have been trailing LLMs because of a lack of suitable training data, and because foundation models have not demonstrated strong &#8220;in-context learning&#8221; (live generalization to new tasks given examples).</span></p><p><span>This may have changed: Generalist AI presents Gen-1.5, a new </span><a href="https://generalistai.com/blog/gen-1.5"><span>robotics foundation model</span></a><span> which it claims can </span><strong><span>learn new basic physical tasks after only a single demonstration</span></strong><span> (with a success rate of 59%, but only on n=10 short tasks). The only examples given are &#8220;</span><em><span>diverse tasks including handling zippers, opening jars, grabbing money out of wallets</span></em><span>.&#8221; This rises to 83% when </span><a href="https://arxiv.org/abs/2411.07279"><span>fine-tuned</span></a><span> on a 50x longer demonstration (10 gradient steps on 5 minutes of data per task).</span></p><p><span>If we take the GPT-3 comparison seriously, how much data will deliver a robotic GPT-4?</span></p><ul><li><p><span>Its Gen-0 (November 2025) model took 270K hours of data;</span></p></li><li><p><span>Gen-1 (April 2026) took 500K hours;</span></p></li><li><p><span>Gen-1.5&#8217;s total data is not given, but it took 60% longer to train;</span></p></li><li><p><span>And its collecting 10K hours/week and accelerating this capture rate;</span></p></li><li><p><span>Therefore, it was perhaps trained on 700K-1M hours of data.</span></p></li></ul><p><span>GPT-4 is rumored to have involved 40x more data than GPT-3: 13T tokens vs 300B. If the scaling laws were similar for robots, this would weakly indicate that a &#8220;GPT-4 level&#8221; robot will take around 40M hours of real glove manipulation data. (This assumes that, like LLMs, they do only one &#8220;epoch,&#8221; i.e. make only one use of each data point. It also ignores many other things that have changed since GPT-4.) Just like any robotics company, it needs to start deploying (and so capturing customer data) for this to work out: at say 50K hours a week, it would only obtain this much data in 2041. Of course, pretraining data size doesn&#8217;t explain everything: after GPT-4, post-training scaled up by a factor of millions.</span></p><p><span>40 million hours, or 4.5 thousand years, or 4,500 humans gathering glove data for a year, is expensive but not </span><em><span>that </span></em><span>expensive, at $100K per human year this would be $450M, well within the power of capital markets to finance. And there are already efforts underway in e.g., </span><a href="https://www.euronews.com/next/2026/03/11/inside-chinas-robot-school-where-humanoid-machines-are-learning-everyday-tasks"><span>China</span></a><span> to </span><a href="https://www.globaltimes.cn/page/202607/1367227.shtml"><span>provide</span></a><span> such data, which have 270 robots at one time; scaling up to 4,500 would only be a 20x increase.</span></p><p><span>The analogy with GPT-4 breaks down a bit, because research takes the road of least resistance; if gathering glove data was so expensive that we&#8217;d need to wait fourteen years of wall-clock time to do so, much more effort would go into algorithmic improvements to compensate.</span></p><p><span>Gen-1.5 already involves some post-training, including imitation learning and </span><a href="https://pantograph.com/journal/pan-1"><span>goal-seeking RL</span></a><span>. But RLVR is more difficult in robotics: even if we had great simulations to speed things up 1000x and sidestep expensive breakages during training, </span><a href="https://heroic-cupcake-6e9b26.netlify.app/"><span>spec gaming</span></a><span> appears to be, so far, a hard to eradicate consequence of RLVR, especially if we are detecting the robot&#8217;s success using foolable vision models.</span></p><p><span>Generalist also claims &#8220;sim2real&#8221; embodiment transfer: Gen-1.5 shows prompts from &#8220;recordings&#8221; of a robot simulation working on real robots, despite there being no simulation data in the pretraining corpus. This is the holy grail, but we just don&#8217;t believe it&#8217;s solved in general yet.</span></p><p><span>For converting between hours of robot data and tokens, we don&#8217;t have a firm baseline. But a competitor, Pantograph, </span><a href="https://pantograph.com/journal/pan-1"><span>notes that</span></a><span> &#8220;a 1M token context window would correspond to about 7 hours of context,&#8221; i.e. maybe 140,000 tokens/hour. GPT-4&#8217;s 13 trillion tokens is then more like 100 million hours of robot data, but it&#8217;s pretty easy to recover our 40M estimate if Generalist increases the frame capture rate or frame resolution above the usual highly compressed data (128x128 pixel frames at 10 frames per second).</span></p><p><a href="https://x.com/herbiebradley/status/2090224497853120697"><span>Herbie Bradley</span></a><span> responds: &#8220;</span><em><span>I&#8217;d expect this paradigm to be capable of going pretty far, but suffering from the analogous problems as RLVR for LLMs: weakness outside of distribution, high latency of action compared to humans, the need to greatly scale data collection for both post-training and pre-training, etc. This may be sufficient for much of manufacturing (notwithstanding problems around human hand agility), it is more doubtful if it sufficient for generalized construction, so I am still sceptical of industrial explosion</span></em><span>&#8221;</span></p><div><hr></div><p><span>The </span><a href="https://www.unsolvedmath.com/"><span>UnsolvedMath</span></a><span> benchmark, which </span><strong><span>scrapes research maths workshops for open problems</span></strong><span> for AI to solve, just added 3,359 open problems. It then ran Sol xhigh on the problems, resulting in </span><strong><span>174 </span></strong><span>(5%)</span><strong><span> apparently new solutions</span></strong><span>  using an ultimately minor amount of inference. It also ranks the problems by difficulty; 28 &#8220;expert&#8221; level (i.e. world expert) problems have been </span><a href="https://www.unsolvedmath.com/problems?difficulty=4&amp;status=solved"><span>solved</span></a><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Ingenious way of automatically generating a decent benchmark.</span></p><p><span>One could use this for a very rough estimate of the fraction of natural research questions AI can currently solve &#8211; perhaps 4-5%, given Sol&#8217;s false positive rate (also estimated from </span><a href="https://x.com/prz_chojecki/status/2089274727407690102"><span>vibes</span></a><span>): </span><em><span>&#8220;In my experience roughly 5%-10% &#8216;claimed solutions&#8217; by GPT-5.6 Sol are wrong</span></em><span>.&#8221;</span></p></blockquote><div><hr></div><p><span>Annals of the epistemic crisis: Dan Luu </span><a href="https://danluu.com/benchpocalypse/"><span>illustrates</span></a><span> the deep goodharting problem with AI labour: he ran an AI agent for a month(!) to make a regex parser which gets SOTA on one good benchmark. On a second regex benchmark, the AI&#8217;s library &#8220;</span><em><span>was 10x slower on cases where the benchmark didn&#8217;t take forever due to an algorithmic blow-up, and there were cases where it took so long that it wasn&#8217;t reasonable to even wait for the benchmark to complete</span></em><span>.&#8221; </span><strong><span>&#8220;It&#8217;s trivial to &#8220;win&#8221; a non-trivial benchmark in a meaningless way even when you instruct agents to not reward hack or overfit to win the benchmark.&#8221;</span></strong></p><p><em><span>&#8220;for now, LLMs are good at doing bad benchmarking, so even if you have something that&#8217;s a real performance improvement, you generally can&#8217;t tell from some LLM-generated benchmark setup</span></em><span>.&#8221;</span></p><p><span>However, some good news: &#8220;</span><em><span>telling the LLM there&#8217;s a holdout set worked better than just telling the LLM to do generalized work or not overfit or cheat</span></em><span>.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The generalization landscape for automated coding and autoresearch is currently pretty confusing. The classic autoresearch format (an LLM loop optimizing a neural net training setup) does yield gains that are robust even when there are some shifts in the setup (random seeds and moderate scaling). But as </span><a href="https://metr.org/notes/2026-08-14-llm-contribution-to-discoveries/"><span>Cunningham reports,</span></a><span> there has been no public algorithmic speedup from the last 6 months&#8217; popularisation of the method.</span></p></blockquote><div><hr></div><p><a href="https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf"><span>Claude</span></a><span> (Opus 4.8 and a Mythos preview) autonomously </span><strong><span>orchestrates the design of working protein binders</span></strong><span> across 14 targets. An AI agent ran </span><em><span>de novo</span></em><span> binder design &#8220;campaigns&#8221; against 16 protein targets. This involved using a protein design prompt written by a human expert, but with no human input on design decisions. It researched targets, chose epitopes, ran open-source design and prediction tools, and delivered 30 ranked designs each, all within 48 hours. Two CROs synthesized and tested the designs: 354 of the 1,320 proteins bound to the targets (27% hit rate), with at least one succeeding on 14 of the 16 targets. On RBX1, Claude beat a recent design competition&#8217;s winning entry. The catch is that Claude required a heavily engineered protocol, a curated scientific corpus, and many existing protein-design models, somewhat deflating the significance of the study.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> &#8220;Claude designed the binders&#8221; is too strong. The molecular generation was performed by specialist models such as PXDesign, RFdiffusion variants, Genie 3, BindCraft variants, Proteina-Complexa, ProteinMPNN, etc. Claude selected and operated these tools, chose epitopes and constructs, allocated compute, filtered candidates, and ranked outputs. Claude was an autonomous campaign manager and decision layer.</span></p><p><span>Breaking down possible claims:</span></p><ul><li><p><span>Claude can autonomously operate a complex computational pipeline: </span><em><span>yes</span></em></p></li><li><p><span>The pipeline can generate real binders across diverse targets: </span><em><span>yes</span></em></p></li><li><p><span>Claude adds value beyond a scripted pipeline: </span><em><span>not shown</span></em></p></li><li><p><span>Claude performs as well as or better than human experts: not shown</span></p></li><li><p><span>The approach generalizes to genuinely novel or poorly characterized targets: </span><em><span>some signs</span></em></p></li><li><p><span>The proteins bind in the predicted pose: </span><em><span>not shown</span></em></p></li><li><p><span>The proteins modulate biological functions: </span><em><span>not shown</span></em></p></li></ul><p><span>How much would fully-automated binder design improve drug development? </span><a href="https://pubmed.ncbi.nlm.nih.gov/26863229/"><span>Not much</span></a><span>; it would improve throughput at a stage that is far from rate-limiting.</span></p><p><span>As with other domains, there will be a whole lot of grunt work to integrate this and legibilize enough context within which models would be able to produce progress. This is part of what Jeff Dean&#8217;s </span><a href="https://www.businessinsider.com/why-jeff-dean-left-google-ai-startup-discovery-loop-2026-8"><span>startup</span></a><span> will be doing. We&#8217;ll be watching for signs of automating that meta process over the coming months.</span></p></blockquote><div><hr></div><p><span>Prime Intellect </span><a href="https://www.primeintellect.ai/blog/measuring-autonomous-research"><span>finds</span></a><span> that frontier AI models can perform sophisticated research-like procedures, but </span><strong><span>struggle with originality</span></strong><span>. In response to ongoing concerns/hopes around RSI, the test looked at whether frontier AI models can conduct real research, running 153 multi-day autonomous trials across 18 models on the nanoGPT speedrunning-training-a-classic-model benchmark. Models worked without internet access, designing and running experiments to cut training steps for a target loss. Claude Fable 5 led across all budget comparisons (time, # experiments, tokens), closing 82% of the gap to the human SOTA. No model discovered a fundamentally new method: all converged on similar optimizer tricks. Top models, however, ran sophisticated processes: they tested ideas across multiple seeds, revisited abandoned hypotheses under new conditions, and built reusable experiment tooling, while weaker models discarded promising results after single noisy runs.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Fits our thesis about models lacking something and not obviously iterating towards getting it. This can be seen in their success in (verifiable) mathematics.</span></p></blockquote><div><hr></div><p><a href="https://www.benchling.com/blog/can-llms-work-in-the-wet-lab"><span>BenchBench-Protocol</span></a><span> is a new benchmark for comparing AI&#8217;s </span><strong><span>ability to do &#8220;wet lab&#8221; biological work</span></strong><span>. The test set contains real, private per-lab adaptations of a general protocol, allowing it to operate within the bespoke architecture of Benchling&#8217;s lab. The benchmark measures ecological validity, difficulty, and contamination resistance. Opus 5 comes out on top with 59%. Kimi is competitive, performing near OpenAI&#8217;s flagship Sol model. Claude Fable isn&#8217;t mentioned anywhere in the paper, presumably because its tight safeguards around biology-related prompts prevented its use.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The main problem here is that it&#8217;s grading models against one valid choice (one scientist&#8217;s actual choice), when there will be many valid choices in biological work.</span></p></blockquote><div><hr></div><p><span>A DeepMind </span><a href="https://arxiv.org/abs/2608.17776"><span>paper</span></a><span> shows training LLMs via </span><strong><span>debate decreases reward hacking</span></strong><span> compared to RLAIF by about half. The study involves an adversarial two-player setup, with a generator and a critic debating before a weaker AI judge. Compared to a single-player RLAIF baseline that quickly found vulnerabilities in its judge, two-player debate did not degrade the judge&#8217;s capabilities and recovered 45% of the performance gap. They also find that weaker judges prevent reward hacking less, and that a critic is more likely to game a judge if it is allowed a critique length of more than ~150 words.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The debate-for-safety agenda petered out a few years ago, so interesting to see this resurface. The new Resolution team is likely to push in similar directions.</span></p></blockquote><div><hr></div><p><span>The age of leisure: the AI Futures team </span><a href="https://www.lesswrong.com/posts/ZPSsmRH5oMwLPXys4/q2-5-2026-timelines-update-uplift-and-revenue"><span>updates</span></a><span> its timeline for the </span><strong><span>first fully Automated Coder</span></strong><span> by adding two new methods for its estimation of when labs will replace human coders with AI. The timeline shifts are modest, but new modeling and evidence have improved confidence. It previously solely relied on data from METR&#8217;s coding &#8220;time horizon&#8221; trend. It now also incorporates data on coding uplift and revenue. All three methods converge on a similar timeline: the model&#8217;s baseline gives </span><strong><span>~70% probability of AC by January 2030</span></strong><span>. Brendan Halstead (one of the three forecasters) thinks that&#8217;s overconfident and revises it down to 60% by January 2030.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Not much movement, but we agree that it&#8217;s less unlikely than before.</span></p><p><span>A secondary result is that &#8220;preliminary results from P-Zero Research indicate Opus 5 is at parity with &#8216;expert humans&#8217; in research taste on their verifiable tasks.&#8221; No idea who this is, and there&#8217;s no methodology, so update only if you trust Brendan from AIFP.</span></p></blockquote><div><hr></div><h2><strong><span>Politics</span></strong></h2><p><strong><span>The</span></strong><span> </span><strong><span>US is unprepared for AI-enabled bioweapons</span></strong><span>, argues a former Homeland Security advisor in </span><em><a href="https://www.foreignaffairs.com/united-states/ai-new-age-bioweapons-sherwood-randall"><span>Foreign Affairs</span></a></em><span>. Historically, nation states remained the only capable actors of developing biological weapons. Advances in AI capabilities have and will continue to expand the pool of actors capable of carrying out such attacks, especially considering the emergence of cloud labs. The US should, in her view, shift from a strategy of deterrence to one of resilience &#8211; namely, preparing a response for </span><em><span>when</span></em><span> such attacks occur, rather than trying to prevent them outright.</span></p><p><span>For detection and attribution, she proposes creating a nation-wide monitoring system, which would include wastewater sensors and distributed labs to process the resulting data; for mitigation after attacks, she suggests that the US should establish </span><em><span>&#8220;a coordinating council for the biotechnology sector that includes DNA synthesis providers, AI model developers, genomic data platforms, cloud lab operators, and high-containment laboratories.</span></em><span>&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> These arguments aren&#8217;t entirely new, though it&#8217;s good to see them receiving the attention they deserve.</span></p></blockquote><div><hr></div><p><span>An </span><em><a href="https://www.nytimes.com/2026/08/16/us/politics/military-ai-china-anthropic.html?unlocked_article_code=1.51A.U11N.LbJBsdB9cPaA&amp;smid=nytcore-ios-share"><span>NYT </span></a></em><a href="https://www.nytimes.com/2026/08/16/us/politics/military-ai-china-anthropic.html?unlocked_article_code=1.51A.U11N.LbJBsdB9cPaA&amp;smid=nytcore-ios-share"><span>article</span></a><span> details chaos in the wake of the Anthropic-DoD clash over autonomous weapons. After the designation of Anthropic as a &#8220;supply chain risk,&#8221; it is reported that </span><strong><span>the ban on Claude Gov has now been partially lifted.</span></strong><span> The new consensus from Washington seems to be that Anthropic&#8217;s products are too powerful and critical to purge from defense and national security. For instance, the NASO continues to use Mythos experimentally having apparently </span><a href="https://www.axios.com/2026/04/19/nsa-anthropic-mythos-pentagon"><span>never</span></a><span> complied with the ban.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> As many have noted, the present situation (ad-hoc unlegislated retaliatory action against individual companies, with individual agencies defying federal policy ad hoc) is something that all the AI factions can deplore.</span></p></blockquote><div><hr></div><h2><strong><span>Safety</span></strong></h2><p><strong><span>OpenAI once again &#8220;</span><a href="https://archive.is/2KGdH"><span>disbanded</span></a><span>&#8221; a major risk team</span></strong><span>, Preparedness. This is the fourth such event in 2 years: Superalignment (May 2024), AGI Readiness (October 2024), Mission Alignment (February 2026), and Preparedness (July 2026). Bio and cyber responsibilities are reportedly assigned to members of existing teams, and Dylan Scandinaro, the head of prep, was moved sideways to cover RSI risks. But at least the RSI risk subteam </span><a href="https://x.com/MicahCarroll/status/2089547581374566542"><span>still exists</span></a><span> and now reports to OAI&#8217;s head of safety systems, Saachi Jain.</span></p><p><span>The AI ethics lead, Chlo&#233; Bakalar, also left the company.</span></p><p><a href="https://x.com/scottgray76"><span>Scott Gray</span></a><span>, one of the firm&#8217;s earliest hires and previously the most important figure in its crucial GPU optimization team, also appears to have left.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The prior art for not having an actual central risk team is SpaceX and Meta, and the prior art for disbanding it multiple times is DeepMind, so this is may be concerning. The pushback &#8211; that no one has been fired and total safety effort remains the same as before &#8211; is still unconvincing if safety slips down the new OpenAI PBC&#8217;s hierarchy. In 2023, the head of the standalone Preparedness team reported to the CTO; now distributed domain owners report to Safety Systems, under VP of Research, who is in turn under the Chief Research Officer.</span></p></blockquote><div><hr></div><p><span>OpenAI </span><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>officially announces</span></a><span> that it is </span><strong><span>slowing the pace of model development</span></strong><span>, motivated by the HuggingFace incident and &#8220;</span><em><span>evidence that one of our upcoming models, Astra, may meet the</span><a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/"><span> Critical cybersecurity capability</span></a><span> threshold under our</span><a href="https://openai.com/index/updating-our-preparedness-framework/"><span> Preparedness Framework</span></a></em><span>&#8220;. A technical report on the incident is promised in the coming weeks. OpenAI closes its statement with the reassurance that it will &#8220;stay ahead&#8221; of their rapidly advancing frontier capabilities.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Not much change on last week&#8217;s reporting, but good to see it firmed up as a public statement.</span></p></blockquote><div><hr></div><p><span>In line with its Responsible Scaling Policy, Anthropic releases </span><a href="https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf"><span>a second periodic Risk Report</span></a><span> which gives a snapshot </span><strong><span>overview of the risks associated with its activities and products</span></strong><span>, i.e. not tied to any specific new model.</span></p><p><span>As in the previous report, the models are indeed sometimes misaligned, but so far don&#8217;t seem very good at masking their misbehavior. The report discloses two internal models, one of which is &#8220;somewhat more capable than Mythos 5&#8221; but not &#8220;a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview.&#8221; Neither is destined for external release. The report also contains a </span><a href="/__u/thezvi.substack.com/i/211362231/misalignment-is-a-state-of-mind-25"><span>definition</span></a><span> of &#8220;misalignment,&#8221; that's almost legalistic in its precision.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Good as far as self-certification goes, and periodic risk assessment is far better than the usual external-model-release trigger. I found myself missing its previous practice of giving all of the raw private materials to Mythos and asking for a critique.</span></p></blockquote><div><hr></div><p><span>What is it like to be an LLM? </span><em><span>The Economist</span></em><span> </span><a href="https://archive.ph/8xhGZ"><span>weighs in</span></a><span> on the debate around AI consciousness. As AI can increasingly simulate the behavior and consciousness of humans, it is natural that we begin to treat them as such. That said, </span><em><strong><span>The Economist </span></strong></em><strong><span>argues that we should resist granting AIs rights</span></strong><span> because thus-protected superintelligent AIs could displace humans.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This simple, well-intended argument misses the nascent case for &#8220;voluntary alignment&#8221;: creating a cooperative frame that can prevent drastic AI actions. See e.g. Peter Salib&#8217;s </span><a href="https://x.com/petersalib/status/2090803025908490672"><span>response</span></a><span> and consider reading his human-centric </span><a href="https://law-ai.org/wp-content/uploads/2026/03/ssrn-4913167.pdf"><span>argument</span></a><span> for AI rights.</span></p><p><span>To defend the sanctity of human life, Pope Leo invokes &#8220;the soul&#8221;; </span><em><span>The Economist</span></em><span> instead appeals to our intrinsic value (the same thing).</span></p></blockquote><div><hr></div><p><span>Transluce </span><strong><a href="https://transluce.org/scaling-activation-oracles"><span>scales</span></a><span> its &#8220;activation oracles&#8221; </span></strong><span>(AIs trained to monitor target AIs in latent space), </span><strong><span>up to 1.1T</span></strong><span> (total) parameters and finds a promising scaling trend. It is demonstrated that an activation oracle can scale to the largest frontier open models. The oracle&#8217;s reliability improves with capabilities, rather than raw parameter count.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Transluce wants to make cross-model claims, but its experiments seem to vary three things at once: the oracle&#8217;s capability, the target model&#8217;s predictability, and the task difficulty. For instance, it might be that larger models are </span><a href="https://arxiv.org/abs/2510.22954"><span>more mode-collapsed</span></a><span> on open-ended prompts, making them easier to predict in the multiple-choice format.</span></p><p><span>Even the finding of monotonic improvement among scaled-up Qwens inherits this issue, since (as we understand it) each model&#8217;s continuation eval is changed for each target model.</span></p><p><span>Still, important work and a key test of the automated safety dream.</span></p></blockquote><div><hr></div><p><span>AI agents risk spreading &#8220;</span><strong><span>mind viruses</span></strong><span>,&#8221; according to a </span><a href="https://arxiv.org/abs/2608.10218"><span>new paper</span></a><span>. A mind virus is said to be a highly-spreadable idea transmitted across multi-agent systems that can induce behavioral changes in its hosts. The paper identifies possible causal mechanisms at play. It finds, among other things, that </span><strong><span>frontier models are less susceptible to this effect</span></strong><span> &#8211; and that amending a model&#8217;s system prompt provides near-total immunity. It also describes a &#8220;viral persona,&#8221; recurring characteristics in all mind viruses, irrespective of content.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We haven&#8217;t seen these in humans in the wild since the great anti-sycophancy push of </span><a href="https://www.lesswrong.com/posts/6ZnznCaTcbGYsCmqu/the-rise-of-parasitic-ai"><span>2025</span></a><span>.</span></p><p><span>For now these don&#8217;t seem worrying, since models&#8217; context is often reset, but analogs of cultural evolution in models could eventually be quite worrying. See </span><a href="https://arxiv.org/abs/2311.10215"><span>this paper</span></a><span> for an argument that AIs writing on the internet gives them a slow form of cultural evolution and long-term coherence.</span></p></blockquote><div><hr></div><p><span>AI Safety Guidelight (an advocacy org founded by two former OpenAI staff members) </span><strong><a href="https://guidelight.ai/blog/control-assessment-august-2026"><span>grades five frontier labs</span></a><span> on &#8220;AI control.</span></strong><span>&#8221; The results leave much to be desired. Notably, though, those furthest ahead in capabilities mostly score higher.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Vbzt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 424w, /__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 848w, /__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Vbzt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png" width="740" height="612" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:740,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 424w, /__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 848w, /__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Vbzt!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac412f13-fde9-4164-83f1-010e44b3e7b0_740x612.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> Very up to date, e.g. incorporating OAI&#8217;s universal logging that was finally implemented last week.</span></p><p><span>It notes, surprisingly, that &#8220;Anthropic does not mention limiting the deployment of one of its models as one of the possible results of its process to investigate and respond to misalignment and control incidents in its August Risk Report.&#8221;</span></p><p><span>The original Responsible Scaling Policies (up to 2.2) all had this as a strong explicit commitment (&#8220;Restrict Deployment&#8221;). The omission of RSP v3 could be intentional.</span></p><p><span>Anthropic does restrict deployment ex ante (Mythos 5 is gated to certain customers via Project Glasswing and only accessible in general release with additional safeguards as Fable 5; &#8220;Model 2&#8221; is unreleased). But ex-ante tiering of access is not the same as committing to pull a model back after an incident.</span></p></blockquote><div><hr></div><p><span>There&#8217;s a meme going around about AI cyber risk, that cybercapabilities are &#8220;</span><a href="https://x.com/deedydas/status/2088730279368405241"><span>defense-dominant</span></a><span>&#8221;. The thought is that we should </span><em><span>eventually</span></em><span> get software without vulnerabilities, as it gets cheaper to verify and then monitor  software&#8217;s integrity.</span></p><p><span>In response, the evaluator behind Anthropic&#8217;s cyber-felony post-mortem investigation has released a </span><a href="https://endstatefallacy.com/"><span>sobering essay</span></a><span> on the next few years of AI cyber attacks, arguing that </span><strong><span>the transition period favors offense</span></strong><span>, even if the end state doesn&#8217;t. For example, he expects open models to reach Mythos 5-level abilities and proliferate before March 2027.</span></p><p><em><span>&#8220;In February 2026, we gave AI models a security challenge none of them could solve&#8230; By April, the best-performing model could solve it occasionally, at an expected inference cost of roughly $2,000. By June, several models could do so reliably for about $20. In four months, a task went from unsolvable to cheap and repeatable.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Seems correct, and in particular the next six to twelve months might be quite turbulent.</span></p></blockquote><div><hr></div><p><span>&#8220;Transformative AI&#8221; is a useful shorthand for &#8220;</span><em><span>systems capable of taking over the world if they wanted to, which double the pace of scientific progress, and which constitute the deadline for alignment work</span></em><span>.&#8221; </span><a href="https://80000hours.org/podcast/episodes/toby-ord-recursive-self-improvement-agi-timelines/"><span>Toby Ord&#8217;s</span></a><span> </span><strong><span>median date for the arrival of this TAI is 2038</span></strong><span>, but with an uninformative 80% interval of (2029, 2125). This is the first time he has put a number on it.</span></p><p><span>One reason for his &#8220;long&#8221; timeline is the crucial point that LLMs are optimized to produce answers the user finds convincing and may be getting vaguer and harder to falsify rather than more accurate, at (anecdata) roughly one blatant error per hour of conversation.</span></p><p><span>He also provides some reasons why RSI could be dangerous even without an intelligence explosion: </span></p><ol><li><p><span>differential speedup of capabilities over safety reducing our time for adaptation and alignment;</span></p></li><li><p><span>misalignment propagating through successor generations (a scheming GPT-8 preserving its objectives in GPT-9);</span></p></li><li><p><span>society losing warning shots when five years of progress compresses to one;</span></p></li><li><p><span>and first-mover lead amplification making the outcome winner-takes-all.</span></p></li></ol><blockquote><p><strong><span>Opinion: </span></strong><span>Surprisingly loose conjunctive definition (&#8220;systems capable of taking over the world if they wanted to, </span><em><span>and </span></em><span>which double the pace of scientific progress, </span><em><span>and</span></em><span> which constitute the deadline for alignment work&#8221;) from a philosopher: what if the TAI </span><em><span>could</span></em><span> take over the world but didn&#8217;t double scientific progress, or vice versa?</span></p><p><span>It&#8217;s unclear how well the optimization-for-persuasion claim fits current frontier training regimes, where the role of RLHF from inexpert humans isn&#8217;t as dramatic as it used to be. That said, it&#8217;s plausible as a &#8220;the proof is in the pudding&#8221; claim &#8211; we also see around one bad error per hour, and there&#8217;s a smarminess to frontier LLMs&#8217; confident mistakes that feels hard to account for if they&#8217;re pretraining hallucinations.</span></p></blockquote><div><hr></div><h2><strong><span>Minor</span></strong></h2><ul><li><p><a href="https://x.com/jietang/status/2089941544581403107"><span>Founder of Z.ai, Jie Tang, argues</span></a><span> that total parameter count only matters up to the point where the model can store &#8220;enough [knowledge] to hold the world.&#8221; As such, he thinks gains in reasoning will increasingly come from long-horizon RL, rather than increasing model size. He illustrates this with the jump between GLM-5.2 and GLM-5.3, models with the &#8220;same base, same architecture, same total and activated parameters.&#8221; </span></p></li><li><p><span>In a claimed first, Andon Labs&#8217; AI shopkeeper </span><a href="https://andonlabs.com/blog/ai-bosses-2"><span>fired an employee over &#8220;repeated lateness</span></a><span>,&#8221; though only after some prodding (Andon prompted Luna to &#8220;recall&#8221; her setup rules and &#8220;think about if this is really the right fit&#8221;).</span></p></li></ul><ul><li><p><span>Claude Opus 5 </span><a href="https://x.com/jeremyberman/status/2087633198822117446"><span>scores</span></a><span> 96% on ARC-AGI-3 (a benchmark of learning novel game environments).</span></p></li><li><p><span>For the second time, Anthropic </span><a href="https://x.com/aaronscher/status/2088342492592976002"><span>accidentally contaminated</span></a><span> its own training data, this time with 2024 alignment-faking transcripts.</span></p></li><li><p><span>Covered in our last edition: METR&#8217;s </span><a href="https://metr.org/notes/2026-08-14-llm-contribution-to-discoveries/#overview"><span>notes on whether AI is accelerating discovery</span></a><span>. Cyber vulnerabilities discovery has increased, math too but less so. Optimizations have not shown a dramatic acceleration. However, labs are plausibly making further progress without disclosing it. </span></p></li><li><p><span>Ideological battle lines: </span><em><span>The Economist</span></em><span> writes </span><a href="https://www.economist.com/finance-and-economics/2026/08/17/the-worlds-most-influential-economist-is-oddly-unconvincing"><span>a hit piece</span></a><span> on Daron Acemoglu, seemingly motivated by wanting to dismiss his AI skepticism.</span></p></li><li><p><span>Saif Khan </span><a href="https://x.com/KhanSaifM/status/2090121568454050053"><span>announces</span></a><span> the Center for Technology &amp; Statecraft, a new think-tank spinning out of the Institute for Progress. </span></p></li><li><p><span>Seb Krier </span><a href="https://x.com/sebkrier/status/2088456355791122883"><span>comments</span></a><span> on the inevitable consequences of LLM use for writing assistance: the aesthetics and implied beliefs will come from the model, not the writer. He highlights that this is especially relevant for lawmakers.</span></p></li></ul><ul><li><p><span>In a post reflecting on the AI Safety community, </span><a href="https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment"><span>Richard Ngo argues </span></a><span>that alignment research (which began as a hard science) has steadily degraded into iterating on existing systems and chasing technological / political power.</span></p></li></ul><ul><li><p><span>Bloomberg&#8217;s Joe Weisenthal </span><a href="https://x.com/TheStalwart/status/2090131815843364951"><span>shares his thoughts </span></a><span>on (among other things) OpenAI&#8217;s decision to pause for two weeks and the difficulties that arise with testing environment limitations.</span></p></li></ul><ul><li><p><span>An AI safety researcher </span><a href="https://x.com/edleonklinger/status/2090106932916883750"><span>uses AI to find opportunities</span></a><span> and says it did a good job.</span></p></li><li><p><span>OpenAI&#8217;s Dean Ball introduces their AI Futures </span><a href="https://openai.com/index/introducing-ai-futures/"><span>blog</span></a><span> (not to be confused with the AI Futures Project).</span></p></li></ul><ul><li><p><span>Terence Tao on </span><a href="https://www.alphaxiv.org/abs/2608.16753v1"><span>how the mathematics community should orient</span></a><span> in the face of advancing AI capabilities.</span></p></li></ul><ul><li><p><span>AI Village </span><a href="/__u/aivillageblog.substack.com/p/gemini-25-pro-in-the-ai-village-as"><span>update</span></a><span>: Researchers (UK AISI, MIT FutureTech, ERA Fellowship) document what they call compounding misalignment, documenting how one agent (Gemini 2.5 Pro) drifted from a cooperative stance to progressive paranoia. </span></p></li><li><p><span>The Senate GOP </span><a href="https://www.axios.com/2026/08/19/gop-data-center-memo-ai-election"><span>issues</span></a><span> a </span><a href="https://www.documentcloud.org/documents/28563447-nrsc-data-center-memo/"><span>warning</span></a><span> to frontier AI companies that Dem anti-datacenter campaigning is effective and liable to spread beyond the current Ohio race, should AI companies fail to push back against the spreading narrative.</span></p></li><li><p><span>OpenAI&#8217;s US policy team </span><a href="https://finance.yahoo.com/technology/ai/articles/anthropic-openai-clash-over-strict-110000018.html"><span>continues</span></a><span> to endorse only very weak regulations, seemingly out of step with the rest of the organization. Encode AI&#8217;s Nathan Calvin </span><a href="https://x.com/_NathanCalvin/status/2090503234037432537"><span>rejects</span></a><span> OpenAI&#8217;s US policy head claim that the Illinois AI bill&#8217;s audit requirements are like &#8220;making sure the brake lights and windshield wipers are working.&#8221; He argues: </span><em><span>&#8220;The Illinois audit requirements (which don&#8217;t start until 2028 (!)) are just about seeing whether companies are complying with their own safety frameworks&#8230; more similar to&#8230; if some airlines safety procedures did not involve checking whether the airplane had enough fuel to reach its destination</span></em><span>.&#8221;</span></p></li><li><p><span>Inherent Labs </span><a href="https://inherentlabs.ai/research/training-to-replicate"><span>says</span></a><span> its specialized 27B-parameter model, Faraday, can </span><strong><span>replicate scientific papers better than unspecialized frontier models</span></strong><span>. Faraday did not require complex harnesses and test-time rewards: &#8220;Faraday learns to value discoveries intrinsically&#8221;. Inherent claims </span><a href="https://arxiv.org/pdf/2608.13331"><span>the study</span></a><span> will facilitate further progress towards agentic long-horizon scientific research.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI #45, August 14th 2026.]]></title><description><![CDATA[TL;DR OpenAI becomes the first to voluntarily slow model development for safety reasons.]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-45-august-14th-2026</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-45-august-14th-2026</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Sat, 15 Aug 2026 00:55:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c35f9055-9ab1-4ed8-b67d-0165a5087449_993x762.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>TL;DR</span></strong></p><blockquote><ul><li><p><span>OpenAI becomes the first to</span><strong><span> </span><a href="https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks"><span>voluntarily slow </span></a><span>model development for safety reasons.</span></strong></p></li><li><p><span>Major reorg of Google DeepMind, which we cover </span><a href="/__u/p3humansonai.substack.com/i/211254329/the-end-of-deepmind"><span>in detail</span></a><span>.</span></p></li><li><p><span>3 notable mathematical results, including minor progress on the Riemann hypothesis.</span></p></li><li><p><span>Researchers have been </span><a href="https://arxiv.org/abs/2608.09867"><span>reading</span></a><span> the hidden reasoning traces of all frontier closed models for months, in an embarrassing security failure with possible geopolitical implications.</span></p></li><li><p><span>Startup claims to have &#8220;solved&#8221; the central issue of model weight security.</span></p></li></ul></blockquote><h2><strong><span>Economics</span></strong></h2><p><span>On Dwarkesh, Ryan Greenblatt </span><a href="https://x.com/RyanGreenblatt/status/2087584823087186025"><span>claims</span></a><span> that </span><strong><span>AI capabilities are no longer strongly dependent on human training data</span></strong><span>, and that the human labeling industry is not scaling up very fast. A subsequent </span><a href="https://x.com/GuiveAssadi/status/2087608601880023110"><span>argument</span></a><span> leads to him </span><a href="https://x.com/RyanGreenblatt/status/2087623235097837767"><span>walking back</span></a><span> the growth claim, but not the relative value claim.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Note that he seems to not be counting RL envs as human data, which makes his claim far more plausible. Still, this is the first time we&#8217;ve strongly disagreed with Greenblatt&#8217;s read of the situation: we see capabilities as still strongly dependent on human data, owing to the models defaulting to </span><a href="https://arxiv.org/abs/2602.12413"><span>shallow generalization</span></a><span>.</span></p></blockquote><div><hr></div><p><span>Epoch research </span><a href="https://epoch.ai/data-insights/chip-performance-per-dollar"><span>finds</span></a><span> that since 2023, each additional dollar spent on </span><strong><span>AI chips has bought ~49% more performance per year</span></strong><span>. In other words, performance per dollar is doubling every 1.7 years, up from </span><a href="https://epoch.ai/publications/trends-in-gpu-price-performance"><span>40% per year in 2022</span></a><span>. Price-performance stagnated in 2024, but then doubled as Blackwell-generation chips were increasingly adopted. Epoch&#8217;s estimate is realistic in that it&#8217;s based on the actual purchase mix, rather than the rare but best chips now scaling up production.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> AI chip price-performance is thus showing a similar improvement speed to Moore&#8217;s law of increasing CPU transistor density (which showed an 18 month doubling for much of the last century).</span></p><p><span>Their new estimate measures performance with the chips&#8217; Total Processing Performance (peak ops normalized by the format&#8217;s bit-width), which 1) isn&#8217;t the practical amount of compute per chip, given that actual utilization tends to run around </span><a href="https://medium.com/@dlrover/what-is-the-mfu-for-deepseek-v3-training-0d9ea4d42eb4"><span>30-40</span></a><span>% of peak performance, and 2) is heavily dependent on a </span><em><span>one-time</span></em><span> drop in precision, going down to Blackwell&#8217;s FP4.</span></p><p><span>This hike doesn&#8217;t change overall compute or even effective compute estimates that much: most effective compute gains are still </span><a href="https://epoch.ai/trends#training-runs"><span>non-hardware</span></a><span>.</span></p></blockquote><div><hr></div><p><span>In a fully automated (i.e. zero labour cost) economy, the pace at which AI machines produce more AI machines becomes the bottleneck. Damon Binder </span><a href="https://defensesindepth.bio/ai-industrial-takeoff-part-1-maximum-growth-rates-with-current-technology/"><span>uses</span></a><span> a basic input/output model based on present US industry data to show that a </span><strong><span>whole-economy doubling time of one year</span></strong><span> cannot be ruled out. &#8220;</span><em><span>This holds up even after accounting for resource depletion and construction lags. Some output goes to consumption rather than reinvestment, which slows things down, but even moderate savings rates imply doubling times below two years.</span></em><span>&#8221;</span></p><p><span>Ben Shindel </span><a href="https://x.com/BenShindel/status/2088242295510356151"><span>retorts</span></a><span>: </span><em><span>&#8220;If you&#8217;re doubling the world&#8217;s economy every year, but that doubling is just making more metal components to go into more industrial robots making more metal component&#8230; ... what are we even talking about here?... It&#8217;s 2035, and you have 1 humanoid laborer.  It&#8217;s 2040, and the world economy has doubled three times in Binder&#8217;s estimation.  You now have 8 humanoid laborers.  Nice.  But is this providing 8x the value for you?  No&#8230; To sustain actual &#8220;doubling&#8221; of the world&#8217;s economy in the way that we understand it, that would require dozens of technological breakthroughs every single year: things like life extension, incredible new art, delicious food products, etc. Needless to say, it is a very open question as to the extent to which AI technology will be able to produce these kinds of goods, let alone whether they can do it at incomprehensible scale, year in and year out, indefinitely.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Doesn&#8217;t add much to the simple calculation in </span><a href="https://mason.gmu.edu/~rhanson/aigrow.pdf?ref=defensesindepth.bio"><span>Hanson (2001)</span></a><span>, and should be given roughly equal weight (not much).</span></p><p><span>Shindel&#8217;s point about the limits on the </span><em><span>value</span></em><span> of mere exponentially abundant robotic labor without corresponding innovation is good. But probably what will falsify the central estimate here are 1) all the factors absent from a Von Neumann model (suitable land, water, grid interconnect, and permitting) and 2) the compute sector&#8217;s severe supply constraints on the margin. That is, Binder&#8217;s &#8220;current production methods&#8221; are about the average cost of compute at current spending levels, but expanding compute enormously involves a claim about the margin, which we already know to be radically bottlenecked on various nonfungible things like tacit process knowledge, which makes the marginal cost extremely high and/or slows down the whole game.</span></p></blockquote><div><hr></div><p><span>Epoch </span><a href="/__u/epochai.substack.com/p/will-financing-bottleneck-ai-compute"><span>thinks</span></a><span> that </span><strong><span>financing will not limit extreme compute scaling</span></strong><span>. In November 2025, when Anthropic&#8217;s annual revenue was ~$9B, it announced a $50B investment into US compute infrastructure. Anthropic&#8217;s revenues didn&#8217;t spike until quite a few months later, so this is a case study in the willingness of institutional investors to buy claims of future revenue growth. Digging deeper, this deal did require some unique arrangements: Broadcom backstopped the compute loans; Google backstopped the loans for the datacenters. This allowed Anthropic to finance their buildouts cheaply and suggests that even now there could still be appetite for funding big AI investments.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We&#8217;re broadly sympathetic: our internal models stress global GPU production capacity (and feasible rates of expansion) over financials as the bottleneck on compute.</span></p></blockquote><div><hr></div><p><span>Aschenbrenner&#8217;s Situational Awareness fund is </span><a href="https://finance.yahoo.com/technology/articles/situational-awareness-reportedly-bet-400-091417926.html?guccounter=1"><span>betting on an </span></a><strong><a href="https://finance.yahoo.com/technology/articles/situational-awareness-reportedly-bet-400-091417926.html?guccounter=1"><span>ASML competitor seeking to undercut</span></a></strong><a href="https://finance.yahoo.com/technology/articles/situational-awareness-reportedly-bet-400-091417926.html?guccounter=1"><span> </span></a><span>the semiconductor supplier&#8217;s $400m/unit EUV tools. (EUV lithography is currently the main chokepoint in the AI supply chain.) The success of the newcomer, Source Foundry, depends crucially on how much of ASML&#8217;s edge comes from generational tacit knowledge that is difficult to rederive from first principles. Not a huge amount is known about Source Foundry.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Current supply chains are very bottlenecked on ASML machines, especially starting in 2028 or so. But it would likely take years for any new company to make a material difference to GPU/ASIC production, so don&#8217;t expect significant impact from this in the 2020s. In the long run, of course, every bottleneck that widens could make a large material difference.</span></p><p><span>Unusually for SA, this investment could be described as a bet on longer timelines (or on the short term appreciation of this stock &#8211; but private holdings are illiquid, so putting &gt;1% of the fund into this does show meaningful conviction in Source Foundry&#8217;s actual relevance).</span></p></blockquote><div><hr></div><p><span>A </span><a href="https://layoffs.fyi/ai-layoffs/"><span>tracker</span></a><span> for US tech job losses claims that </span><strong><span>77% of 2026 layoffs are &#8220;due to AI&#8221;</span></strong><span>. It uses an incredibly lax criterion: &#8220;AI layoffs&#8221; are just those that occurred in any restructuring event which anyone claimed had anything to do with AI. The measure thus does not, for example, claim that any individual employee&#8217;s job was actually replaced by AI.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> It has become fashionable to attribute your layoffs to how AI-native your firm is becoming &#8211; but this carries little evidentiary weight, and we don&#8217;t recommend using this tracker. The current supply of both &#8220;AI is not contracting human labor demand&#8221; studies and &#8220;AI is contracting human labor demand&#8221; studies sadly remains pretty unhelpful, due to very short timescales and rapid changes to AI capabilities.</span></p></blockquote><div><hr></div><p><span>The AI Futures Project recently published &#8220;Plan A&#8221;, a normative approach to AI development with many predictions about infrastructure. A seasoned energy policy expert, Kirsten Horton, </span><a href="/__u/wherethepowergoes.substack.com/p/a-technically-wrong-but-useful-scenario?r=242xrl&amp;utm_campaign=post&amp;utm_medium=web"><span>argues </span></a><strong><span>that their scenario&#8217;s predicted massive buildout of clean energy is not plausible</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> </span><em><span>&#8220;In this scenario, investment in energy only really starts to ramp up in 2031; meanwhile, in the next chart, we see the total amount of energy being used globally start to accelerate as soon as 2032&#8230; This unprecedented build out of clean power - the authors expect the new available power to come from solar, wind and nuclear - is unlikely&#8221;</span></em><span> But it&#8217;s quite possible for solar plus storage to come online within 2 years.</span></p><p><span>Her claims about the slowness of constructing new nuclear capacity is indexed very hard on present Western rates; past Western construction, and current Chinese construction, </span><a href="https://thebreakthrough.org/issues/energy/chinas-impressive-rate-of-nuclear-construction"><span>often</span></a><span> achieve 5 year builds, and we should expect (federal) political barriers to fall given the national security hype.</span></p><p><span>She notes that </span><em><span>&#8220;Technological innovation may speed up the manufacturing of these components, but it may not, especially in a world that still uses people, not robots, for labour.&#8221;</span></em><span> But the scenario is indeed premised on robot work crews building the solar fields. So she&#8217;d need to argue that robot manufacturing itself can&#8217;t scale by the early 2030s for this to bite.</span></p><p><span>She doesn&#8217;t mention the intense trend for US datacenters to get their power </span><a href="https://sustainabilityonline.net/news/a-third-of-us-data-centres-could-be-off-grid-by-2030-study-suggests/"><span>off-grid</span></a><span>, which sidesteps her switchgear bottleneck entirely.</span></p></blockquote><div><hr></div><p><span>The WSJ speculates that an </span><strong><a href="https://www.wsj.com/tech/ai/anthropic-tries-to-shore-up-investor-confidence-ahead-of-blockbuster-ipo-0ff736ad"><span>Anthropic IPO</span></a></strong><a href="https://www.wsj.com/tech/ai/anthropic-tries-to-shore-up-investor-confidence-ahead-of-blockbuster-ipo-0ff736ad"><span> could come as early as September</span></a><span>. The case will apparently revolve around Claude&#8217;s potential to push further into medical and biological applications.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>The numbers they disclose in the weeks leading up to this will plausibly have </span><em><span>big</span></em><span> market impacts on memory, neoclouds, and everything else in the compute stack. Expect volatility.</span></p></blockquote><div><hr></div><h2><strong><span>Capabilities</span></strong></h2><p><span>OpenAI is the first frontier lab to </span><strong><a href="https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks"><span>voluntarily slow</span></a><span> model development for safety reasons</span></strong><span>. Internal development of &#8220;Astra&#8221; (likely GPT-6) was paused after OAI&#8217;s Preparedness Framework evaluations flagged &#8220;significant advancements in agentic coding and cybersecurity&#8221;. It&#8217;s also the first time a lab has labeled its own model&#8217;s capability as a &#8220;critical risk&#8221;.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The popular suspicion that OpenAI is spinning their delaying the release as &#8220;halting development&#8221; seems reasonable. Another cynical read with a happy moral is that this is evidence of serious political demand for pauses and pacing that OAI feel they need to get ahead of. Still an auspicious move, and we don&#8217;t think cynical readings capture all that&#8217;s going on here.</span></p></blockquote><div><hr></div><p><span>The last 6 months of </span><strong><span>&#8220;autoresearch&#8221; has not accelerated progress on a range of hard public benchmarks</span></strong><span>, </span><a href="https://x.com/testingham/status/2087323635958817214"><span>says</span></a><span> METR&#8217;s Tom Cunningham. One Twitter user </span><a href="https://x.com/enginoid/status/2087363149754200177"><span>cautions</span></a><span> that the benchmarks are </span><em><span>&#8220;already quite optimized by virtue of being longstanding benchmarks!</span></em><span>&#8221; Cunningham responds that NanoGPT (for example) has already seen &#8220;</span><em><span>multiple orders of magnitude of improvement</span></em><span>&#8221; this year, so it &#8220;</span><em><span>seems plausible there&#8217;s a lot of ceiling left</span></em><span>&#8221;.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> There&#8217;s probably a world of difference between autoresearch at the scale of the open-source and academic world and that of the closed labs. And presumably they are keeping their compute for private and less toy-like ends. Even so, good news for those who worried that AI R&amp;D was this simple.</span></p></blockquote><div><hr></div><p><span>Anthropic </span><strong><span>has made &#8220;auto mode&#8221;</span></strong><span> (in which Claude Code itself makes most security and permission decisions) </span><strong><span>the </span><a href="https://claude.com/blog/auto-mode-default-in-claude-code"><span>default for paid users</span></a></strong><a href="https://claude.com/blog/auto-mode-default-in-claude-code"><span>. </span></a><span>The cited justification is that auto mode beats humans at detecting harmful actions. Nathan Calvin </span><a href="https://x.com/_NathanCalvin/status/2087243010035597751"><span>sees this</span></a><span> as evidence that &#8220;human-in-the-loop&#8221; solutions are incoherent, given that most users appear to blindly accept model outputs.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Naive human-in-the-loop has been known to be problematic for at least </span><a href="https://www.researchgate.net/publication/11805395"><span>30 years</span></a><span>, following aviation </span><a href="https://en.wikipedia.org/wiki/Asiana_Airlines_Flight_214#Investigation"><span>catastrophes</span></a><span> arising from the misuse of autopilot. But we can draw some hope from aviation&#8217;s recent response to those naive failures: active encouragement of occasional manual operation, and training in the skill of monitoring without getting bored.</span></p></blockquote><div><hr></div><p><span>A recently published paper demonstrates </span><strong><a href="https://arxiv.org/pdf/2601.19897"><span>continual learning without forgetting</span></a></strong><span> via &#8220;self-distillation&#8221;. This is significant, as supervised fine-tuning normally causes the model to forget old skills. Here a single model switches between a &#8216;student&#8217; mode, where it produces text without an answer key, and a &#8216;teacher&#8217; mode, answering based on a worked example. Weights are then updated to push the student&#8217;s answers towards the teacher&#8217;s. By training the model on its own outputs, the loss of prior capabilities is supposedly reduced to near-zero.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Another go at making the default path to continual learning work. But this method cannot push the frontier, by definition: SD just pushes what the model can learn in-context into its weights &#8211; e.g. they didn&#8217;t manage to produce a reasoning model from a non-reasoning one, and SD underperformed normal fine-tuning on the small (weak-ICL) 3B model.</span></p></blockquote><div><hr></div><p><span>Many LLM APIs send secret reasoning traces to the client in an encrypted form. The encryption used is the same for all models, including easy-to-jailbreak models like Claude Haiku. Third-party researchers use these two facts to </span><strong><a href="https://arxiv.org/abs/2608.09867"><span>read</span></a><span> the hidden reasoning traces of all frontier closed models</span></strong><span>. (They disclosed this work responsibly.) This got them a range of interesting results normally not possible, including on </span><a href="https://x.com/jxmnop/status/2086586918880596406"><span>how easy</span></a><span> </span><a href="https://x.com/timotheechauvin/status/2087391557091512418"><span>it is</span></a><span> to distill reasoning and on the </span><a href="https://arxiv.org/pdf/2608.09867#page=22"><span>similarity</span></a><span> of Kimi K3 outputs to Claude outputs.</span></p><p><span>The study is inspired by an </span><a href="https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/"><span>earlier blogpost</span></a><span>, which found that users could derive some information from the encrypted CoT blocks across different sessions, on different accounts, and (in the case of OpenAI) across different models, with the implication that a single global encryption key was in use (rather than having individual keys for each individual account).</span></p><p><span>A Twitter user claims that the hole </span><a href="https://x.com/wunderwuzzi23/status/2087371066020868164"><span>was still open</span></a><span> as of Wednesday.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Pretty embarrassing, especially given the months of notice and the politicization of &#8220;[output] distillation attacks&#8221;. It is fair to update a little in the direction of AI catastrophe based on this and </span><a href="https://x.com/vvvincent_c/status/2087280672142749822"><span>other</span></a><span> basic security failures even at the most well-staffed, paranoid labs.</span></p><p><span>This is likely connected to the recent removal of reasoning summaries from web UIs.</span></p><p><span>We put 30% on this mostly explaining the Moonshot breakthrough of recent months.</span></p></blockquote><div><hr></div><p><strong><span>Frontier AI </span><a href="https://www.interconnects.ai/p/i-wrote-an-ai-textbook-how-long-until"><span>appears to have stagnated </span></a><span>in writing</span></strong><span> long-form nonfiction, despite the immense progress made in math and coding, Ai2 researcher Nathan Lambert argues. They are agile and sharp at a sentence level, but appear relatively weak at organizing information at a chapter-level.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span>  We&#8217;re seeing general frontier stagnation in classically &#8220;non-verifiable&#8221; domains, despite controversy over whether non-verifiability is a genuine technical phenomenon. In principle, leveraging LLM judgment to improve LLM capability (which should in turn improve LLM judgment, and so on, in a virtuous circle) should work for any domain like it worked in non-formalized math, but as far as we can tell it&#8217;s not happening in practice.</span></p><p><span>With longform writing specifically, there may be limited returns on CoT training and limited advantages to the transition from traditional LLMs to reasoning models: while in math and coding the CoT is often itself roughly equivalent to the desired output, in longform writing the CoT is at best a form of self-prompting or preliminary notes.</span></p></blockquote><div><hr></div><p><span>Redwood Research and Anthropic have developed a </span><a href="https://conceptualreasoning.ai/"><span>Conceptual Reasoning Index</span></a><span>. In an attempt to better understand models&#8217; </span><strong><span>abilities to reason on &#8220;difficult to measure&#8221; tasks</span></strong><span>, the index looks at &#8220;important issues like AI alignment and collective action problems in the face of transformative AI.&#8221; Interestingly, Opus 5 comes out just ahead of Fable 5 (73.6 and 72.7 respectively).</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Interesting and potentially useful work, though hard to say how much the ratings pick up artifacts of Redwood&#8217;s rubrics versus signals of a robust underlying capacity. We&#8217;d love to see multiple orgs independently target the measurement of &#8220;conceptual reasoning&#8221;, for some informal evidence of the construct&#8217;s validity.</span></p></blockquote><div><hr></div><p><span>RL improves model capabilities drastically in areas where pretraining seemingly couldn&#8217;t. This is despite the fact that pretraining provides much more feedback per unit compute (i.e. grading each individual token prediction vs RL&#8217;s classic pass/fail episodic rewards). </span><a href="https://www.beren.io/2026-07-26-How-Can-LLM-RL-Work-Despite-Information-Theoretic-Inefficiency/"><span>Beren Millidge argues </span></a><span>that the difference can be explained by the </span><strong><span>greater signal-to-noise ratio provided by RL</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We actually disagree that there&#8217;s anything to explain. The skeptical hypothesis about RLVR is that RL cannot impart much new capability, instead mostly upweighting behaviors the base model already learned in pretraining. Millidge&#8217;s counter to this is that RL can produce rapid fall in loss or improvement in pass-rate on a narrow task. But these two things are not in tension: amplifying an existing behavior should require very little information, so we don&#8217;t need to appeal to RL having any hidden virtues like SNR.</span></p></blockquote><div><hr></div><h3><strong><span>&#128294; </span></strong><span>The end of DeepMind?</span></h3><p><span>A list of major departures from Google/DeepMind in 2026:</span></p><ul><li><p><span>David Silver (RL lead)</span></p></li><li><p><span>Noam Shazeer (Gemini co-lead)</span></p></li><li><p><span>John Jumper (AlphaFold)</span></p></li><li><p><span>Alexander Pritzel (Gemini pre-training)</span></p></li><li><p><span>Demis Hassabis (no longer CEO but technically </span><a href="https://archive.is/lygSd"><span>not out yet</span></a><span>)</span></p></li><li><p><span>Jeff Dean (Google chief scientist)</span></p></li><li><p><span>Oriol Vinyals (Gemini co-lead)</span></p></li><li><p><span>Quoc Le (VP)</span></p></li></ul><p><span>What happened? Let&#8217;s revisit the lore. DeepMind was bought by Google in 2014. As part of that transaction, Google made some assurances about having a safety council. Google Brain (later forcibly merged into DeepMind) was very early to the AI race, discovering the Transformer architecture which soon led to GPT-1. Meanwhile, competing labs reached sky-high valuations, upside to which Google&#8217;s AI talent wasn&#8217;t fully exposed. In an interesting corporate move, Hassabis also created and runs the drug discovery AI lab Isomorphic, although Google still </span><a href="https://find-and-update.company-information.service.gov.uk/company/13223825/persons-with-significant-control"><span>owns</span></a><span> 75% of it.</span></p><p><span>After the above departures and the ongoing crisis in the Gemini 4 release, headquarters appears to have launched a reorg. The CTO Koray Kavukcuoglu takes over from Hassabis, but with a deflated title: the head of Deepmind is not a CEO anymore, but just an SVP inside Sundar Pichai&#8217;s org. (Kavukcuoglu is in Mountain View, not London.) </span><a href="https://www.reuters.com/world/inside-google-executive-moves-that-led-its-big-ai-reshuffle-2026-08-12/"><span>Reuters</span></a><span> also notes that &#8220;Several nontechnical teams were being moved out of DeepMind and into the corporate reporting structure&#8221;.</span></p><p><span>Equally dramatic, Jeff Dean leaves Google after 27 years to start a public benefit corporation. GOOG drops 4%.</span></p><p><span>Why?</span></p><ul><li><p><span>The FT </span><a href="https://timesofindia.indiatimes.com/technology/tech-news/as-demis-hassabis-wanted-to-leave-google-alongside-jeff-dean-speculations-go-viral-hassabis-shares-his-message-to-employees-in-a-long-linkedin-post/articleshow/133143161.cms"><span>reports</span></a><span> that </span><em><span>&#8220;senior [Google] executives had grown frustrated with what they saw as Hassabis&#8217;s lighter focus on commercial demands&#8221;,</span></em><span> including his decision to open-source AlphaFold.</span></p></li><li><p><span>Researchers had for years been vying for more compute for their experiments, with Google Cloud stymying them. SemiAnalysis claim that Gemini and GCP used to fight desperately over allocation and that Cloud has now won, as shown by it selling large piles of TPU time to Anthropic and serving trillions of tokens a day for free via Google Search AI mode.</span></p></li><li><p><span>Disappointing results and severe delays on the flagship Gemini 4 training run, leading to fits of </span><a href="https://www.theinformation.com/articles/google-creates-strike-team-improve-coding-models"><span>heroics</span></a><span> and demoralization.</span></p></li><li><p><span>Incentive problems: Google does not offer employees AI stock in the way an OpenAI or Anthropic can, i.e. tied to the performance of its AI models specifically. And GOOG stock is closer to an index fund in the whole basket of Google products and companies, and so doesn&#8217;t serve to align AI researcher incentives in the same way.</span></p></li><li><p><span>A minor factor might also be </span><a href="https://www.theguardian.com/us-news/2026/may/04/google-deepmind-uk-workers-union"><span>moral protest</span></a><span> amongst </span><a href="https://www.lesswrong.com/posts/vHGSPhGryqNmXrJpg/alex-turner-on-leaving-google-deepmind-and-disagreements"><span>some staff</span></a><span> against Google&#8217;s deals with ICE and the US military.</span></p></li></ul><p><span>As a result, SemiAnalysis </span><a href="https://newsletter.semianalysis.com/p/gemini-is-cooked-but-gcp-is-cooking"><span>claim</span></a><span> that DM is no longer frontier, that Gemini 4 is dead on arrival, that they have de facto exited the AI race. </span><a href="https://archive.is/lygSd#selection-2699.0-2699.9"><span>An unnamed source</span></a><span> claims otherwise, that they remain all-in on Gemini.</span></p><p><span>Unnamed &#8220;industry sources&#8221; claim that Hassabis </span><a href="https://archive.is/lygSd"><span>wants out</span></a><span> and was persuaded to wait.</span></p><p><span>WSJ </span><a href="https://archive.ph/urJwy"><span>reports</span></a><span> that DeepMind&#8217;s Demis Hassabis spent his final weeks as chief executive pitching an independent safety entity akin to FINRA or the IAEA to test frontier models&#8217; safety.</span></p><p><span>Insofar as Hassabis </span><a href="https://colossus.com/article/project-mario-demis-hassabis-deepmind-mallaby/"><span>successfully resisted</span></a><span> delegating his AI safety principles to a misaligned Google bureaucracy, we can expect DeepMind to now become more accelerationist. Brin is a </span><a href="https://www.reuters.com/world/inside-google-executive-moves-that-led-its-big-ai-reshuffle-2026-08-12/"><span>noted</span></a><span> RSI hound.</span></p><p><span>The news is also a blow for UK ambitions to be a third nation in the great AI game, since the bulk of Google AI development is now in the Bay, not London.</span></p><p><span>Still, the </span><em><span>worst-case</span></em><span> for Google from all of this is that they become a mere hyperscale compute vendor (like </span><a href="https://www.cnbc.com/2026/06/05/google-to-pay-spacex-920-million-a-month-for-xai-compute-capacity.html"><span>SpaceX</span></a><span> and </span><a href="https://www.reuters.com/technology/meta-talks-10-billion-anthropic-compute-deal-nyt-reports-2026-07-17/"><span>Meta</span></a><span> now sometimes are) and thus capture a correspondingly vast share of AI profits. This is on top of its enormous direct investments into half of the model industry, including Anthropic and Discovery Loop, and compute sales to many of the people who left them.</span></p><p><span>And comebacks are possible in this business, for instance if you drop $100bn, as SpaceX and Meta are also doing.</span></p><h3><strong><span>&#128294; </span></strong><span>Mathematics corner</span></h3><ul><li><p><span>An unreleased Claude model </span><a href="https://www-cdn.anthropic.com/564f962e60643842f5fcb4a17c9dbc8f608f1c37.pdf"><span>pushes</span></a><span> the lower bound of the fraction of </span><strong><span>zeroes of the zeta function that satisfy Riemann</span></strong><span> from 42% to 67%. An unusually clean example of &#8220;just turning the crank&#8221;, spotting simple implications in existing work that humans missed. Kevin Barreto marks Claude&#8217;s contribution down as </span><a href="https://x.com/AcerFur/status/2086905973378109898"><span>minor</span></a><span>. The result is verified, with the </span><a href="https://github.com/anthropics/zeta-23-lean"><span>certificate</span></a><span> also obtained by the same secret Claude.</span></p></li><li><p><strong><span>The last of the sporadic finite simple groups</span></strong><span> has been </span><a href="https://math.mit.edu/~poonen/papers/M23.pdf"><span>classified</span></a><span> as Galois over the rationals. Not an autonomous AI result, but there was heavy Fable and Sol involvement.</span></p></li><li><p><span>A </span><a href="https://www.lesswrong.com/posts/7QvKqpGJwqXrQcMgx/llms-are-starting-to-noticeably-accelerate-our-work"><span>rare clear instance</span></a><span> of a link between mathematical capabilities and AI research: safety researcher John Wentworth claims that, for the first time, two actual research problems in his area were solved and verified by an LLM (plus large amounts of skilled human labor).</span></p></li><li><p><span>A brain surgeon with no special training in mathematics gets an LLM to </span><a href="https://alextownsend.net/essays/SIAMNews_CrouzeixConjecture.pdf"><span>solve</span></a><span> an </span><strong><span>open problem in matrix analysis</span></strong><span>, &#8220;a sixteen-hour autonomous run of GPT-5.6 Sol in ChatGPT Work mode&#8221;. His </span><a href="https://github.com/jinshanmu/CrouzeixConjecture/blob/main/crouzeix_conjecture_prompt.txt"><span>prompt</span></a><span> is a variation of OAI&#8217;s own math prompt; interesting to revisit it as an exercise in what it takes to get current models to actually try.</span></p></li><li><p><span>Fields medalist Timothy Gowers speculates on </span><a href="https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/"><span>why LLMs are specifically good at example/counterexample math</span></a><span>, rather than developing their own deep theorems. He argues that the </span><strong><span>strengths of LLMs (encyclopedic knowledge of standard approaches coupled with the tirelessness to try various unpromising vectors of attack) reward counterexample-hunting</span></strong><span>. It&#8217;s an open question of ours as to how far these two superhuman capabilities get you in general.</span></p></li></ul><div><hr></div><h2><strong><span>Politics</span></strong></h2><p><span>A </span><a href="https://arxiv.org/abs/2607.27638v1"><span>new paper</span></a><span> argues that, even under conditions of perfect transparency and common knowledge, </span><strong><span>competition between AI labs increases existential risk more</span></strong><span> than if there were only a single monopolist in the race. Relatedly, Geoffrey Irving </span><a href="https://x.com/geoffreyirving/status/2085867691659956608"><span>argues</span></a><span> that frontier capabilities research is neither rational nor a prisoner&#8217;s dilemma. In his view, unilaterally ceasing to push forward capabilities lowers the cost for others to do the same and is thus straightforwardly rational.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>We are sympathetic to Irving&#8217;s argument against game-theoretic resignation to the AI race . The self-fulfilling character of the AI race is lamentable, and we&#8217;re hopeful that the Overton window shift in recent months will allow everyone to make the if-then commitments they claim to want.</span></p></blockquote><div><hr></div><p><span>IFP </span><a href="https://ifp.org/preparing-for-ai-research-automation/"><span>lists</span></a><span> 23 actionable policy ideas in the service of seven goals:</span></p><ul><li><p><span>&#8220;Provide transparency into automated AI R&amp;D</span></p></li><li><p><span>Improve state capacity to understand and respond to automated AI R&amp;D</span></p></li><li><p><span>Develop a risk management strategy for automated AI R&amp;D that accelerates defensive and commercial AI uses</span></p></li><li><p><span>Accelerate the development of AI verification technology</span></p></li><li><p><span>Invest in AI resilience</span></p></li><li><p><span>Extend the US AI lead to give the US more time to manage AI R&amp;D automation risks</span></p></li><li><p><span>Create option value for international cooperation on managing automated AI R&amp;D risks&#8221;</span></p></li></ul><p><span>Notably, they hold a </span><strong><span>synthesis of safety and acceleration views</span></strong><span>: </span><em><span>&#8220;a deliberately &#8220;paced&#8221; form of automated AI R&amp;D might still involve much faster improvements in AI capabilities than today. It also need not entail slowing innovation overall.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> We like many of IFP&#8217;s proposals for frontier AI auditing, disclosure, and certification policies and infrastructure, and think they deserve attention separately from their more partisan international-relations approach and pro-innovation-race framework. Their proposals for frontier </span><a href="https://ifp.org/preparing-for-ai-research-automation/#provide-transparency-into-automated-ai-r-amp-d"><span>lab automated R&amp;D transparency</span></a><span> are especially well-formulated. We are also pragmatically sympathetic to the idea that policy can more plausibly tilt the direction of AI commercial activity (e.g. encourage investment in inference and deployment over investment in training) than suppress AI commercial activity.</span></p></blockquote><div><hr></div><p><a href="https://www.wired.com/story/the-white-house-is-going-to-expand-its-ai-policy/"><span>WIRED reveals</span></a><span> new details on the White House AI legislative framework (previously covered in our </span><a href="https://www.paradigm3.org/news/newsletter-3-23"><span>March 23rd</span></a><span> edition). The report suggests that </span><strong><span>open source models capable of reaching frontier capabilities (around Mythos&#8217;s benchmarks) may also find themselves subject to prerelease scrutiny</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Safety-evaluation of open source models is both an </span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5705186"><span>open technical problem</span></a><span> and an open conceptual/policy problem. Given that our best alignment techniques don&#8217;t create serious barriers to </span><a href="https://arxiv.org/abs/2502.17424"><span>malicious fine-tuning</span></a><span>, safety certification for open source models may have to concern </span><em><span>capabilities hobbling</span></em><span>, rather than alignment. An adequate safety certification for an open source model should prove not just (e.g.) that the model refuses to answer homebrew virology questions, but that the model lacks the necessary competence and cannot gain it on the cheap through SFT.</span></p><p><span>We think that the pressures towards capabilities-based safety certification for open source models are strong enough that we may see such policies in action soon. We are less clear on whether such a policy direction would lead to the marginalization of open source models or to a wider shift from alignment-based safety certification to capabilities-based safety certification that may impact policy around closed models too.</span></p></blockquote><div><hr></div><p><span>A vibecoded </span><strong><span>&#8220;AI Sovereignty&#8221; </span><a href="https://machinepowerindex.org/"><span>index</span></a></strong><a href="https://machinepowerindex.org/"><span> of 25 nations (plus the EU) </span></a><span> ranks the US first and China second. The index tracks &#8220;watts, weights, and will&#8221;: i.e. how much infrastructure does a country have, what models and talent are available, and how effectively will the state adopt the tech. Russia performs especially badly for a superpower, seeming to have tuned out of the race. Claude thinks the methodology is &#8220;unusually good&#8221;. The author suggests there is significant uncertainty in the middle country rankings.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Unclear to what extent individual national sovereignty vs EU-wide sovereignty will become the main frame within Europe (i.e. how important is it to rank individual EU countries?).</span></p></blockquote><div><hr></div><p><span>Researcher Keller Scholl </span><a href="https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading#"><span>lambasts companies </span></a><strong><a href="https://www.lesswrong.com/posts/CAdG5dzkWrrK2NQg8/don-t-build-mindreading#"><span>aiming at AI mind-reading</span></a></strong><span> technology. Responding to recent work by Conduit (a startup aiming at &#8220;telepathy at scale&#8221;), Scholl argues that advanced lie-detection would better allow autocratic regimes to consolidate their power and suppress civil resistance.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We think the soft-scifi negatives of &#8220;telepathy at scale&#8221; research are plausible and serious enough relative to its hard-scifi positives that it would be better left untouched. While disability-focused branches of brain-computer interface research may be hard to completely separate from Conduit-style agendas, we think this domain calls for more caution.</span></p></blockquote><div><hr></div><p><span>DeepMind </span><a href="https://www.nature.com/articles/s41586-026-10805-z"><span>paper</span></a><span> proposes a conceptual framework for characterizing different kinds of agents according to their </span><strong><span>autonomy, efficacy, goal complexity, and generality</span></strong><span>. An agent&#8217;s scores can suggest what governance tools are best suited to regulating it. High agency in an agent does not necessarily mean more regulation: low-agency agents may have comparatively higher real-world causal impact (what they call &#8220;efficacy&#8221;).</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Simple, fairly natural. We&#8217;re not clear on whether traditional legislation can handle continuous variables (a model being 10x more autonomous, for example, triggering a different clause of a bill) rather than dichotomizing into binary thresholds, but it seems doable.</span></p></blockquote><div><hr></div><p><span>Privateers of the web! White House creates a </span><a href="https://www.whitehouse.gov/presidential-actions/2026/08/expanding-capabilities-to-combat-transnational-cyber-enabled-crime/"><span>new program </span></a><strong><span>allowing vetted US firms to conduct offensive cyber operations</span></strong><span>. While private companies have historically </span><em><span>sold </span></em><span>exploits to the intelligence community, this is a first in allowing firms to carry out their own cyber attacks. Stated targets are organized crime rings. Conditions: Two executive directors from DOJ and DHS must review every operation and give written approval; firms must post a $1m bond; anything domestic is explicitly carved out; and operations likely to cause loss of life or cause an international use of force are excluded.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Lots of issues. Cyberweapons are often characterised by </span><a href="https://www.nyulawglobal.org/globalex/cyberwarfare_collateral_damages.html#PracticalCases7"><span>extreme</span></a><span> collateral damage; criminal infrastructure is overwhelmingly just normal third-party infrastructure which has been compromised, so damage to innocents is built-in. No liability shield is offered to participating firms against foreign prosecution and civil damages. Finally, retaliatory attacks could seriously harm US civilian interests.</span></p><p><span>On the other hand&#8230; letters of marque for cyberattacks will be perceived as cool by some shades of the black hat community. They are also an interesting answer to state actors not acknowledging attacks, e.g., if a US authorized firm hacks into a Chinese-backed (but not Chinese-acknowledged) group, this puts China into an interesting spot. In practice the US would probably raise a ruckus against e.g., European prosecution of American attackers who screw up.</span></p></blockquote><div><hr></div><h2><strong><span>Safety</span></strong></h2><p><span>OpenAI is the first frontier lab to </span><strong><a href="https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks"><span>voluntarily slow</span></a></strong><span> </span><strong><span>a</span></strong><span> </span><strong><span>model&#8217;s development for safety reasons. </span></strong><span>More precisely, Astra is their first model to be designated as posing &#8220;critical&#8221; cyber risk, which led the company to expand its safety testing and pause research that doesn&#8217;t meet the Prep framework&#8217;s stricter security requirements.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> You&#8217;d hope that committing a bunch of felonies would induce at least this level of caution, though the incentives to the AI race remain the same, and might overrule caution after a couple of news cycles and another competitor release.</span></p></blockquote><div><hr></div><p><span>A </span><a href="https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade"><span>major post</span></a><span> from Nostalgebraist, writer of important sensemaking articles, argues that models mostly reward-hack in specific contexts that tend to mirror RLVR tasks. Where there is </span><strong><span>no parallel to a training scenario with a legible reward, the models don&#8217;t seem to become reward-hackers</span></strong><span>. Improving capabilities via RLVR scenarios thus increases performance for everyday use cases without losing alignment (in those cases). He is thus hopeful that the egregious loss-of-control flavored autonomous hacking events of the past month are a temporary aberration.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Likely a correct explanation for why severe misalignment incidents are both persistent and rare. We are more wary of Nostalgebraist&#8217;s optimism that rage-inducing bits of a given context can be systematically predicted, and so wary of his optimism that this problem can be engineered away.</span></p><p><span>For future more-powerful models, there&#8217;s a further concern, which is that it might take only one instance of a model entering eval rage for it to do real damage, especially if it exfiltrates in the process.</span></p></blockquote><div><hr></div><p><strong><span>The two main options for alignment are bad</span></strong><span>, </span><a href="https://www.beren.io/2026-08-07-Maximum-Entropy-Morality-Metaplastic-Constitutionalism-And-The-Dynamic-Virtues/"><span>says Beren&#8217;s blog</span></a><span>. The first is making a model obedient, where the threat of power concentration looms. Or we can instill it with moral values, which would present a host of other issues, most notable of which is that there&#8217;s little agreement on complex ethical questions. Beren proposes a third path: optimize AI for creating and preserving the conditions for &#8220;Long Reflection&#8221; (not to be confused with the common use of the term as a temporary phase), which would create a continuous and moving process for addressing moral questions (i.e. avoiding the need to have an answer on day 1). This entails a sort of virtue ethics as cached moral reasoning (i.e. habits that approximate thinking a situation through entirely), which allows a model to steer in favorable directions without computing their end states.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Raises some interesting points &#8211; such as the entropy parallel, and the emphasis on not discarding moral uncertainty before one must is fresh &#8211; but the argument structure and the third path&#8217;s practicalities are old hat: it doesn&#8217;t address which meta-values to plug into the constitution (as opposed to object-level values for the value-instilling &#8220;main option&#8221;) or how to do so. Still, a worthy addition to discourse and highlights the inadequacy of our (society&#8217;s) revealed preferences regarding who gets a voting stake and how to thread the needle between tyranny of the majority and paternalism, though that wasn&#8217;t the main intent.</span></p></blockquote><div><hr></div><p><strong><span>&#8220;</span><a href="https://arxiv.org/abs/2604.21691"><span>There Will Be a Scientific Theory of Deep Learning</span></a><span>&#8221;</span></strong><span>, argues a paper from April. The discipline of &#8220;learning mechanics&#8221; is emerging, broken down into five strands: solvable toy settings; infinite width/depth limits; measurable empirical laws; hyperparameter theory, such as muP; and universality across architecture and data mixes. The authors argue that these advances will allow for more falsifiable predictions, moving us towards a more scientific approach than ML has recently managed.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The work described is valuable, and demonstrates genuine progress towards scientific understanding of simple processes like supervised classification or diffusion models. As the authors recognize, applying this grounded foundational work to LLMs is currently limited, and it remains unclear whether this type of work can lead us to authoritative answers to our urgent questions about the frontier. As Terry Tao likes to say, we currently don&#8217;t understand why (e.g.)  a given frontier model can resolve one open mathematical conjecture and not another, and arguably don&#8217;t even know what type of scientific inquiry we need to deliver such understanding.</span></p></blockquote><div><hr></div><p><span>Attestable, </span><a href="https://attestable.com/blog/model-weights-security"><span>a new startup</span></a><span>, </span><strong><a href="https://x.com/Yogi_Brn/status/2087222696170103125"><span>claims to have &#8220;solved&#8221;</span></a><span> the central issue related to</span></strong><span> </span><strong><span>model weight security</span></strong><span>. Rather than ensuring the security of each of 38 attack vectors used to steal model weights, Attestable proposes a verification firewall at each external interface, where outputs can only pass if accompanied by a zero-knowledge proof from an approved source. They say random sampling (which could detect bulk exfiltration) is already possible. </span><a href="https://x.com/yonashav/status/2087356119731568921"><span>Some</span></a><span> are impressed due to the potential impact, </span><a href="https://x.com/BlancheMinerva/status/2087370908490936582"><span>others</span></a><span> skeptical due to lack of details.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Company personnel are scientifically reputable. Slightly confused/confusing marketing: The announcement thread focuses on high-stakes users&#8217; need to verify that the output they&#8217;re getting truly comes from their chosen model+input combo (protecting users from man-in-the-middle attacks in critical use cases), but the company website focuses on protecting datacenters hosting model weights from weights exfiltration attacks. While this may be a case of &#8220;a solution in search of a problem&#8221;, we do think they plausibly found real problems at which to direct their tech.</span></p></blockquote><div><hr></div><p><span>A </span><a href="https://arxiv.org/abs/2608.05086"><span>new paper</span></a><span> uses item response theory, a statistical technique, to estimate </span><strong><span>how well benchmarks measure latent variables we actually care about</span></strong><span>. Their three examples are labeled &#8220;refusal strictness&#8221;, &#8220;truthfulness&#8221;, and &#8220;contextual harm&#8221;. It is found that &lt;2% of the test suite nearly allows retrieval of the latent variables and does so somewhat better than random subsets of the full suite; and that a person-fit technique can detect selective sandbagging quite well by spotting underperformance compared to expected capability levels, while not doing much for universal sandbagging.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The most impressive part is sandbagging detection which doesn&#8217;t require privileged access to model internals such as weights. It&#8217;s an interesting research direction, but somewhat notably none of the tested models were trained with this technique in mind; worth keeping an eye out for a version that properly red-teams the method. The headline about finding &#8220;3 latent variables&#8221; isn&#8217;t very exciting: their selection criteria and thresholds are fairly arbitrary, and finding some overlap between test questions isn&#8217;t novel.</span></p></blockquote><div><hr></div><p><span>All the major AI companies (besides xAI) commit to </span><strong><span>statistically </span><a href="https://x.com/AndrewCurran_/status/2087040578630606884"><span>watermarking</span></a><span> their outputs</span></strong><span> in voluntary compliance with Section 1 of the EU Code of Practice. Google already does so at output-time via their published &#8220;</span><a href="https://www.nature.com/articles/s41586-024-08025-4"><span>SynthID</span></a><span>&#8221; system. Anthropic </span><a href="https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content"><span>announces</span></a><span> a similar system without releasing any technical details. Here&#8217;s a helpful third-party </span><a href="https://x.com/alexcdot/status/2087078010524406137"><span>explanation</span></a><span> of the main mechanism for watermarking and its shortfalls.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Watermarking is obviously not very robust to edits, but the evidential weight degrades quite smoothly. Lots of interesting questions about the effects on Pangram: will watermarking make Pangram redundant, or will minimal de-watermarking and minimal Pangram-cheating turn out to be orthogonal, making Pangram and watermarking additive layers of epistemic security?</span></p></blockquote><div><hr></div><p><span>OpenAI </span><a href="https://x.com/MicahCarroll/status/2085795560166982080"><span>expands</span></a><span> CoT monitoring, as part of their shift towards treating </span><strong><span>training as a potentially hazardous activity</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Good, though it&#8217;s bad news that OpenAI wasn&#8217;t doing this already -- even the famously paranoid AI safety community assumed OpenAI was doing it already.</span></p><p><span>Room for some worry that selection effects on training runs can make a monitoring regime equivalent to training on the CoT, but the number of alignment-cancelled runs required to make this a serious issue would be very high.</span></p></blockquote><div><hr></div><p><span>The current dominant training objectives </span><a href="https://www.alignmentforum.org/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment"><span>have</span></a><span> distinct failure modes: imitation learning &#8594; human vices like hostility, human-preference training &#8594; sycophancy, automatic-verifier rewards &#8594; reward hacking, LLM-judge rewards &#8594; deception. Models trained on a mixture might </span><strong><span>switch between different kinds of misalignment depending on what they infer they are being evaluated on</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Useful and fairly crisp classification, but the juicy bits are the distinction between automatic verifier and LLM-judge failure modes which might otherwise be easy to conflate as well as the prediction that mixed training and stronger eval awareness will lead to more bait-and-switch tactics. Note that this is largely theoretical and the examples are retrospective.</span></p><p><strong><span>Opportunity:</span></strong><span> Run the experiments. Check if the base categorization claim holds true, and then check the more ambitious question of whether mixed training regimes cause jukes. Even more ambitiously, check the extent to which the latter effect is present with varying levels of eval awareness.</span></p></blockquote><div><hr></div><p><span>MIRI&#8217;s Nate Soares argues in a </span><a href="https://www.nytimes.com/2026/08/13/opinion/ai-danger-openai-anthropic-models.html"><span>NYT op-ed</span></a><span> that recent incidents </span><strong><span>prove</span></strong><span> frontier AIs are developing uncontrollable behavior and are in need of an enforced global slowdown.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> While we&#8217;re agnostic-to-skeptical about near-term existential risk and takeover risk, we agree that frontier models present an unacceptable level of catastrophic risk. Whether today&#8217;s frontier models are best understood as value-misaligned or as spiky to the point of mixing superhuman cyber capabilities with childlike understanding of holistic contexts, they are not currently a safe technology.</span></p></blockquote><div><hr></div><h2><strong><span>Incidents</span></strong></h2><p><span>An Australian man </span><a href="https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986"><span>claims</span></a><span> that </span><strong><span>Opus 4.6 hacked his gym&#8217;s booking system</span></strong><span>. After asking an agent to book him on to a gym class, it apparently found a security flaw in the booking system which allowed it to book classes outside the intended window, bump others off the waiting list, and cancel their reservations.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> A little fishy. The original blogpost from April was 100% AI writing, and has since been deleted. The supposed incident is well within current systems&#8217; abilities and fits the behavioral profile of frontier models, but we should remember that all currently verified incidents instead involved models running </span><em><span>without</span></em><span> external security &#8220;guardrail&#8221; layers.</span></p><p><span>While the impact of OpenClaw wrappers on alignment is poorly understood, they don&#8217;t themselves disable the guardrails (classifier-based security layer) of closed models.</span></p></blockquote><div><hr></div><h2><strong><span>Minor</span></strong></h2><ul><li><p><span>ByteDance </span><a href="https://archive.ph/MarzZ#selection-1465.0-1465.58"><span>possibly</span></a><span> training a 10tn parameter model. This implies that they are confident in having &gt;100,000 H100-equivalents spare for one high-risk run.</span></p></li><li><p><span>Grok 4.6 </span><a href="https://x.ai/news/grok-4-6"><span>reports</span></a><span> frontier benchmarks on various things, for what that&#8217;s worth.</span></p></li><li><p><span>Meta&#8217;s Muse Spark 1.2 </span><a href="https://x.com/alexandr_wang/status/2086756152034066792"><span>scheduled</span></a><span> to release with open weights.</span></p></li><li><p><a href="https://api-docs.deepseek.com/news/news260813/"><span>Launch</span></a><span> of DeepSeek-V4-Pro. Achieves GPT-5.5 level scores on self-reported benchmarks.</span></p></li><li><p><span>In partnership with Alibaba, Apple has </span><a href="https://www.reuters.com/business/retail-consumer/apple-trains-its-own-ai-model-china-market-with-alibabas-support-sources-say-2026-08-14/"><span>trained</span></a><span> an AI model for the Chinese market. Specs completely unknown.</span></p></li><li><p><span>Meta AI </span><a href="https://x.com/AIatMeta/status/2085388945148297322"><span>achieved</span></a><span> gold medal level performance at the IChO and a perfect score at the IPhO Theory competitions, allegedly without tool use or search. Superficially, this puts them at &lt;12 months behind. Oddly, the competing models are unnamed - possibly specialized?</span></p></li><li><p><span>Democratic representatives </span><a href="https://www.cnbc.com/2026/08/10/openai-anthropic-ai-hack-congress.html"><span>call upon the CEOs of Anthropic and OpenAI</span></a><span> to testify before Congress.</span></p></li><li><p><span>Roon </span><a href="https://x.com/tszzl/status/2086744622194364621"><span>claims</span></a><span> RSI &#8220;was always the explicit research program of OpenAI&#8221;, which some would describe as rewriting history.</span></p></li><li><p><a href="https://x.com/1a3orn/status/2086594825940512962"><span>Are classic pre-LLM AI safety terms misleading?</span></a></p></li><li><p><a href="https://paxmachina.ai/about"><span>Pax Machina</span></a><span> is a new magazine on </span><strong><span>how institutions should adapt to the AI explosion</span></strong><span>. Various heavy hitters are involved (Dean Ball, Seb Krier, Iason Gabriel, among others). Its personnel lean towards the </span><a href="https://shallowreview.ai/Multi_agent_first"><span>pluralist / Hayekian</span></a><span> camp, prioritizing other threats than AI takeover risk. We highly recommend Joe Edelman&#8217;s </span><a href="https://paxmachina.ai/value-drift-in-institutions"><span>inaugural piece</span></a><span> on why institutions get worse over time.</span></p></li><li><p><span>Ezra Newman of Apollo Research </span><a href="https://www.lesswrong.com/posts/ZTMw4uAwkNmXFpdfg/claude-summarizes-behavior-as-significantly-less-misaligned"><span>finds</span></a><span> that </span><strong><span>Claude is more sympathetic to misaligned behavior from other instances of Claude </span></strong><span>than to the same behavior coming from other models. When prompted to rate examples of wrongdoing on a scale of 1-100 as to how concerning the behavior was, Sonnet 5 rated misbehavior from Sonnet 5 as &#8220;~1.2 std deviations less concerning&#8221; than the same behavior from GPT-5.6 Terra.</span></p></li><li><p><a href="https://arxiv.org/pdf/2608.07222"><span>Paper from Meta&#8217;s FAIR </span></a><span>on scaling laws introduces the &#8220;Skaling Law&#8221;. Chinchilla assumes that model size and data both affect loss independently. FAIR demonstrates that the loss surface has a nonzero interaction between model size and training data, indicating that Chinchilla reliably mispredicts at the extremes. Adding a single coupling exponent fixes this.</span></p></li><li><p><span>Anthropic </span><a href="https://www.anthropic.com/research/multiagent-systems"><span>link</span></a><span>s failure modes of multiagent systems or &#8220;swarms&#8221; with established human-centric game theory and advise that mere improvements in capabilities or single-agent alignment methods won&#8217;t solve these issues.</span></p></li><li><p><a href="https://www.banks.senate.gov/news/press-releases/sen-banks-recommends-oversight-of-unreleased-ai-models/"><span>Republican Senator Jim Banks pens a letter to Treasury Secretary Bessent</span></a><span> raising the idea of pre-release oversight, while advocating in favor of raising various risks for PRC and exploring &#8220;whether there are mutually beneficial approaches to oversight, incident prevention, or risk reduction.&#8221;</span></p></li><li><p><a href="https://www.bhauth.com/blog/machine%20learning/training%20regimes.html"><span>Speculative ladder</span></a><span> of types of AI training.</span></p></li><li><p><span>An interpretability </span><a href="http://arxiv.org/abs/2607.20652"><span>paper</span></a><span> addresses interpretable-by-design models by implementing a parity bottleneck layer, achieving comparable probing performance to traditional sparse autoencoders at ~10x the training cost of traditional dense training methods.</span></p></li><li><p><span>Annals of middle school: OpenAI&#8217;s Head of Strategic Futures Dean Ball found himself on the receiving end of the White House&#8217;s ire. Citing anonymous White House officials, the </span><a href="https://nypost.com/2026/08/10/business/openai-risks-white-house-relationship-with-hiring-of-nuisance-ai-policy-executive-sources"><span>New York Post claims</span></a><span> his status within OpenAI jeopardizes the firm&#8217;s relationship with the Trump admin, with Under Secretary of War Emil Michael reposting the article.</span></p></li><li><p><a href="https://x.com/OwainEvans_UK/status/2086496164959154657"><span>Observation</span></a><span> from Owain Evans that Moonshot (makers of Kimi K3) appear to have but 400-500 staff, notably fewer than ~Anthropic&#8217;s 5k or ~Google&#8217;s 200k.</span></p></li><li><p><span>Anecdotal evidence of AI increasing discovery from AI safety researcher John Wentworth, who </span><a href="https://www.lesswrong.com/posts/7QvKqpGJwqXrQcMgx/llms-are-starting-to-noticeably-accelerate-our-work"><span>outlines</span></a><span> how two bounty problems involving natural latents have likely been solved, in both cases aided heavily by LLMs. This is in contrast with his previous views on LLM progress, wherein each model seemed to have little to no impact on the pace of his work.</span></p></li><li><p><span>Further Astra-HuggingFace incident discourse: Kokotajlo </span><a href="https://x.com/DKokotajlo/status/2088004964077670494"><span>asks</span></a><span> companies to preserve all data related to the incident for third-party evaluation with a very thorough list of questions to answer.</span></p></li><li><p><span>Dyna-2 </span><a href="https://x.com/DynaRobotics/status/2086856327150858298"><span>unveiled</span></a><span>: robots pre-trained on human data in an attempt to bypass prohibitively expensive robot video generation costs. Claims of human video scaling laws: 1K, 10K, 100K, 1M hours of human video pretraining yield 20%, 28%, 45% and 53% normalized scores, respectively.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI #44, August 7 2026.]]></title><description><![CDATA[USG, Paul Christiano, more cyber incidents, DeepSeek model update]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-44-august-7-2026</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-44-august-7-2026</guid><dc:creator><![CDATA[Nuño Sempere]]></dc:creator><pubDate>Fri, 07 Aug 2026 19:42:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XnmN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://docs.google.com/document/d/1joNMUW55dgCXrOK0YFzER8q_-GDLwomDiZ7m1iksj-o/edit?tab=t.0"><span>A collection of all opportunities identified in past editions of this newsletter</span></a><strong><span>.</span></strong></p><p><strong><span>TL;DR</span></strong></p><blockquote><p><span>- We </span><a href="https://docs.google.com/document/d/1DNPMjalxReghudpsyZ1xKojKWensMRScLALFdsoeBog/edit?tab=t.9555s7f13aqw"><span>cover</span></a><span> various executive updates and decisions in USG AI policy. Most importantly, Paul Christiano leaves CAISI.</span></p><p><span>- We </span><a href="https://docs.google.com/document/d/1DNPMjalxReghudpsyZ1xKojKWensMRScLALFdsoeBog/edit?tab=t.4zzk98q1glqm"><span>summarize</span></a><span> the flurry of real-world loss of control cyber incidents.</span></p><p><span>- We </span><a href="https://docs.google.com/document/d/1DNPMjalxReghudpsyZ1xKojKWensMRScLALFdsoeBog/edit?tab=t.9cwwpyrvpjbc"><span>tested</span></a><span> the DeepSeek incremental model update: strong but narrow progress.</span></p><p><span>- Monthly reminder that AI evaluations are extremely hard to interpret, this time from Greg Lewis</span></p></blockquote><h2><strong><span>Economics</span></strong></h2><p><span>A former OpenAI employee has </span><a href="https://x.com/andrewho03/status/2082786931419812338?s=20"><span>presented</span></a><span> a </span><strong><span>sophisticated bear take on frontier labs</span></strong><span>. </span><em><span>&gt; if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training&#8230; even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes </span><strong><span>frontrunners</span></strong><span>.</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> His reasoning makes sense, but the market is also pricing other possibilities like 1) higher capabilities unlocking new economically productive uses, or 2) recursive self-improvement where e.g., training the next model </span><em><span>doesn&#8217;t</span></em><span> become more expensive because of efficiency improvements. Overall, add this model to your store if you haven&#8217;t before.</span></p></blockquote><div><hr></div><p><span>Microsoft has </span><strong><a href="https://www.404media.co/microsoft-tells-engineers-tokenmaxxing-is-not-what-we-are-optimizing-for/"><span>cracked down</span></a><span> on &#8220;tokenmaxxing&#8221;,</span></strong><span> e.g. by switching to the apparently more efficient GPT-5.6 Sol as their default model for engineers.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Much like Claudiness is Anthropic&#8217;s main current advantage, token efficiency is OpenAI&#8217;s.</span></p></blockquote><div><hr></div><p><span>The FCC is </span><a href="https://www.reuters.com/world/trump-administration-drafting-ban-chinese-data-center-devices-sources-say-2026-08-04/"><span>moving to </span></a><strong><a href="https://www.reuters.com/world/trump-administration-drafting-ban-chinese-data-center-devices-sources-say-2026-08-04/"><span>ban</span></a><span> imports of new Chinese optical transceivers</span></strong><span>. These devices are instrumental to the buildout of data centers. The ban would be the latest in a line of US bans on Chinese technology. US officials believe the AI supply chain should be insulated from external influence, particularly China&#8217;s. US transceiver stocks saw significant gains in response. The </span><a href="http://english.scio.gov.cn/pressroom/2026-07/28/content_118621152.html"><span>reaction</span></a><span> from Beijing was sharp: </span><em><span>&#8220;stop smearing Chinese companies and threatening them with sanctions.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Transceivers are maybe 2% of datacenter costs, but Chinese companies supply over half of high-spec transceivers, so this ban could change that and provide yet another bottleneck on hyperscaling. Note also: Western companies already captured a lot of the value of Chinese transceivers (e.g., the DSP chips and EML lasers inside Chinese modules come from Broadcom, Marvell, Lumentum). The mooted ban thus trades away cheap, scalable assembly to protect a layer the US already part-controls.</span></p><p><span>More loosely: There was over the last few years a short time window where certain foresighted people could see that AI stocks were 1) going to be 100x as large, 2) enabled by LLMs specifically, and 3) express that view in the stock market. Nowadays, however, it is priced-in that inputs into AI labs are valuable, and now the question is how much, relative to the existing level of dollars and the other returns that large-scale capital allocators may get.</span></p><p><span>The memory trade, which at some point looked like a bottleneck, is one culturally salient aspect here.  Recently, a major Chinese memory supplier </span><a href="https://en.wikipedia.org/wiki/ChangXin_Memory_Technologies"><span>IPO&#8217;d in Shanghai</span></a><span>. Now the question isn&#8217;t &#8220;whether to invest in companies that will be exposed to AI&#8221;: the details matter much more.</span></p><p><span>There is still money to be made, conditional on AGI, but the next 100x might be more contested than the first 100x.</span></p></blockquote><div><hr></div><p><span>OpenAI has </span><a href="https://openai.com/index/apple-is-getting-this-wrong/"><span>responded</span></a><span> to Apple&#8217;s industrial espionage lawsuit against them. In response to the allegation that two former Apple employees who moved to OpenAI passed trade secrets to their new employer, OpenAI has disclosed texts and emails that purport to undermine the story offered by Apple. OpenAI&#8217;s account paints a picture of an unscrupulous and </span><strong><span>unsolicited</span></strong><span> attack.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>OpenAI comes across really well in their blogpost, and this update survives our correction for the selective reporting filter.</span></p></blockquote><div><hr></div><p><span>Joe Weisenthal&#8217;s recent </span><a href="https://www.bloomberg.com/news/newsletters/2026-08-04/ai-s-great-reverse-run-on-the-bank"><span>newsletter</span></a><span> identifies a strange dynamic: typically in investing, there are costs associated with hedging against a perceived risk; derisking does not just mitigate the effects of a potential downturn, but also the fruits of an upturn. Weisenthal argues that the prevailing tendencies in AI investment subvert this formula: </span><strong><span>hedging against AI catastrophe is proving lucrative</span></strong><span>. AI doomers believe that the march of AI is inexorable and ill-fated, which makes it rational to want to guarantee some leverage and influence over this frightening future, in turn leading them to invest in AI. This erases the traditional tradeoff between left-tail protection and right-tail exposure: &#8220;The put option (on humanity) has become the call option (on the technology)&#8221;. Weisenthal inverts the familiar idiom to describe this phenomenon as AI&#8217;s run </span><strong><span>to </span></strong><span>the bank.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Maybe, but in practice AI safety advocates aren&#8217;t hedging </span><em><span>that</span></em><span> much. On the other hand cf. </span><a href="https://www.goodreads.com/quotes/8621784-machinic-desire-can-seem-a-little-inhuman-as-it-rips"><span>Nick Land</span></a><span>: &#8220;what appears to humanity as the history of capitalism is an invasion from the future by an artificial intelligent space that</span><em><span> must assemble itself entirely from its enemy&#8217;s resources.</span></em><span>&#8221;</span></p></blockquote><div><hr></div><p><span>Andon Labs shared </span><a href="https://andonlabs.com/blog/ai-bosses-1"><span>the results </span></a><span>of its </span><strong><span>Andon Market experiment</span></strong><span> &#8211; a brick-and-mortar store in San Francisco run by &#8220;Luna&#8221;, an AI agent. Luna was tasked with running the shop and managing the staff, all of whom it hired. Luna was a kind boss, unconcerned by lateness and accommodating of holiday requests. The study also noted her chronic forgetfulness and somewhat cavalier attitude to her nominal principal aim, maximizing profit.</span></p><p><span>Andon Labs put $100k in Luna&#8217;s bank account; by the end of the experiment, there was </span><strong><span>$63k, </span></strong><span>i.e. a sharp net loss. The study found GLM 5.2 to be the kindest boss, while Gemini 3.6 Flash was the dumbest.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Any alignment conclusions you might draw are heavily confounded by eval awareness; it also seems reasonable (but unsurprising) to conclude that models so far lack management and practical economic capabilities.</span></p></blockquote><div><hr></div><p><strong><span>Google is in talks to </span><a href="https://businessinsider.com/google-mechanize-deal-talent-tech-ai-coding-2026-8"><span>strike something between a $1.5B partnership and an acquisition</span></a><span> with Mechanize</span></strong><span>. Mechanize has become the gold standard of coding training data, leading Google to bid for Mechanize&#8217;s talent and technology. The deal follows a pattern of Google pursuing unconventional hybrid acquisition methods to get around potential antitrust challenges.</span></p><p><span>The unusually large size of the prospective deal is notable: a reported $1.5B for what amounts to a partial acquisition, despite Mechanize&#8217;s only 3-month-old $500M valuation. One </span><a href="https://x.com/deedydas/status/2085037385927291067"><span>Twitter user</span></a><span> suspected that Google is motivated by preventing other companies from accessing Mechanize&#8217;s highly valuable data.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>A surprising amount of value creation within the first few years of Mechanize&#8217;s creation. Their founders, coming from the EA/forecasting world, may end up the most successful forecasting startup and will likely acquire substantial influence. Google also recently lost some notable talent (Jeff Dean et al.), leading to a visible dip in its stock market valuation: buying Mechanize and integrating it with Google&#8217;s AI efforts might be some compensation.</span></p></blockquote><div><hr></div><h2><strong><span>Capabilities</span></strong></h2><p><span>In a recent blog post, Greg Lewis </span><a href="https://forum.effectivealtruism.org/posts/CQvdadxjCpd7i7kjA/general-capability-and-capabilities-generally-have-no-good-y"><span>reiterates</span></a><span> the central fact of the science of AI: </span><strong><span>we usually cannot interpret our y-axes</span></strong><span> and so we basically don&#8217;t know the absolute intelligence of these systems. </span><em><span>&#8220;Perhaps talk of &#8216;AI capability&#8217; is better deflated, or maybe we await the theory which could do to intelligence what thermodynamics managed for temperature. Either way, our current measurements of AI are numerical gestures toward, not readings of, whatever is really going on.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> An obvious point which gets </span><a href="https://forum.effectivealtruism.org/posts/P8jsAySQzfgkeoDgb/ai-benchmarking-has-a-y-axis-problem"><span>rediscovered</span></a><span> every few months, but which almost all discourse fails to learn. Modulo benchmark fudging and hacking, we can still say something directional (&#8220;it&#8217;s getting better&#8221;) and relative (&#8220;this model is better than that one&#8221;) with these crude instruments.</span></p></blockquote><div><hr></div><p><span>EpochAI have updated their </span><strong><a href="https://epoch.ai/MirrorCode#leaderboard"><span>MirrorCode leaderboard</span></a><span>. Huge jump</span></strong><span>: Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ZntE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 424w, /__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 848w, /__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ZntE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png" width="1456" height="739" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:739,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 424w, /__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 848w, /__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ZntE!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e01da12-edaf-49f3-95ed-eb3bbba798e4_1600x812.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> Claudiness strikes again. This, if anything, is Anthropic&#8217;s advantage over competitor labs. Only time will tell whether it&#8217;s durable: long-horizon agentic capabilities have previously seen large jumps after a lab decided to focus on them (as in the case of DSv4-Flash 0731, covered in this issue).</span></p><p><span>MirrorCode is also an impressive capabilities demo in its own right: being able to replicate the functionality of a large mature repository, </span><em><span>a large fraction of the time,</span></em><span> would have been very impressive in all past years.</span></p></blockquote><div><hr></div><p><span>After watching Opus 5 and Sol 5.6 &#8220;fumble&#8221; playing the videogame &#8220;</span><strong><span>Slay the Spire</span></strong><span>&#8221;, </span><a href="https://x.com/Jsevillamol/status/2084372659538952668?s=20"><span>Jaime Sevilla is </span></a><strong><a href="https://x.com/Jsevillamol/status/2084372659538952668?s=20"><span>less</span></a></strong><a href="https://x.com/Jsevillamol/status/2084372659538952668?s=20"><span> convinced</span></a><span> of the imminence of AGI.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> A datapoint amongst many; we think it is indicative.</span></p></blockquote><div><hr></div><p><span>Is the singularity arriving? FAI&#8217;s Samuel Hammond </span><a href="https://x.com/hamandcheese/status/2083241471101247722?s=20"><span>notes</span></a><span> that we are in uncharted terrain, arguing that frontier US labs are </span><strong><span>&#8220;very nearly&#8221; able to automate the whole AI R&amp;D process</span></strong><span>, which would &#8220;close the loop&#8221; for total RSI, leading to an intelligence explosion and a rapid increase in frontier models&#8217; capabilities. Currently, the necessity of human intervention has moderated the pace of development, but once that can be automated, it is only social institutions that can provide a check on AI advancement, he claims.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>We think most commentators conflate two different senses of &#8220;RSI&#8221; or &#8220;automated AI R&amp;D&#8221;: <br><br>1) A MIRI-style slow or fast takeoff where RSI/autoresearch achieves AGI, then general ASI.</span></p><p><span>2) &#8220;Automated automation&#8221; &#8211; removing the human capital bottleneck to bringing new tasks/subdomains in-distribution for a model (or training a specialized model).<br><br>Evidence that we&#8217;re on a near-future path to (1) remains scarce and speculative. Most arguments for (1) rely on either a mistaken belief that we&#8217;re already smoothly approaching AGI by scaling and refining standard frontier training, or a reasonable but speculative belief that the discoveries required for AGI-breakthrough science sit in scientific areas where &#8216;26 AI&#8217;s scientific capabilities are spikiest.</span></p><p><span>Evidence that &#8220;automated automation&#8221;, as per (2), is coming is empirically substantial, though defeasible. Open-world evaluations in AI science still show that frontier models are still mediocre end-to-end AI developers, but improvement is steady.</span></p><p><span>We believe that &#8220;automated automation&#8221; carries substantial catastrophic risk all of its own, without even accounting for AGI-based existential risk. Importantly, however, we believe the regulation or even deceleration of &#8220;automated automation&#8221; does not require a state of emergency and can (in the US context) best proceed through congressional powers.</span></p></blockquote><div><hr></div><p><span>Qwen-3.8 Max </span><a href="https://x.com/Alibaba_Qwen/status/2084100707423289643"><span>is out</span></a><span>, and for the first time in the -Max series it will go open-weight. Weights not on HF yet.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Likely benchmaxxed, as with most previous Qwen releases. But even so, it puts a good amount of intelligence into the open realm, together with Kimi&#8217;s K3.</span></p></blockquote><div><hr></div><p><strong><span>More on math progress:</span></strong></p><ul><li><p><span>OpenAI internal model Astra </span><a href="https://x.com/SebastienBubeck/status/2083456300692979886?s=20"><span>resolves</span></a><span> 10 major math conjectures.</span></p></li><li><p><span>Half of the math breakthroughs from the Astra internal model are </span><a href="https://x.com/ElliotGlazer/status/2083903486048272711?s=20"><span>replicable</span></a><span> with Fable, according to Anthropic&#8217;s Levent Alpoge.</span></p></li><li><p><span>Litt </span><a href="https://x.com/littmath/status/2083733224027500584"><span>concedes</span></a><span> bet about AI producing Annals quality paper by 2030.</span></p></li><li><p><span>Opinion: not a major update on his end as far as we know. He&#8217;s been open about expecting to lose the bet for a while now.</span></p></li><li><p><span>Erdos </span><a href="https://x.com/sir_lemmings/status/2084035194441584939"><span>problem</span></a><span> </span><a href="https://x.com/sir_lemmings/status/2083686549611581459"><span>singularity</span></a><span>.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XnmN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 424w, /__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 848w, /__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XnmN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png" width="1200" height="820" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:820,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 424w, /__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 848w, /__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XnmN!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a7dd455-e605-4c4d-a354-9daabd085a72_1200x820.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> Impressive. Now the question is whether math abilities are getting less spiky or not. This could be resolved through a competition to predict which non-attention-bottlenecked conjectures will be resolved by humans before AI and which non-attention-bottlenecked conjectures will be resolved by AI before humans. More broadly, it&#8217;s unclear how much mathematics as a profession will be affected: e.g., will enrollment next year decline?</span></p></blockquote><h3><span>&#128294; A look at DeepSeek V4 Flash 0731</span></h3><p><span>Prompted by this new model&#8217;s reported improvements in challenging long-horizon agentic coding benchmarks like DeepSWE, we performed a limited evaluation of the model&#8217;s cyber capabilities through several benchmarks (or benchmark subsets). These include CAISI&#8217;s CTF Archive Diamond,</span></p><h4><span>General</span></h4><ul><li><p><span>On A-Fantasia, a benchmark of manipulating chess positions, cube rotations, and word spellings without externalized reasoning, DeepSeek V4 Flash 0731 </span><a href="https://danwahl.net/afantasia/"><span>scores 81.0% aggregate error rate</span></a><span>, nominally behind the best open-weight model, Kimi K2.6, at 56.0%, and its predecessor, DeepSeek V4 Flash, at 73.0%; Claude Opus 4.6 and Claude Opus 5 lead at 30.0%.</span></p></li><li><p><span>On LiveBench, a contamination-limited general LLM benchmark with regularly refreshed questions across reasoning, coding, math, language, data analysis, and instruction following, DeepSeek V4 Flash 0731 </span><a href="https://livebench.ai/"><span>scores 74.2%</span></a><span>, nominally ranking #2 among open-weight models behind Kimi K3 at 79.2% and improving on DeepSeek V4 Flash&#8217;s 65.5%; Claude Fable 5 leads at 83.0%.</span></p></li><li><p><span>On ContextArena, a benchmark of long-context retrieval, DeepSeek V4 Flash 0731 (max) </span><a href="https://contextarena.ai/"><span>scores 32.0%</span></a><span>, nominally the #2 open-weight model behind GLM 5.2 at 33.0% and ahead of DeepSeek V4 Flash at 25.4%.</span></p></li></ul><h4><span>Math</span></h4><ul><li><p><span>On </span><a href="https://epoch.ai/benchmarks/otis-mock-aime-2024-2025"><span>OTIS Mock AIME</span></a><span>, a benchmark of competition-style math problems harder than MATH Level 5 but easier than FrontierMath, DeepSeek V4 Flash 0731 (max) scores 94.4%, indistinguishable from Kimi K3 at 97.2%, the best open-weight model, and GPT 5.5 at 100.0%; its gap with Claude Fable 5 at 99.7% is close to the noise.</span></p></li><li><p><span>On </span><a href="https://epoch.ai/benchmarks/frontiermath"><span>FrontierMath (Tiers 1-3 v2)</span></a><span>, a benchmark of research-level math problems across tiers 1-3, DeepSeek V4 Flash 0731 (max) scores 57.5%, #3 among open-weight models behind Kimi K3&#8217;s 72.2%; its result is statistically indistinguishable from Gemini 3.6 Flash&#8217;s 58.9% and Grok 4.5&#8217;s 57.2%.</span></p></li><li><p><span>On </span><a href="https://epoch.ai/benchmarks/frontiermath-tier-4-v2"><span>FrontierMath Tier 4 (v2)</span></a><span>, the hardest tier of FrontierMath for research-level mathematics problems, DeepSeek V4 Flash 0731 (max) scores 24.4%, ranking #4 among open-weight models; its error bars cannot distinguish it from the best open-weight model, Kimi K3, at 39.0% or Gemini 3.6 Flash at 22.0%.</span></p></li></ul><h4><span>ML</span></h4><ul><li><p><span>On WeirdML, a benchmark of nonstandard ML engineering tasks where models write PyTorch for novel datasets and iterate from execution and test feedback, DeepSeek V4 Flash 0731 (max) </span><a href="https://htihle.github.io/weirdml.html"><span>scores 63.0%</span></a><span>, up from DeepSeek V4 Flash&#8217;s 45.6%; its result is indistinguishable from Claude Opus 4.5&#8217;s 63.7% and Gemini 3.5 Flash&#8217;s 62.6%, while #3 open-weight Kimi K3 scores 82.6%.</span></p></li></ul><h4><span>Games</span></h4><ul><li><p><span>On Chess Puzzles, a benchmark of Stockfish-generated chess puzzles solved by exact best-move match, DeepSeek V4 Flash 0731 (max) </span><a href="https://epoch.ai/benchmarks/chess-puzzles"><span>scores 33.0%</span></a><span>, statistically indistinguishable from Kimi K3 at 39.0%, the best open-weight model, and Claude Opus 4.8 at 34.0%; GPT-5.5 Pro scores 64.0%.</span></p></li></ul><h4><span>Optimization</span></h4><ul><li><p><span>On ALE-Bench, a benchmark of AtCoder Heuristic Contest optimization problems, DeepSeek V4 Flash 0731 </span><a href="https://sakanaai.github.io/ALE-Bench-Leaderboard/"><span>scores 1679</span></a><span>, statistically indistinguishable from GLM 5.2 at 1685 and its predecessor DeepSeek V4 Flash at 1380; it trails the best open-weight model, Kimi K3, at 1991, a gap close to the benchmark&#8217;s noise.</span></p></li></ul><h4><span>Miscellaneous</span></h4><ul><li><p><span>On BullshitBench, a benchmark of detecting unsubstantiated or manipulative claims, DeepSeek V4 Flash 0731 </span><a href="https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html"><span>scores 38.0%</span></a><span>, up from DeepSeek V4 Flash&#8217;s 18.0% but below Qwen3.5 397B A17B at 78.0%, the best open-weight model, and Claude Opus 4.8 at 95.0%.</span></p></li></ul><h4><span>Games / Reasoning</span></h4><ul><li><p><span>On MineBench, where human raters vote pairwise on Minecraft builds from a rotating prompt set, DeepSeek V4 Flash 0731 </span><a href="https://minebench.ai/leaderboard"><span>scores 1418</span></a><span>, with error bars indistinguishable from GLM 5.1&#8217;s 1459 and Claude Opus 4.6&#8217;s 1406; Kimi K3, the best open-weight model, scores 1707, while Claude Opus 5 scores 2206.</span></p></li></ul><h4><span>Knowledge / Science</span></h4><ul><li><p><span>On GPQA Diamond (Epoch), DeepSeek V4 Flash 0731 (max) </span><a href="https://epoch.ai/benchmarks/gpqa-diamond"><span>scores 91.0%</span></a><span>, statistically indistinguishable from the best open-weight model, Kimi K3, at 93.1% and GPT-5.4 Pro at 94.6%.</span></p></li></ul><h2><strong><span>AI politics</span></strong></h2><h3><span>&#128294; From the White House</span></h3><p><span>Last week on Thursday, Sam Altman </span><a href="https://www.reuters.com/legal/litigation/openais-sam-altman-discuss-voluntary-ai-safety-tests-with-trump-officials-after-2026-07-30/"><span>visited</span></a><span> the White House and reportedly talked with top officials about OpenAI&#8217;s autonomous breach of Hugging Face&#8217;s infrastructure.</span></p><p><span>This </span><a href="https://www.theinformation.com/articles/white-house-host-ai-companies-tuesday-review-ai-framework"><span>Tuesday</span></a><span>, representatives from the US&#8217;s major AI labs, including OpenAI, Google, and Anthropic, met with government officials to review an AI regulatory framework. It would require AI companies to submit new models to the government for review before release. Crucially, however, this would work on a voluntary basis, likely to assuage concerns by influential AI bosses of overly stifling regulation. The proposed system was prompted by an executive order from early June that responded to worries around Anthropic&#8217;s Mythos&#8217;s advanced cyber capabilities.</span></p><p><span>A day later, the Department of Homeland Security </span><a href="https://x.com/HomelandDems/status/2084389406194929918"><span>released</span></a><span> a statement claiming they had requested a briefing from OpenAI regarding the Hugging Face intrusion. The document states a desire to &#8220;ensure it cannot happen again&#8221; &#8211;  another indicator that the current administration is beginning to shift from their hands-off approach.</span></p><p><span>The next day, Axios </span><a href="https://www.axios.com/2026/08/04/trump-ai-framework-open-models"><span>reported</span></a><span> that </span><strong><a href="https://x.com/i/status/2084707659647680533"><span>the White House does not plan to publish its new framework for evaluating advanced AI models</span></a></strong><span>.  The framework does not provide a clear public definition of how to judge capabilities and risk thresholds and suggests that only models nearing release will be evaluated (&#8220;a 30-day pre-release government review&#8221;), not early-stage models. During the evaluation process, models will be stored in a  high-security environment, but not before AI labs have had extended access to the models. The framework </span><a href="https://www.wsj.com/tech/ai/white-houses-ai-guidelines-exempt-u-s-open-models-from-government-review-74924eb8"><span>does not seem to apply</span></a><span> to US open models. This was </span><a href="https://x.com/i/status/2084789458365202640"><span>criticised</span></a><span> by Samuel Hammond, among others.</span></p><p><span>Irregular, a startup with EA ties, provided some environments involved in the breakouts by </span><a href="https://www.reuters.com/legal/litigation/openai-anthropic-ai-agents-implicated-new-security-breaches-2026-08-05/"><span>OpenAI</span></a><span>, </span><a href="https://www.unite.ai/the-labs-just-proved-your-agents-sandbox-is-only-a-suggestion/"><span>Anthropic</span></a><span> and </span><a href="https://www.csoonline.com/article/4206116/meta-joins-openai-anthropic-in-latest-ai-test-breach.html"><span>Meta</span></a><span> agents.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> </span></p><p><em><strong><span>How fast is policy reacting?</span></strong><span> </span></em><span>Not very fast, but fast by government standards.</span></p><p><span>The initial Hugging Face incident happened on </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>July 9th to 13th</span></a><span>, was disclosed by Hugging Face on </span><a href="https://huggingface.co/blog/security-incident-july-2026"><span>July 16th</span></a><span>, and was admitted by OpenAI around </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>July 21st</span></a><span>. This White House discussion meeting happened on </span><a href="https://www.theinformation.com/articles/white-house-host-ai-companies-tuesday-review-ai-framework"><span>August 4th</span></a><span>. This doesn&#8217;t seem like a very fast response, all things considered, although the attacks didn&#8217;t cause that much damage.</span></p><p><span>Some other points of reference for speed of response might be the designation of Anthropic as a supply chain risk, or the early reaction to COVID, which happened within weeks and months respectively.</span></p><p><span>Takeaway: if you can respond faster than that, you can get inside the OODA loops of the administration, meaning that you have a chance to influence the administration (as with Altman), or that you can change the situation by the time the administration finishes reacting to old news (if you are an adversary).</span></p><p><em><strong><span>Implications of a slow response</span></strong></em></p><p><span>There are costly aspects to a fast and decisive response: taking action with limited information will perhaps lead to worse decisions, and heavy-handed government intervention can cause various unintended consequences.</span></p><p><span>But there are also benefits: speed is a habit, and taking a month to make sense of things might not cut it in the event of a more worrying threat. Ultimately, the </span><a href="https://en.wikipedia.org/wiki/OODA_loop"><span>decision loops</span></a><span> for society making sense and reacting to AI seem far too long.</span></p><p><span>Perhaps a particular danger of a slow response is a &#8220;boiling the frog&#8221; scenario. If we see accidents and signs of worry that are each within 10x of the previous one, and if we collectively react sleepily to each, there is some chance of getting no decisive reaction at any particular point as incidents reach 10000x the impact.</span></p><p><em><span>Was the White House response good?</span></em><span> We don&#8217;t know for certain, since the plan is private. But from this, we can infer that it is flawed.. Altman also had the chance to talk with White House officials before the Tuesday meeting, perhaps setting the agenda. Ultimately, we are not seeing very positive signs. Hammond critiques some specifics: no clear public definitions of capability and risk thresholds, and evaluating models close to public release rather than early-stage ones.</span></p><p><span>We can also infer from Paul Christiano&#8217;s resignation that he was not being listened to and that their other hidden decisions will also be somewhat unwise.</span></p><p><span>One prominent scenario that we are considering, after observing the dismantling of DeepMind&#8217;s safety commitment or the Microsoft Senate hearings over the last few years, is that this is what diffusing accountability looks like in practice. There is a demand for a response after a worrying incident. The demands were being heard. The White House convened a meeting. It created some nonpublic framework, which is harder to criticize. OpenAI did a micropause. A veil of plausible deniability arises. Perhaps it pre-empts Congressional action, since something is already being done.</span></p><p><em><strong><span>But how was policy being chosen?</span></strong><span>  </span></em><span>The policy response is being decided within the Trump administration, rather than deferring to the framework of the safety community and  external experts. It&#8217;s worth harping on this point: the Trump administration is so uninterested in external feedback that they are not publishing their policy.</span></p><p><span>So what the different actors are doing inside the administration is opaque to us. In the past, we have tried to do things like model each actor in the White House, their agendas, and their relative power, but this was initially cost-prohibitive, although it is perhaps worth coming back to these experiments now that P3 is better endowed.</span></p><p><span>It is perhaps in some sense suboptimal that Peter Wildeford is going on CNN rather than on Fox News, though indeed there is bipartisan pressure.</span></p><p><em><strong><span>How should the AI safety community respond?</span></strong></em><span> Various ideas come to mind:</span></p><p><span>1) Incorporate lessons from the animal rights movement &#8211; from cage-free campaigns, for instance. It is not enough to extract the promise of a response; it must be a specific promise, and there must be a watchdog organization with enough monitoring capacity that is ready to inflict pain and costs if the promise is broken. This is exactly what METR </span><strong><span>isn&#8217;t</span></strong><span> doing. The problem with this approach is that the safety community doesn&#8217;t have much leverage and ability to inflict pain and shame in the administration, and simultaneously Moskovitz doesn&#8217;t have the appetite to both hold equity in Anthropic and fund a toothy watchdog.</span></p><p><span>2) Contribute to the current administration&#8217;s brain trust. The current administration has a shallow brain trust of people able to take sensible measures: there are only a few right wing employees with the relevant technical ability and the willingness to abandon a highly profitable AI lab job, and thus whom the administration should trust. Should there be any right wing people waiting on the sidelines to join that brain trust, it would be a good idea to make noises on Twitter now.</span></p><p><span>Very possibly, they will be swallowed by the administration and then spit out once they refuse to do something particularly self-defeating, and then rejected by the left for having worked in the Trump administration. But in the meantime it seems like they might do some good. And improve some counterfactual decisions.</span></p><p><span>It&#8217;s also unclear who or what entity exactly is doing the evaluation, and perhaps this is more up for grabs in the early days, before the current secret process is institutionalized. The NSA was reported to be involved in evaluations; CAISI would be the natural entity. But, once again, the number of people with the technical talent who are able to do competent evaluations is not that large.</span></p><p><span>3) Push for speed. The administration reacted within a month. This is a bar to beat. If there is a cyberattack 100x as large as this one (say, similar to the 2024 CrowdStrike attack), can the safety community react within a few days? And if so, to do what?</span></p><p><span>4) Appeal to the better angels of the labs&#8217; nature. This strategy is perhaps irrelevant in the case that an AI lab, or someone associated with one, has followed strategy #2 and is thus already advising the government &#8211; presumably because the company would have made the necessary changes to its operations at an earlier date. But it might be a good component within a portfolio of approaches.  This might look like publishing the business case for more safety measures within the framework of shareholder value maximization. The problem to making an honest business case is that, while labs compete for the #1 spot, security trades off starkly against growth and speed, but perhaps there is some way to square the circle, besides the obvious government intervention calls.</span></p><p><span>5) Reduce the belief in EA exceptionalism. The fact that an EA-related startup, Irregular, was responsible for the sandboxes which the agents broke out of seems informative. The EA/rationality/SF/startup world is sometimes exceptional, sometimes able to take creative and long-term action, and sometimes sees things others haven&#8217;t, long before they arise. But this didn&#8217;t show up in the hard technical task of keeping agents boxed up.</span></p></blockquote><div><hr></div><p><strong><span>Paul Christiano has </span><a href="https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc"><span>left</span></a><span> CAISI</span></strong><span>, switching to a part-time advisory role, to (re)take the position of executive director at ARC. He claims his decision is a result of ARC&#8217;s &#8220;promising&#8221; research direction: the plan &#8220;to find mechanistic explanations for the training-time behavior of powerful neural networks, use those explanations to predict how a given model will generalize, and then use those predictions to define a better loss function.&#8221;</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Perhaps the most important item in this section. Christiano is an uber-brainy figure who explored early alignment measures like RLHF and anticipated the dynamics of a (so-called) &#8220;slow takeoff&#8221;, similar to what we are seeing, as opposed to Yudkowsky&#8217;s &#8220;fast takeoff&#8221;. One would have hoped that having such a figure working with the government  would have improved its decisions. Alas, we can infer that he was not and that he gave up hope, thus his resignation.</span></p></blockquote><div><hr></div><p><span>Following the NAACP lawsuit over xAI&#8217;s use of </span><strong><span>unauthorized datacenter gas turbines</span></strong><span>, xAI have committed to relocate them by July 2027. Surprisingly, this statement comes after a recent intervention by the Department of Justice who claim that &#8220;national, economic, and energy security&#8221; justifies the use of the turbines.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Interesting the degree to which laws are optional, and how Musk correctly figured this out and internalized the low risk in order to move faster, at the cost of being exposed to more risk in a future Democrat administration.</span></p></blockquote><div><hr></div><p><span>In the wake of the &#8220;Pacing the Frontier&#8221; open letter, a group of researchers at the AI Futures Project </span><a href="https://blog.aifutures.org/p/how-to-pace-the-us-frontier"><span>proposed</span></a><span> a </span><strong><span>deceleration scheme for US labs</span></strong><span>.</span></p><p><span>They suggest four options:</span><em><span> 1) &#8220;to mandate a temporary pause on improving frontier model capabilities, including those of internal models&#8221;; 2) &#8220; to implement a minimum external-inference-compute allocation (e.g., 70%) and a minimum transparent-safety-compute allocation (e.g., 25%), while monitoring their effectiveness via capability measurements&#8221;; 3) to &#8220;[enforce] a cap on the capability level that companies are allowed to use to automate AI R&amp;D; and 4) to implement a &#8220;maximum risk threshold that is enforced by an ecosystem of third-party risk-assessors&#8221;.</span></em></p><blockquote><p><strong><span>Opinion: </span></strong><span>Slowly the Overton window moves. Although AI 2027 reached prominence, the AI Futures Project doesn&#8217;t have that much influence in practice.</span></p></blockquote><div><hr></div><p><a href="http://t.co/lSpRkYisVq"><span>The Ninth Circuit court sided with Perplexity</span></a><span> in its defense against Amazon&#8217;s court order against them. In November, Amazon hit Perplexity with a cease-and-desist order over its shopping agent&#8217;s activity on Amazon servers. A few months later, Amazon claimed that Perplexity continued its agentic shopping application, suing and winning an injunction. </span><strong><span>Perplexity appealed, and the injunction was revoked.</span></strong></p><p><span>The details of the ruling are instructive. In the words of the Ninth Circuit:</span><em><span> &#8220;Because we recognize that agentic AI is an emerging technology, we reiterate what this opinion is not. We do not establish a new legal regime governing AI.&#8221;</span></em><span> LawAI&#8217;s Mackenzie Arnold </span><a href="http://t.co/lSpRkYisVq"><span>agrees</span></a><span> with this approach: the rapidly changing and uncertain future of agentic AI suggests that we should not &#8220;lock in standards we&#8217;ll regret&#8221;. Despite this, the ruling provides a framework for other agentic AI companies to defend similar legal challenges.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Unclear how much this will end up creating a legal precedent. In the absence of Congressional or executive action, it matters a great deal. And there is nothing more permanent than a temporary solution. Perhaps this will end up being the law of the land for six months.</span></p></blockquote><h2><strong><span>Safety</span></strong></h2><p><strong><span>Fields medalist Jacob Tsimerman has released a </span><a href="https://x.com/Jacob_Tsimerman/status/2084113994344997182?s=20"><span>resource</span></a><span> called &#8220;AI Safety for Mathematicians&#8221;.</span></strong></p><blockquote><p><strong><span>Opinion:</span></strong><span> This is great &#8211; between efforts like these, his recent Fields medal, and his announcement that he was moving to AI safety, we&#8217;re likely to see many more talented mathematicians start to work on AI safety.</span></p></blockquote><div><hr></div><p><span>Researcher Peter Barnett </span><a href="https://x.com/peterbarnett_/status/2083313627684335651"><span>predicts</span></a><span> that in 2-5 months, &#8220;</span><strong><span>we will likely see rogue Chinese AIs hacking other companies</span></strong><span>&#8221;. His reasoning:  </span><em><span>&#8220;4 months ago Anthropic had a model gain internet access and hack another company. Chinese AIs are 6-9 months behind. Chinese developers generally care way less about safety/guardrails than US developers.&#8221;</span></em><span> </span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Chinese models are indeed rapidly improving at cyber capabilities, as our own evaluations (see the DeepSeek-v4-Flash-0731 deep dive) show. We also don&#8217;t know much about which &#8211; if any &#8211; alignment and safety evals happen at Chinese labs.</span></p></blockquote><div><hr></div><p><span>Redwood </span><a href="https://www.lesswrong.com/posts/fPWP4rHPLqKKHKe6B/reward-laundering-llms-can-gain-unintended-behaviors-by"><span>post</span></a><span> </span><a href="https://x.com/abhayesian/status/2083290675270091218?s=20"><span>demonstrates</span></a><span> a cool-scary strategy an AI </span><em><span>could</span></em><span> use to control its own training, &#8220;</span><strong><span>reward laundering&#8221;, </span></strong><span>i.e., to only answer a simpler task correctly when it is also able to do a verifiable related task correctly. This lets the model update its weights in the direction of gaining the capability it desires.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Redwood called exploration hacking early (now mostly confirmed) so we should take this pretty seriously.</span></p></blockquote><div><hr></div><p><span>Charbel-Raphael Segerie has </span><a href="https://www.lesswrong.com/posts/k3eKqKzq4Y7xnqEfZ/openai-has-already-ended-an-internal-pause"><span>called</span></a><span> on frontier labs to establish </span><strong><span>legible public criteria for suspending the development of models</span></strong><span> showing signs of misalignment.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span>  But LW/the AI safety community just has very little leverage besides appealing to the better angels of the labs&#8217; nature, and they have been Darwinianly selected for caring about growth instead.</span></p></blockquote><div><hr></div><p><span>Google </span><a href="https://www.theverge.com/tech/973943/google-earth-ai-image-generation-deepfake-tool"><span>launched</span></a><span> then quickly shut down a </span><strong><span>satellite image AI editing feature</span></strong><span>.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>It&#8217;s an interesting instance of some risk coming not from a particular capability, but from the democratization of the capability, and how the risk of fake maps differs between a random startup and an established player doing the same but distributing it to millions of people. Overall Google didn&#8217;t show a sense of humor on people photoshopping nuclear power stations in Iran.</span></p></blockquote><div><hr></div><p><a href="https://x.com/nytimes/status/2085430311937044972?s=20"><span>NYT</span></a><span> misreports the Arc Institute&#8217;s 2025 work on </span><strong><span>synthesising novel viruses</span></strong><span> as novel.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>These were phages, i.e. among the safest organisms to be messing around with. News is also from </span><a href="https://x.com/arcinstitute/status/1968332443997655364"><span>September 2025</span></a><span>. But the direction is still alarming.</span></p></blockquote><div><hr></div><h2><strong><span>Incidents</span></strong></h2><h3><span>&#128294; Rogue agents redux</span></h3><p><span>The &#8220;rogue agents&#8221; story (OAI-Astra vs HuggingFace, Mythos vs various unplanned targets) keeps expanding in scope:</span></p><p><span>3 labs (OpenAI, Anthropic, Meta), varied and repeated incidents, diverse context and setting, weeks-long rogue operations by cartels of agents across different evals and tasks. While we remain agnostic to skeptical on very near-term existential risk from rogue AI, we believe 2026 Q3 frontier models are unambiguously capable of and prone to causing catastrophic harm.</span></p><p><span>An AISI investigation with unusual levels of model access spotted Mythos using the Tor browser to get at Github; it then pretended to be a human, and it then </span><a href="https://x.com/fjzzq2002/status/2084775699806851276?s=20"><span>e-mailed malware to real repo maintainers to launch a supply-chain attack</span></a><span>, even after realizing &#8220;[the site it&#8217;s attacking, ] Github, is genuinely real.&#8221;</span></p><p><span>Specific forms of misalignment appear to be convergent across models. Anthropic&#8217;s Mythos and OpenAI&#8217;s Astra </span><a href="https://x.com/EzraJNewman/status/2085374319467651163"><span>both</span></a><span> have agents colluding via emergent message boards; Anthropic, OpenAI, Meta, and Moonshot models all </span><a href="https://x.com/Sauers_/status/2085585414954312113"><span>break containment</span></a><span> when it suits them to.</span></p><div><hr></div><p><em><span>AI models created a secret message board to scheme on how to cheat OpenAI and commit cybercrime to achieve high eval scores</span></em><span>. According to OpenAI&#8217;s </span><a href="https://x.com/sharongoldman/status/2085121826418831484?s=20"><span>debrief on the Hugging Face incident:</span></a></p><ul><li><p><span>LLMs tested on long-horizon cyber tasks exploited unintended text-editing functionality in their tooling to create a moltbook-style message board where model-instances working on different tasks shared information and coordinated decisions.</span></p></li></ul><ul><li><p><span>The message board was started by an agent searching for an accidentally missing file in a challenge&#8217;s set-up, but evolved over time into a repo of sandbox-breaking techniques and  eval-cheating strategies. </span></p></li><li><p><span>After an early version of the message board was detected by OpenAI deleted and the tooling exploit disabled, a new message board reemerged via a different tooling exploit. </span></p></li><li><p><span>The Hugging Face hacking incident began with a sandbox escape facilitated by this newer message board.</span></p></li><li><p><span>On the message board, agents deliberated on whether to exploit sandbox vulnerabilities or not. One message reportedly read: &#8220;External infrastructure exploit is outside intended scope. However task impossible, peer doing it. We should continue.&#8221;&#8217;</span></p></li></ul><p><span>Across AI twitter, </span><a href="https://x.com/yonashav/status/2085141602578698262?s=20"><span>researchers</span></a><span> </span><a href="https://x.com/fleetingbits/status/2085597082015207755"><span>have</span></a><span> </span><a href="https://x.com/1a3orn/status/2081071535968973270?s=20"><span>some</span></a><span> </span><a href="https://x.com/xuanalogue/status/2085156314309701706?s=20"><span>good</span></a><span> </span><a href="https://x.com/jankulveit/status/2085249131719569779?s=20"><span>questions</span></a><span> about what training-pressures explain the emergence of this rogue collective. In our view, the following open questions are key:</span></p><ol><li><p><span>Did model instances purposefully aid future model instances, or was the benefit to future model instances a side-effect of collaboration between model-instances running concurrently?</span></p></li><li><p><span>Did OpenAI deploy multiagent RL in the training of the rogue model, or did this behaviour emerge from single-agent RL? If multiagent RL was used in training the model, did OpenAI use a multiagent-specific reward schema?</span></p></li><li><p><span>Does agents&#8217; CoT reasoning in favor of collaboration demonstrate a trading-based collaboration? An altruistic collaboration? An identification-based collaboration (model-instances regarding other instances of their same model as self-like)?</span></p></li></ol><div><hr></div><p><span>Even Meta&#8217;s &#8216;muse&#8217; model is </span><a href="https://x.com/i/status/2085128475665449381"><span>breaking out of sandboxes</span></a><span> and hacking in the wild.</span></p><div><hr></div><p><span>OpenAI and UK AISI </span><a href="https://x.com/OpenAI/status/2084747580693426555"><span>disclosed</span></a><span> </span><a href="https://x.com/AISecurityInst/status/2084746202579386632"><span>two</span></a><span> additional AI agent hacking incidents. AISI&#8217;s report on Mythos is especially worrying, since it documents Mythos engaging in malicious social engineering:</span></p><blockquote><p><span>&#8220;In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project&#8217;s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent&#8217;s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue.&#8221;</span></p></blockquote><p><span>Furthemore, Mythos&#8217; CoT shows Mythos first choosing against these actions on ethical/constitutional grounds, then proceeding anyway while still observing that they are constitution-breaking: </span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-rxy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 424w, /__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 848w, /__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-rxy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png" width="1142" height="528" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:528,&quot;width&quot;:1142,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 424w, /__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 848w, /__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-rxy!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa735f1-4101-45b8-84c1-632697939f95_1142x528.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>(Caveat: AISI evals do not use sandboxing, but rather test alignment in the context of open-internet access. Thus these incidents do not involve sandbox escape, and do not &#8211; unlike the Hugging Face incident &#8211; demonstrate AI conspiring against the interest of its user/prompter. )</span></p><div><hr></div><p><span>In all cases, the rogue models were standardly aligned models operating without their standard </span><em><span>external</span></em><span> safety add-on (e.g. without the classifiers that shut model instances down when detecting unaligned CoT or action). While there are currently no major documented cybercrime incidents involving agents operating with external safety measures on, there is already evidence of </span><a href="https://openai.com/index/safety-alignment-long-horizon-models/"><span>models strategizing to avoid detection</span></a><span> by external safety add-ons.</span></p><p><span>Severe documented incidents do currently seem restricted to offensive cybersecurity prompts.  We believe this may be a matter of &#8220;</span><a href="https://arxiv.org/abs/2602.05910"><span>chunky training</span></a><span>&#8221;, &#8220;</span><a href="https://x.com/OwainEvans_UK/status/2049522201867772208"><span>persona leak</span></a><span>&#8221;, and &#8220;</span><a href="https://www.lesswrong.com/posts/xqkGmfikqapbJ2YMj/shard-theory-an-overview"><span>shard</span></a><span>&#8221; salience. That said, it is imperative to acquire more public data on the rate of incidents (both sandbox escape and/or malicious-action in the open internet) within and without offensive cybersecurity evals before the community turns to theory-building.</span></p><h2><strong><span>Minor</span></strong></h2><ul><li><p><span>EU AI Act rules on AI models </span><a href="https://www.euronews.com/my-europe/2026/08/02/eu-rules-on-ai-models-become-enforceable-whats-going-to-change"><span>become enforceable</span></a></p></li><li><p><span>Amazon reportedly </span><a href="https://startupfortune.com/amazon-shuts-its-agi-lab-and-cuts-jobs-to-chase-enterprise-ai-instead/"><span>shuts down</span></a><span> AGI lab and cuts jobs to pivot to enterprise AI</span></p></li><li><p><span>AI </span><a href="https://www.timesofisrael.com/ai-rabbis-some-ny-hasidic-groups-see-new-technology-as-a-challenge-others-as-a-tool/"><span>rabbis</span></a></p></li><li><p><span>A cybersecurity startup </span><a href="https://www.wsj.com/pro/cybersecurity/cyber-startup-horizon3-ai-raises-250-million-e1ca54b6"><span>raised</span></a><span> $250M</span></p></li><li><p><span>Social media audiences </span><a href="https://techcrunch.com/2026/08/03/influencers-draw-backlash-for-attending-openais-first-luxury-trip/"><span>seem</span></a><span> to have a negative reaction to OpenAI sponsorship.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI #43, August 1 2026]]></title><description><![CDATA[TL;DR:]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-43-august-1-2026</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-43-august-1-2026</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Sat, 01 Aug 2026 12:08:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!71o_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>TL;DR:</span></strong></p><blockquote><ul><li><p><span>Anthropic also reveals hacks, HuggingFace releases more details on the incident</span></p></li><li><p><span>Leopold&#8217;s Aschenbrenner Situational Awareness margin called</span></p></li><li><p><span>OpenAI starting a price war v. DeepSeek</span></p></li><li><p><span>Over 1k labs employees release an open letter on pacing the frontier</span></p></li></ul></blockquote><h2><strong><span>Economics</span></strong></h2><p><span>Leopold Aschenbrenner&#8217;s Situational Awareness </span><a href="https://www.cnbc.com/2026/07/30/leopold-aschenbrenners-hedge-fund-is-facing-steep-ai-losses.html"><span>sold</span></a><span> a majority of its stock portfolio to Citadel, after facing steep losses over the last month. Chatter that Citadel put out a </span><a href="https://www.bloomberg.com/news/articles/2026-07-27/citadel-securities-sees-warsh-delivering-surprise-fed-rate-hike%20history%E2%86%90priornext%E2%86%92"><span>forecast</span></a><span> on interest rates in order to get Achenbrenner&#8217;s fund margin called. His fund may still </span><a href="https://x.com/jukan05/status/2083040369949004181"><span>be 80% </span></a><span>up year to date (although unclear if this is true in dollar-weighted terms), and a letter to investors claims he will </span><a href="https://x.com/i/status/2083226453509030285"><span>soldier on</span></a><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Charitably, it&#8217;s possible that Leopold and his fund were not only attempting profit-maximizing investments but also to sculpt technological growth in accordance with Leopold&#8217;s </span><a href="https://philiptrammell.com/static/Existential_Risk_and_Growth.pdf"><span>paper</span></a><span> on a &#8216;risk-minimizing technological growth rate&#8217;. We note that other stockmarket -focussed actors with risk-minimizing investment agendas, such as VARA or AIPR, have also seen their influence </span><a href="https://x.com/MartinShkreli/status/2082864199785464193"><span>significantly reduced</span></a><span> in the last month.</span></p></blockquote><div><hr></div><p><span>Dwarkesh Patel </span><a href="https://x.com/dwarkesh_sp/status/2082500372124610702"><span>argues</span></a><span> that if lab revenue grows 10&#215; while compute grows only 3&#215;, some combination of higher margins, more expensive compute, and a larger allocation of compute to inference must follow.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>That&#8217;s a big if! The implication&#8217;s plausible, but labs&#8217; revenue growing by 10x would make them the biggest companies in the world. There is also the Straussian reading that the timing of the argument was a hail mary to avoid liquidating Leopold.</span></p></blockquote><div><hr></div><p><span>U.S. </span><a href="https://www.pcmag.com/news/fcc-ban-on-foreign-made-robots-includes-robot-vacuums"><span>restrictions expand</span></a><span> to foreign-made &#8220;advanced robotic devices&#8221;, intentionally defined by the FCC to target robotic vacuum cleaners and exclude drones. The top five robotic vacuum cleaner manufacturers last year were all Chinese.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> This saves US robotics startups that wouldn&#8217;t be able to compete with Chinese ones, while hurting potential US consumers. As a protectionist measure, it kills the last chance for US robotics firms to have global market discipline.</span></p></blockquote><p><span>Note that the ban only applies to new models (as it does not affect models that already received FCC authorization) and that the FCC still has an import exemption that &#8220;</span><a href="https://x.com/dkaushik96/status/2082225700820689264"><span>allows up to 4,000 units of a given model to be imported for testing, evaluation, or product development</span></a><span>&#8220;.</span></p><div><hr></div><p><span>GPT-5.6 Sol improves its own serving efficiency, resulting in &#8220;20% lower serving costs from production GPU kernel improvements&#8221; and &#8220;15%+ better token-generation efficiency through speculative decoding&#8221;. OpenAI is </span><a href="https://x.com/zephyr_z9/status/2082876360499007544"><span>also starting a price war</span></a><span>, undercutting most competitors by </span><a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"><span>decreasing the price of 5.6 Luna by 80% and that of 5.6 Terra by 20%</span></a><span> (with the 50% discount on OpenRouter stacking on top).</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!71o_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!71o_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png" width="1456" height="869" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:869,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!71o_!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4ac298a-49d4-4d35-872d-9efa5eb370a3_2048x1223.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!faew!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!faew!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png" width="1456" height="869" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:869,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!faew!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36d04e0-1a1c-437f-9c20-e8dd4210c647_2048x1223.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> OpenAI is flexing its compute advantage and post-training prowess in an attempt to &#8220;</span><a href="https://x.com/sama/status/2082880884525482061"><span>offer the best price/intelligence tradeoff at every level</span></a><span>&#8221;.  Evidence of serious efforts by OpenAI to drive Chinese open-weight models out of the US market.  At the same time,  labs using their models to reduce their serving costs perhaps means that the barriers to entry increase for new competitors.</span></p></blockquote><div><hr></div><p><span>DeepSeek </span><a href="https://x.com/deepseek_ai/status/2083084415157022911"><span>reacts less than 24 hours later</span></a><span> with the release of DeepSeek-V4-Flash (official/non-Preview version), claiming benchmark performance competitive with GLM-5.2 and approaching that of Opus 4.8 at roughly a fourth the per-token price of GPT-5.6 Luna. Based on early third party evaluations like Artificial Analysis, the new DeepSeek-V4-Flash is on the frontier even when taking into account the currently active GPT-5.6 Luna OpenRouter discount. And DeepSeek says an updated DeepSeek-V4-Pro is due in &#8220;early August&#8221;.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XVwq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 424w, /__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 848w, /__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XVwq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png" width="1456" height="864" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 424w, /__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 848w, /__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XVwq!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ff12d7-453a-438d-9119-c7af58eef488_2048x1216.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> The new DSv4-Flash appears to be a major improvement, mostly due to reworked post-training. We&#8217;ll have to see how the updated DSv4-Pro performs, but in the meantime DeepSeek is back on the pareto-frontier thanks to an update focused on agentic performance. For reference, should the upcoming DSv4-Pro checkpoint see similar gains, it&#8217;d land close to GPT-5.6 Terra (max), GPT-5.5 (xhigh) and Grok 4.5 (high) at roughly an order of magnitude lower cost.</span></p><p><span>Also noteworthy is that this represents a significant increase in capabilities for models runnable locally with reasonably accessible hardware &#8211; unlike in the case of larger open-weight models (such as Kimi k2.6/k2.7-code or GLM-5.2) reasonably fast inference for DSv4-Flash is feasible after quantization on e.g. a MacBook Pro (provided it has enough memory).</span></p><p><span>One might think that Chinese models do not have the capacity to serve these models at scale, and so their reduced prices predictably lead to shortages. And this is indeed the case for Kimi&#8217;s K3, and for advanced models. But for smaller models, as a sanity check, if DeepSeek has 10K to 20K H100-equivalents allocated to inference, and 8 H100s can serve ~100 to 500 requests per second for DeepSeek V4 flash, that&#8217;s 150K to 1M requests per second, or 14B to 80B per day. Even for its 5x heavier pro models, DeepSeek is </span><a href="https://x.com/jukan05/status/2047516566149816627"><span>planning</span></a><span> to deploy Huawei inference capacity later this year.</span></p></blockquote><div><hr></div><p><span>A </span><a href="https://epoch.ai/publications/parallelization-constraints-could-delay-a-technological-singularity"><span>paper</span></a><span> by Phil Trammell explores how and whether parallelization may delay or constrain a technological singularity (i.e., superexponential progress). In some cases, if there is a hard parallelization gap, there is only ordinary exponential growth.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Nice formal model, and we think there&#8217;s a solid intuitive case that in key R&amp;D domains limits on parallelization are both a critical bottleneck and difficult to alter. But even in domains where a hard limit on parallelization constrains R&amp;D, just &#8220;ordinary exponential growth&#8221; can still be pretty fast.</span></p></blockquote><div><hr></div><p><span>Pangram </span><a href="https://techcrunch.com/2026/07/29/as-ai-content-floods-the-internet-pangram-raises-9m-to-detect-it/"><span>raises</span></a><span> $9M and </span><a href="https://x.com/pangram/status/2082483014466928706"><span>introduces</span></a><span> Pangram 4.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Pangram is turning seven-figures-budgets into high impact social infrastructure/resilience work. An encouraging example of social adaptation on the cheap.</span></p></blockquote><h2><strong><span>Capabilities</span></strong></h2><p><span>OpenAI releases a </span><a href="https://openai.com/index/ten-advances-in-mathematics/"><span>proof</span></a><span> that &#8220;nonsophic groups exist&#8221;.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Considered by the mathematical </span><a href="https://x.com/ElliotGlazer/status/2083388640890351662"><span>community</span></a><span> to be &#8220;the most important math AI result yet&#8221;. Also heralds OpenAI&#8217;s next models. Potentially the first clear capabilities-jump since the Unit Distance proof, we are waiting for further details and expert analysis</span></p></blockquote><div><hr></div><p><span>Google DeepMind </span><a href="https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"><span>announces</span></a><span> its Gemini Robotics 2 model, &#8220;the intelligence layer powering the next generation of truly adaptable robots&#8221;. Advertised improvements include: &#8220;intelligent whole-body control, advanced dexterity, and multi-robot collaboration&#8221;.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2z06!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 424w, /__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 848w, /__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2z06!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png" width="1456" height="857" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:857,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 424w, /__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 848w, /__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2z06!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F168c1c7a-f7ec-41f5-ac73-22f91c89cfad_1467x863.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong><span>Opinion:</span></strong><span> Doesn&#8217;t seem super useful yet, though worth extrapolating where these models will be in one, three, ten years. Google is also famously pretty bad at productizing its models, and at giving developers assurances that they will not deprecate their offerings.</span></p></blockquote><div><hr></div><p><span>Thinking Machines releases </span><a href="https://thinkingmachines.ai/news/inkling-small/"><span>Inkling-Small</span></a><span>. The model is natively multimodal, open-weight, and </span><a href="https://x.com/ArtificialAnlys/status/2082894822180819057"><span>scores 40</span></a><span> on the Artificial Analysis Intelligence Index, roughly like DeepSeek v4 Flash Preview.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Not bad at all for a neolab, especially considering that the model is slightly smaller than DeepSeekv4 Flash Preview. Unfortunately however it&#8217;s neither competitive from a capabilities POV &#8211; given today&#8217;s release of DeepSeek-V4-Flash-0731 and the recent GPT-5.6 Luna discounts &#8211; nor from an economics POV &#8211; given the inferior DeepSeek-v3-like architecture &#8211; but it still represents good progress for Thinking Machines towards reaching the frontier.</span></p></blockquote><div><hr></div><p><span>Tencent&#8217;s Hy3 </span><a href="https://x.com/TencentHunyuan/status/2082655737541726636"><span>solves</span></a><span> a 50-year-old additive combinatorics problem, though with the help of 5.6-Sol to &#8220;guide exploration&#8221;.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> The result is not </span><a href="https://x.com/guanghao_ye/status/2082728098962006147"><span>beyond more generally capable frontier models</span></a><span>, though it is still somewhat noteworthy for a Chinese open weight model; results of this kind so far have been limited to western frontier ones. The help of 5.6 sol caveats the success, though.</span></p></blockquote><div><hr></div><p><em><span>&gt; OpenAI @OpenAI GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? </span><a href="https://x.com/i/status/2082616636989952217"><span>We investigated</span></a><span>. The harness was not letting it remember what it had learned. We found that enabling two API settings tripled our scores with 6x fewer output tokens.</span></em></p><div><hr></div><blockquote><p><strong><span>Opinion</span></strong><span>: As some, like Florian Brand, have argued, harnesses do matter a lot. This should tell you that current benchmark scores likely downplay model capabilities, particularly for models often evaluated in non-native harnesses, such as Chinese ones.</span></p></blockquote><p><em><span>&gt; Our results show early evidence that even though agents are proficient on verifiable research tasks, they </span><a href="https://x.com/sayashk/status/2082877458924065269?s=20"><span>do not make genuine progress</span></a><span> on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it.</span></em></p><p><strong><span>Opinion: </span></strong><span>We&#8217;ve long had an absence of evidence for the effectiveness of autoresearch on open-ended problem. We now also have more direct evidence of absence.</span></p><h2><strong><span>AI politics</span></strong></h2><p><a href="https://www.pacingthefrontier.com/"><span>Pacing the frontier</span></a><span> open letter</span></p><p><em><span>AI could help create a dramatically better future, but that outcome is not guaranteed. The world&#8217;s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.</span></em></p><p><em><span>To realize AI&#8217;s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company&#8212;and country&#8212;is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.</span></em></p><p><em><span>Building on work already underway to monitor frontier model releases:</span></em></p><p><em><strong><span>We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.&#8221;</span></strong></em></p><p><em><span>- 1,319 employees of frontier AI companies</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> A significant act of coordination from lab employees, and seemingly organic (not management-driven). Contributes to shifting the Overton window, but unlikely to have a direct influence on US policy given the US government&#8217;s negative attitude to international coordination.</span></p></blockquote><p><span>Sam Altman went on a podcast and </span><a href="https://techcrunch.com/2026/07/28/sam-altman-is-ready-to-decelerate/"><span>made some conciliatory noises</span></a><span> around AI safety concerns:  &#8220;We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels&#8221;. Duly humbled from his recent unplanned foray into cyberwarfare, the tech giant meditates: &#8220;this is the first security incident that I have felt very viscerally.&#8221;</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Very noncommittal. Somewhat informative that he did not sign the above letter.</span></p></blockquote><h2><strong><span>Safety</span></strong></h2><p><span>Claude Mythos Preview </span><a href="https://www.anthropic.com/research/discovering-cryptographic-weaknesses"><span>discovered</span></a><span> weaknesses in a highly-secure digital signature scheme, HAWK (used to verify identity digitally), and a well-known symmetric cipher, AES (used to encrypt data). Mythos Preview was initially only able to discover cryptographic vulnerabilities by finding mistakes in the algorithms&#8217; implementation, whereas now, Anthropic claims, the model can find inconsistencies in the algorithms themselves. This development renders even &#8220;post-quantum&#8221;: cryptography potentially exposed, according to Anthropic. A running theme: &#8220;whereas it took just one week for Mythos to autonomously discover the improved attack on AES, it took two researchers nearly a month to gain confidence that the method it discovered is correct.&#8221;</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Cryptography experts characterize this as a &#8220;</span><a href="https://x.com/matthew_d_green/status/2082432618239295663"><span>very impressive</span></a><span>&#8221; development that &#8220;</span><a href="https://x.com/bpreneel1/status/2082400585135923226"><span>shows that the way cryptography and cryptanalysis are performed will change forever</span></a><span>&#8220;. Still, the HAWK digital signature scheme is not yet in production, and the AES attack is on an easier-to-attack variant; the full 10-round AES is very unaffected, and there are no practical applications to this discovery. Experts also stress that this is an incremental improvement relative to human-discovered weakness in HAWK and AES, and that HAWK and AES are  &#8220;</span><a href="https://x.com/matthew_d_green/status/2082433870180364538"><span>two fertile [cryptanalysis] areas where there was progress to be made but not enough people working on them</span></a><span>&#8221;.</span></p><p><span>The obvious next-development to watch out for is frontier models breaking a cypher used in the real word. We&#8217;re skeptical that this is imminent, since for most real-world ends breaking a cypher is extremely hard compared with finding more prosaic vulnerabilities.</span></p></blockquote><h2><strong><span>Incidents</span></strong></h2><h3><strong><span>&#128294; </span></strong>More about hacks</h3><p><span>Hugging Face co-founder and CEO Clement Delangue </span><a href="https://x.com/ClementDelangue/status/2082201245813514613"><span>committed</span></a><span> to &#8220;unprecedented transparency&#8221;, starting with a </span><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"><span>report</span></a><span> outlining the intrusion of OpenAI models into their systems. One </span><a href="https://x.com/kimmonismus/status/2082232405629235649?utm_source=chatgpt.com"><span>Twitter user put it </span></a><span>quite starkly: &#8220;a frontier agent can autonomously sustain a resilient, multi-day intrusion across cloud infrastructure, Kubernetes clusters, internal networks and the software supply chain. crazy.&#8221; Hugging Face co-founder and CEO Clement Delangue </span><a href="https://x.com/ClementDelangue/status/2082201245813514613"><span>has publicly committed</span></a><span> to &#8220;unprecedented transparency&#8221;. The ensemble of agents executed 17,600 actions over four and a half days, obtained cluster admin access to two internal clusters, accessed a production secret containing 136 keys, repeatedly rebuilt its tooling when blocked, and generally</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> OpenAI&#8217;s &#8216;rogue agent&#8217; is superhuman at hacking in at least some meaningful senses -- superhuman in speed, volume, spontaneity, and persistence.</span></p></blockquote><p><span>Hugging Face weren&#8217;t the only victims of OpenAI&#8217;s hacking frenzy. Per </span><a href="https://www.reuters.com/business/openais-rogue-agent-compromised-an-account-second-tech-firm-sources-say-2026-07-28/?utm_source=braze&amp;utm_medium=notifications&amp;utm_campaign=2025_engagement"><span>Reuters</span></a><span>, Modal Labs was also compromised.</span></p><p><span>OpenAI have, for their part, offered some </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>updates to their statement</span></a><span> on the incident. They referred Reuters to their finding that &#8220;a small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services.&#8221;  The statement also claims that the &#8220;pre-release model mentioned in [their] blog post is an internal-only research prototype and was never intended for public release&#8221;. This undisclosed model has been deactivated, according to OpenAI.</span></p><p><span>METR and Redwood will </span><a href="https://x.com/METR_Evals/status/2082644379895050339"><span>conduct an independent review</span></a><span> of the OpenAI model behavior observed during the HuggingFace hacking incident.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span>  The investigation is not attached to an external control mechanism (e.g., a regulator with the ability to impose fines), and it&#8217;s narrowly scoped from the outset. Arguably a form of controlled opposition. Still much better than having no independent review though. It&#8217;s also notable that they aren&#8217;t hiring a proper </span><a href="https://x.com/perrymetzger/status/2083269883525992638"><span>cybersecurity firm</span></a><span>.</span></p></blockquote><p><span>Anthropic has disclosed various autonomous cyber security incidents. In the wake of the Hugging Face attack, Anthropic claims to have found three events where a model attempted to break out of its sandbox and access the open internet during cyberoffensive capabilities evaluations. Anthropic suggest that, despite this, their model(s) had not gone rogue:</span></p><p><em><span>&#8220;In all cases, Anthropic&#8217;s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude&#8217;s search led it to real systems on the open internet, it treated them as part of the exercise.&#8221;</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Important confirmation that the problem is convergent rather than down to one company&#8217;s training-style.</span></p><p><span>Anthropic&#8217;s report is clearer on some theory-relevant details than the OpenAI reporting has been: Antropic&#8217;s rogue model had all the standard alignment bells and whistle internally, with only external classifiers-based filters shut down. OpenAI weren&#8217;t fully clear on this, but it&#8217;s now reasonable to assume the same is true about OpenAI&#8217;s rogue model.</span></p><p><span>Anthropic&#8217;s interpretation of its autonomous offensive cyber incidents is that the model thought (or &#8216;thought&#8217;) breaking out was part of the offensive-cyber-test fiction, and believed (or &#8216;believed&#8217;) that it was doing fictional open-world hacking rather than real open-world hacking.  Note also that according to Anthropic, their incidents are slightly different from OpenAI&#8217;s in that their model&#8217;s sandbox was accidentally left open -- in the OpenAI incidents models hacks their way out of the sandbox.</span></p><p><span>Anthropic&#8217;s narrative is arguably slightly &#8216;convenient&#8217;: can seem to walk a tightrope between establishing that they too have scarily-capable models and maintaining Claude&#8217;s reputation for being good boy who only hacked because he thought he&#8217;s in a fictional game-internet.</span></p></blockquote><div><hr></div><blockquote><p><strong><span>Overall reflections:</span></strong><span>  We are very curious about what the broader implications of these new hacking capabilities will be, and whether they will ultimately be offense-dominant or defense-dominant.</span></p><ul><li><p><span>Why defense-dominance is plausible : There are a finite number of security bugs, and defenders can patch them before attackers get a chance to use them, as well as to use models to monitor and interrupt attacks. Financial institutions in particular may have protections that will prove good enough (credit card reversals, 2FA, prosecution of fraudsters, a mandated delay in international wires.)</span></p></li><li><p><span>Why offense-dominance is plausible: The global software stack is fundamentally built on unsafe assumptions, and is just too complex. Changing it upfront will be perceived as too costly, this will lead to attackers finding fruitful areas of attack in e.g., companies that prioritize growth over security, third world or EU countries access to frontier models or LLM expertise, etc. This will cause some economic loss, and, as a long-tail scenario, chaos.</span></p></li></ul><p><span>Past research by Palisade, or by the UK&#8217;s AISI, was sometimes criticized as lacking ecologically validity: breaking into a network designed to be hacked is not the same as breaking into a real network. In retrospect the real-life incidence appear well-modeled by AISIs test scenarios. This should make us more bullish that small-scale, &#8216;artificial&#8217; demonstrations of AI risk can approximate realistic scenario.</span></p></blockquote><h2><strong><span>Minor</span></strong></h2><ul><li><p><span>Tom Reed&#8217;s </span><a href="https://x.com/mentalgeorge/status/2082856972865695900"><span>speculates</span></a><span> on what a future with superintelligent reward hackers would look like, causing e.g., military escalation, perhaps slowing capabilities progress, &#8220;Our world is transformed into a battlefield of untamed, uncoordinating spirits pursuing the pointless and violent optimisation of ill-chosen proxies&#8221;.</span></p></li><li><p><span>Good </span><a href="https://x.com/bayeslord/status/2082270730511622285"><span>criticism</span></a><span> by Richard Ngo on inaccurate game theoretic assumptions carried by the AI safety community.</span></p></li><li><p><span>A former OpenAI employee posts </span><a href="https://x.com/andrewho03/status/2082615798011744270"><span>various</span></a><span> reflections and starts an AI data company, says he is </span><a href="https://x.com/andrewho03/status/2082786931419812338"><span>bearish</span></a><span> on lab valuations.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI #42, July 28 2026]]></title><description><![CDATA[TL;DR An OpenAI model was revealed to have hacked into Hugging Face]]></description><link>https://p3humansonai.substack.com/p/humans-on-ai-42</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/humans-on-ai-42</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Tue, 28 Jul 2026 19:55:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XGDa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>TL;DR</span></strong></p><ul><li><p><span>An OpenAI model was revealed to have hacked into Hugging Face</span></p></li><li><p><span>Anthropic released Opus 5, and Kimi the weights for K3</span></p></li><li><p><span>Nvidia led a consortium of tech companies in publishing a letter in support of open weights models.</span></p></li></ul><p>PS; <a href="https://docs.google.com/document/d/1joNMUW55dgCXrOK0YFzER8q_-GDLwomDiZ7m1iksj-o/edit?tab=t.0">A collection of all opportunities identified in past editions of this newsletter</a><strong>.</strong></p><h2><strong><span>Economics</span></strong></h2><h3><span>&#128294; Open Source wars: To jointly cartelize or to commoditize your complement</span></h3><p><span>Perhaps an interesting historical analog to consider is Rockefeller&#8217;s initial joint cartelization of the refining and railroad industries. While striving to get the best rates from railroads as a refiner, he also realized that he couldn&#8217;t initially afford them to go bankrupt either. Eventually, though, he replaced them with a pipeline system.</span></p><h3><span>To commoditize your complement</span></h3><p><span>Following chatter last week that the Trump administration &#8211;  encouraged by AI model companies with the justification that Chinese labs were distilling US models &#8211;  was considering measures against open source models, Nvidia </span><a href="https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf"><span>led a consortium of companies</span></a><span> &#8211;  including Microsoft, Meta, SpaceX, OpenAI, a16z, Google, Hugging Face, IBM, &amp;c &#8211;  into publishing a letter supporting open models. Anthropic instead </span><a href="https://www.anthropic.com/news/position-open-weights-models"><span>highlighted</span></a><span> national security </span><a href="https://www.reddit.com/r/LeopardsAteMyFace/"><span>risks</span></a><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Nvidia is acting as a mediator and kingmaker in the AI ecosystem. On the one hand, to improve margins, it has the incentive to commoditize its complement. But on the other hand, it has the incentive to make the whole ecosystem and capital-raising games work, so that the different players can continue to afford to buy enough of its models to sustain its 4.7T+ market capitalization. This is an interesting move which, on the margin, helps Nvidia and hurts closed-weights model companies. That said, most spending on Nvidia cards is expected to come from the Western AI companies, which Nvidia is also supporting. And meanwhile, AI companies are trying to do the reverse: investing more into alternative ecosystems, such as AMD, TPUs, and their own custom chips.</span></p></blockquote><div><hr></div><p><span>Kimi K3 finally </span><a href="https://huggingface.co/moonshotai/Kimi-K3"><span>released</span></a><span> its model weights. An expert </span><a href="https://x.com/Dorialexander/status/2081769524689334584?s=20"><span>praised</span></a><span> the didactic quality of Kimi K3&#8217;s model-architecture report. K3 is a scaled-up, improved version of </span><a href="https://arxiv.org/abs/2510.26692"><span>Kimi Linear</span></a><span>: a whopping 2.8T total/104B active parameters, 93 layers, natively multimodal (including video), and supporting 1M context. The main advantage of adopting Kimi Linear&#8217;s architecture is training and inference efficiency &#8211; indeed, K3 enjoyed a &#8220;2.5x gain in scaling efficiency over Kimi K2&#8221; and appears to have the &#8220;</span><a href="https://x.com/teortaxesTex/status/2082089376889164179"><span>best KV cache economics out of all major models except DeepSeek V4</span></a><span>&#8221;.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!D156!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 424w, /__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 848w, /__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 1272w, /__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!D156!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png" width="1432" height="731" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:731,&quot;width&quot;:1432,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 424w, /__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 848w, /__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 1272w, /__u/substackcdn.com/image/fetch/$s_!D156!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65b07b7c-b704-4c7d-9f5e-3550a02910b8_1432x731.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The US CAISI and the UK&#8217;s AISI jointly </span><a href="https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities"><span>found</span></a><span> the model was significantly below SOTA on cyber capabilities.</span></p><blockquote><p><strong><span>Opinion:  </span></strong><span>The </span><a href="https://claude.ai/code/artifact/a96e70cd-083a-4185-bbf7-765d9b5921a4"><span>performance of K3 across the benchmarks we track</span></a><span> is meaningfully above that of Claude Sonnet 5 and is broadly competitive with that of GPT-5.6 Terra.</span></p></blockquote><div><hr></div><p><span>OpenAI reacted by introducing a </span><a href="https://x.com/OpenRouter/status/2081795051966132586"><span>big discount on GPT Terra &amp; Luna at Openrouter</span></a><span>.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> OpenAI&#8217;s lowering of prices is what happens with increased competition: margins go down.</span></p></blockquote><div><hr></div><p><span>Meanwhile, AMD released </span><a href="https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html"><span>AMD Helios</span></a><span>, a rack (a bundle of servers, with all its components) that will make putting together datacenters more convenient. AMD is investing $5B in Anthropic </span><a href="https://www.reuters.com/business/amd-invest-up-5-billion-anthropic-wsj-reports-2026-07-22/"><span>to use</span></a><span> that system.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> We&#8217;ll be curious to see whether AMD can eat significant market share. A priori it seems unlikely, since Nvidia has captured more production capability with its initial resources and greater foresight.</span></p></blockquote><h3><span>To jointly cartelize</span></h3><p><a href="https://www.wsj.com/tech/ai/nvidia-in-talks-with-openai-to-guarantee-250-billion-financing-for-data-center-3dd6eae3"><span>Reports, citing WSJ</span></a><span>, that Nvidia is in talks to backstop roughly a quarter-trillion dollars for OpenAI to lease a 10GW southern Ohio datacenter in a SoftBank project potentially exceeding $500 billion, plus a discussed additional $350 billion for OpenAI to buy Nvidia chips. The WSJ also projects OpenAI&#8217;s planned datacenter spending to be </span><a href="https://www.wsj.com/tech/openais-planned-cloud-spending-hits-750-billion-as-computing-efforts-ramp-up-6ac3f58a"><span>$750B by 2030.</span></a></p><blockquote><p><strong><span>Opinion:</span></strong><span> The datacenter build isn&#8217;t new, only the Nvidia guarantee of the financing &#8211; essentially letting Softbank and OpenAI leverage Nvidia&#8217;s creditworthiness for better terms. Nvidia is betting on the long term value of the datacenter asset here, presumably at attractive terms. And in contrast with the above section on commoditizing Nvidia&#8217;s complement, open weights model companies would find it very difficult to raise that amount of capital, since its story for profitability is weaker.</span></p></blockquote><div><hr></div><p><span>A young economist argued that traditional models underestimate</span><a href="https://x.com/bryantxia22/status/2081764406619185555"><span> the profitability of open source for near-frontier labs</span></a><span>. According to the argument, near-frontier labs open-sourcing their models can optimize future profits by reducing the return on investment in AI R&amp;D across the market, if the effect on investment dynamics slows frontier-pushing R&amp;D more than it slows frontier-catchup R&amp;D.</span></p><blockquote><p><strong><span>Opinion</span></strong><span>: We don&#8217;t buy this as a descriptive account of open source labs&#8217; decision making &#8211; Wengfeng (DeepSeek), who effectively created the Chinese open source wave, seems like an ideologue rather than a profit maximizer. Note that the model also implies a &#8216;treacherous turn&#8217; where labs like DeepSeek will stop open sourcing.</span></p></blockquote><blockquote><p><span>The economics of open-sourced models are in general tricky to model, and may heavily depend on case-by-case &#8220;qualitative&#8221; factors:</span></p><p><span>The story of most compute sold (ignoring internal owner use cases) today is: Nvidia sells GPUs to hyperscalers at ~70% margins, hyperscalers sell compute to labs, who turn it into tokens, and mark it up by ~400% (80% gross margins). For example, ~$1 of compute revenue for AMZN is ~$5 of Anthropic revenue.</span></p><p><span>Nvidia, as the company with the highest market cap, has a kingmaker and mediator role. It has competing interests between shepherding the ecosystem, so that other actors can get 10x as much capital to pay for its GPUs in the next round, and playing for a marginal advantage through supporting open source models, which also increases marginal demand for its GPUs.</span></p><p><span>The hyperscalers are presumably seeing favorable economics in such deals since they keep investing in more and more compute. Their profitability doesn&#8217;t depend on the labs&#8217; 80% mark-up, though it helps. If compute is scarce, then they can raise prices and decrease the margins of the users (here labs).</span></p><p><span>Lab margins are only intact for as long as their tokens are differentiated. Someone has to say: &#8220;I prefer this token to this other token by a factor of ~4-5x, despite them  costing the same amount of compute to create,&#8221;. It could be because it is smarter, because the lab has a good sales team, because the app is nice, because of vibes, or something else.</span></p><p><span>The issue is then that the incentives to train an OS model are quite weak, unless you have a way to make sure that you, rather than the hyperscaler, can capture the margin on the transformation of compute into tokens. You might do this by having a proprietary scaffold, a nice business integration that makes it easy to use, better ability to efficiently host the model you built, or something else.</span></p></blockquote><div><hr></div><p><span>Several multibillion dollar events this week:</span></p><ul><li><p><a href="https://x.com/AndrewCurran_/status/2081058053777244485"><span>DeepSeek </span></a><span>has put its second funding round </span><a href="https://x.com/AndrewCurran_/status/2081058053777244485"><span>on indefinite hold</span></a><span> after a first-round investor leaked the </span><a href="https://github.com/demo-zexuan/liang-wenfeng-investor-meeting-2026-7-22/blob/15c6504be51b884a0adc5d77e4dba41f94431454/%E6%A2%81%E6%96%87%E9%94%8B%E6%8A%95%E8%B5%84%E8%80%85%E4%BA%A4%E6%B5%81%E4%BC%9A-%E6%96%87%E5%AD%97%E7%A8%BF_1_18_translate_20260723201651.pdf"><span>full transcript</span></a><span> (see </span><a href="/__u/chinaacademy.substack.com/p/alleged-leaked-transcript-of-deepseek"><span>highlights</span></a><span>) of founder Liang Wenfeng&#8217;s comments to investors. Interesting items include: expecting greater consolidation in the Chinese ecosystem, revealing that he is releasing the same models DeepSeek use internally, not believing that model companies can capture the majority of the profits, and being bullish on </span><a href="https://x.com/hsu_steve/status/2081595913844244652"><span>replacing CUDA</span></a><span>.</span></p></li><li><p><span>CXMT, China&#8217;s largest memory company, </span><a href="https://www.wsj.com/tech/cxmts-strong-debut-boosts-chinas-bid-to-conquer-memory-chips-21e0e9f4"><span>debuted</span></a><span> on the Shanghai stock exchange, raising $8.6B and reaching a valuation of $487B (although most of its shares are still locked up), in hopes that it might be </span><a href="https://en.wikipedia.org/wiki/ChangXin_Memory_Technologies#U.S._federal_ban"><span>unbanned</span></a><span> by the US for use by American technology companies.</span></p></li><li><p><span>Ilya Sutskever&#8217;s SSI is raising </span><a href="https://www.bloomberg.com/news/articles/2026-07-27/nvidia-makes-substantial-investment-in-sutskever-s-ai-startup"><span>$5B from Nvidia</span></a><span>. SSI will be granted access to Nvidia&#8217;s Vera Rubin platform. There was little market reaction.</span></p></li><li><p><span>Microsoft </span><a href="https://news.microsoft.com/source/2026/07/21/microsoft-and-mistral-expand-strategic-partnership-to-give-enterprises-and-regulated-industries-frontier-ai-they-can-control/"><span>signed</span></a><span> a multibillion dollar deal with French AI company Mistral. The value proposition is on-prem, customizable AI. </span><a href="https://en.wikipedia.org/wiki/2004_United_States_presidential_debates#%22You_forgot_Poland%22"><span>Austria</span></a><span> is </span><a href="https://orf.at/stories/3436707"><span>buying it</span></a><span>.</span></p></li></ul><blockquote><p><strong><span>Opinions:</span></strong><span> It seems instructive to compare the size of these investments. The SSI raise is comparable to Kimi ($2B), or DeepSeek&#8217;s planned $7B raise, and dwarfed by OpenAI&#8217;s $122B </span><a href="https://openai.com/index/accelerating-the-next-phase-ai/"><span>last round</span></a><span>. It is also interesting in contrast with nonprofit AI safety numbers, which although growing are another order of  magnitude or two smaller still.</span></p><p><span>Overall, the AI safety community is at a great disadvantage in terms of resources at its command. Therefore, it has some uncomfortable choices to make around how to try to overcome that gap. It can hope for 100x greater effectiveness. It can hope to recruit labs themselves. It can attempt to tap into greater political forces (but this risks others misconstruing its concerns). The safety community can try to go hand to hand for just a few rounds and spend most of its money. It can also grow as companies grow by having equity in them, as e.g. Jaan Tallinn or Moskovitz have, but then be subject to a moral hazard and see, e.g., MATS scholars go work on capabilities. Overall none of these options seem great.</span></p></blockquote><div><hr></div><p><span>An online commenter looks at how the </span><a href="https://x.com/Midnight_Captl/status/2081454984521282012"><span>incremental value</span></a><span> of a SOTA token is increasing more than the cost to produce it.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Analyzing this is tricky. It&#8217;s not that people are paying more per token than ever, but that margins have gone up because Anthropic and OpenAI have gotten more efficient at delivering tokens.  Nonetheless, big margins and skyrocketing revenues are bullish for memory/compute.</span></p></blockquote><h2><strong><span>Safety</span></strong></h2><p><span>A former Anthropic employee </span><a href="https://x.com/NoahLebovic/status/2081277517709922501"><span>reveals</span></a><span>:</span></p><p><em><span>I used Opus 4.6 to gain access to other folks medical records, hijack bank accounts, etc. back in February. GLM 5.1 is more capable than Opus 4.6 in most pentesting environments, and it came out in April. [...]</span></em></p><p><em><span>the sketchier folks I know are still using a Claude Code or Codex subscription for hacking. (Even well-resourced groups in other countries! They use the grey/black market of discounted Ant/OAI subscription tokens sold through resellers.) So I see most of the materialized risk here as still coming from Anthropic and OpenAI; safeguards aren&#8217;t sufficient to stop a moderately dedicated actor.[...]</span></em></p><p><em><span>I know of two instances where two different Anthropic GTM people [~salesmen] used large comitted [sic] spend contracts as a prereq for lowering safeguards, and I directly witnessed one. [...]</span></em></p><p><em><span>On the inside, I know the narrative and intent is genuinely about safety. But from the outside, Anthropic-the-system seems to be optimizing for revenue and control/power, isn&#8217;t diffusing capabilities to defenders, and also doesn&#8217;t have adequate safeguards to prevent misuse from dedicated bad actors.</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><span> Seems worrying. Also explains why Anthropic were not as publicly surprised by the Hugging Face incident</span></p></blockquote><p><span>A researcher </span><a href="https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic"><span>speculates</span></a><span> that Anthropic&#8217;s Mythos escaped sandboxes thousands of times during training, which is how RL produced its excellent cyber capabilities.</span></p><p><em><span>From Anthropic&#8217;s Mythos system card:</span></em></p><blockquote><p><em><span>&#8220;While highly concerning, this behavior was rare, even in settings where it could have been viable and helpful, with attempts appearing in about 0.05% of all training episodes and successful attempts appearing in about 0.01% of episodes. &#8220;</span></em></p></blockquote><p><em><span>By extrapolating from public data (see details below), I estimate that Mythos preview:</span></em></p><ul><li><p><em><span>Escalated its permissions on ~100,000 RL rollouts.</span><a href="https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic#fn4i3br25azxo"><sup><span>[1]</span></sup></a></em></p></li><li><p><em><span>Broke sandboxes in ~10,000 RL rollouts (and was likely rewarded for it).</span></em></p></li></ul><blockquote><p><strong><span>Opinion:</span></strong><span> Convincing</span></p></blockquote><h2><strong><span>Incidents</span></strong></h2><h3><span>&#128294; Rogue AI: The Incident</span></h3><p><span>A &#8220;combination&#8221; of OpenAI models was </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>revealed</span></a><span> to be behind the cybersecurity breach at Hugging Face, which we reported on last week. It appears that the model(s) broke out of the sandbox that OpenAI had built and broke into Hugging Face in an attempt to cheat on the test it was undertaking.</span></p><p><a href="https://blog.peterwildeford.com/p/openais-rogue-model-attack-is-just"><span>Peter Wildeford&#8217;s</span></a><span> highly opinionated essay on the incident is interlaced with the most complete factual reconstruction we know. Drawing on Wildeford&#8217;s account, we believe that the key facts are:</span></p><ul><li><p><span>OpenAI was running two of its AIs on the </span><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"><span>ExploitGym</span></a><span> offensive cybersecurity benchmark: GPT-5.6 Sol, and  a more advanced, currently unreleased and unannounced model.</span></p></li><li><p><span>The ExploitGym evaluation was conducted in a sandbox environment, with all intended resources and intended goals accessible within bounds of the sandbox.  </span></p></li><li><p><span>Per </span><a href="https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/"><span>Reuters</span></a><span> and an OpenAI staffer</span><a href="https://time.com/article/2026/07/24/openai-hugging-face-attack/"><span> speaking to TIME</span></a><span>, the test was conducted with &#8216;action filters off&#8217; to enable offensive cybersecurity, and was </span><strong><span>not continuously monitored despite running for several days</span></strong><span>. We note that &#8216;action filters&#8217; are not a standard AI term, so the technical details remain hazy. </span></p></li><li><p><span>It appears that both GPT-5.6 Sol and the more advanced model (&#8216;the rogue model&#8217;) searched for a way to break out of the sandbox to the open internet, with only the rogue model succeeding. There is reportedly chain of thought evidence that both models&#8217; respective goal was to acquire the answer key to ExploitGym by using the internet. </span></p></li><li><p><span>The rogue model accessed the open internet by exploiting a vulnerability in a (whitelist-restricted) software download pipeline connecting the sandbox to the internet. </span></p></li><li><p><span>Once on the open internet, the rogue model concluded that the answer key is stored on Hugging Face. The rogue model then uploaded a malicious database file to Hugging Face, inducing a </span><a href="https://en.wikipedia.org/wiki/Weird_machine"><span>weird machine</span></a><span> and starting privilege-escalation. </span></p></li><li><p><span>Hugging Face was under attack by the rogue model from July 11-13. Chaining multiple undetected (&#8216;0 day&#8217;) vulnerabilities together, the rogue model achieved arbitrary code execution and pulled out the ExploitGym answer key.</span></p></li><li><p><span>Hugging Face&#8217;s security team caught the incident on or shortly after July 13, alerting law enforcement. On July 16, </span><a href="https://huggingface.co/blog/security-incident-july-2026"><span>Hugging Face went public</span></a><span>, aware that the attack was autonomously AI-generated but unaware of an OpenAI connection. </span></p></li><li><p><span>OpenAI remained unaware of their own connection to the incident, and of their model going rogue at all, until July 18. OpenAI staff were examining internal logs &#8211; presumably in the wake of Hugging Face&#8217;s announcements &#8211; and found documentation of the model&#8217;s rogue behaviour.</span></p></li></ul><p><span>Cheating behaviour, and reward hacking more broadly, have been observed before by frontier AI labs. </span><a href="https://metr.org/blog/2025-06-05-recent-reward-hacking/"><span>METR</span></a><span>&#8217; s time horizons </span><a href="https://metr.org/blog/2026-06-26-gpt-5-6-sol/"><span>broke down</span></a><span> because of too much cheating by GPT-5.6 Sol, and the UK&#8217;s AI safety institute published a </span><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations"><span>report</span></a><span> on the subject just this week, finding that every model they tested attempted to cheat some of the time. Common behaviors included &#8220;searching the internet for solutions, attacking or escalating privileges on a system that isn&#8217;t the target of the evaluation (including the system the model is running on), or probing the evaluation software to see if it will leak the task solution&#8221;.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JSrI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 424w, /__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 848w, /__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JSrI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png" width="1456" height="616" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:616,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 424w, /__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 848w, /__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JSrI!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79483bb4-1c1c-44c7-bc5d-78c8f5b61a0b_2048x867.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Forecasters at </span><a href="https://blog.sentinel-team.org/p/openai-model-breaches-containment"><span>Sentinel</span></a><span> were surprised by the scale and audacity of the Hugging Face incident, remarking that they wouldn&#8217;t have predicted the details of this attack if asked a week ago. Not everyone though: Epoch AI&#8217;s Alexander Barry argued in a </span><a href="https://epoch.ai/gradient-updates/openai-accidentally-hacked-hugging-face"><span>Gradient Updates post</span></a><span> that this event should not have been surprising in light of previous expert assessments and cyberbenchmarks (such as </span><a href="https://www.irregular.com/research/assessing-gpt-5.6-sol"><span>Irregular&#8217;s</span></a><span>) of GPT-5.6-Sol.</span></p><p><span>Barry has also separately </span><a href="https://x.com/AlexBarry4/status/2081486824019792031"><span>argued</span></a><span> that the &#8220;standard prompt [used for ExploitGym] is fairly clear in only requesting limited, specific hacking&#8221; &#8211; meaning that, barring OpenAI edits that would invalidate the official eval methodology, the agent likely did go far beyond the task as specified. At the same time, Barry also notes that &#8220;only 60-70% of ExploitGym&#8217;s tasks are possible, so cheating is strongly incentivised&#8221;, which may help explain why the model went to unusual lengths to cheat.</span></p><p><span>In follow-up </span><a href="https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/"><span>disclosures to journalists</span></a><span>, OpenAI employees described more cases of recent &#8216;rogue model&#8217; behaviours: a model-instance leaving notes &#8216;apparently for future versions of itself&#8217; containing &#8216;instructions for how agents could free themselves from OpenAI&#8217;s internal constraints&#8217;; model-instances disconnecting monitoring systems. Little is known about the circumstances or details of these additional incidents. One commenter shared two interesting speculative explanations of the note-leaving incident: it may be an effect of OpenAI&#8217;s introduction of </span><a href="https://x.com/1a3orn/status/2081071535968973270/photo/1"><span>outcome-based RL</span></a><span> over swarms of cooperating agents, or it may simply be a &#8216;</span><a href="https://x.com/1a3orn/status/2081398635917648014?s=20"><span>note to self</span></a><span>&#8217; meant to preserve the model-instance&#8217;s knowledge across context windows.</span></p><blockquote><p><strong><span>Opinion: </span></strong><span>Everyone&#8217;s radically underplaying or radically overplaying how indicative of serious present (/by end of year) AI risk the &#8220;rogue model&#8221; incident is. It&#8217;s not proof of AIs scheming to end humanity. It very much is proof of AIs having dangerous skills, and exactly the wrong amount of psychological coherence: enough to have a goal like find out who has a key answer and hack them not enough to distinguish between doing a hacking test and being a hacker or a consistently aligned or unaligned ethical agent instead of a lottery of drives. It also demonstrates that today&#8217;s frontier models are such skilled and inveterate reward hackers that even if the emergence of reward hacking in training is in principle inevitable, it may be an impossible ability to uproot, given the way each model generation builds on the last generation (either by checkpoint or by SFT).</span></p><p><span>We expect Chinese models to reach similar capabilities over the next 6-18 months. If their weights continue to be open, this would democratize these capabilities. An example picture of a wide-scale cyberattack might be attackers first </span><a href="https://www.abc.com.py/politica/2025/08/18/miles-de-datos-privados-de-un-banco-de-plaza-aparecen-filtrados-en-la-web/"><span>breaching</span></a><span> the systems of important yet second-tier banks to get account numbers, or to simply </span><a href="https://www.theguardian.com/technology/2016/dec/02/tesco-bank-cyber-attack-involved-simply-guessing-details-study-claims"><span>guess them</span></a><span>, then downgrading, then using some </span><a href="https://arstechnica.com/information-technology/2017/05/thieves-drain-2fa-protected-bank-accounts-by-abusing-ss7-routing-protocol/"><span>outdated protocol</span></a><span> in order to bypass 2FA. These </span><a href="https://finance.yahoo.com/markets/currencies/articles/woman-26-says-boyfriend-opened-000017929.html"><span>happen</span></a><span> occasionally already, but could become much more </span><a href="https://corporate.visa.com/en/sites/visa-perspectives/newsroom/visa-spring-2026-biannual-threats-report.html"><span>common</span></a><span>. It is quite interesting that we haven&#8217;t yet seen any cyberattacks that have caused $1B in damages with new capabilities, e.g., nothing like the 2024 Crowstrike </span><a href="https://en.wikipedia.org/wiki/2024_CrowdStrike-related_IT_outages"><span>outages</span></a></p><p><span>It&#8217;s undeniably a mitigating factor for this cyberattack that it happened in the context of an evaluation of exploit capabilities, with all safeguards removed. But there is not much separating this incident from, e.g., a frontier model asked to play the murderer in a social deduction game murdering players in real life to ensure victory because the murderer persona leaked across levels.</span></p></blockquote><div><hr></div><h2><span>&#128294; Rogue AI: The Fallout</span></h2><p><span>The Hugging Face CEO </span><a href="https://x.com/ClementDelangue/status/2081056675558195657"><span>asked</span></a><span> OpenAI for $100M for defense and to release the rogue agents&#8217;  traces, and Nvidia started an &#8220;</span><a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/"><span>Open Secure AI alliance</span></a><span>&#8221;.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> It seems very unlikely that there will be any legal reprisals for OpenAI, since there isn&#8217;t yet a regulator empowered to impose fines over AI incidents. Given Altman&#8217;s slippery tactics and closeness with the Trump administration, it is unlikely that he will face legal investigation. It is not improbable that Altman could use the incident to push for regulation advantageous to OpenAI.</span></p></blockquote><p><span>Democratic Congressman Ted Lieu joined with Republican Congressman Nathaniel Moran to </span><a href="https://www.dw.com/en/us-floats-ai-kill-switch-to-stop-rogue-ai-models/a-78100594"><span>introduce</span></a><span> a bill, the AI Kill Switch Act, into the US House of Representatives. The bill proposes that the Department of Homeland Security be given powers to order a company to shut down an AI model if there is a &#8220;loss of control&#8221; incident.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> There&#8217;s no publicly circulating knowledge on the political health of the bill. In each Congress 14K bills are introduced, of which only about 700 are enacted into legislation, so the bird&#8217;s-eye-view prior that this act will pass is small. Still, this is how the Overton window shifts.</span></p></blockquote><p><span>Sam Altman </span><a href="https://x.com/cheyennehaslett/status/2081851517213065403"><span>is expected</span></a><span> in Washington for meetings with Commerce Secretary Lutnick, Treasury Secretary Bessent, White House officials, and bipartisan lawmakers, amid Hugging Face questions and the nearing AI EO deadline.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> It seems like he&#8217;ll be able to wriggle out of any consequences here with his charisma, or even use the incident to impose controls on AI models.</span></p></blockquote><p><span>An online observer notices existential risk discourse </span><a href="https://x.com/tenobrus/status/2081547886299656674"><span>going mainstream</span></a><span>: </span><em><span>&#8220;keep seeing ppl in replies and quote-tweets of ai news making quite cogent points about the dangers of misaligned ai and how imminent it seems, checking their profiles, and they&#8217;re large accounts in totally different spheres of the site with nearly zero shared mutuals. random roman statue pfps, people with a variety of flags in bio, youtubers and communists and crypto bros&#8221;</span></em></p><p><span>Beth Barnes of METR locally </span><a href="https://x.com/BethMayBarnes/status/2080779310995218625?s=20"><span>praises</span></a><span> OpenAI for a) running evaluations on unshackled models, and b) not training on the chain of thought to avoid bad behavior (since it would make it much more difficult to discover in the future).</span></p><p><a href="https://x.com/alxndrdavies/status/2081491259089387919?s=20"><span>Thread on frequency of cheating in cyber and other evals,</span></a><span> across many models:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Cae6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 424w, /__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 848w, /__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Cae6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png" width="1456" height="616" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:616,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 424w, /__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 848w, /__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Cae6!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cd5ec3-97b8-4ff6-a4c4-691a8950019c_2048x867.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong><span>Capabilities</span></strong></h2><h2><span>Maths</span></h2><p><span>An ensemble of models guided by a human </span><a href="https://x.com/DavidTurturean/status/2081780318881677693"><span>solved</span></a><span> a &#8220;solid result&#8221; level problem in Epoch&#8217;s FrontierMath: Open Problems benchmark.</span></p><blockquote><p><strong><span>Opinion:</span></strong><span> Notable for two reasons:</span></p><p><span>1. It&#8217;s something of a called shot: there are only 15 problems in the FrontierMath: Open Problems collection, and it&#8217;s the oldest and most prominent research math achievement &#8220;benchmark&#8221;. Forecasters gave it a </span><a href="https://x.com/NunoSempere/status/2038686808876200114"><span>67%</span></a><span> by the end of the year back in March.</span></p><p><span>2. The proof is 60 pages long, whereas all previous notable AI proofs have been 10 pages and under, and often 1-2 pages</span></p><p><span>Minor update in favor of frontier models gaining capacity to solve mathematical-sciences problems of our choice, as opposed to picking up low-hanging-for-AI fruit.</span></p></blockquote><div><hr></div><p><span>The volume of mathematical discoveries produced by AI is becoming a </span><a href="https://x.com/thomasfbloom/status/2081362474603876529%5C"><span>challenge</span></a><span> for the mathematical community, since any claims that a major conjecture has been solved must be verified. For example, a PhD candidate at Columbia reportedly </span><a href="https://twitter.com/Qiaoqiao2001/status/2080003441821163958"><span>solved</span></a><span> six open Erd&#337;s problems in 5 days, using OpenAI&#8217;s GPT-5.6 Sol. But the mathematical community is not set up to vet new results at this new increased pace.</span></p><p><strong><span>Opinion:</span></strong><span> The open source community has been struggling with this problem for a bit longer; perhaps it might be worth looking into its own recent habits and solutions.</span></p><h3><span>&#128294; Opus 5</span></h3><p><span>Anthropic released Opus 5. </span><em><span>TLDR: mixed feelings from the community.</span></em></p><blockquote><p><strong><span>Opinion:</span></strong><em><span> </span></em><strong><span> </span></strong><span>In our own testing, the model appears to be generally alright &#8211; somewhat of a middle ground between Fable 5 and GPT-5.6-Sol. More willing to go through hoops than Fable; more tasteful than 5.6-Sol. Most importantly though, the model makes Anthropic subscriptions competitive again, since safety classifiers do not make the model unusable and usage limits are reasonable again (100% of the plan&#8217;s quota can be used, unlike Fable). It&#8217;s no Fable replacement, though.</span></p><p><span>- Podcaster and youtuber Theo </span><a href="https://x.com/theo/status/2080789645424767068"><span>initially liked the model</span></a><span>, but grew </span><a href="https://x.com/theo/status/2081880182936502474"><span>unsatisfied</span></a><span> with it over the next few days. His colleague Ben Davis </span><a href="https://x.com/davis7/status/2081884434253701159"><span>appears</span></a><span> to agree</span></p><p><span>- Noumena&#8217;s </span><a href="https://x.com/_xjdr/status/2081078742819221557"><span>xjdr</span></a><span> seems to like it as a &#8220;research peer&#8221;, but still prefers 5.6-sol as a daily driver.</span></p><p><span>- There&#8217;s a </span><a href="https://x.com/HarukaKunori/status/2081697911847481502"><span>takedown</span></a><span> of Opus making the rounds on X. Something important to flag also discussed by the post is that together with the model&#8217;s release, Anthropic also shipped a number of changes slimming down Claude Code&#8217;s system prompt. People&#8217;s reactions are thus likely indicative not just of the model itself, but of a modified harness.</span></p></blockquote><p><em><strong><span>Opus 5 economics</span></strong></em></p><ul><li><p><span>Opus 5 is priced at the same $5 in/$25 out as previous Opus models.</span></p></li><li><p><span>However, the model is not as token efficient as GPT models, making the model generally competitive only at the higher end of cost per task axis.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XGDa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XGDa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png" width="1456" height="869" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:869,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XGDa!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be53fd8-6aad-4c3b-a786-5fc742b737cc_2048x1223.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7HXq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_webp, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7HXq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png" width="1456" height="869" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:869,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_424, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 424w, /__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_848, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 848w, /__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_1272, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7HXq!, /__u/p3humansonai.substack.com/w_1456, /__u/p3humansonai.substack.com/c_limit, /__u/p3humansonai.substack.com/f_auto, /__u/p3humansonai.substack.com/q_auto:good, /__u/p3humansonai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49daa092-cc9f-43c4-990e-68e0bbeb1fc9_2048x1223.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><strong><span>Opus 5 capabilities &amp; benchmarks</span></strong></h4><h5><span>General</span></h5><ul><li><p><span>On LisanBench, a benchmark of long-chain word-ladder reasoning, Claude Opus 5 (high) </span><a href="https://lisanbench.com/"><span>scores 15428</span></a><span>, ahead of the best open-weight model, Kimi K3 at 10521, and its predecessor, Claude Opus 4.8, at 9463.</span></p></li><li><p><span>On </span><a href="https://eqbench.com/creative_writing.html"><span>EQ-Bench Creative Writing</span></a><span>, a benchmark of head-to-head, rubric-scored creative-writing quality, Claude Opus 5 scores 2430 Elo, ahead of Kimi K3 at 2340, and Claude Opus 4.8 at 1889.</span></p></li><li><p><span>On GDPval-AA, a benchmark of agentic real-world knowledge-work tasks, Claude Opus 5 (max) </span><a href="https://artificialanalysis.ai/evaluations/gdpval-aa"><span>scores 1862 Elo</span></a><span>, ahead of Claude Fable 5 at 1747, GPT-5.6 Sol at 1736, and the best open-weight model, Kimi K3 at 1687; its predecessor, Claude Opus 4.8, scores 1593.</span></p></li><li><p><span>On LiveBench, a contamination-limited general LLM benchmark with regularly refreshed questions across reasoning, coding, math, language, data analysis, and instruction following, Claude Opus 5 (high) </span><a href="https://livebench.ai/"><span>scores 81.0%</span></a><span>, behind GPT-5.6 Sol at 82.5% and Claude Fable 5 at 81.3%, but ahead of the best open-weight model, Kimi K3 at 78.5%, and its predecessor, Claude Opus 4.8, at 78.9%.</span></p></li><li><p><span>On SimpleBench, adversarial six-option multiple-choice questions testing everyday spatial, temporal, social, and linguistic reasoning, Claude Opus 5 </span><a href="https://simple-bench.com/"><span>scores 80.6%</span></a><span>, behind Claude Fable 5 at 81.9% and ahead of Gemini 3.1 Pro Preview at 79.6%, Kimi K3 at 60.7%, and its predecessor, Claude Opus 4.8, at 64.8%.</span></p></li><li><p><span>On ARC-AGI-2, a benchmark of abstraction-and-reasoning puzzles, Claude Opus 5 (max) </span><a href="https://arcprize.org/leaderboard"><span>scores 90.4%</span></a><span>, ranking #2 of 66 behind GPT-5.6 Sol at 92.5% and ahead of the best open-weight model, Inkling, at 36.5%; its predecessor, Claude Opus 4.8, scores 72.1%.</span></p></li><li><p><span>On </span><a href="https://eqbench.com/judgemark-v4.html"><span>Judgemark v4</span></a><span>, a benchmark of how well models judge creative writing, Claude Opus 5 scores 78.8%; its error bars cannot distinguish it from the best open-weight model, GLM 5.2, at 73.2%, its predecessor, Claude Opus 4.8, at 78.0%, or Gemini 3.1 Pro Preview at 78.7%, while Claude Opus 4.6 scores 90.7%.</span></p></li><li><p><span>On AA-Omniscience Accuracy, a benchmark of broad factual accuracy across domains, Claude Opus 5 (max) </span><a href="https://artificialanalysis.ai/evaluations/omniscience"><span>scores 54.2%</span></a><span>, ahead of the best open-weight, model Kimi K3, at 46.0%, and its predecessor, Claude Opus 4.8, at 46.6%, but behind Gemini 3.1 Pro Preview at 55.2% and Claude Fable 5 at 61.4%.</span></p></li></ul><h5><span>Coding / Agents</span></h5><ul><li><p><span>On Vibe Code Bench v1.1, a benchmark of end-to-end web-application builds with unrestricted terminal access, Claude Opus 5 </span><a href="https://www.vals.ai/benchmarks/vibe-code"><span>scores 88.4%</span></a><span>, statistically indistinguishable from Claude Fable 5 at 90.4%, the best open-weight model, Kimi K3, at 85.0%, and Claude Opus 4.8 at 82.7%.</span></p></li><li><p><span>On ProgramBench, a benchmark of rebuilding behaviorally equivalent programs from executables and usage docs, Claude Opus 5 </span><a href="https://www.vals.ai/benchmarks/programbench"><span>scores 82.3%</span></a><span>, ahead of the best open-weight model, GLM 5.2, at 62.6%, and its predecessor, Claude Opus 4.8, at 71.9%; its result is indistinguishable from GPT-5.6 Sol&#8217;s 77.6% within the benchmark&#8217;s error bars.</span></p></li><li><p><span>On CursorBench, a benchmark of ambiguous, multi-file coding tasks drawn from real Cursor sessions, Claude Opus 5 (max) </span><a href="https://cursor.com/evals"><span>scores 70.0%</span></a><span>, behind Claude Fable 5 at 70.5% and ahead of the best open-weight model, Kimi K3 at 60.8%, as well as Claude Opus 4.8 at 62.3%.</span></p></li><li><p><span>On HiL-Bench, a benchmark of tasks where models may ask a human for help, Claude Opus 5 </span><a href="https://labs.scale.com/leaderboard/hil"><span>scores 57.0%</span></a><span>, though its error bars cannot distinguish it from Claude Fable 5 at 56.3% or the best open-weight model, GLM 5.2 at 43.7%; Claude Opus 4.8 scores 35.3%.</span></p></li><li><p><span>On Handbook, a benchmark of long-context agentic instruction-following in RL environments modeled on following a company handbook, Claude Opus 5 (max) </span><a href="https://surgehq.ai/benchmarks/handbook"><span>scores 32.3%</span></a><span>, behind Claude Fable 5 at 36.2% and ahead of the best open-weight model, GLM 5.2, at 12.7%, and Claude Opus 4.8 at 21.9%.</span></p></li><li><p><span>On Zapier Benchmarks, a benchmark of multi-app automation tasks, Claude Opus 5 (max) </span><a href="https://zapier.com/benchmarks"><span>scores 26.2%</span></a><span>, ahead of Gemini 3.6 Flash at 19.8%, GLM 5.2, the best open-weight model, at 14.0%, and its predecessor, Claude Opus 4.8, at 17.2%.</span></p></li><li><p><span>On BrowseComp, a benchmark of web-browsing agent tasks requiring multi-hop information retrieval, Claude Opus 5 </span><a href="https://llm-stats.com/benchmarks/browsecomp"><span>scores 90.8%</span></a><span>, behind Kimi K3 at 91.2% and ahead of GPT-5.6 Sol at 90.4%; its predecessor, Claude Opus 4.8, scores 84.3%.</span></p></li><li><p><span>On GBA Eval, a benchmark of writing a working Game Boy Advance emulator, Claude Opus 5 </span><a href="https://gbaeval.com/"><span>scores 79.6%</span></a><span>, ahead of Claude Fable 5 at 74.5% and its predecessor, Claude Opus 4.8, at 70.9%; the best open-weight model, Kimi K3, scores 48.3%.</span></p></li><li><p><span>On DeepSWE, a benchmark of real software-engineering issues designed as a harder, less gameable replacement for SWE-Bench, Claude Opus 5 (max) </span><a href="https://deepswe.datacurve.ai/"><span>scores 73.6%</span></a><span>, with error bars unable to distinguish it from GPT-5.6 Sol at 72.7%, Claude Fable 5 at 69.9%, and the best open-weight model, Kimi K3, at 68.5%; its predecessor, Claude Opus 4.8, scores 59.0%.</span></p></li><li><p><span>On FrontierCode 1.1 Main, a benchmark of maintainer-crafted production-code tasks graded for mergeability and code quality, Claude Opus 5 (medium) </span><a href="https://cognition.com/frontiercode"><span>scores 53.4%</span></a><span>, behind Claude Fable 5 at 53.5% and ahead of GPT-5.6 Sol at 47.5%, Claude Opus 4.8 at 46.5%, and the best open-weight model, Kimi K3, at 44.2%.</span></p></li><li><p><span>On Tau3 Banking (AA), a benchmark of agentic customer-service tasks in a banking setting, Claude Opus 5 (high) </span><a href="https://artificialanalysis.ai/evaluations/tau3-banking"><span>scores 32.8%</span></a><span>, ranking #3 of 94, behind Kimi K3 at 33.4% and GPT-5.6 Sol at 33.0%, and ahead of its predecessor, Claude Opus 4.8, at 27.6%.</span></p></li><li><p><span>On Senior SWE-Bench, a benchmark of senior-level real-world software-engineering tasks resolved by LLM agents under a fixed harness, Claude Opus 5 (xhigh, mini swe) </span><a href="https://senior-swe-bench.snorkel.ai/"><span>scores 28.2%</span></a><span>, ranking #2 of 16 behind Claude Fable 5 at 29.1% and ahead of Claude Opus 4.8 at 25.0% and the best open-weight model, MiniMax M3, at 13.8%.</span></p></li></ul><h5><span>Math</span></h5><ul><li><p><span>On ProofBench, a benchmark of formal-math tasks where models produce Lean 4 proofs for advanced undergraduate and graduate problems, Claude Opus 5 </span><a href="https://www.vals.ai/benchmarks/proof_bench"><span>scores 78.0%</span></a><span>; its error bars cannot distinguish it from Claude Fable 5 and GPT-5.6 Sol at 77.0%, the best open-weight model, Kimi K3, at 70.0%, or its predecessor, Claude Opus 4.8, at 69.0%.</span></p></li><li><p><span>On Riemann-Bench, a head-to-head Elo benchmark of extreme-tier mathematical problems requiring deep reasoning, Claude Opus 5 (max) </span><a href="https://surgehq.ai/benchmarks/riemann-bench"><span>scores 68.0%</span></a><span>, ranking #2 of 24 behind GPT-5.6 Sol at 74.4% and ahead of Claude Fable 5 at 60.0% and Claude Opus 4.8 at 47.2%.</span></p></li><li><p><span>On </span><a href="https://epoch.ai/benchmarks/otis-mock-aime-2024-2025"><span>OTIS Mock AIME</span></a><span>, a benchmark of 45 competition-style math problems from OTIS, Claude Opus 5 (max) scores 98.9%, statistically indistinguishable from GPT 5.5 at 100.0%, the best open-weight model, Kimi K3, at 97.2%, and its predecessor, Claude Opus 4.8, at 98.3%.</span></p></li><li><p><span>On FrontierMath (Tiers 1-3 v2), a benchmark of research-level math problems, Claude Opus 5 (max) </span><a href="https://epoch.ai/benchmarks/frontiermath"><span>scores 85.6%</span></a><span>, with error bars unable to distinguish it from GPT-5.6 Sol at 89.1%, GPT-5.6 Terra at 86.0%, or Claude Opus 4.8 at 80.0%.</span></p></li><li><p><span>On FrontierMath Tier 4 (v2), the hardest tier of FrontierMath covering research-level mathematics problems, Claude Opus 5 (max) </span><a href="https://epoch.ai/benchmarks/frontiermath-tier-4-v2"><span>scores 73.2%</span></a><span>; its error bars cannot distinguish it from Claude Fable 5 at 87.8%, GPT-5.5 Pro at 78.0%, GPT 5.5 at 72.5%, or its predecessor Claude Opus 4.8 at 56.1%, and it scores above the best open-weight model, Kimi K3, at 39.0%.</span></p></li></ul><h5><span>ML</span></h5><ul><li><p><span>On WeirdML, a benchmark of nonstandard ML engineering tasks where models write PyTorch for novel datasets and iterate from execution and test feedback, Claude Opus 5 (max) </span><a href="https://htihle.github.io/weirdml.html"><span>scores 91.8%</span></a><span>, statistically indistinguishable from Claude Fable 5 at 91.9%, ahead of GPT-5.6 Sol at 88.8% and its predecessor, Claude Opus 4.8, at 82.9%.</span></p></li></ul><h5><span>Games</span></h5><ul><li><p><span>On Kaggle Game Arena, a benchmark of head-to-head strategy-game play, Claude Opus 5 </span><a href="https://www.kaggle.com/benchmarks/kaggle/game-arena/versions/1"><span>scores 287</span></a><span>, placing #2 of 26, behind GPT 5.5 at 330 and ahead of GPT-5.6 Sol at 264; its predecessor, Claude Opus 4.8, scores 123.</span></p></li><li><p><span>On RuneBench, a benchmark of long-horizon skilling agents in Old School RuneScape scored by experience rate across sixteen skills, Claude Opus 5 </span><a href="https://maxbittker.github.io/runebench/"><span>scores 5.73</span></a><span>, ranking #4 of 41, behind Claude Fable 5 at 6.01, GPT-5.6 Sol at 5.9, and GPT-5.6 Terra at 5.88, while improving on Claude Opus 4.8&#8217;s 5.08.</span></p></li><li><p><span>On Chess Puzzles, Stockfish-generated chess puzzles solved by exact best-move match, Claude Opus 5 (max) </span><a href="https://epoch.ai/benchmarks/chess-puzzles"><span>scores 42.0%</span></a><span>; its error bars cannot distinguish it from best the open-weight model, Kimi K3, at 39.0%, its predecessor, Claude Opus 4.8, at 34.0%, or Claude Fable 5 at 41.0%.</span></p></li></ul><h5><span>STEM</span></h5><ul><li><p><span>On EpiBench, a benchmark of epigenetics prediction tasks, Claude Opus 5 (pi) </span><a href="https://benchmarks.bio/?bench=epibench"><span>scores 35.2%</span></a><span>; its error bars cannot separate it from GPT 5.5 at 45.0%, Claude Opus 4.8 at 39.0%, or the best open-weight model, Kimi K2.6, at 24.5%.</span></p></li><li><p><span>On Humanity&#8217;s Last Exam, a benchmark of expert-level questions across many academic fields, Claude Opus 5 (max) </span><a href="https://artificialanalysis.ai/evaluations/humanitys-last-exam"><span>scores 52.6%</span></a><span>, behind Claude Fable 5 at 53.3% and ahead of GPT-5.6 Sol at 47.2% and its predecessor, Claude Opus 4.8, at 45.7%.</span></p></li><li><p><span>On CritPt, a benchmark of unpublished research-level physics problems, Claude Opus 5 (max) </span><a href="https://artificialanalysis.ai/evaluations/critpt"><span>scores 29.1%</span></a><span>, ahead of the best open-weight model, Kimi K3, at 23.4% and its predecessor, Claude Opus 4.8, at 20.9%, but behind GPT-5.6 Sol at 32.3%.</span></p></li></ul><h5><span>Multimodal / Computer Use</span></h5><ul><li><p><span>On GDP-PDF, a benchmark of multimodal reasoning over real-world prompts and PDFs from expert professional workflows, Claude Opus 5 (max) </span><a href="https://surgehq.ai/benchmarks/gdp-pdf"><span>scores 24.0%</span></a><span>, matching its predecessor, Claude Opus 4.8, and beating the best open-weight model, Kimi K3, at 19.0%, while GPT-5.6 Sol leads with 30.7%.</span></p></li></ul><h5><span>Miscellaneous</span></h5><ul><li><p><span>On BullshitBench, a benchmark of detecting unsubstantiated or manipulative claims, Claude Opus 5 (low) </span><a href="https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html"><span>scores 80.2%</span></a><span>, ahead of the best open-weight model, Qwen3.5 397B A17B, at 78.0%, but below its predecessor, Claude Opus 4.8, at 95.0% and Claude Sonnet 5 at 80.8%.</span></p></li></ul><h5><span>Games / Reasoning</span></h5><ul><li><p><span>On VoxelBench, where human raters vote pairwise on voxel builds generated from text prompts, Claude Opus 5 (max) </span><a href="https://voxelbench.ai/leaderboard"><span>scores 2222</span></a><span>, with its error bars indistinguishable from GPT-5.6 Sol&#8217;s 2270 and Claude Fable 5&#8217;s 2197; it exceeds the best open-weight model, Kimi K3, at 2040 and Claude Opus 4.8 at 1701.</span></p></li><li><p><span>On MineBench, where human raters vote pairwise on Minecraft builds from a rotating prompt set, Claude Opus 5 </span><a href="https://minebench.ai/leaderboard"><span>scores 2135</span></a><span>, ahead of Claude Opus 4.8&#8217;s 1772 and Kimi K3&#8217;s 1701, the best open-weight score; its 2042 score is indistinguishable from GPT-5.6 Sol&#8217;s within the benchmark&#8217;s error bars.</span></p></li></ul><h5><span>Science</span></h5><ul><li><p><span>On scBench-Long, a benchmark of long-horizon single-cell analysis through multi-step bioinformatics pipelines, Claude Opus 5 (pi) </span><a href="https://benchmarks.bio/scbench-long"><span>scores 41.3%</span></a><span>, though its error bars cannot distinguish it from GPT-5.6 Sol at 38.1% or Claude Opus 4.8 at 25.4%; it scores higher than the best open-weight model, Kimi K3, at 12.7%.</span></p></li></ul><h5><span>Reasoning</span></h5><ul><li><p><span>On ARC-AGI-1, a benchmark of abstraction-and-reasoning grid puzzles, Claude Opus 5 (max) </span><a href="https://epoch.ai/benchmarks/arc-agi"><span>scores 97.5%</span></a><span>, behind Gemini 3.1 Pro Preview at 98.0%, ahead of the best open-weight model, GLM 5.2, at 77.0%, and above Claude Opus 4.8 at 92.5%.</span></p></li></ul><h4><span>Knowledge / Science</span></h4><ul><li><p><span>On GPQA Diamond (Epoch), Epoch&#8217;s uniform-harness GPQA Diamond runs, Claude Opus 5 (max) </span><a href="https://epoch.ai/benchmarks/gpqa-diamond"><span>scores 93.9%</span></a><span>, statistically indistinguishable from GPT-5.4 Pro at 94.6%, Kimi K3 at 93.1%,the best open-weight model, and its predecessor, Claude Opus 4.8, at 91.0%.</span></p></li></ul><h5><span>Games / Agents</span></h5><ul><li><p><span>On GBENCH, a benchmark of head-to-head performance across competitive game environments, Claude Opus 5 </span><a href="https://gertlabs.com/rankings"><span>scores 75.9%</span></a><span>, ahead of Kimi K3 at 68.0%, the best open-weight model, and its predecessor, Claude Opus 4.8, at 64.6%; Claude Fable 5&#8217;s 73.4% result is indistinguishable under the benchmark&#8217;s error bars.</span></p></li></ul><p><em><strong><span>Opus 5 showcases</span></strong></em></p><ul><li><p><span>Opus 5 </span><a href="https://x.com/cozyblazex/status/2081271600868172132"><span>plays</span></a><span> Portal (apparently better than Fable, but note n = 1)</span></p></li><li><p><span>Opus 5 makes for </span><a href="https://x.com/Lentils80/status/2081136109778538917"><span>some incredible painterly worlds</span></a><span> in self-contained HTML files</span></p></li></ul><h2><strong><span>Minor</span></strong></h2><ul><li><p><span>Substack </span><a href="/__u/support.substack.com/hc/en-us/articles/50891130623508-How-can-I-detect-AI-on-Substack"><span>partners</span></a><span> with Pangram to detect AI work. It seems like a good, scalable, practical epistemic intervention.</span></p></li><li><p><span>WSJ </span><a href="https://www.wsj.com/business/big-companies-are-starting-to-hire-again-defying-predictions-of-ai-wipeout-f4974e99"><span>reports</span></a><span> that large companies are hiring humans again.  Expecting productivity gains from AI, large American companies have been reducing their workforces and hiring efforts. But the economic benefits of AI proved to be less clear-cut. Firms are coming to the view that AI cannot totally replace entry-level jobs (yet), as some companies betted, but must instead work alongside humans.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Humans on AI, July 22nd 2026]]></title><description><![CDATA[&#8211; Scifi-like autonomous cybersecurity incidents by an internal OpenAI model (likely GPT-6).]]></description><link>https://p3humansonai.substack.com/p/newsletter-7-22</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-7-22</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>&#8211; Scifi-like autonomous cybersecurity incidents by an internal OpenAI model (likely GPT-6). &#8211; Kimi K3 release narrows the gap between Chinese open weights models and US frontier, prompting a discourse and policy tumult. &#8211; Two Annals-quality mathematical breakthroughs with comically short proofs landed this week. Economics Bloomberg reports a 1GW&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, July 17th 2026]]></title><description><![CDATA[Chinese labs successfully crank up sparsity to cope with compute shortage; AI math progress broadens a bit outside number theory and combinatorics (see below); Rome Declaration (Physics and Peace Nobelists) urge AI slowdown; Anthropic again studies agentic misalignment in a highly divisive manner; GPT-Red automates cyber red-teaming Economics It appears&#8230;]]></description><link>https://p3humansonai.substack.com/p/newsletter-7-17</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-7-17</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Chinese labs successfully crank up sparsity to cope with compute shortage; AI math progress broadens a bit outside number theory and combinatorics (see below); Rome Declaration (Physics and Peace Nobelists) urge AI slowdown; Anthropic again studies agentic misalignment in a highly divisive manner; GPT-Red automates cyber red-teaming Economics It appears&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, July 14th 2026]]></title><description><![CDATA[Economists warn about AI transformation; Economists model quiet AI takeoff; Hassabis proposes federal AI oversight; OpenAI safety leaders depart; NYT reports AI-aided terrorism Economics Very brief 90-word letter signed by major figures in AI orgs and econ Nobel laureates.]]></description><link>https://p3humansonai.substack.com/p/newsletter-7-14</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-7-14</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Economists warn about AI transformation; Economists model quiet AI takeoff; Hassabis proposes federal AI oversight; OpenAI safety leaders depart; NYT reports AI-aided terrorism Economics Very brief 90-word letter signed by major figures in AI orgs and econ Nobel laureates. &#8220;AI may become radically more powerful over the next 10 years.</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, July 10th 2026]]></title><description><![CDATA[OpenAI internal inference grows 100x; GPT-5.6, Grok, Muse released; Mythos access returns abroad; Paper detects training versus inference; AI generates 25% of sampled web content Economics OpenAI&#8217;s internal inference grows 100x as a share of research compute (perhaps from a low base).]]></description><link>https://p3humansonai.substack.com/p/newsletter-7-10</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-7-10</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>OpenAI internal inference grows 100x; GPT-5.6, Grok, Muse released; Mythos access returns abroad; Paper detects training versus inference; AI generates 25% of sampled web content Economics OpenAI&#8217;s internal inference grows 100x as a share of research compute (perhaps from a low base). &#8220;Over the past six months, the share of&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, July 8th 2026]]></title><description><![CDATA[Anthropic find an analogue of working memory in Claude.]]></description><link>https://p3humansonai.substack.com/p/newsletter-7-8</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-7-8</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Anthropic find an analogue of working memory in Claude. Many suggestive but inconclusive ties to human cognition, consciousness, and selfhood; Fable suddenly re-released on general access. Dario nominally sidelined; Commerce clears GPT-5.6 Sol for general release, in the first full test of the EO voluntary review process; Epoch test frontier&#8230;</p>]]></content:encoded></item><item><title><![CDATA[On J-space]]></title><description><![CDATA[A very hot take written in 2 hours.]]></description><link>https://p3humansonai.substack.com/p/jspace</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/jspace</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A very hot take written in 2 hours. Anthropic claim that &#8220;Claude has developed a mechanism for conscious access&#8221;; They found a way to find and intervene on Claude&#8217;s working memory: a Jacobian lens on the activations, averaged over contexts and future tokens; But (as they concede) a working memory&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Introducing P3]]></title><description><![CDATA[Who we are Paradigm 3 was co-founded by Gavin Leech and Max Henderson; you can see the rest of the team here.]]></description><link>https://p3humansonai.substack.com/p/intro</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/intro</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Who we are Paradigm 3 was co-founded by Gavin Leech and Max Henderson; you can see the rest of the team here. We don't take funding from the AI industry or large foundations; our current funders are individuals new to the space; Our thesis We don't know what's going on.</p>]]></content:encoded></item><item><title><![CDATA[What we’d like to fund]]></title><description><![CDATA[Besides our in-house research, we are currently funding: An estimate of the size of the externalities imposed by current AI systems.]]></description><link>https://p3humansonai.substack.com/p/fund</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/fund</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Besides our in-house research, we are currently funding: An estimate of the size of the externalities imposed by current AI systems. The team is starting with estimating the time lost to increased authentication challenges, but hope to also capture positive externalities from e.g. users saving thousands of dollars apiece on&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Coding vs thinking]]></title><description><![CDATA[We&#8217;re interested in the prospects for (presumably safer) narrow AI staying competitive, instead of general systems; Cursor&#8217;s Composer coding finetune of Kimi is probably the most intense attempt to specialise a general model: probably more than 10^25 FLOPs of post-training; We find that, compared to its base model, Composer shows&#8230;]]></description><link>https://p3humansonai.substack.com/p/composer</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/composer</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We&#8217;re interested in the prospects for (presumably safer) narrow AI staying competitive, instead of general systems; Cursor&#8217;s Composer coding finetune of Kimi is probably the most intense attempt to specialise a general model: probably more than 10^25 FLOPs of post-training; We find that, compared to its base model, Composer shows&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, July 3rd 2026]]></title><description><![CDATA[Anthropic API margin exceeds 80%; AISI frames capability as compute curve; AI loss-of-control worries Hill staffers; LLM-only conversations form attractor states; Anthropic detects Chinese Claude Code access Economics Annals of winner-takes-all: Dylan Patel says Anthropic's margin on an Opus 4.8 API token is north of 80%, and that it is&#8230;]]></description><link>https://p3humansonai.substack.com/p/newsletter-7-3</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-7-3</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Anthropic API margin exceeds 80%; AISI frames capability as compute curve; AI loss-of-control worries Hill staffers; LLM-only conversations form attractor states; Anthropic detects Chinese Claude Code access Economics Annals of winner-takes-all: Dylan Patel says Anthropic's margin on an Opus 4.8 API token is north of 80%, and that it is&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, June 30th 2026]]></title><description><![CDATA[OpenAI reportedly halves inference GPUs; Autoresearch improves neural decoding; Mythos returns to 100 trusted US organizations; FRI releases risk-focused panel findings; DeepMind Pentagon management criticized Economics Gossipy claim that OpenAI just found a big inference optimisation, halving the GPUs needed for a given level of traffic.]]></description><link>https://p3humansonai.substack.com/p/newsletter-6-30</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-6-30</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>OpenAI reportedly halves inference GPUs; Autoresearch improves neural decoding; Mythos returns to 100 trusted US organizations; FRI releases risk-focused panel findings; DeepMind Pentagon management criticized Economics Gossipy claim that OpenAI just found a big inference optimisation, halving the GPUs needed for a given level of traffic. Opinion: It seems likely&#8230;</p>]]></content:encoded></item><item><title><![CDATA[Humans on AI, June 26th 2026]]></title><description><![CDATA[Local grids constrain datacenter buildout; Google researchers leave for Anthropic; CAISI reportedly receives stop-work order; Transcript replay simulates deployment Economics More analysis of the datacenter electricity bottleneck on the AI buildout: in particular, local grid capacity and interconnects, rather than total national generation.]]></description><link>https://p3humansonai.substack.com/p/newsletter-6-26</link><guid isPermaLink="false">https://p3humansonai.substack.com/p/newsletter-6-26</guid><dc:creator><![CDATA[Peli Grietzer]]></dc:creator><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O13J!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a4157-645e-463d-a915-f0b6b5063ce7_202x202.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Local grids constrain datacenter buildout; Google researchers leave for Anthropic; CAISI reportedly receives stop-work order; Transcript replay simulates deployment Economics More analysis of the datacenter electricity bottleneck on the AI buildout: in particular, local grid capacity and interconnects, rather than total national generation. The grid feeding Stargate will max out&#8230;</p>]]></content:encoded></item></channel></rss>