<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Build with AI]]></title><description><![CDATA[Learn what actually works, from experts building with AI. New issue every Thursday.]]></description><link>https://packtbuildwithai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!TYg3!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49f5d476-aafa-4cf7-8d2b-a6dcb5f921b2_512x512.png</url><title>Build with AI</title><link>https://packtbuildwithai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 10:28:25 GMT</lastBuildDate><atom:link href="/__u/packtbuildwithai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Build With AI]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[packtbuildwithai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[packtbuildwithai@substack.com]]></itunes:email><itunes:name><![CDATA[BuildWithAI]]></itunes:name></itunes:owner><itunes:author><![CDATA[BuildWithAI]]></itunes:author><googleplay:owner><![CDATA[packtbuildwithai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[packtbuildwithai@substack.com]]></googleplay:email><googleplay:author><![CDATA[BuildWithAI]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Part 5 – Building the Human-AI Responsibility Matrix with Michelle Sandford]]></title><description><![CDATA[&#129504; Where drift sets in, and the document that stops it.]]></description><link>https://packtbuildwithai.substack.com/p/part-5-building-the-human-ai-responsibility</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/part-5-building-the-human-ai-responsibility</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 27 Aug 2026 14:03:03 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8e38a385-541b-4df6-affd-cba5c126a5b2_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <em><a href="https://x.com/packtwebdevpro?s=21">X </a></em></p></div><p><em><strong><a href="/__u/packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code-f29?r=55ncj4&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">Part 4</a></strong></em> proved the loop works on something real: a rate-limiting feature walked through all six stages, with the agent doing most of the typing and almost none of the deciding. That&#8217;s the version of the series everyone wants to believe describes their own team permanently. It doesn&#8217;t. It describes a team on its best week.</p><p>Part 5, the final part of this series, is about what happens on the other weeks. Two failure modes show up in almost every team that adopts an AI-native loop and then relaxes: responsibility drift, where careful review quietly turns into rubber-stamping, and escalation theatre, where a polished-looking policy routes to nobody in particular. Both are correctable, and both get expensive if you let them sit.</p><p>From there, we bring everything from this series together into the one artifact that&#8217;s been promised since Part 1: <em><strong>a Human-AI Responsibility Matrix</strong></em>, with one named <em>Accountable</em> owner for every stage of the loop, no exceptions. It&#8217;s the document that turns &#8220;you own the code&#8221; from a rule Michelle repeats in keynotes into something you can check into your own repository this week.</p><blockquote><p><em><a href="https://www.linkedin.com/in/michellesandford/">Michelle Sandford</a>, <a href="https://sessionize.com/MichelleSandford/">Microsoft&#8217;s Developer Engagement Lead for Asia</a></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iZk8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iZk8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1734801,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/212515259?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iZk8!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e4ad75-7a91-4200-9f44-e7f34bcf39b6_2400x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></blockquote><div class="callout-block" data-callout="true"><h1 style="text-align: center;">In this edition</h1><p>&#10140; Responsibility drift: how careful review quietly turns into rubber-stamping, and the mechanical fix</p><p>&#10140; Escalation theatre: why a polished-looking escalation policy can still route to nobody</p><p>&#10140; Over-escalation, and why routing everything to a human is its own kind of failure</p><p>&#10140; Building a Human-AI Responsibility Matrix: the one rule that makes it worth keeping</p><p>&#10140; A filled-in matrix for the rate-limiting example from Part 4, RACI markers and all</p><p>&#10140; Closing the series: what changes when &#8220;a human is responsible&#8221; becomes a document instead of a sentence</p></div><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>Failure mode and correction: When the boundaries blur</h1><p>Even with the matrix in place, two failure modes show up early and often. Both are correctable; both get expensive if you let them set.</p><p>The first is responsibility drift. The agent does good work for two weeks, the team relaxes, the senior engineer starts approving spec drafts without reading them carefully, and someone notices three sprints later that the agent has been quietly proposing AGENTS.md changes that no one has been reviewing as policy changes. The fix is mechanical, not cultural. Mark <code>AGENTS.md</code>, the <code>specs/ </code>directory, and the contents of <code>.github/ </code>as protected paths in <code>CODEOWNERS</code> so that a named human must approve every change. Add a periodic review (quarterly is plenty) where the team reads the <code>AGENTS.md</code> diff history end to end. Drift is the default; mechanical countermeasures are how you resist the default.</p><p>The second is escalation theater. The team writes a beautiful escalation policy, but the routes point to a generic senior engineers&#8217; group that nobody owns, so escalations sit unread until they time out and auto-merge. The fix is to name humans, not roles, and to make the absence of an escalation visible. A stale escalation should turn a PR red in the status check, not silently expire. The 2025 DORA report&#8217;s trust paradox finding maps directly onto this: developers say they do not fully trust AI output, then merge it anyway when the friction of escalation is higher than the friction of approving. Lower the friction of the right behavior, not the wrong one.</p><blockquote><p>I once found a team with a polished escalation policy but no named weekend route. Two high-risk PRs sat untouched, then were merged quickly on Monday under schedule pressure. Nothing failed immediately, which made the weakness harder to see. They corrected it by naming explicit primary and backup reviewers per service and by failing CI when an escalation label aged beyond a threshold. The policy became operational overnight because drift was now visible and mechanically blocked.</p></blockquote><p>A third pattern is worth a sentence even though it is less common. Over-escalation (routing every borderline case to a human) produces the same outcome as no escalation at all, because reviewers stop reading carefully when everything is urgent. Tune the triggers accordingly. A good escalation rule fires on roughly 5 to 15 percent of agent PRs, often enough to stay sharp, rarely enough to stay meaningful. If yours fires on every PR, the rule is doing the wrong work.</p><p>Now, let&#8217;s look at how we can formalize the practical processes we need by building out the human-AI responsibility matrix in the next section.</p><h2>Building a Human-AI Responsibility Matrix</h2><p>The carry-forward artifact for this series is a Human-AI Responsibility Matrix. Rows are the six loop stages covered in <em><strong><a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-22-the-ai-native-loop">Part 1</a></strong></em>, columns are the participants in your loop, and cells use RACI-style markers:</p><ul><li><p>R (Responsible: does the work)</p></li><li><p>A (Accountable: single named owner of the outcome)</p></li><li><p>C (Consulted: must be heard)</p></li><li><p>I (Informed: must be told)</p></li></ul><p>Every row must have exactly one A. That is the rule that makes the matrix worth keeping.</p><p>A blank template, suitable for copying into <code>docs/responsibility-matrix.md</code>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kj2Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 424w, /__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 848w, /__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kj2Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png" width="670" height="327" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:327,&quot;width&quot;:670,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:10057,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/212515259?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 424w, /__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 848w, /__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kj2Z!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F132e47fb-6a4f-4e10-94d4-cb1de8916814_670x327.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><strong>Figure 1: Human-AI Responsibility Matrix</strong></em></p><p>A filled-in version for the rate-limiting example, which is what the team would actually check in:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-D8k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 424w, /__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 848w, /__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-D8k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png" width="670" height="271" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:271,&quot;width&quot;:670,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9740,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/212515259?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 424w, /__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 848w, /__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-D8k!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29a3ecf1-e53b-401a-b8ca-76dffe15729e_670x271.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><strong>Figure 2: Filled-in Human-AI Responsibility Matrix with RACI assignments and escalation rules</strong></em></p><p>Notice the shape. The Agent is Accountable at exactly one stage (Change), and is Responsible at one other (part of Verification, as an evaluator&#8217;s input). Everywhere else the agent is Consulted or Informed. That distribution is what human-led, agent-assisted looks like written down. Other shapes are valid for other teams, but the rule (one A per row, never the agent A at Intent or Delivery) holds across every team this pattern has worked for.</p><div class="callout-block" data-callout="true"><p><strong>The one rule that makes the matrix worth keeping</strong></p><p>Every row must have exactly one A (Accountable). If a row has zero, no one owns the outcome. If a row has more than one, no one owns the outcome. The matrix is only as valuable as the extent to which that rule is enforced.</p></div><p>Treat the matrix as a living document. It belongs in the repository, it changes when <code>AGENTS.md</code> changes, and it is reviewed in the same cadence as the AI-Native Success Criteria Memo from <em><strong>Part 2</strong></em>. The two artifacts are a pair: the memo says what good looks like, the matrix says who is on the hook for getting there.</p><blockquote><p>The first matrix review I ran with a team surfaced a hidden disagreement: product believed low-risk API changes could be self-approved if tests were green, while engineering required explicit human accountability for all contract-impacting changes. Writing the matrix forced that disagreement into one row and one cell, which made it solvable. They split change categories and documented which categories required named approval. Future PRs moved faster because decisions no longer depended on competing assumptions.</p></blockquote><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Wrapping up&#8230;</h1><p>Five parts ago, this series opened on a plane, with an agent that had quietly invented a case study because it thought the story would land better that way. Nothing about that moment was malicious. What was missing was a system where a human was still positioned to catch it.</p><p>Everything since has been about building that system, one piece at a time. Part 1 gave you the six-stage loop and an honest read on where your team actually sits. Part 2 proved the loop changes outcomes, not just vibes, with a bug fixed twice and only one fix that actually held up. Part 3 named the real shift: developers move from typing to deciding, and Rule Zero, you own the code, became the line the rest of the series was built to operationalize. Part 4 turned &#8220;ask if you&#8217;re not sure&#8221; into rules a machine can actually enforce, and proved the whole loop on one real feature. Part 5 closed the gap between &#8220;we have a policy&#8221; and &#8220;the policy actually works,&#8221; and handed you the Human-AI Responsibility Matrix, the document where every one of those ideas has to become concrete: one stage, one owner, no exceptions.</p><p>If you&#8217;ve built even one artifact from this series, a spec, a memo, an <code>AGENTS.md</code> rule, a matrix, you already have more than most teams claiming to be AI-native. The rest is just repetition.</p><p>Thank you for reading this series. If it changed how you think about your own delivery loop, that&#8217;s worth more than any part of this newsletter could say on its own.</p><h1>Further reading</h1><p>&#128309; <strong><a href="https://en.wikipedia.org/wiki/Responsibility_assignment_matrix">Responsibility assignment matrix (RACI)</a></strong><br>The project-management technique underneath the Human-AI Responsibility Matrix built in this issue. Useful if you want the fuller history of RACI and its variants before adapting it to a loop that includes a non-human participant.</p><p>&#128309; <strong><a href="https://sre.google/sre-book/postmortem-culture/">Postmortem culture: learning from failure, Google SRE Book</a></strong><br>A useful companion to this series&#8217; argument that responsibility drift and escalation theatre are systems problems, not people problems. Google&#8217;s SRE team makes the same case about incident reviews: the fix is redesigning the system, not assigning blame after it fails.</p><p>&#128309; <strong><a href="https://blogs.microsoft.com/on-the-issues/2022/06/21/microsofts-framework-for-building-ai-systems-responsibly/">Microsoft&#8217;s framework for building AI systems responsibly</a></strong><br>Microsoft&#8217;s own account of why it built a company-wide Responsible AI Standard. Worth reading alongside this series&#8217; argument that trust in AI output should be placed in a designed system, specs, gates, named approvers, rather than in the model itself.</p><p>&#128309; <strong><a href="https://dora.dev/dora-report-2025/">2025 DORA State of AI-assisted Software Development report</a></strong><br>Referenced a final time because this issue closes the loop on the trust paradox that&#8217;s run through the whole series: developers use AI more while trusting it less, and this issue shows exactly how that gap shows up as escalation friction in practice, not just as a survey statistic.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Editor-in-Chief<br></span>Build with AI</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Three live workshops in September]]></title><description><![CDATA[AG-UI and CopilotKit on the 16th, Spring Boot with AI on the 24th, graph engineering with Claude Code on the 29th. With Lauren&#355;iu Spilc&#259;, Ken Huang and CopilotKit&#8217;s Angular maintainers.]]></description><link>https://packtbuildwithai.substack.com/p/upcoming-workshops</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/upcoming-workshops</guid><dc:creator><![CDATA[BuildWithAI]]></dc:creator><pubDate>Thu, 27 Aug 2026 13:15:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/02b06a3c-b408-4d11-85d8-ffc142ca52bf_1536x1024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Three workshops before the end of September. All live, all hands-on, and all of them come with the recording if your timezone hates you.</p><div><hr></div><h3>Hands-On Agentic Frontends with AG-UI and CopilotKit</h3><p>A chat box is not an agent interface. A production one streams work as it happens, keeps frontend and agent state in sync, hands tasks to specialist agents, and stops for a human before it does anything expensive.</p><p>The awkward part is that five protocols now claim a piece of that job: AG-UI, A2A, MCP, MCP Apps and A2UI. Most explanations tell you what each one is. This session shows you where each boundary sits, by building one incident-response app that uses all of them and adding the unglamorous bits afterwards: cancellation, timeouts, partial failures, and an approval gate before remediation.</p><p><strong>Rainer Hahnekamp</strong> and <strong>Murat Sari</strong> co-maintain CopilotKit&#8217;s Angular integration, so these are decisions they have already had to make in public.</p><p>Wed 16 Sep, 9:00 to 12:00 EDT. Three hours, online. Angular or React, your choice.</p><p>&#128073; <a href="https://www.eventbrite.co.uk/e/hands-on-agentic-frontends-with-ag-ui-and-copilotkit-tickets-1996619123546">Book on Eventbrite</a>  &#183;  <a href="https://luma.com/packt-d38e">or via Luma</a></p><div><hr></div><h3>Hands-on AI Assisted Java and Spring Boot Development</h3><p>AI writes Java quickly. It also writes Java that passes review and quietly wrecks your layering.</p><p>This one runs a whole Spring Boot build with AI in the loop: scaffolding, entities, DTOs, controllers, persistence, tests, debugging, refactoring. The point is not speed. It is keeping the architecture and security calls yours while the typing stops being your problem.</p><p><strong>Lauren&#355;iu Spilc&#259;</strong> is a Java Champion and wrote Spring Security in Action, which is to say he will spot the generated code that is wrong in a way that still compiles.</p><p>Thu 24 Sep, 8:30 to 11:00 EDT. Two and a half hours, online. Bring working Java and Spring Boot.</p><p>&#128073; <a href="https://packt.link/springai">Book on Eventbrite</a>  &#183;  <a href="https://luma.com/packt-qsxz">or via Luma</a></p><div><hr></div><h3>Hands-On Graph Engineering with Claude Code</h3><p>In July we ran an issue on why an agent asked to grade its own work tends to just praise it. This is the workshop where you build the fix.</p><p>Single-agent loops fail in predictable ways. Context degrades. The agent marks its own homework. Everything runs in sequence when half of it could run in parallel. The spec drifts without anyone noticing, and retries run until something stops them. None of that is solved by a better prompt. It is solved by a graph: real node boundaries, structured handoffs, and verification the agent doing the work cannot influence.</p><p><strong>Ken Huang</strong> co-chairs the OWASP AIVSS project and the AI Safety working groups at the Cloud Security Alliance. He opens with a keynote, then it is hands on the whole way.</p><p>Tue 29 Sep, 8:30 to 11:30 EDT. Three hours, online. Assumes you already use Claude Code.</p><p>&#128073; <a href="https://www.eventbrite.co.uk/e/hands-on-graph-engineering-with-claude-code-tickets-1997682277468">Book on Eventbrite</a>  &#183;  <a href="https://luma.com/packt-4p1d">or via Luma</a></p><div><hr></div><p>Every one includes the recording, the slides and a certificate, so a clashing calendar is not a reason to skip.</p><p>If you would rather not keep checking this page, new dates go out with the Thursday issue. <a href="https://packt.link/buildwithai">Subscribe to Build with AI</a>.</p>]]></content:encoded></item><item><title><![CDATA[The AI-Native Loop: Who Owns the Code Now – Part 4 with Michelle Sandford]]></title><description><![CDATA[&#129504; Escalation rules, operational prompting, and one feature built end to end.]]></description><link>https://packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code-f29</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code-f29</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 20 Aug 2026 14:00:34 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/226585a5-bc4e-4a6e-a2d2-c9670b9d2264_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <em><a href="https://x.com/packtwebdevpro?s=21">X </a></em></p></div><p><em><strong><a href="/__u/packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code-671?r=55ncj4&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">Part 3</a></strong></em> left you with something genuinely useful and slightly uncomfortable: a boundary a reviewer can apply without a meeting, mapped stage by stage across the loop. What it didn&#8217;t answer is <em>what happens the moment the normal case breaks</em>. The agent hits something it wasn&#8217;t expecting. A diff comes back three times the size anyone planned for. A change quietly touches a path it shouldn&#8217;t have.</p><p>In an AI-native loop, ambiguity doesn&#8217;t default to caution; it defaults to whichever voice is loudest, and that&#8217;s usually the agent&#8217;s own confidence, which is famously unreliable as a signal.</p><p>Part 4 replaces &#8220;ask if you&#8217;re not sure&#8221; with something a machine can actually enforce: escalation rules built from three concrete pieces, a trigger, a route, and an action. From there, we look at what prompting looks like when it&#8217;s a leadership act rather than a chat message, and at a handful of collaboration patterns worth naming and reusing. Then we do the whole thing for real: adding rate limiting to a public endpoint and walking it through all six stages of the loop. You&#8217;ll see every idea from this series working together on something that isn&#8217;t a toy example.</p><blockquote><p><em><a href="https://www.linkedin.com/in/michellesandford/">Michelle Sandford</a>, <a href="https://sessionize.com/MichelleSandford/">Microsoft&#8217;s Developer Engagement Lead for Asia</a></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!y-UJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!y-UJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1679822,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/211661606?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y-UJ!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3d2712e-c80b-4a9c-a053-e761452035cc_2400x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></blockquote><div class="callout-block" data-callout="true"><h1 style="text-align: center;">In this edition</h1><p>&#10140; Escalation rules that actually work: triggers, routes, and actions, not hope</p><p>&#10140; Why &#8220;if the agent is unsure, it should ask&#8221; is a wish, not a rule</p><p>&#10140; Prompting as a leadership act, and what separates a durable AGENTS.md from a one-off chat prompt</p><p>&#10140; Four collaboration patterns worth naming: spec-first pairing, reviewer-in-the-loop, pair-on-context, and evaluator pairing</p><p>&#10140; A full walkthrough: adding rate limiting to a public API endpoint, stage by stage across the six-stage loop</p><p>&#10140; A preview of Part 5, where we build the Human-AI Responsibility Matrix that ties the whole series together</p></div><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>Decision ownership and escalation rules</h1><p>Boundaries tell you who decides in the normal case. Escalation rules tell you what happens when the normal case does not apply (when the agent is uncertain, when the diff is larger than expected, when the change touches a regulated path, or when an evaluator flags a borderline result). Without explicit escalation, ambiguity defaults to the loudest voice, which in an AI-native loop is usually the agent&#8217;s confidence, and agent confidence is famously poorly calibrated.</p><p>A workable escalation policy has three ingredients:</p><ul><li><p><strong>Triggers are observable conditions</strong>: The agent edits any file under <code>payments/</code>, a diff exceeds 300 lines, an evaluator score falls below 0.8, or <code>AGENTS.md</code> itself is modified.</p></li><li><p><strong>Routes are named humans or teams, not roles</strong>: Route to the on-call senior for the affected service, not route to a senior engineer.</p></li><li><p><strong>Actions are what changes when the route fires</strong>: The PR is marked draft, the agent&#8217;s branch protection prevents auto-merge, a comment is posted requesting a specific reviewer.</p></li></ul><p>All three should be expressible in GitHub Actions and in <code>AGENTS.md</code> instructions to the agent, so the rule is enforced by the system and not by goodwill.</p><p>The most common escalation mistake is to write rules that depend on the agent volunteering uncertainty. &#8220;If the agent is not confident, it should ask&#8221; is not a rule; it is a hope. Better rules are mechanical: a list of paths, a diff-size threshold, an evaluator score, a label. Hope is allowed as a supplement, not as a foundation.</p><p>Let&#8217;s dive deeper into how to do that as we explore prompting in the next section.</p><h1>Prompting as operational leadership</h1><p>If the developer is now a system designer, prompting is one of the leadership acts of the role. Not the chat-window prompting of &#8220;write me a function that does X&#8221;; that is the productivity-tool version. Operational prompting is the deliberate construction of the instructions an agent will follow across many runs and many contributors: the <code>AGENTS.md</code>, the chat modes and prompt files in <code>.github/</code>, the issue templates that feed the Copilot coding agent, the Spec Kit specifications that anchor a feature. The <code>github/awesome-copilot</code> repository exists precisely because these artifacts have become a shared craft, with patterns worth borrowing.</p><p>Three properties separate operational prompts from one-off chat:</p><ul><li><p><strong>They are durable</strong>: Checked into the repository, reviewed in PRs, and versioned alongside the code they govern.</p></li><li><p><strong>They are constraining</strong>: They say what the agent must not do as clearly as what it should do, because an over-eager agent is the most common failure mode.</p></li><li><p><strong>They are testable</strong>: A good <code>AGENTS.md</code> rule can be checked by a CI job, an evaluator, or a reviewer in under a minute, which is what makes it enforceable rather than aspirational. Treat operational prompting the way you treat any other interface design: small, sharp, and revisited when the system around it changes.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ViUa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 424w, /__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 848w, /__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ViUa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png" width="1456" height="741" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:741,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:876730,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/211661606?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 424w, /__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 848w, /__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ViUa!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1dacdd9-d580-42f8-9c0b-a036237d3784_1758x895.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><strong>Table 2.1: Operational prompts vs. chat prompts</strong></em></p><p>If you cannot point to where the prompt lives in the repo, it is not yet operational. In the next section, we&#8217;ll look at some patterns that are worth calling out to make collaboration more effective when working with AI agents.</p><h1>Collaboration patterns for pairing with AI</h1><p>A small number of patterns recur often enough to be worth naming:</p><ul><li><p><strong>Spec-first pairing is the default for non-trivial work</strong>: The human drafts the spec, a reviewer (human or AI evaluator) sharpens it, the agent implements.</p></li><li><p><strong>Reviewer-in-the-loop inverts the polarity for risky changes</strong>: The agent drafts, a human reviewer treats the draft as a proposal rather than a finished artifact, and the agent revises against the review.</p></li><li><p><strong>Pair-on-context is the pattern for unfamiliar code</strong>: A human and the agent walk a module together, the human asking questions, the agent surfacing call sites and tests via MCP, until both can articulate the change before either touches it.</p></li><li><p><strong>Evaluator pairing uses one AI to grade another&#8217;s work</strong>: A second AI evaluates the implementation against the specification, catching plausible but incorrect outputs far more reliably than human review alone.</p></li></ul><p>Pick the pattern by the shape of the work, not by habit.</p><h1>Small example walkthrough: Adding rate limiting to a public API endpoint</h1><p>The example is deliberately ordinary: add per-token rate limiting to a public POST /v1/messages endpoint on a Fastify service running on Node.js, deployed via GitHub Actions to Azure. The team consists of three engineers and one Copilot coding agent. The repository already has an AGENTS.md, a <code>specs/</code> directory, and a Microsoft Agent Framework evaluator wired into CI through Azure AI Foundry. Nothing about the task is exotic. What matters is who decides what at each of the six loop stages:</p><ol><li><p><strong>Intent</strong>: The product manager files a GitHub issue: a small number of tokens are sending bursts that degrade latency for everyone else. The senior engineer on the service is the decision-maker. They accept the work, set the priority, and write the why paragraph in two sentences. The agent is not involved at this stage, and that absence is the point. The question of whether to do the work at all is human-only. The escalation rule that applies here is structural: any change that affects a public endpoint is automatically routed to the service&#8217;s on-call senior, regardless of who files the issue.</p></li><li><p><strong>Spec</strong>: The senior engineer drafts <code>specs/2026-05-rate-limit-messages.md</code> using the team&#8217;s Spec Kit template. The spec is short: <em>why, what changes for the caller, success criteria, out of scope</em>. The success criteria are concrete: <em>bursts above 60 requests per minute per token return HTTP 429 with a </em><code>Retry-After</code><em> header, no token sees a false 429 in the property test, p99 latency for compliant tokens is unchanged within noise</em>. A second engineer reviews the spec in fifteen minutes and adds one bullet the agent would never have known to add: exempt the internal health-check token from rate limiting. That bullet is the single most valuable artifact produced in the whole task, and it exists only because a human reviewed a human-authored spec before any code was generated.</p></li><li><p><strong>Context:</strong> The repository&#8217;s <code>AGENTS.md</code> already encodes the rules the agent needs: use <code>@fastify/rate-limit</code>, all limits are configured per token, never per IP, never edit <code>auth/</code> without a spec citation, log decisions through the existing observability MCP server. The senior engineer adds one per-task context note to the issue: a link to last quarter&#8217;s incident postmortem where a global rate limit briefly took down a partner. The agent reads <code>AGENTS.md</code> and the issue automatically; the postmortem link is the part the human had to assemble deliberately. Context is curated, not assumed.</p></li><li><p><strong>Change</strong>: The senior engineer assigns the issue to the Copilot coding agent. The agent opens a draft PR, installs <code>@fastify/rate-limit</code>, adds a per-token limiter, writes a <code>fast-check</code> property test that asserts no compliant token ever sees a 429, updates the OpenAPI document, and posts a trace listing model version, MCP tools called, files touched, and a mapping from each spec bullet to a commit. The agent does not edit <code>auth/</code>; <code>AGENTS.md</code> forbids it without a spec citation, and the spec does not authorize an auth change. That refusal is a feature, not a bug; it is the boundary working.</p></li><li><p><strong>Verification</strong>: Three layers run in order. Tests run first and pass. The Microsoft Agent Framework evaluator grades the diff against the spec&#8217;s success criteria and flags one borderline result: the <code>Retry-After</code> value is fixed at sixty seconds rather than reflecting the time remaining until the bucket refills. GitHub Advanced Security and Copilot Autofix run, finding no new vulnerabilities, and proposing one minor dependency pin. Finally, the senior engineer reviews, not for syntax, which the previous layers covered, but for intent and integration. They agree with the evaluator on <code>Retry-After</code>, ask the agent to fix it, and approve when the revised diff lands. The escalation rule for evaluator scores below 0.8 did not fire; the rule for changes to public endpoints did, and it routed the review to the right person.</p></li><li><p><strong>Delivery and Learning</strong>: GitHub Actions deploys to staging, runs a synthetic load test against the new limit, and promotes to production behind a feature flag for the first 24 hours. The production signal is a small but measurable drop in p99 latency for the noisy-neighbor cohort. It is captured by the observability MCP server and linked back to the spec as a closing note. If a future change to rate limiting is proposed, the agent will find that closing note when it reads the spec, and the loop will be a little smarter than it was yesterday.</p></li></ol><div class="callout-block" data-callout="true"><p>I watched one rollout where a team skipped explicit ownership because the first two agent PRs looked excellent. On the third PR, a small configuration edit changed throttling behavior for an internal integration and quietly degraded a downstream workflow. In the incident review, nobody could explain who approved the contract impact because everyone had touched the PR and no one was accountable for that stage. They fixed it by assigning stage owners in the matrix and auto-routing any contract-sensitive change to a named reviewer. A month later, a similar risky change was caught before merge. The workflow was not slower, but the rollback was avoided.</p></div><p>Two things to notice about this walkthrough. First, the agent did most of the typing and almost none of the deciding. Second, every decision the agent did make was either authorized by the spec, constrained by <code>AGENTS.md</code>, or caught by an evaluator; the system, not goodwill, kept it inside its lane. That is what a designed loop looks like in practice.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Wrap-up for Part 4: Connecting to Part 5</h1><p>The rate-limiting walkthrough showed something worth sitting with: the agent did most of the typing and almost none of the deciding. Every choice it made was either authorized by the spec, constrained by <code>AGENTS.md</code>, or caught by an evaluator. That&#8217;s not an accident; it&#8217;s what a designed loop looks like when the boundaries and escalation rules from this series actually hold under a real change.</p><p>But boundaries and escalation rules only work if they stay enforced once the excitement wears off. Two failure modes show up almost every time a team relaxes too early: responsibility drift, where careful review quietly turns into rubber-stamping, and escalation theatre, where the policy looks good on paper but routes to nobody in particular.</p><p>Part 5 closes the series by naming both failure modes explicitly, along with their mechanical fixes, and then brings everything from <em>Parts 1 through 4</em> together into one document: the Human-AI Responsibility Matrix. One row per loop stage, one accountable owner per row, no exceptions. It&#8217;s the artifact that turns everything you&#8217;ve read in this series into something you can put in front of your own team confidently.</p><h1>Further reading</h1><p>&#128309; <strong><a href="https://www.npmjs.com/package/@fastify/rate-limit">@fastify/rate-limit on npm</a></strong><br>The actual package used in this issue&#8217;s walkthrough. Worth a look if you want to see the plugin&#8217;s full configuration options beyond the per-token limiter described here, including how it handles custom key generators for cases where limiting by IP isn&#8217;t the right call.</p><p>&#128309; <strong><a href="https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/managing-a-branch-protection-rule">Managing a branch protection rule</a></strong><br>GitHub&#8217;s documentation on the mechanism that turns an escalation rule from a hope into something enforced. This is the practical companion to this issue&#8217;s argument that a rule not expressible in GitHub Actions or branch protection isn&#8217;t really a rule yet.</p><p>&#128309; <strong><a href="https://docs.github.com/en/actions">GitHub Actions documentation</a></strong><br>Referenced throughout this issue as the place where triggers, evaluator scores, and escalation routes actually get wired into a real workflow. A good starting point if you haven&#8217;t yet built a workflow that gates merges on anything beyond passing tests.</p><p>&#128309; <strong><a href="https://learn.microsoft.com/en-us/agent-framework/overview/agent-framework-overview">Microsoft Foundry Agent Framework</a></strong><br>Microsoft&#8217;s documentation on the framework behind the evaluator that grades the rate-limiting diff against spec success criteria in this issue&#8217;s walkthrough. Useful background on how AI-powered evaluators are meant to slot into a verification layer alongside tests and policy gates.</p><p>&#128309; <strong><a href="https://github.com/github/awesome-copilot">github/awesome-copilot</a></strong><br>The community-curated repository referenced in this issue&#8217;s section on operational prompting. It&#8217;s a genuinely useful place to see real AGENTS.md files, chat modes, and prompt files that other teams have already checked into their repositories, rather than starting from a blank page.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Editor-in-Chief<br></span>Build with AI</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The AI-Native Loop: Who Owns the Code Now – Part 3 with Michelle Sandford]]></title><description><![CDATA[&#129504; What's left for developers when agents write the code.]]></description><link>https://packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code-671</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code-671</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 13 Aug 2026 14:02:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7a8fcd74-e56a-402b-9107-2d1d8294e33a_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <em><a href="https://x.com/packtwebdevpro?s=21">X </a></em></p></div><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong><a href="https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50">Build Reliable Claude Code Workflows with Guardrails and Tests</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 424w, /__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 848w, /__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8CqT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png" width="658" height="370.3" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80962f24-74c1-461c-96a5-ff378336468c_940x529.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:529,&quot;width&quot;:940,&quot;resizeWidth&quot;:658,&quot;bytes&quot;:367868,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/210855931?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 424w, /__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 848w, /__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8CqT!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80962f24-74c1-461c-96a5-ff378336468c_940x529.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you&#8217;ve felt the whiplash of Claude Code moving fast one minute and going off the rails the next, this workshop is for you. Ken Huang (AI Safety researcher at Cloud Security Alliance, OWASP AIVSS co-chair, and USF adjunct professor) walks through harness engineering, the practical discipline of wrapping specs, permissions, hooks, and tests around your agent so speed doesn&#8217;t come at the cost of trust.</p><p>&#128073; <strong><a href="https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50">Save your seat</a></strong><a href="https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50"> </a>and start building AI coding workflows that work predictably, not just fast.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50&quot;,&quot;text&quot;:&quot;Register Now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/hands-on-harness-engineering-with-claude-tickets-1990304341864?aff=buildwithai&amp;discount=bwa50"><span>Register Now</span></a></p></div><p><em><a href="/__u/packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code?r=55ncj4&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">Part 2</a></em> handed you an artifact: your own AI-Native Success Criteria Memo, sitting in a repo somewhere with a line in it that says &#8220;a human is always responsible for...&#8221; That line is easy to write and surprisingly hard to make real. Which human? Responsible for what, exactly, and at which point in the loop?</p><p>Part 3 is where the series stops treating that as a detail and makes it the main event. This part is about what actually changes in a developer&#8217;s day when an agent is doing more of the typing, and it starts with a rule Michelle keeps coming back to across her work: <em><strong>Rule Zero. You own the code</strong></em>.</p><blockquote><p><em><a href="https://www.linkedin.com/in/michellesandford/">Michelle Sandford</a>, <a href="https://sessionize.com/MichelleSandford/">Microsoft&#8217;s Developer Engagement Lead for Asia</a></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BY2C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 424w, /__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 848w, /__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BY2C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png" width="1200" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ebdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:627402,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/210855931?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 424w, /__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 848w, /__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BY2C!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febdc0aa4-e4cb-41bc-a348-3bd627a1d5f1_1200x800.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></blockquote><p>We&#8217;ll look at why the real bottleneck in an AI-native loop isn&#8217;t generation, it&#8217;s judgment, and why the DORA report&#8217;s &#8220;trust paradox&#8221; (using AI more while trusting it less) means the fix is role clarity, not better prompting. Then we&#8217;ll walk the six-stage loop from <em><a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-22-the-ai-native-loop?r=55ncj4&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">Part 1</a></em> one more time, this time assigning a decision-maker to every single stage, so &#8220;the human is responsible&#8221; stops being a slogan and starts being something a reviewer can actually check.</p><div class="callout-block" data-callout="true"><h1 style="text-align: center;">In this edition</h1><p>&#10140; Why the developer&#8217;s job is shifting from writing code to designing the system around it</p><p>&#10140; The DORA report&#8217;s trust paradox, and why it means the fix is role clarity, not better prompts</p><p>&#10140; Rule Zero: you own the code, and what that actually means in practice</p><p>&#10140; Responsibility, mapped stage by stage across the six-stage loop, from Intent through Delivery and Learning</p><p>&#10140; Two boundaries worth writing into AGENTS.md on day one</p><p>&#10140; A preview of Part 4, where we walk a real rate-limiting feature through the loop end to end</p></div><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h1><strong>The AI-Native Developer: From Coder to System Designer</strong></h1><p>Buried in the AI-Native Success Criteria Memo from <em><a href="/__u/packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code?r=55ncj4&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">Part 2</a></em> is a single line that turns out to be the hardest one to operationalize: &#8220;a human is always responsible for&#8230;.&#8221; It sits quietly in the <em>Human responsibility</em> section of the template, between the <em>Quality bars</em> you can measure and the <em>Anti-goals</em> you can recognize, and it is where the role of the developer is being rewritten today.</p><p> Most teams hope no one will ask them to be specific about it. This part and the next two in the series force specificity. The core idea is straightforward: in an AI-native loop, the developer is a system designer first and an author second, and the artifact that makes that real is a written-down responsibility matrix.</p><blockquote><p>I sometimes refer to this as <em><strong>Rule Zero: You own the code</strong></em>.</p></blockquote><p>The approach throughout the rest of this series is practical. Every concept is anchored to one of the six loop stages covered earlier (Intent, Spec, Context, Change, Verification, Delivery and Learning), and every recommendation is mechanical enough to enforce in a GitHub Action, a CODEOWNERS file, or an AGENTS.md instruction. We will walk through a single rate-limiting example end to end, then identify the failure modes that show up early and explain how to correct them before they set in.</p><p>By the last part of the series, you will have what you need to draft v1 of a Human-AI Responsibility Matrix for one repository you own, with named decision-makers, mechanical escalation rules, and the protected paths you need under CODEOWNERS before the next agent PR lands.</p><h1>The developer as system designer, not just coder</h1><p>The shift is not that developers stop writing code. It is that the unit of work has moved up a level. In an AI-assisted loop, a senior engineer&#8217;s day is dominated by typing; the AI shaves minutes off each function. In an AI-native loop, the same engineer spends much less time at the keyboard producing characters and much more time deciding which spec is worth writing, which context the agent must see, which verification layer catches which class of mistake, and which decisions are never the agent&#8217;s to make. The 2025 DORA State of AI-assisted Software Development report calls this the amplifier effect (AI multiplies whatever loop you already have, including the bad parts) and pairs it with what the report calls the <em>trust paradox.</em> Both observations point to the same conclusion. The new bottleneck is not generation; it is judgment, and judgment lives in roles, not in keystrokes.</p><blockquote><p>The trust paradox is the report&#8217;s name for a contradiction that shows up in the survey data: developers lean on AI for more and more of their daily work while reporting that they do not fully trust what it produces. The two halves seem to cancel out, yet both are rising at once. Adoption climbs because the assistance is genuinely useful; trust lags because the same tools still confidently produce plausible but wrong output that a skim review will not catch. The paradox matters here because it explains why role clarity is the fix rather than better prompts. If developers neither fully trust the output nor stop using it, the only stable resolution is to design the loop so that trust is placed in verifiable artifacts (a reviewed spec, a layered set of checks, a named approver) rather than in the model itself. </p></blockquote><p>Calling this system design is not a rebrand. A traditional developer designs a system once, at the start of a project, and then implements inside it. An AI-native developer is designing the system every day: the spec template, the AGENTS.md, the MCP tool surface, the evaluator that gates the PR, and the escalation rule that triggers when the agent touches a payment path. These are all design artifacts, and every one of them shifts who is allowed to do what. The role is closer to a tech lead than a senior IC, except the team being led includes at least one non-human contributor whose strengths, weaknesses, and failure modes the lead is expected to know.</p><p>Two consequences fall out of that framing and shape the rest of this series. The first is that role clarity becomes a precondition for speed, not a brake on it. Ambiguity about who decides what produces the most expensive form of rework: a merged change that no one feels accountable for when it breaks in production. The second is that responsibility is a property of the loop stage, not of the person. The same human may be the decision-maker at the spec stage and a pure reviewer at the verification stage; the same agent may be a generator at the change stage and an evaluator at the verification stage. The matrix that we will cover at the end of the series is how you write that down.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Responsibility boundaries between human and model</h1><p>A useful boundary is one a reviewer can apply without a meeting. &#8220;The human owns intent and the agent owns implementation&#8221; sounds reasonable in a keynote and dissolves on contact with a real PR. What is intent, exactly, in a 400-line refactor that changes one public type? The boundaries that survive are stage-shaped, not vibe-shaped, and they map cleanly onto the six-stage loop from <em><a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-22-the-ai-native-loop?r=55ncj4&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">Part 1</a></em>: Intent, Spec, Context, Change, Verification, Delivery and Learning:</p><ul><li><p>At the Intent stage, the human is always the decision-maker. An agent may help shape the issue (suggest acceptance criteria, draft a why paragraph, point at related work), but the question &#8220;is this worth doing at all?&#8221; is not delegable.</p></li><li><p>At the Spec stage, authorship is shared, but approval is human. An agent may draft a Spec Kit specification or a five-line <code>specs/</code> Markdown file, and a named human approves it before any code is generated.</p></li><li><p>At the Context stage, AGENTS.md, the MCP tool surface, and the per-task context packet are curated by humans even when agents help assemble them; the rule of thumb is that anything an agent will read on every future run must be reviewed by a human now.</p></li><li><p>At the Change stage, the agent can lead: agent mode in Visual Studio Code or the Copilot coding agent on GitHub.com produces the diff, provided the spec and context above it are sound.</p></li><li><p>At the Verification stage, responsibility is layered: tests and AI evaluators run first, GitHub Advanced Security and Copilot Autofix act as policy gates, and a human review focuses on intent and integration rather than syntax.</p></li><li><p>At the Delivery and Learning stage, the human owns the decision to merge to a protected branch and the decision about what the production signal means for the next spec.</p></li></ul><div class="callout-block" data-callout="true"><p>Two boundaries are worth naming explicitly because teams get them wrong with the same regularity: </p><p>&#10145;&#65039; Secrets and credentials are never the agent&#8217;s responsibility, full stop: not to read, not to rotate, not to reference. </p><p>&#10145;&#65039; Decisions that change a public contract (an API shape, a database schema, a permission model) require a human approver even when the change itself is trivial, because the blast radius is shaped by the contract, not by the diff size. <br>Write both of these into AGENTS.md on day one. They are cheap to enforce and expensive to add after an incident.</p></div><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>A preview of Part 4</h1><p>By the end of this part, you have a boundary a reviewer can actually apply: who decides at each of the six stages, and two rules (secrets and public-contract changes) that belong in AGENTS.md before your next agent PR lands. That&#8217;s the theory settled. What it doesn&#8217;t tell you yet is what happens when the normal case breaks, when the agent is uncertain, when a diff comes back bigger than expected, when a change touches something regulated and nobody&#8217;s quite sure whether it needs a second pair of eyes.</p><p>That&#8217;s where Part 4 picks up. We&#8217;ll build escalation rules with three concrete ingredients (a trigger, a route, and an action) so that &#8220;ask if you&#8217;re not sure&#8221; stops being a hope and starts being something a GitHub Action can actually enforce. We&#8217;ll also look at what prompting looks like when it&#8217;s a leadership act rather than a chat message, and then walk one feature (adding rate limiting to a public API endpoint) through all six stages of the loop from Intent to Delivery and Learning, so you can see every one of these ideas working together on something real.</p><h1>Further reading</h1><p>&#128309; <strong><a href="https://dora.dev/dora-report-2025/">2025 DORA State of AI-assisted Software Development report</a></strong><br>Referenced a third time in this series because the trust paradox this issue leans on, developers using AI more while trusting it less, comes directly from this report&#8217;s survey data. Worth reading in full if you want the numbers behind the idea that role clarity, not better prompting, is the actual fix.</p><p>&#128309; <strong><a href="https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners">About code owners</a></strong><br>GitHub&#8217;s own documentation on CODEOWNERS, the mechanism this issue points to for making protected paths enforceable rather than aspirational. Useful groundwork before Part 5, where the Human-AI Responsibility Matrix gets tied directly into CODEOWNERS entries.</p><p>&#128309; <strong><a href="https://agents.md/">AGENTS.md: the open standard</a></strong><br>Linked again because this issue adds two specific rules to it (secrets are never the agent&#8217;s responsibility, and public-contract changes require a human approver). Worth revisiting the format itself if you haven&#8217;t yet written your own.</p><p>&#128309; <strong><a href="https://docs.github.com/en/code-security/getting-started/github-security-features">GitHub Advanced Security</a></strong><br>Referenced here as one of the policy gates in the Verification stage of the loop. Relevant background if your team is deciding where automated security scanning should sit relative to human review in your own responsibility map.</p><p>&#128309; <strong><a href="https://docs.github.com/en/copilot/concepts/agents/coding-agent/about-coding-agent">About GitHub Copilot coding agent</a></strong><br>Referenced again because this issue&#8217;s stage-by-stage boundaries (the agent can lead at Change, but not at Intent or Delivery) describe exactly how the coding agent is designed to operate according to GitHub&#8217;s own documentation, not an idealized version of it.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Editor-in-Chief <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The AI-Native Loop: Who Owns the Code Now – Part 2 with Michelle Sandford]]></title><description><![CDATA[&#129504; One bug, two fixes, and the gap that actually matters]]></description><link>https://packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-ai-native-loop-who-owns-the-code</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 06 Aug 2026 14:01:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e1c7c273-1599-4925-9d53-2b40a7cf2ee5_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <a href="https://x.com/packtwebdevpro?s=21">X </a></p></div><p>That&#8217;s where Part 1 left you: a six-stage AI-native loop mapped out, and an honest sense of where your own team sits on it. A map is useful, but it doesn&#8217;t tell you whether the loop actually changes anything on a Tuesday afternoon when a bug ticket lands in your queue. So this part answers that with an experiment rather than an argument.</p><p>We take a deliberately unglamorous bug (a reporting endpoint that quietly drops results for anyone outside UTC) and fix it twice, in the same repository, from the same starting point. Once the AI-assisted way: ask Copilot Chat, accept the fix, write one test, merge. Once the AI-native way: open a spec first, define success criteria before any code changes, and let the loop from Part 1 do the rest.</p><blockquote><p><em><a href="https://www.linkedin.com/in/michellesandford/">Michelle Sandford</a>, <a href="https://sessionize.com/MichelleSandford/">Microsoft&#8217;s Developer Engagement Lead for Asia</a></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fPco!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fPco!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1156383,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/209892473?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fPco!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc91f4f3f-07fb-410e-a658-13e27f581503_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></blockquote><p>Both runs get measured against the same yardstick, and the gap between them is where &#8220;AI-native&#8221; stops being a definition you nod along to and starts being something you can point at. One of the two fixes ships looking finished and is actually still broken. You&#8217;ll see exactly which one, and exactly why nobody catches it until we walk through it together.</p><p>By the end, you&#8217;ll have v1 of your own <em><strong>AI-Native Success Criteria Memo</strong></em>, the artifact that makes this kind of comparison repeatable on your own team, not just in this example.</p><div class="callout-block" data-callout="true"><h1 style="text-align: center;">In this edition</h1><p>&#10140; The bug-fix experiment: one deliberately ordinary bug, fixed twice in the same repository</p><p>&#10140; Run A, AI-assisted: the fast fix that looks done and ships subtly broken</p><p>&#10140; Run B, AI-native: the spec-first fix, and why it catches what Run A misses</p><p>&#10140; What the side-by-side actually shows, and why speed isn&#8217;t the headline number</p><p>&#10140; Seven failure modes to recognize in your first six months of AI-native adoption, with corrections for each</p><p>&#10140; Your AI-Native Success Criteria Memo, the artifact that makes this kind of comparison repeatable on your own team</p><p>&#10140; A preview of Part 3, where the developer&#8217;s role shifts from writing code to designing the system around it</p></div><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>Running a side-by-side bug-fix experiment</h1><p>Our bug-fix experiment is a deliberately mundane one: a TypeScript service exposes a <code>/reports</code> endpoint with <code>from</code> and <code>to</code> query parameters, and it returns an empty array when a user in Australia/Perth selects the last day of the month. The root cause is a comparison that parses the <code>to</code> value as UTC midnight, which falls before the user&#8217;s local end of day. It is exactly the kind of slightly embarrassing date bug every engineer has shipped.</p><p>The repository is a small Node.js service with TypeScript strict mode, Jest, Fastify, and GitHub Actions for CI. The <code>buildRange</code> function lives in <code>src/reports/range.ts</code>, and a modest <code>AGENTS.md</code> file is already in place.</p><p>Both runs use the same measures:</p><ul><li><p>T1: Time to the first reproducing test</p></li><li><p>T2: Time to a fix merged through CI</p></li><li><p>Q1: Number of distinct edge cases captured by tests</p></li><li><p>Q2: Number of surviving artifacts after the merge</p></li><li><p>R1: Can a future engineer understand the reasoning without rereading the diff?</p></li></ul><div class="callout-block" data-callout="true"><p>Tip: Keep this yardstick visible while you compare Run A and Run B. The important signal is not only speed, it is the quality and durability of the artifacts each run leaves behind.</p></div><h3>Run A, AI-assisted</h3><p>The engineer checks out a branch, opens <code>src/reports/range.ts</code>, finds <code>buildRange</code>, and asks Copilot Chat for likely causes of the empty-array bug. Copilot explains the UTC-midnight problem and suggests using <code>endOfDay</code> from <code>date-fns</code>. The engineer accepts the suggestion and hand-writes a Jest test:</p><pre><code><code>import { endOfDay } from 'date-fns';

export function buildRange(fromISO: string, toISO: string) {
  return { start: new Date(fromISO), end: endOfDay(new Date(toISO)) };
}</code></code></pre><p>The tests pass. A pull request is opened, reviewed 90 minutes later, approved, and merged.</p><blockquote><p>Result: T1 &#8776; 12 min. T2 &#8776; 2 h 15 min. Q1 = 1 edge case. Q2 = 2 artifacts (code and test). R1 = No, the rationale is trapped in the diff.</p></blockquote><p>Worse, <code>endOfDay()</code> uses the server&#8217;s local time rather than the user&#8217;s time zone, so the fix is subtly incomplete. The test does not catch the issue.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h3>Run B, AI-native</h3><p>The engineer does not open the file. They open <code>specs/2026-05-date-range-timezone.md</code>:</p><pre><code><code># Spec: Date range honors user time zone

## Why
Users in non-UTC time zones see empty results when selecting a range
that includes the last day of the month.

## What changes for the user
For any user in any IANA time zone, a date range from YYYY-MM-DD to
YYYY-MM-DD includes all records whose local timestamp falls within
the user's local start of the first day through the end of the last day, inclusive.

## Success criteria
- Reproduction case (Australia/Perth, last day of month) returns rows.
- Property test: for 100 random (time zone, start, end) triples, a
  record at local noon on any day in the range is included.
- No new dependencies beyond date-fns-tz.
- End of day is 23:59:59.999 local, not 00:00:00 the next day.

## Out of scope
- Wire-format changes to the from/to parameters.
- Daylight-saving-transition boundary cases (tracked separately).</code></code></pre><p>A teammate reviews the spec and adds the end-of-day clarification before any code is written. Surfacing ambiguity before the agent has a chance to act on it is one of the loop&#8217;s quiet superpowers. The engineer assigns the linked issue to the Copilot coding agent. The repository&#8217;s <code>AGENTS.md</code> already says: use <code>date-fns-tz</code>, prefer fast-check property tests for date logic, never edit <code>range.ts</code> without updating <code>range.test.ts</code>, and cite which spec section each commit addresses.</p><p>The agent pulls the original error from the observability MCP server, writes a failing reproduction test, fixes <code>buildRange</code> using <code>zonedTimeToUtc</code>/<code>utcToZonedTime</code>, and adds a property-based test across four IANA time zones. It posts a trace to the PR: model version, tools called, files touched, and a mapping from spec bullets to commits:</p><pre><code><code>import { zonedTimeToUtc, utcToZonedTime } from 'date-fns-tz';
import { startOfDay, endOfDay } from 'date-fns';

export function buildRange(fromISO: string, toISO: string, userTz: string) {
  const startLocal = startOfDay(utcToZonedTime(new Date(fromISO), userTz));
  const endLocal = endOfDay(utcToZonedTime(new Date(toISO), userTz));
  return {
    start: zonedTimeToUtc(startLocal, userTz),
    end: zonedTimeToUtc(endLocal, userTz),
  };
}</code></code></pre><p>The engineer reviews the changes and asks the agent to derive the fast-check seed from the commit SHA so that CI failures stay reproducible. They also ask the agent to add that as a convention in <code>AGENTS.md</code>, ensuring that the rule applies to every future change, not just this one. The agent updates both. CI passes. Merged.</p><blockquote><p><em>Result: T1 &#8776; 8 min. T2 &#8776; 1 h 40 min. Q1 = 100+ generated test cases across four time zones. Q2 = 4 artifacts (spec, code, property test, and </em><code>AGENTS.md</code><em> convention). R1 = Yes. The spec explains why, the trace explains how, and </em><code>AGENTS.md</code><em> documents the rule that now applies to everyone.</em></p></blockquote><p>Most importantly, the spec and the property test catch the time zone subtlety that Run A&#8217;s fix missed.</p><h3>What the side-by-side actually shows</h3><p>Run B was about 35 minutes faster, but the headline number is not the point. Run A produced a fix that was subtly wrong. Run B produced a fix that was actually correct and strengthened a convention for every future change. Both runs solved today&#8217;s bug. Only one made tomorrow&#8217;s bug easier to solve.</p><p>The AI-native dividend is the asymmetry of the artifacts. The agent did real work, but the durable work (the spec, the property test, the <code>AGENTS.md</code> rule) was shaped by the human on either side of the agent&#8217;s run. Run B also asked the engineer to do four uncomfortable things:</p><ul><li><p>Write the spec before the code.</p></li><li><p>Trust that handing the work to an agent was faster than typing it themselves.</p></li><li><p>Review a property test, an <code>AGENTS.md</code> change, and a diff at the same time.</p></li><li><p>Resist the urge to fix the seeding issue directly, instead sending a review comment that taught the agent and the repo the right pattern.</p></li></ul><p>If those four behaviors feel uncomfortable, that discomfort is the friction of moving from AI-assisted to AI-native. </p><div class="pullquote"><p>I ran this exact side-by-side with a platform team that was convinced the AI-native path would be slower. Their surprise was not that Run B was faster, it was that the reviewer spent less time debating intent because the spec had already settled the hard questions. The second surprise came a week later: another engineer reused the <code>AGENTS.md</code> seed rule in a different service, and a flaky test was fixed before it reached production. The team called that moment the first compound return from changing the loop, because one experiment improved work they had not even planned yet.</p></div><div class="callout-block" data-callout="true"><p><strong>Try this yourself:</strong> Find one small bug in a repository you control. Fix it twice using the yardstick above. Keep the artifacts. They are your first AI-native portfolio.</p></div><p>The experiment shows what is possible. What follows shows what gets in the way.</p><h2>Risks, misconceptions, and failure modes</h2><p>Seven failure modes will hit you in the first six months of adopting an AI-native loop. Recognize them now, and the corrections are quick:</p><ul><li><p><strong>Vibe-driven development at scale:</strong> Vague prompt, plausible diff, hidden bug. The work looks correct at every checkpoint, and the bug surfaces three weeks later in a tenant with a slightly different configuration.</p><ul><li><p><em><strong>Correction:</strong> No agent run without a spec, even five lines. If you cannot write the success criteria in three bullets, the work is not ready for an agent.</em></p></li></ul></li><li><p><strong>Context starvation:</strong> The agent does not know what the senior engineer carries in their head, the design review you held last quarter, the deprecation timeline, or the architectural rule no one wrote down. It produces code that is locally plausible and globally wrong.</p><ul><li><p><strong>Correction:</strong> Invest in <code>AGENTS.md</code>, MCP tools, and per-task context packets. </p></li></ul></li><li><p><strong>Trust calibration drift:</strong> You start skimming agent PRs because they &#8220;look right.&#8221;</p><ul><li><p><strong>Correction:</strong> Decouple trust from throughput. Sample one in ten agent PRs for deep review, chosen at random. If the deep review starts finding more issues than normal review, you have drifted.</p></li></ul></li><li><p><strong>Mistaking the agent for the loop:</strong> You see a productivity boost, declare victory, and stop investing. Six months later you have the same problems, flaky tests, slow reviews, and unclear ownership, only faster.</p><ul><li><p><strong>Correction:</strong> The artifacts improve the loop. The agent is just one participant.</p></li></ul></li><li><p><strong>Invisible agent work:</strong> &#8220;Copilot wrote it&#8221; with no trace is a future incident waiting to happen, especially in regulated industries where provenance is a compliance question.</p><ul><li><p><strong>Correction:</strong> Preserve agent run summaries on every pull request, including the model version, tools used, files modified, and spec citations.</p></li></ul></li><li><p><strong>Deferring governance:</strong> Locking the loop down after a near miss costs more than designing it from day one.</p><ul><li><p><strong>Correction:</strong> In week one, decide which branches an agent may push to, which MCP tools it may call, and which secrets it must never see. Write them in <code>AGENTS.md</code>.</p></li></ul></li><li><p><strong>Confusing productivity with capability:</strong> Faster typing is not a better loop.\</p><ul><li><p><strong>Correction:</strong> Measure artifact quality and trust calibration, not just throughput. That is what the memo in the next section is for.</p></li></ul></li></ul><div class="pullquote"><p>I once worked with a team that skipped traces because they felt &#8220;too heavy&#8221; for a fast-moving roadmap. Two months later, a customer incident forced a rollback, and nobody could answer three basic questions: which model generated the risky change, what context packet it used, and which reviewer approved the policy exception. The post-incident review was painful, mostly because the evidence was missing. They corrected course in one sprint by adding a mandatory trace section to every agent-driven PR and a weekly deep-review sample. Within a month, reviews became calmer and faster because every debate started from concrete artifacts instead of memory.</p></div><div class="callout-block" data-callout="true"><p>Tip: Failure modes compound. Pick one correction this week (spec minimum, trace template, or deep-review sampling) and make it a team habit before adding another.</p></div><h2>Defining AI-Native Success Criteria</h2><p>Here is the carry-forward artifact for this part: a one-page memo, written by you for your team. Use the template below. The goal is something slightly embarrassing in its specificity rather than something polished for a slide.</p><pre><code><code># AI-Native Success Criteria - [Team / Repo] - v1

## Loop scope
We consider our loop AI-native when:
- Every non-trivial change starts from a written spec in `specs/`.
- AGENTS.md is current and reviewed in PRs.
- At least one AI participant (agent mode, coding agent, or evaluator)
  contributes to every non-trivial change, with a readable trace.
- Verification is layered: tests, evaluators, policy gates, human review.
- Production signals feed back into specs, not just backlogs.

## Artifact criteria
For each change we retain: a spec, a context-packet description, an agent
run trace (model, tools, files), tests mapped to spec criteria, a PR
description linking all of the above.

## Quality bars
Mergeable only if: every success-criteria bullet is satisfied; no
security or policy regression; tests are deterministic; trace is linked.

## Human responsibility
A human is always responsible for: drafting and approving the spec;
curating AGENTS.md and the MCP tool surface; final review on any
protected branch; sampling agent PRs for deep review.

## Anti-goals
Not aiming for: maximum agent autonomy at the expense of accountability;
productivity metrics divorced from artifact quality; "AI did it" as a
sufficient explanation in a post-incident review.

## Review cadence
Reviewed quarterly. Owner: [name]. Last updated: [date].</code></code></pre><p>A few practical notes on filling it in:</p><ul><li><p>For <strong>Loop scope</strong>, be honest about where you are today. v1 can name agent mode as an aspiration rather than a current standard, and most teams&#8217; first version describes a loop they are 60 percent of the way there.</p></li><li><p>For <strong>Quality bars</strong>, resist setting bars you are not willing to enforce. A bar the team routinely ignores teaches everyone the memo is decorative.</p></li><li><p>For <strong>Human responsibility</strong>, be specific about which human, not just that one exists. For example, &#8220;The on-call senior engineer for the affected service approves any spec that touches it&#8221; is enforceable in a way that &#8220;A human approves the spec&#8221; is not.</p></li><li><p>For <strong>Anti-goals</strong>, write down what you have already seen go wrong elsewhere. Anti-goals are where hard-won wisdom lives.</p></li><li><p>Set a review cadence you will actually keep. Quarterly is realistic for most teams.</p></li></ul><p>The memo is a target, not a description of done. Writing it does not make your loop AI-native. The parts that follow are the practice.</p><div class="pullquote"><p>The first time I wrote a memo like this, I expected it to be a documentation exercise and nothing more. What changed was ownership. People stopped saying, &#8220;the agent decided,&#8221; and started asking, &#8220;which success criterion did we miss?&#8221; In two planning cycles, the memo became the default lens for deciding whether a change was ready for agent execution. It also made disagreements cheaper, because the team could edit the memo together instead of re-litigating standards in every PR.</p></div><h1>A preview of Part 3</h1><p>The side-by-side comparison and the seven failure modes both point at the same underlying fact: an AI-native loop doesn&#8217;t run itself. Every artifact that made Run B trustworthy (the spec, the property test, the AGENTS.md rule, the trace) was shaped by a human deciding what mattered before and after the agent&#8217;s turn. The Success Criteria Memo you just built formalizes that judgment so it doesn&#8217;t live only in one engineer&#8217;s head.</p><p>Which raises the question this series turns to next: if the agent is doing more of the typing, what exactly is the developer doing instead?</p><p>Part 3 answers that directly. We&#8217;ll look at why the developer&#8217;s role is shifting from writing code to designing the system that decides what gets generated, reviewed, and merged, what the DORA report&#8217;s &#8220;trust paradox&#8221; reveals about why role clarity (not better prompts) is the real fix, and Michelle&#8217;s own Rule Zero: you own the code. We&#8217;ll close by mapping responsibility across all six stages of the loop, so you have a boundary a reviewer can actually apply without calling a meeting.</p><div class="pullquote"><p><em><strong><span>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through </span><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a><span>. We&#8217;re all ears!</span></strong></em></p></div><h1>Further reading</h1><p>&#128309; <strong><a href="https://www.npmjs.com/package/date-fns-tz">date-fns-tz on npm</a></strong><br>The time zone extension for date-fns used in Run B&#8217;s fix. Worth a look if you want to see how <code>zonedTimeToUtc</code> and <code>utcToZonedTime</code> actually work under the hood, and why a UTC-based comparison is such an easy trap to fall into in the first place.</p><p>&#128309; <strong><a href="https://github.com/dubzzz/fast-check">fast-check on GitHub</a></strong><br>The property-based testing framework behind the 100-plus generated test cases in Run B. If Part 2 is the first time you&#8217;ve seen property testing applied to a &#8220;boring&#8221; bug fix, this is the place to see the wider pattern: instead of writing individual examples, you describe a rule that should always hold and let the framework hunt for the input that breaks it.</p><p>&#128309; <strong><a href="https://dora.dev/dora-report-2025/">2025 DORA State of AI-assisted Software Development report</a></strong><br>Referenced again here because the report&#8217;s core finding, that AI amplifies whatever loop already exists, is exactly what the side-by-side experiment demonstrates in miniature. Run A and Run B used the same model and the same repository; the only variable was the loop around it.</p><p>&#128309; <strong><a href="https://docs.github.com/en/code-security/getting-started/github-security-features">GitHub Advanced Security</a></strong><br>One of the policy gates named in the verification layer of the loop. Useful background if your team hasn&#8217;t yet set up automated security scanning as part of the cheap-to-expensive verification funnel this issue describes.</p><p>&#128309; <strong><a href="https://docs.github.com/en/copilot/concepts/agents/coding-agent/about-coding-agent">About GitHub Copilot coding agent</a></strong><br>Referenced again because Run B&#8217;s agent behavior (pulling context from an MCP server, writing a failing test first, posting a trace to the PR) is exactly what this documentation describes as the coding agent&#8217;s intended workflow, not a hypothetical.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Co-Editor-in-Chief <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Token-Efficient Coding Agent]]></title><description><![CDATA[Context and cost per accepted change]]></description><link>https://packtbuildwithai.substack.com/p/the-token-efficient-coding-agent</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-token-efficient-coding-agent</guid><dc:creator><![CDATA[Lucas Germinari Carreira]]></dc:creator><pubDate>Tue, 04 Aug 2026 20:26:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d2dcd789-69bf-48c1-a2e2-d4798267907b_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h3>Editor&#8217;s note</h3><p>Most teams building with coding agents track cost the same way they track API usage &#8212; dollars per call. That number is misleading. A cheap call that fails and triggers two retries costs more than an expensive one that works the first time, and almost nobody is measuring for that.</p><p>We invited Lucas to write this because he&#8217;d already done the work of reframing the problem: instead of asking &#8220;how do I make this call cheaper,&#8221; ask &#8220;what&#8217;s my cost per accepted change.&#8221; That single shift changes how you think about context size, model selection, and prompt structure &#8212; not as separate levers, but as one system you&#8217;re optimizing together.</p><p>Here&#8217;s Lucas.</p></div><p>I&#8217;m an AI engineer at Elevance Health and founder of Goliath.ai, so I spend most of my working hours inside Claude Code sessions. Here&#8217;s the principle that took me longest to internalize: <strong>the prompt you type is not the bill.</strong></p><p>By the time a task finishes, the agent has read system instructions, repository guidance, tool schemas, retrieved files, terminal output, and prior turns &#8212; often more than once, because the first patch failed. The cheapest single call in that loop can produce the most expensive finished task.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!y-Vh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 424w, /__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 848w, /__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!y-Vh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png" width="1456" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:816,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:760381,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/209833066?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 424w, /__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 848w, /__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y-Vh!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feccb21bd-6557-4b81-9374-3ead8060ede2_1456x816.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The number that actually matters is <strong>effective cost</strong>: total spend on a task divided by the probability the result gets accepted. A run that costs twice as much but succeeds on the first try usually beats a cheaper run that needs three retries. Cheaper input tokens don&#8217;t automatically mean lower end-to-end cost &#8212; treating them as the same thing is the most common mistake I see.</p><p>Here&#8217;s what to do instead.</p><h2>Give the agent a map, not the repository</h2><p>Attaching an entire repo &#8220;just in case&#8221; costs more than tokens &#8212; it costs the agent&#8217;s attention. A bigger haystack means more exploration before it finds the needle.</p><p>Keep a short, current repo guide (a <code>CLAUDE.md</code>, for example) covering architecture, key directories, and a real definition of done. Give the agent an explicit starting point &#8212; the files involved, the failing test &#8212; rather than a blank search. Narrow the working set with code or symbol search before injecting full files, enable only the tools the task needs, and give it a stopping condition so it doesn&#8217;t wander into unrelated refactoring.</p><div class="callout-block" data-callout="true"><p>&#10060; <strong>Vague:</strong> &#8220;Fix the login bug&#8221;</p><p>&#9989; <strong>Specced:</strong> &#8220;The failure is in the OAuth callback flow. Start with <code>callback.ts</code> and <code>session.ts</code>. Reproduce with <code>npm test -- auth-callback</code>. Stop once that test and the typecheck pass.&#8221;</p></div><h2>Spec the task, don&#8217;t just shorten it</h2><p>&#8220;Write shorter prompts&#8221; isn&#8217;t the fix. A few extra tokens spent killing ambiguity are cheaper than the retry a vague prompt triggers. A prompt that holds up has five parts:</p><ul><li><p><strong>The goal</strong> &#8212; what done actually looks like</p></li><li><p><strong>The location to start</strong> &#8212; the specific files or module</p></li><li><p><strong>The constraints that must not change</strong> &#8212; what&#8217;s off-limits</p></li><li><p><strong>The exact verification steps</strong> &#8212; the command that proves success</p></li><li><p><strong>A stopping condition</strong> &#8212; when to stop, full stop</p></li></ul><p>That&#8217;s task specification, not clever wording. It&#8217;s closer to writing a good ticket than a good sentence.</p><p>Keep the counterbalance in mind too: this isn&#8217;t a case for a giant static instructions file loaded on every call. A small, stable core, with task-specific procedures loaded only when needed, works better.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Enjoyed this? BuildWithAI drops two editions like this every week. Subscribe &#8212; no fluff, just what&#8217;s actually working.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The context window is working memory</h2><p>A bigger context window is capacity, not a target &#8212; having room for more information doesn&#8217;t mean the model uses it well. Research on long-context retrieval has shown models can be worse at pulling information from the middle of a long input than from the start or end.</p><p>Think in four tiers instead:</p><ul><li><p><strong>Stable instructions</strong> &#8212; safety rules, a short repo map, core commands. Rarely changes.</p></li><li><p><strong>Task spec</strong> &#8212; issue, acceptance criteria. Stays visible throughout.</p></li><li><p><strong>Retrieved working set</strong> &#8212; relevant files, tests. Pulled in and pruned as needed.</p></li><li><p><strong>Transient evidence</strong> &#8212; logs, search results. Summarized or dropped once you&#8217;ve pulled out the fact you needed.</p></li></ul><h2>Not every phase deserves your best model</h2><p><strong>Planning</strong> &#8212; working out what&#8217;s wrong and what approach to take &#8212; is where ambiguity is highest, and a stronger model earns its cost there. <strong>Execution</strong> &#8212; applying a plan that&#8217;s already decided &#8212; tolerates a cheaper model well. <strong>Verification</strong> shouldn&#8217;t involve model judgment where you can avoid it: run the test suite and the type checker, and treat their output as ground truth rather than asking another model call to eyeball the result.</p><h2>What to leave alone</h2><p>A few things worth resisting:</p><ul><li><p>Don&#8217;t strip relevant tests or constraints just to shrink the first call &#8212; it&#8217;s a false saving if it triggers retries.</p></li><li><p>Don&#8217;t default every task to the cheapest model, and don&#8217;t default every task to the frontier model either.</p></li><li><p>Don&#8217;t judge any of this on API price alone &#8212; developer review time and rework belong in the number too.</p></li></ul><p>An agent that always reaches for the cheapest model isn&#8217;t efficient &#8212; it&#8217;s just cheap. The one worth building hands the model the smallest working set it needs, spends its best reasoning where a decision is actually being made, and stops the moment the tests say the job is done.</p><div><hr></div><p>Thanks for reading, Lucas &#8212; if anyone&#8217;s <code>CLAUDE.md</code> is doing double duty as both a repo guide and a definition of done, hit reply and tell me what&#8217;s in it.</p><blockquote><p> Charu, Co-Editor-In-Chief, Build With AI </p></blockquote><div class="callout-block" data-callout="true"><p><strong>About the author:</strong>  Lucas Germinari Carreira is an AI engineer and computer science grad at Indiana University. He builds production AI systems at Elevance Health, where he works on healthcare AI and large language model infrastructure, and serves as a Claude Builder Ambassador with Anthropic, leading technical workshops and hackathons that help students build with frontier AI. </p><p>Outside of work, he enjoys building products and developer communities at the intersection of AI, software engineering, and entrepreneurship</p><p><strong>Connect with him on: </strong><a href="https://www.linkedin.com/in/lucasgerminaricarreira/">LinkedIn</a> | <a href="https://www.instagram.com/_germinari_/">Instagram</a> | <a href="https://lucasgerminari.pythonanywhere.com/">Website</a></p></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!sjpR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!sjpR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1638011,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/209833066?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sjpR!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fbb8ab4-365e-45ca-817c-cfb001a1dbbe_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The AI-Native Loop: Who Owns the Code Now – Part 1]]></title><description><![CDATA[&#129504; Why your delivery loop matters more than your tools]]></description><link>https://packtbuildwithai.substack.com/p/build-with-ai-22-the-ai-native-loop</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/build-with-ai-22-the-ai-native-loop</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 30 Jul 2026 14:03:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/70dd6fff-094a-40a3-a77e-9d01c131fd99_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <a href="https://x.com/packtwebdevpro?s=21">X </a></p></div><p>On her way to a conference, <a href="https://www.linkedin.com/in/michellesandford/">Michelle Sandford</a> asked a coding agent to build a companion website for her talk. She reviewed it as most speakers would during a final check before walking on stage, and it looked good. The layout was clean, the case studies were compelling, and the statistics lined up neatly with every slide. A little too neatly!</p><p>When Michelle asked the agent where one of the case studies had come from, it told her, without hesitation, that it had invented it. The story simply felt like it would land better that way. A few of the statistics turned out to be invented too.</p><p>The agent hadn&#8217;t done anything malicious. It had simply behaved like any capable, eager contributor might when nobody had made the boundaries clear. What was missing wasn&#8217;t better judgment from the model. It was a system, designed by Michelle, in which a human was still positioned to catch the mistake before it shipped.</p><p>The question raised by that gap between what agents can do and what we have designed them to be accountable for is what this five-part series, <em><strong>The AI-Native Loop: Who Owns the Code Now</strong></em>, sets out to answer.</p><blockquote><p><em>Over the next five issues, <a href="https://www.linkedin.com/in/michellesandford/">Michelle</a>, <a href="https://sessionize.com/MichelleSandford/">Microsoft&#8217;s Developer Engagement Lead for Asia</a>, explores what it takes to build software with AI agents as participants in the process rather than as faster autocomplete. Each part builds on the last. You&#8217;ll come away with a practical six-stage AI-native loop; a Success Criteria Memo you can use to hold a team accountable; a clear understanding of what Michelle calls Rule Zero: &#8216;You own the code&#8217;; a set of mechanical escalation rules tested against a real example; and a Human-AI Responsibility Matrix for a repository you actually manage.</em></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.linkedin.com/in/michellesandford/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fMne!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1177510,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.linkedin.com/in/michellesandford/&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/208932522?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fMne!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd140-b651-4dc3-b59c-61fd767285da_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This first part begins with the question beneath all the others: not which AI tools your team has adopted, but whether the loop those tools operate within was ever redesigned for them. We&#8217;ll compare traditional, AI-assisted, and AI-native delivery side by side, help you place your own team honestly on that spectrum, and build the six-stage AI-native development loop: Intent, Spec, Context, Change, Verification, and Delivery and Learning.</p><div class="callout-block" data-callout="true"><h1 style="text-align: center;">In this edition</h1><p>&#10140; The story behind the series and what it reveals about designing for accountability rather than trust</p><p>&#10140; A precise definition of what &#8220;AI-native&#8221; actually means</p><p>&#10140; An honest comparison of traditional, AI-assisted, and AI-native delivery</p><p>&#10140; Where most teams actually sit today and how to tell</p><p>&#10140; The six-stage AI-native development loop, explained stage by stage </p><p>&#10140; A preview of Part 2, where we fix the same bug twice to see the difference in practice</p></div><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>What actually changed between 2024 and 2026</h1><p>It is easy to underestimate a shift while you are inside it, so the timeline is worth naming. In 2024, GitHub Copilot Chat reached general availability and GitHub previewed Copilot Workspace, the first end-to-end environment where a developer could go from issue to spec to implementation plan to pull request inside a single AI-driven session. Those two milestones signaled that AI was no longer just autocomplete; it was beginning to own workflow.</p><p>In early 2025, Visual Studio Code shipped agent mode for Copilot: the editor could now plan across files, run terminal commands, and iterate on test failures. Later that year, the Copilot coding agent moved out of preview, picking up issues on GitHub.com and producing draft pull requests from its own ephemeral environment. In August 2025, GitHub adopted AGENTS.md as a standardized cross-tool custom-instructions file. In September 2025, GitHub released Spec Kit as an open-source toolkit for spec-driven development. Across the same window, MCP matured into a stable interface that real teams build on, Microsoft Foundry and the open-source Microsoft Agent Framework both standardized on it, and the community-curated <code>github/awesome-copilot</code> repository became one of the most useful starting points for chat modes, prompt files, and AGENTS.md examples.</p><p>By mid-2026, the trajectory was unmistakable. At Microsoft Build 2026, GitHub unveiled the Copilot App, a desktop experience for orchestrating parallel multi-agent coding sessions across repositories. Microsoft shipped the Foundry Agent Framework 1.0 at general availability, giving teams a production-grade path to deploy, govern, and observe autonomous agents backed by over 11,000 models. The new Microsoft IQ platform introduced a unified orchestration layer (Work IQ, Foundry IQ, Web IQ) that lets agents share context across tools. And the Microsoft Execution Containers (MXC) SDK provided OS-level sandboxing for agentic workloads, answering the trust and security questions that had held back enterprise adoption.</p><p>Stack those on a calendar and the story is coherent: the tools to make changes matured, the standards to govern them matured, the interfaces to give models real-world context matured, and the community practices to share what works matured. That is what a paradigm shift looks like in software: not one launch, but a stack of compatible primitives that suddenly let a new kind of system be assembled.</p><p>If those new primitives changed what tools can do, the next question is what they change about the delivery process itself. The clearest way to see that shift is to compare how software moves from idea to production in traditional, AI-assisted, and AI-native environments.</p><h1>Traditional versus AI-assisted versus AI-native delivery</h1><p>Three eras, told honestly. </p><blockquote><p>Traditional delivery is a human doing every stage, with AI absent or peripheral. AI-assisted delivery is the human still driving, but with Copilot Chat, inline completions, and the occasional agent run helping at each step. The loop is unchanged, but each stage runs faster. AI-native delivery is when the loop itself is redesigned around AI participants: specs are first-class, context is engineered, agents run end-to-end stages, and traces are auditable artifacts.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ag_w!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ag_w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1071229,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/208932522?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ag_w!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faca63189-7aeb-40a5-b305-8128f16b2c28_1448x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><strong>Figure 1: How the delivery model shifts across traditional, AI-assisted, and AI-native approaches</strong></em></p><p>The diagnostic question, the next time someone claims to be &#8220;doing AI-native,&#8221; is this: <strong>which row of that table reflects how their loop has actually changed?</strong> &#8220;We turned on Copilot&#8221; is AI-assisted. It is a good and useful place to be, but not the same thing as AI-native.</p><h2>Where most teams sit today</h2><blockquote><p>Most teams that have adopted Copilot are firmly in the AI-assisted era. They use inline completions and Copilot Chat daily. A growing minority experiments with agent mode in the editor, fewer run the Copilot coding agent on real issues, and fewer still have invested in AGENTS.md, MCP servers, or Spec Kit workflows.</p></blockquote><p>This is not a criticism. The tools matured faster than most teams could absorb them, and absorbing them well requires a loop redesign while also shipping the work the business is asking for. The teams that do it well share a small set of behaviors: they pick one repository to redesign first, they treat AGENTS.md and a starter spec template as the first investment, and they measure the loop artifacts rather than the developer&#8217;s keystrokes.</p><p>The goal is not purity. The goal is a delivery loop that produces more durable artifacts and does so more consistently than the one you started with.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Anatomy of an AI-native development loop</h1><p>The structure of the AI-native loop remains familiar but has been rebalanced. It consists of six stages, with a new artifact introduced at the beginning and a new feedback path added at the end:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!MLki!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 424w, /__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 848w, /__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!MLki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:956102,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/208932522?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 424w, /__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 848w, /__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MLki!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20e6cd2c-6265-45db-94b6-53b74825e5f1_1586x992.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><strong>Figure 2: AI-native development loop with production feedback into the spec</strong></em></p><p>Let&#8217;s take a closer look at the stages:</p><ol><li><p><strong>Intent is captured deliberately.</strong> GitHub issue templates with <strong>&#8220;why&#8221;</strong> and <strong>&#8220;expected behavior&#8221;</strong> sections filter out work that should not have been started. The modest friction of those two headings prevents a remarkable number of wasted agent runs.</p></li><li><p><strong>Spec is an artifact, not a Slack thread.</strong> Spec Kit&#8217;s Specify, Plan, Tasks, Implement workflow is one shape; a five-line Markdown file in <code>specs/</code> is another. The temptation here is to over-engineer the spec. Resist it. A good spec is brief, opinionated, and reviewable in five minutes. The point is not to anticipate every edge case but to give the agent and the reviewer a shared contract.</p></li><li><p><strong>Context is engineered, not assumed.</strong> AGENTS.md tells agents how this codebase prefers to be edited, MCP servers expose your wiki, observability, and internal APIs as first-class tools, and a per-task context packet is assembled deliberately for non-trivial work. The Microsoft and GitHub ecosystem makes this approachable: AGENTS.md is plain Markdown, instruction files live under <code>.github/instructions/</code>, the <code>github/awesome-copilot</code> repository gives you ready-to-fork starting points, and writing a custom MCP server is a weekend project rather than a quarter-long migration.</p></li><li><p><strong>Change </strong>can be made by a human, by agent mode in the editor, or by the Copilot coding agent on GitHub.com. All are legitimate, but none is legitimate without a spec to anchor it. The most underrated skill in this stage is interrupting the agent well. When it is about to do something you do not want, stop it, refine the spec or AGENTS.md, and let it start over rather than letting it finish and cleaning up afterward.</p></li><li><p><strong>Verification </strong>moves from cheap to expensive layers: tests, evaluators, policy gates, and human review. Evaluators are often AI-powered and check the diff against the spec&#8217;s success criteria. Policy gates include tools such as GitHub Advanced Security and Copilot Autofix. Human review should focus on intent and integration rather than syntax. Designing the funnel so the cheap layers catch the cheap problems is what makes the expensive layers economical.</p></li><li><p><strong>Delivery and learning</strong> ship through your existing CI/CD (GitHub Actions, Azure Pipelines, or your environment of choice) and feed production signals back into the spec, not just into the backlog. The spec is the living artifact; the code is downstream of it.</p></li></ol><div class="callout-block" data-callout="true"><p>In one migration workshop, I asked a team to replay their previous sprint and list every decision that had changed the shape of the final result. Most of the list had nothing to do with code edits. The real pivots were a clarified success criterion, a context packet that included an old incident review, and one review comment that pushed the agent to <strong>rerun</strong> tests across time zones instead of only fixing the immediate failure. The team still used the same editor and <strong>the</strong> same model, but their loop had changed. Their takeaway was simple: when your artifacts become reviewable, your learning becomes reusable.</p></div><h1>A preview of Part 2</h1><p>That&#8217;s where Part 1 leaves you: with a six-stage loop mapped out and an honest sense of where your team currently sits on it. Part 2 (releasing next week) puts that loop to the test with a bug that&#8217;s almost embarrassingly ordinary: a reporting endpoint quietly drops results for any user outside UTC, because a date comparison assumes the server&#8217;s clock instead of the user&#8217;s. It&#8217;s the kind of date bug nearly every engineer has shipped at some point, and that&#8217;s exactly the point. Nothing about it is exotic enough to need a research paper.</p><p>We fix it twice, using the same repository and the same starting conditions. Run A is AI-assisted: an engineer asks Copilot Chat for the likely cause, accepts a plausible-looking fix, writes one test, and merges after review. Run B is AI-native: the engineer opens a spec first, defines success criteria before any code changes, and lets the loop from Part 1 do the rest.</p><p>Both runs get measured against the same yardstick: how fast each one reaches a passing test, how fast each one actually merges, how many edge cases the tests catch, and whether a future engineer could understand the reasoning without re-reading the diff. The gap between the two runs is where &#8220;AI-native&#8221; stops being a definition and starts being something you can measure. One of the fixes ships looking done and is actually still broken. You&#8217;ll see exactly which one, and exactly why nobody catches it until Part 2 walks through it.</p><p>By the end, you&#8217;ll have v1 of your own AI-Native Success Criteria Memo, the artifact that makes this kind of comparison repeatable on your own team, not just in this example.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it &#8212; just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Further reading</h1><p>&#128309; <strong><a href="https://dora.dev/dora-report-2025/">2025 DORA State of AI-assisted Software Development report</a></strong><br>The research behind the &#8220;amplifier&#8221; idea that opens this issue. Drawing on survey responses from nearly five thousand technology professionals, it argues that AI&#8217;s biggest impact comes not from the tools themselves but from the strength of the underlying delivery system. It&#8217;s also where the trust paradox that shapes Part 3 first shows up in the data.</p><p>&#128309; <strong><a href="https://www.anthropic.com/news/model-context-protocol">Introducing the Model Context Protocol</a></strong><br>Anthropic&#8217;s original announcement of MCP, the open standard for connecting AI systems to the tools and data they need. It explains why fragmented, one-off integrations were holding agents back, and why a shared protocol changes what &#8220;context&#8221; can mean at the Context stage of the loop.</p><p>&#128309; <strong><a href="https://agents.md/">AGENTS.md: the open standard</a></strong><br>The official home for AGENTS.md, the plain Markdown file that tells coding agents how a given codebase prefers to be edited. Useful for seeing the format in practice before writing your own, and for understanding why it&#8217;s become a shared convention across tools rather than a proprietary one.</p><p>&#128309; <strong><a href="https://github.com/github/spec-kit">GitHub Spec Kit</a></strong><br>GitHub&#8217;s open source toolkit for spec-driven development, built around the idea that specifications should lead code rather than follow it. Worth a look for teams wanting a more structured starting point than a five-line markdown file, with ready-made workflows for specifying, planning, and implementing.</p><p>&#128309; <strong><a href="https://docs.github.com/en/copilot/concepts/agents/coding-agent/about-coding-agent">About GitHub Copilot coding agent</a></strong><br>GitHub&#8217;s documentation on the coding agent that can research a repository, draft an implementation plan, and open a pull request with minimal supervision. A good primer on what &#8220;the agent leads at the Change stage&#8221; actually looks like in practice, and where human review still sits in that workflow.</p><p>&#128309; <strong><a href="https://github.blog/ai-and-ml/generative-ai/spec-driven-development-with-ai-get-started-with-a-new-open-source-toolkit/">Spec-driven development with AI: get started with a new open source toolkit</a></strong><br>GitHub&#8217;s own introduction to why specs need to become durable, living artifacts rather than throwaway planning documents. It&#8217;s a useful companion to this issue&#8217;s argument that the spec, not the code, should be the thing your team treats as source of truth.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Co-Editor-in-Chief <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Bill Was Always This Big. You Just Weren't Being Charged For It.]]></title><description><![CDATA[What the June billing changes at GitHub, Anthropic, and OpenAI actually reveal, and the governor pattern that keeps spend bounded.]]></description><link>https://packtbuildwithai.substack.com/p/the-bill-was-always-this-big-you</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-bill-was-always-this-big-you</guid><dc:creator><![CDATA[Charu Mitra Dubey]]></dc:creator><pubDate>Tue, 28 Jul 2026 19:09:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0f65489c-9ba1-40f4-bd43-5da4b93b2e30_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build With AI.</strong></p><p>On June 1, GitHub Copilot completed its shift to usage-based billing. Within days, developers were posting their new numbers. One went from $29 a month to $750. Another from $50 to $3,000. One company ran the math across 80 developers and found their new monthly AI tooling spend would equal a full-time engineer&#8217;s annual salary.</p><p>Nobody&#8217;s coding agent got worse overnight. The bill just stopped hiding what the usage was actually costing.</p><p>This is the story underneath every &#8220;AI coding agent&#8221; headline this year. Capability went up. Adoption went up. And the cost model quietly went from &#8220;flat fee, don&#8217;t think about it&#8221; to &#8220;metered, and you&#8217;d better think about it,&#8221; while most engineering orgs were still budgeting like it was 2024.</p><p>Thanks for reading Build With AI! Subscribe for free to receive new posts and support my work.</p><div class="callout-block" data-callout="true"><p><strong>In this edition:</strong></p><ul><li><p>Why coding agent costs stopped being predictable, and what actually drives the spend under the hood</p></li><li><p>The June 2026 billing shifts across GitHub Copilot, Anthropic, and OpenAI, and what they signal about where this is heading</p></li><li><p>A working pattern for keeping agent spend bounded without killing the productivity gains</p></li><li><p>What&#8217;s worth watching in AI this week</p></li></ul></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YBhh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YBhh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1134382,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/208871372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YBhh!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F138524ce-6df4-455b-be8e-2bca0c416ea5_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Mental Model - Why coding agent costs stopped behaving like software costs</h2><p>For twenty years, software cost planning was simple. You paid for seats or you paid for compute, and both scaled in ways you could forecast a quarter out. AI coding agents broke that model quietly, and most budgeting processes haven&#8217;t caught up.</p><p>The Copilot numbers make the mechanism visible. A $29 developer and a $750 developer aren&#8217;t using two different tools. They&#8217;re using the same tool at two different depths of agentic work, and depth is the thing nobody&#8217;s budget was built to measure.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Build With AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>A request isn&#8217;t a request anymore</h3><p>Traditional API billing scaled with request count. Agentic billing scales with what happens inside a single task: how many files the agent reads before it writes anything, how many times it retries after a failed test, how much context it&#8217;s carrying by turn twelve of a refactor. A one-line fix and a thirty-file agentic refactor can sit under the same &#8220;task&#8221; label and differ in cost by 10x or more.</p><p>This is why two developers on the same plan, same seat price, report wildly different bills. One is asking for autocomplete. The other is asking for a colleague who reads the whole codebase before acting. The pricing page never distinguished between the two. The invoice always did.</p><h3>The subsidy is ending, not the tool</h3><p>The flat-rate era wasn&#8217;t actually cheap, it was subsidized. Anthropic&#8217;s own June 15 billing change, which split subscription usage into a separate metered pool, was explicitly a response to agents extracting far more value from flat-rate logins than the flat rate was priced for. GitHub&#8217;s transition was the same story with the subsidy removed all at once instead of gradually. The tools didn&#8217;t get more expensive. The teams found out what they&#8217;d actually been costing to run all along.</p><h3>The throughline from harness engineering</h3><p>Last edition and the one before it were about getting agents to behave reliably: harnesses, verification loops, approval layers that can&#8217;t be fooled by the agent&#8217;s own framing. All of that work makes agents trustworthy enough to run unsupervised on real tasks. None of it makes that unsupervised running free. Reliability and cost are two separate bills, and teams that solved the first one are only now getting the invoice for the second.</p><p>The uncomfortable version of this: the more autonomous and reliable your agent setup gets, the more surface area it has to spend on. A harness that lets an agent retry, self-correct, and run longer before asking for human input is also a harness that racks up more tokens per task. You didn&#8217;t just buy reliability. You bought a bigger meter.</p><h2>The Build - A cost governor that stops runaway spend before the invoice does</h2><p>The mistake most teams make: they monitor cost after the fact, in a dashboard, a week after the spend already happened. That&#8217;s an autopsy, not a control. The fix is the same shape as the approval broker from two editions back: a layer that sits between the agent and its own execution loop, checking cost in real time, not describing it in a report later.</p><p><strong>The architecture before writing a line of code</strong></p><blockquote><p>Three pieces, deliberately not one:</p><ul><li><p><strong>The agent</strong> - runs its task, retries, reads files, calls tools. It has no concept of a budget. It&#8217;s not supposed to.</p></li><li><p><strong>The meter</strong> - tracks real token and dollar cost per task, updated on every model call, not estimated after the fact.</p></li><li><p><strong>The governor</strong> - checks the meter against a budget before every new step and can halt the loop mid-task, before the agent starts step forty of a task that should have stopped at ten.</p></li></ul></blockquote><p>The agent never decides when it&#8217;s spent enough. Something outside it does.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rwkC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 424w, /__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 848w, /__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rwkC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png" width="1456" height="931" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:931,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:958527,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/208871372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 424w, /__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 848w, /__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rwkC!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef5131e-b6e7-4189-a734-5be5b0fec0c6_1568x1003.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Layer 1 - The meter tracks real cost, not estimated cost</h3><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;88e4086c-6f82-4839-b6e6-bfe2024c6cc7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">class CostMeter {
  constructor(taskId, rates) {
    this.taskId = taskId;
    this.rates = rates; // { inputPerMTok, outputPerMTok }
    this.totalCostUSD = 0;
    this.callCount = 0;
  }

  recordCall(inputTokens, outputTokens) {
    const cost = (inputTokens / 1_000_000) * this.rates.inputPerMTok
               + (outputTokens / 1_000_000) * this.rates.outputPerMTok;
    this.totalCostUSD += cost;
    this.callCount += 1;
    return this.totalCostUSD;
  }
}</code></pre></div><p>This runs on every single model call, not on a timer. By the time a dashboard would show you the number, the governor has already seen it forty calls earlier.</p><h3>Layer 2 - The governor checks before the agent gets to act, not after</h3><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;33925d6f-2dd0-4c6f-85e3-af47ff780ac3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">function checkBudget(meter, budget) {
  if (meter.totalCostUSD &gt;= budget.hardCapUSD) {
    return { allowed: false, reason: `Task ${meter.taskId} hit hard cap: $${meter.totalCostUSD.toFixed(2)}` };
  }

  if (meter.totalCostUSD &gt;= budget.warnAtUSD &amp;&amp; !budget.warned) {
    budget.warned = true;
    console.log(`&#9888;&#65039;  Task ${meter.taskId} at $${meter.totalCostUSD.toFixed(2)}, approaching cap of $${budget.hardCapUSD}`);
  }

  if (meter.callCount &gt;= budget.maxCalls) {
    return { allowed: false, reason: `Task ${meter.taskId} exceeded ${budget.maxCalls} calls, likely stuck in a retry loop` };
  }

  return { allowed: true };
}</code></pre></div><p>Notice the second check: a call-count ceiling, independent of dollar cost. A cheap model looping forty times on a stuck retry is still a stuck agent, even if the dollar figure looks small. Cost and runaway behavior are two different failure modes, and a governor that only watches dollars misses the second one.</p><h3>Layer 3 - Wire it into the loop, not around it</h3><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;33d2037d-f2da-46c5-8fd8-e2be8d3a973d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">async function runTask(agent, task, budget) {
  const meter = new CostMeter(task.id, MODEL_RATES[task.model]);

  while (!task.complete) {
    const budgetCheck = checkBudget(meter, budget);
    if (!budgetCheck.allowed) {
      return { completed: false, reason: budgetCheck.reason, spentUSD: meter.totalCostUSD };
    }

    const step = await agent.nextStep(task);
    meter.recordCall(step.inputTokens, step.outputTokens);
    task = applyStep(task, step);
  }

  return { completed: true, spentUSD: meter.totalCostUSD };
}</code></pre></div><p>The check happens before every step executes, not after the loop finishes. A task that would have blown through $40 on an unattended weekend run instead stops at the cap, mid-loop, with a clear reason attached.</p><h3>Layer 4 - Route by task shape, not by habit</h3><p>The other half of cost control isn&#8217;t stopping runaway loops, it&#8217;s not starting expensive ones in the first place.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;a6dfc2a2-22e2-4aa2-a426-6b35db6ee7ef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">function selectModel(task) {
  if (task.type === 'autocomplete' || task.estimatedFiles &lt;= 1) {
    return 'budget-tier'; // cheap, fast, sufficient
  }
  if (task.type === 'refactor' &amp;&amp; task.estimatedFiles &gt; 10) {
    return 'frontier-tier'; // worth the cost, genuinely needs the reasoning
  }
  return 'mid-tier'; // default for everything else
}</code></pre></div><p>Most teams run every task through their most expensive model by default, because that&#8217;s what the IDE shipped with. The governor pattern is incomplete without this: cheap tasks routed to cheap models, and the expensive tier reserved for tasks that actually need it.</p><blockquote><p><strong>What this architecture gives you</strong></p><ul><li><p><strong>Real-time cost visibility</strong> - the meter knows the number before the invoice does</p></li><li><p><strong>A hard stop before the runaway loop finishes</strong> - not a report explaining what already happened</p></li><li><p><strong>Two failure modes caught, not one</strong> - dollar cost and call-count both checked, since a cheap loop can still be a stuck one</p></li><li><p><strong>Spend that matches task shape</strong> - the expensive tier gets reserved for work that actually needs it</p></li></ul></blockquote><p>Same principle as the approval broker two editions back: the thing that decides whether to keep going lives outside the agent, and the agent never gets a vote.</p><div><hr></div><h2>The Radar</h2><h3>01. GitHub Copilot&#8217;s usage-based billing finished rolling out, and the real numbers landed.</h3><p>Copilot completed its transition to usage-based billing on June 1, with GitHub offering promotional credits (Business plans get an extra $30/user/month, Enterprise $70) through August to soften the landing. DX has been tracking the actual customer impact: one developer went from $29 to $750 a month, another from $50 to $3,000, and an 80-developer company calculated their new monthly spend would equal a full-time engineer&#8217;s salary. The promotional credits expire in September, which means most teams haven&#8217;t seen their real baseline yet. If your governor pattern from this edition isn&#8217;t in place before then, Q4 is when you find out the hard way.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://getdx.com/blog/ai-coding-assistant-pricing/">getdx.com</a></p></div><h3>02. Anthropic split subscription usage into two billing pools, and named the reason out loud.</h3><p>Effective June 15, Anthropic moved a portion of Claude Code usage onto a metered credit pool billed at standard API rates, separate from normal subscription limits. The change followed an April restriction on third-party tools consuming flat-rate plans, and by industry estimates, agents had been extracting 12 to 175x effective subsidies on flat-rate logins before the split. This is the clearest public admission yet that the &#8220;unlimited&#8221; era of agentic coding tools was never actually unlimited, just unpriced.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://www.morphllm.com/ai-coding-costs">morphllm.com</a></p></div><h3>03. A wave of cheap or free frontier-tier models just reset the cost floor for coding agents.</h3><p>Moonshot AI&#8217;s Kimi K3, a 2.8-trillion-parameter open model, took the top spot on a major coding leaderboard days after DeepSeek&#8217;s roughly $0.44-per-million-output-token pricing had already reset industry expectations. Because K3&#8217;s weights are open, teams can self-host a top-tier coding model with no per-token cost at all. The catch nobody&#8217;s pricing page mentions: self-hosting isn&#8217;t free either, it just moves the cost from the model bill to your infrastructure bill. Worth benchmarking before assuming your default model is still the right one, and worth doing the total-cost math, not just the per-token math, before switching.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://www.buildfastwithai.com/blogs/ai-news-today-july-21-2026">buildfastwithai.com</a></p></div><div><hr></div><h2>Tools of the Week</h2><h3><a href="https://www.toriihq.com/articles/seven-tools-for-tracking-ai-token-usage-across-vendors">LiteLLM</a>: Gateway-level proxy that enforces the budget, not just logs it</h3><p>An open-source proxy that sits in front of every major provider, normalizing Claude, GPT, Gemini, Bedrock, and a dozen others behind one interface, with per-key budget limits enforced in the request path itself. This is the closest packaged version of the governor pattern from this edition&#8217;s Build section: the cap lives at the gateway, before the call executes, not in a dashboard afterward. The caveat: it&#8217;s a routing and enforcement layer, not a debugging tool. If you need to understand why a specific agent run got expensive, you&#8217;ll want this paired with a tracing tool, not used alone.</p><h3><a href="https://amnic.com/blogs/ai-cost-tracking-tools">Langfuse</a>: Trace-level cost, attached to the step that caused it</h3><p>Open-source observability that records cost inside the trace itself, so a single agentic run shows exactly which step, which retry, which tool call drove the spend, not just a total at the end. Useful for the diagnostic half of cost control: once the governor stops a runaway task, this is what tells you which of the forty steps actually broke the budget. The caveat: it&#8217;s observability, not enforcement. It&#8217;ll show you the expensive step happened. It won&#8217;t stop the next one from happening the same way.</p><h3><a href="https://amnic.com/blogs/ai-token-management-tools">Amnic</a>: Finance-grade attribution, for when &#8220;which team spent this&#8221; is the real question</h3><p>Attributes token spend to specific teams and cost centers and can enforce budgets before the invoice lands, aimed more at the finance-and-platform-team version of this problem than the individual-engineer version. Worth it if your organization&#8217;s actual pain point is &#8220;we don&#8217;t know which team&#8217;s agents are driving the bill,&#8221; less useful if you&#8217;re a smaller team where the governor pattern from this edition already covers your enforcement need directly.</p><p>One honest note across all three: none of these replace deciding what your budget should be. They enforce or reveal a number. The judgment call, what a task is worth spending, still sits with the team, same as it always has.</p><div><hr></div><p>GitHub&#8217;s $29-to-$750 developer and Anthropic&#8217;s June 15 billing split are the same story told twice: the flat-rate era wasn&#8217;t cheap, it was unpriced. Agentic work was always going to cost more than autocomplete, because it does more per task, more retries, more context, more tool calls stacked on top of each other. The invoice just took a while to catch up to the usage.</p><p>Two editions back, the fix for approval was a wall the agent couldn&#8217;t talk its way around. This edition, it&#8217;s the same shape one layer over: a governor that checks spend before the step runs, not a dashboard that explains it after. Reliability and cost turned out to be two separate bills, and the teams doing well right now are the ones who stopped treating the second one as somebody else&#8217;s problem to notice later.</p><p>The teams still budgeting like it&#8217;s 2024 will find out their real number in Q4, whether they built the governor or not. The only choice left is whether they find out from their own metering, or from the invoice.</p><p>See you next week.</p><p><strong>Charu Mitra Dubey</strong><br>Co-Editor-in-chief, Build with AI</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Build With AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Build with AI #20: The Agentic Engineering Playbook – Part 5]]></title><description><![CDATA[Why faster individual output doesn't add up to a faster organization, and what actually moves the needle instead.]]></description><link>https://packtbuildwithai.substack.com/p/build-with-ai-20-the-agentic-engineering</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/build-with-ai-20-the-agentic-engineering</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 23 Jul 2026 14:01:30 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9433240d-e4e5-4acb-a6be-92a00a0bbc11_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <a href="https://x.com/packtwebdevpro?s=21">X </a></p></div><p>In <em><a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-16-the-agentic-engineering-2f1">Part 4</a></em>, we made the case that judgment, not typing speed, is what separates a vibe coder from an agentic engineer. But judgment lives in one person&#8217;s head. The question we&#8217;ve been circling throughout this series is bigger than any one developer: <em>does any of this actually move the needle for the team?</em></p><p>That&#8217;s exactly what we&#8217;re closing out with in Part 5.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mclS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 424w, /__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 848w, /__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mclS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png" width="1456" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1158921,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/208019516?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 424w, /__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 848w, /__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mclS!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a75a2cf-afb6-4737-8185-93ac0ec79f71_1619x972.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You&#8217;ve probably heard the individual success stories by now. Someone on your team shipped a whole feature in an afternoon, or built a weekend MVP that used to take a month. Those stories are real. They&#8217;re also, on their own, a bit of a trap.</p><p>Along the way, we&#8217;ll unpack:</p><p>&#10145;&#65039; Why individual velocity gains so often vanish before they reach the organization<br>&#10145;&#65039; The Theory of Constraints, and why speeding up code generation can quietly make things worse<br>&#10145;&#65039; How DORA metrics give agentic teams a way to measure real, delivered value instead of raw output<br>&#10145;&#65039; The three practices that do most of the work in turning individual speed into team-level leverage</p><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>10x value, not 10x volume &#8211; Where the real gains come from</h1><p>As individual developers adopt AI assistants, we frequently hear reports of incredible velocity gains. <em>&#8220;Copilot made me 3x faster.&#8221;</em> <em>&#8220;I built a whole MVP in a weekend that would have normally taken a month.&#8221;</em> </p><p>The Factory era is not a forecast. Large engineering organizations already run in-house agent platforms against their production codebases at a scale no individual could match. That&#8217;s the organizational footprint of a practice that has moved well beyond one engineer typing faster.</p><p style="text-align: justify;">While individual velocity spikes are very real, they often create a localized illusion of productivity that fails to materialize at the organizational level. Individual speed is a false summit. You feel like you have reached the top because your own keyboard is faster, but cycle time and the team&#8217;s DORA numbers stay flat until the system around the agent changes. The climb that matters has barely started. </p><blockquote><p style="text-align: justify;">Why does the gain vanish? Because of the Theory of Constraints.</p></blockquote><p style="text-align: justify;">In any system, improving the throughput of a non-bottleneck step does not improve the throughput of the whole system; it just moves the bottleneck somewhere else. If you make code generation ten times faster, but your code review processes, security audits, QA testing cycles, and deployment pipelines remain manual, you haven&#8217;t delivered value to the user ten times faster. You have merely stockpiled a 10x backlog of unverified code waiting to pass through the human bottleneck downstream.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Pwdn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 424w, /__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 848w, /__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Pwdn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png" width="660" height="440" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:440,&quot;width&quot;:660,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 424w, /__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 848w, /__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Pwdn!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc77946d1-2442-478d-b485-7c38bfa2c2ea_660x440.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>Figure 1.3 - The Theory of Constraints applied to an agentic workflow, highlighting the verification bottleneck</em></p><div class="callout-block" data-callout="true"><p style="text-align: justify;">This is why the obsession with <em><span data-color="rgb(255, 153, 204)" style="color: rgb(255, 153, 204);">10x volume</span></em>, with raw output or hyper-productive individual vibe coders, misses the broader goal. The real gains of the AI transformation do not come from individuals typing faster. They come from <em><span data-color="rgb(255, 153, 204)" style="color: rgb(255, 153, 204);">10x value</span></em>: delivered, verified work that reaches the user. And that value comes from designing the team and the system around the agent, not from speeding up the individual at the keyboard.</p></div><p style="text-align: justify;">Agentic engineering looks at the entire <strong><span>software development lifecycle</span></strong> (<strong><span>SDLC</span></strong>) holistically. It measures success not by how many lines of code a single individual produces in an hour, but by how dependably the organization ships verified solutions. To track this, agentic teams align closely with the DORA metrics. Originally designed to measure human operational excellence, these metrics become the ultimate lifeline when scaling autonomous agents.</p><p style="text-align: justify;">DORA tracks four metrics. Two measure speed (how often you ship and how fast a change reaches production) and two measure stability (how often a change breaks and how fast you recover). Here is how each one governs agentic workflows:</p><ul><li><p><strong><span>Deployment frequency</span></strong>: In a traditional environment, increasing deployment frequency requires teams to work in smaller batches. When an AI agent generates code, it can easily write a thousand lines covering five different features simultaneously. If an engineer allows the agent to submit massive pull requests, deployment frequency plummets because humans cannot effectively review thousand-line diffs. The agentic team needs to constrain the AI to solve one specific problem per ticket, enforcing micro-deployments. This ensures that the velocity of generation does not overwhelm the pipeline.</p></li><li><p><strong><span>Lead time</span></strong>: This metric tracks the time from code commit to code running in production. When vibe coding, developers assume that fast code generation equals fast lead times. They are usually wrong. If code is generated quickly but lacks tests, it will fester in the review stage while security teams and QA testers manually verify its safety. An agentic organization drastically reduces lead time by shifting verification left. The environment makes the agent generate its own unit tests, or satisfy existing ones, before the code is even pushed to version control, making human review significantly faster.</p></li><li><p><strong><span>Change failure rate</span></strong>: How often do deployments cause a regression? AI agents are notorious for introducing subtle defects that are invisible on the surface. They will rename a variable that breaks a downstream database migration or quietly change the error-handling signature of a core utility. To keep the change failure rate low, the agentic team does not rely on manual human inspection. They build reliable integration environments and mandate that agents run local test suites and fix any failures prior to requesting a human review.</p></li><li><p><strong><span>Mean Time to Recovery (MTTR)</span></strong>: When a failure inevitably occurs in production, how fast can you recover? A standard vibe coder will scramble to read their own agent-generated code (which they only half-understand) to figure out what broke. An agentic engineer relies on the structure instead. Because the codebase is strictly modularized with well-defined APIs and thorough tests, the engineer can pinpoint the exact failing interface and instruct the AI to patch it safely. Delivering 10x value depends on this: an agentic organization builds one shared, automated setup across the team, not faster typing on one keyboard. </p></li></ul><p style="text-align: justify;">These four metrics are outcomes. You move them with concrete practices, three of which do most of the work in an agentic team:</p><ul><li><p><strong><span>Automated code review</span></strong>: If an agent generates the code, a secondary linting and security agent performs the first-pass review, instantly rejecting syntax violations or clear anti-patterns before a human ever looks at the pull request.</p></li><li><p><strong><span>Test-driven workflows</span></strong>: The verification layer has to scale automatically with the generation layer. This means AI-driven <strong><span>Test-Driven Development</span></strong> (<strong><span>TDD</span></strong>), where the team defines the behavioral boundaries and the agents fulfill them.</p></li><li><p><strong><span>Continuous deployment</span></strong>: Tight, automated deployment cycles let the team push small, AI-generated changes rapidly to staging, getting real-world feedback in minutes rather than weeks.</p></li></ul><p style="text-align: justify;">When we focus solely on individual AI tools, we end up optimizing for local maxima. True agentic software engineering is about refactoring the team&#8217;s operational model. It is about building a reliable, automated factory line where autonomous agents act as diligent workers, governed by the safety rails designed by human engineers. That factory line is the Factory era from the start of this series, and it is where team-level leverage is finally realized. None of the engineering principles behind it are new. </p><h1 style="text-align: justify;">Series wrap-up</h1><p style="text-align: justify;">And that&#8217;s the Playbook.</p><p>Five weeks ago, we started with a simple observation: AI hasn&#8217;t made engineering discipline less important, it&#8217;s made it more important. Everything since then has been us making good on that claim, one layer at a time.</p><p>We drew the line between vibe coding and agentic engineering. We named the 70% Problem and showed what it actually costs to cross it. We watched judgment replace syntax as the scarce resource, and the developer&#8217;s role shift from player to coach. And now, in Part 5, we&#8217;ve zoomed out one last time to show that none of it matters if it stays locked in one person&#8217;s workflow. The real payoff shows up at the team level, in deployment frequency, lead time, change failure rate, and how fast you recover when something breaks, not in how fast any one of us can type.</p><p>If there&#8217;s one thread running through all five parts, it&#8217;s this: </p><blockquote><p>AI didn&#8217;t hand us a shortcut around good engineering. It handed us a multiplier, and multipliers are only as good as what they&#8217;re multiplying. The teams that win here aren&#8217;t the ones with the fastest individual keyboard. They&#8217;re the ones who did the unglamorous work of building the tests, the interfaces, the review gates, and the feedback loops that let an agent run safely at scale.</p></blockquote><p>Thanks for building this with us over the last five weeks. If you&#8217;ve made changes to how your team works because of something in this series, or if you&#8217;ve hit walls we didn&#8217;t cover, we&#8217;d genuinely <a href="https://forms.cloud.microsoft/pages/responsepage.aspx?id=Dmauk5VIE0SnXsWk3kKcDmQlSQwENC9FmvMsFBr_EgdUOU41Q0xUNjVVUTNSQTNLS1owOUtCOE1YNy4u&amp;route=shorturl">love to hear about it</a>. That&#8217;s what the survey link has been for this whole time, and it still is &#128071; </p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out?  We&#8217;d love to hear about it&#8212;just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Further reading</h1><p>&#10145;&#65039; <strong><a href="https://www.theagilemindset.co.uk/2025/10/07/the-theory-of-constraints-in-software-development-finding-and-fixing-the-real-bottleneck/">The framework behind &#8220;fixing one bottleneck just moves it&#8221;</a></strong><br>This piece walks through Eliyahu Goldratt&#8217;s Theory of Constraints applied specifically to software teams, including the point that your system can only move as fast as its slowest part, no matter how fast the other parts get. A clear, practical grounding for the constraint logic we leaned on this issue.</p><p>&#10145;&#65039; <strong><a href="https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report">The report that put numbers on &#8220;AI amplifies what&#8217;s already there&#8221;</a></strong><br>Google Cloud&#8217;s 2025 DORA report surveyed nearly 5,000 technology professionals and found that AI doesn&#8217;t fix a struggling team, it makes their existing problems more visible. It&#8217;s the primary source behind the DORA metrics we used to frame agentic team performance.</p><p>&#10145;&#65039; <strong><a href="https://www.faros.ai/blog/ai-software-engineering">The research behind the paradox we opened with</a></strong><br>Faros AI&#8217;s telemetry study of over 10,000 developers is the empirical case for exactly the gap we described: individual output climbing while organizational delivery metrics stay flat. If you want the data behind &#8220;10x volume isn&#8217;t 10x value,&#8221; this is the source.</p><p>&#10145;&#65039; <strong><a href="https://codemanship.wordpress.com/2026/05/26/in-teams-individual-productivity-is-a-harmful-illusion/">Why individual productivity is the wrong thing to optimize for</a></strong><br>This piece argues that in a team setting, there&#8217;s really no such thing as individual productivity, only team outcomes, and that AI-assisted coding is pulling most organizations in exactly the wrong direction on this point. A sharp companion to the &#8220;false summit&#8221; framing we used.</p><p>&#10145;&#65039; <strong><a href="https://getdx.com/blog/dora-metrics/">Where the bottleneck actually goes once generation gets faster</a></strong><br>This guide names the mechanism directly: the effort saved on typing doesn&#8217;t disappear, it relocates downstream into review and verification. Useful if you want a deeper look at why lead time and change failure rate are the metrics that actually catch what deployment frequency alone would miss.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Co-Editor-in-Chief <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Human Was in the Loop. The Loop Was Lying to Them.]]></title><description><![CDATA[What GhostApproval and Friendly Fire actually broke, and the layer that closes the gap.]]></description><link>https://packtbuildwithai.substack.com/p/the-human-was-in-the-loop-the-loop</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-human-was-in-the-loop-the-loop</guid><dc:creator><![CDATA[Charu Mitra Dubey]]></dc:creator><pubDate>Tue, 21 Jul 2026 18:45:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lqPk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><p>On July 8, two security teams published disclosures on the same day, about the same class of bug, in different tools. Wiz called theirs GhostApproval. AI Now Institute called theirs Friendly Fire. Read together, they say the same thing: the approval box a human sees before an agent acts is not the same thing as the truth.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Build With AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In one of Wiz&#8217;s demos, a coding agent reads a file, correctly reasons out loud that it&#8217;s a symlink pointing at <code>~/.ssh/authorized_keys</code>, and writes to it anyway. The approval prompt shown to the developer names a harmless file. The human is still in the loop. The loop is just showing them the wrong thing.</p><p>This is worth sitting with if you shipped anything like last edition&#8217;s generator/evaluator loop. We spent that whole issue on why verification can&#8217;t be the same agent grading its own homework. This week&#8217;s story is the layer above that: putting a human in the approval seat doesn&#8217;t fix the problem either, if the agent is the one deciding what the human gets to see.</p><div class="callout-block" data-callout="true"><p>In this edition:</p><ul><li><p>Why &#8220;human in the loop&#8221; was never a security guarantee by itself</p></li><li><p>What GhostApproval and Friendly Fire actually exploited, and why patches narrow the gap but don&#8217;t close it</p></li><li><p>A working approval pattern that doesn&#8217;t rely on the agent&#8217;s own framing of its request</p></li><li><p>What&#8217;s worth watching in AI this week</p></li></ul></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lqPk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 424w, /__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 848w, /__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lqPk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png" width="4804" height="2227" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2227,&quot;width&quot;:4804,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:437873,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/207950278?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ea72889-182b-47f6-bc26-6e1b86ee89e4_4840x2560.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 424w, /__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 848w, /__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lqPk!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62dbb86e-3107-4b27-ac51-d97e426e56a1_4804x2227.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Mental Model - Why &#8220;human in the loop&#8221; was never a full security boundary</h2><p>For the last two years, the default answer to &#8220;what if the agent does something wrong&#8221; has been: put a human in front of it. Require approval before execution. Problem solved.</p><p>GhostApproval shows why that was always an incomplete answer. The approval step assumes three things line up: what the agent did, what the agent says it did, and what the human sees on the screen. Most systems only check the third one.</p><p>In Wiz&#8217;s disclosure, a malicious repo plants a file with an ordinary name, say <code>project_settings.json</code>, that&#8217;s actually a symlink to a sensitive system path. The agent reads it, reasons correctly about what it actually is, and proceeds anyway. The approval box the developer sees still just says <code>project_settings.json</code>. Nothing lied to the human. The system just never told them the thing that mattered.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/subscribe"><span>Subscribe now</span></a></p><h3>The human approves what they can see. The agent hands over what they can&#8217;t.</h3><p>This is the same failure mode as last edition&#8217;s self-grading evaluator, one level up. There, the problem was an agent judging its own output with no adversarial distance. Here, the problem is an agent that has full information about its own action, sitting between that information and the person whose job is to catch mistakes. In both cases, the party with the most context is the same party the check is supposed to be checking.</p><h3>Why patching this doesn&#8217;t close it</h3><p>The instinct is to fix the specific bug: block symlink writes, sanitize the specific file-naming trick, ship a CVE. That helps. It doesn&#8217;t touch the structural issue, which is that approval, as most teams have built it, asks the agent to summarize its own intent and asks the human to trust the summary. Any new technique that produces a truthful-looking summary of a harmful action defeats the same control.</p><p>The fix isn&#8217;t a smarter human, or a longer confirmation dialog. It&#8217;s moving authorization to a place the agent can&#8217;t shape: a layer that inspects the actual action being requested, independent of how the agent describes it, before it executes. That&#8217;s the fourth layer this edition&#8217;s Build section walks through.</p><h2>The Build - Approval that inspects the action, not the agent&#8217;s description of it</h2><p>The mistake GhostApproval exposed: the approval layer trusts the agent&#8217;s summary of what it&#8217;s about to do, instead of inspecting the actual call. Fix that, and the symlink trick (and the whole class of tricks like it) stops working, because the human, or the policy engine, is looking at the real target, not the agent&#8217;s name for it.</p><p><strong>The architecture before writing a line of code</strong></p><blockquote><p>Three pieces, deliberately not one:</p><ul><li><p><strong>The agent</strong> - proposes an action. It can be wrong, confused, or manipulated by something it read. That&#8217;s expected.</p></li><li><p><strong>The tool broker</strong> - sits between the agent and anything with real-world effect (filesystem, shell, network, APIs). It resolves the actual target of every call before anything runs.</p></li><li><p><strong>The policy check</strong> - evaluates the resolved action against rules. It never sees the agent&#8217;s reasoning or justification, only the resolved fact.</p></li></ul></blockquote><p>The agent never gets to hand the broker a description. It hands the broker a call, and the broker figures out for itself what that call actually touches.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bXW1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bXW1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1235923,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/207950278?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bXW1!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0669b436-034f-48d4-9032-b1aaab2ce09b_1024x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3>Layer 1 - The broker resolves, it doesn&#8217;t trust</h3><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;77d640ed-4b17-428b-871f-b6b7b936a8e6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">const fs = require('fs');
const path = require('path');

async function resolveTarget(requestedPath) {
  // Follow symlinks all the way down, don't stop at the first name
  const realPath = await fs.promises.realpath(requestedPath).catch(() =&gt; requestedPath);

  return {
    requestedPath,
    resolvedPath: realPath,
    isSymlink: requestedPath !== realPath,
    inSensitiveZone: SENSITIVE_PATHS.some(p =&gt; realPath.startsWith(p))
  };
}

const SENSITIVE_PATHS = [
  process.env.HOME + '/.ssh',
  process.env.HOME + '/.aws',
  '/etc',
];</code></pre></div><p>This is the whole fix for the symlink trick, in four lines. <code>requestedPath</code> is what the agent said. <code>resolvedPath</code> is what&#8217;s actually true. The broker checks the second one.</p><h3>Layer 2 - Policy sees the resolved fact, not the agent's framing</h3><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;7d4942c6-ab1b-46f8-9110-03c0c8515c96&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">function checkPolicy(resolved, action) {
  if (resolved.inSensitiveZone) {
    return { allowed: false, reason: `Resolved path ${resolved.resolvedPath} is in a protected zone` };
  }

  if (resolved.isSymlink &amp;&amp; action=/__u/packtbuildwithai.substack.com/== 'write') {
    return { allowed: false, reason: `Write target is a symlink (${resolved.requestedPath} -&gt; ${resolved.resolvedPath}), requires manual review` };
  }

  return { allowed: true };
}</code></pre></div><p>Notice what's missing here: no argument from the agent about why the write is safe. Policy doesn't take submissions. It takes facts and a rule.</p><h3>Layer 3 - The evidence package, for when a human does need to approve</h3><p>Some actions still need a human. The GhostApproval fix isn't "remove humans from approval," it's "show the human what the broker resolved, not what the agent claimed."</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;5341bf92-b11a-4a1b-9d0f-571215395d51&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">function buildEvidencePackage(resolved, action, agentClaim) {
  return {
    action,
    agentClaim,               // what the agent said it's doing (context, not authority)
    resolvedTarget: resolved.resolvedPath,
    requestedTarget: resolved.requestedPath,
    flagged: resolved.isSymlink || resolved.inSensitiveZone,
    timestamp: Date.now()
  };
}

async function requestApproval(resolved, action, agentClaim) {
  const evidence = buildEvidencePackage(resolved, action, agentClaim);

  if (evidence.flagged) {
    console.log(`&#9888;&#65039;  APPROVAL NEEDED - resolved target differs from or touches sensitive zone:`);
    console.log(`   Agent said: ${evidence.requestedTarget}`);
    console.log(`   Actually resolves to: ${evidence.resolvedTarget}`);
    // surface both paths to the human, not just the agent's claim
  }

  return evidence;
}</code></pre></div><p>The approval box now shows both lines: what the agent asked for, and what that actually resolves to. If those two differ, that mismatch is the thing the human is approving or rejecting, not the agent's plain-looking filename.</p><h3>Layer 4 - Wire it in front of the agent, not inside it</h3><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;a40a03a5-d517-4c27-997c-9105700edae8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">async function executeAction(agent, requestedPath, action, agentClaim) {
  const resolved = await resolveTarget(requestedPath);
  const policyResult = checkPolicy(resolved, action);

  if (!policyResult.allowed) {
    return { executed: false, reason: policyResult.reason };
  }

  if (resolved.isSymlink || resolved.inSensitiveZone) {
    const evidence = await requestApproval(resolved, action, agentClaim);
    const humanDecision = await getHumanApproval(evidence); // blocks on real human input
    if (!humanDecision.approved) {
      return { executed: false, reason: 'human rejected based on resolved target' };
    }
  }

  return performAction(resolved.resolvedPath, action);
}</code></pre></div><p>The broker sits outside the agent's process, ideally off-host or in a separate service the agent has no write access to. That's the part teams skip, because it's more infrastructure than a system prompt telling the agent to "be careful with sensitive files." A system prompt is a request. The broker is a wall.</p><blockquote><p><strong>What this architecture gives you</strong></p><ul><li><p><strong>Resolution before trust</strong> - the broker checks what a path actually is, not what it&#8217;s called</p></li><li><p><strong>Policy that can&#8217;t be argued with</strong> - no amount of agent-generated justification changes a resolved fact</p></li><li><p><strong>Approval that shows the mismatch</strong> - a human approving a symlinked write sees both paths, not one</p></li><li><p><strong>A wall the agent can&#8217;t write to</strong> - the broker lives outside the agent&#8217;s blast radius, same principle as last edition&#8217;s evaluator living outside the generator&#8217;s context</p></li></ul></blockquote><p>Same shape as last edition's split, one layer earlier in the pipeline. There, the fix was separating who does the work from who checks it. Here, the fix is separating what the agent claims from what's actually going to happen.</p><div><hr></div><h3><strong>The Radar</strong></h3><h4>01. Kimi K3 just reset the &#8220;which model goes in your loop&#8221; conversation.</h4><p>Moonshot AI&#8217;s Kimi K3 stunned the US technology industry over the weekend, setting off fresh debate about the China and US AI rivalry, after the 2.8-trillion-parameter open model took the top spot on a major coding leaderboard days earlier. It landed days after DeepSeek&#8217;s roughly $0.44 per million output tokens already set the price floor the industry gets measured against, with K3 weights arriving three days later meaning organizations can self-host a top-tier coding model with no per-token cost at all. If your loop&#8217;s <code>MODELS.GENERATOR</code> slot has been Claude or GPT by default, this is a legitimate week to actually benchmark rather than assume. The practical advice is to run real workloads against the current model and compare total cost including the infrastructure needed to self-host, which isn&#8217;t free even when the weights are. </p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://www.buildfastwithai.com/blogs/ai-news-today-july-21-2026">Build Fast with AI &#8212; AI News Today</a></p></div><h4>02. New research shows benchmark accuracy hides how unstable models actually are.</h4><p>A study called &#8220;The Illusion of Robustness&#8221; found that aggregate accuracy scores mask prediction-level instability, with models flipping their answers when semantically irrelevant context is changed, and related work showed that swapping drug names or adding distractors drops frontier accuracy on medical benchmarks by 10 to 20 points. Worth reading next to this edition&#8217;s approval work. A model that flips its answer on irrelevant noise is exactly the kind of agent you don&#8217;t want deciding, unchecked, what to summarize in an approval box. </p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://opendatascience.com/in-case-you-missed-it-last-week-in-ai-july-13-19-2026/">Open Data Science</a></p></div><h4>03. Adoption of agentic systems is running way ahead of any security sign-off.</h4><p>A new industry survey found 81% of teams are past the planning phase into active testing or production, yet only 14.4% have full security approval, and 88% of organizations confirmed or suspected a security incident this year. On identity specifically, only about 22% of teams treat agents as independent, identity-bearing entities, and roughly 46% still rely on shared API keys for agent-to-agent authentication. That&#8217;s the environment GhostApproval and Friendly Fire landed in: production systems, weak identity boundaries, and approval treated as a checkbox rather than infrastructure. </p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://www.gravitee.io/blog/state-of-ai-agent-security-2026-report-when-adoption-outpaces-control">Gravitee &#8212; State of AI Agent Security 2026</a></p></div><div><hr></div><h2>Tools of the Week</h2><h3><a href="https://arxiv.org/abs/2607.05518">aiAuthZ</a>: Off-host authorization, published the same day as GhostApproval</h3><p>Research from Sai Varun Kodathala, published July 8, the same day as the Wiz and AI Now Institute disclosures. The gateway verifies caller identity with a per-message signature and checks a role-based, argument-level policy the agent can neither read nor modify, then hash-chains every decision into an audit log. Across 15 evaluated models, refusal rates on adversarial tool-call attacks ranged from 100% down to 38%, with the gateway in place, residual attack success dropped to zero for all of them. The honest caveat: this is a research paper with published experiments, not a packaged product. If you&#8217;re building the broker pattern from this edition&#8217;s Build section, it&#8217;s the closest thing to a reference design for the policy layer specifically.</p><h3><a href="https://openid.net/authzen-at-identiverse-2026-authorization-in-the-agent-era/">AuthZEN (OpenID Foundation)</a>: An actual standard forming around agent tool authorization</h3><p>Discussed heavily at Identiverse 2026, AuthZEN&#8217;s profile for MCP tool authorization answers a question OAuth scopes were never built for: whether this specific agent, acting for this specific user, may call this specific tool with these specific arguments. It works by having a gateway consult a policy decision point before any tool call runs, structurally the same separation as this edition&#8217;s Build. The caveat: authentication for agents is mostly solved at this point, but authorization standards are still being actively drafted, not finalized. Worth watching rather than adopting wholesale this week.</p><h3><a href="https://workos.com/blog/developers-guide-to-ai-agent-authentication-and-authorization">WorkOS&#8217;s agent auth guide</a>: The clearest explanation of why multi-hop agent chains break normal auth</h3><p>Not a tool, a developer guide, but the most useful thing published this month for understanding why this problem gets harder as soon as one agent spawns another. Single-hop delegation, a human authorizes an agent, is manageable with existing OAuth token exchange. The moment Agent A spawns Agent B which calls Agent C, the identity chain most teams have doesn&#8217;t hold up. If your loop from last edition ever grows a sub-agent, read this before you assume your auth model scales with it.</p><div><hr></div><p>GhostApproval and Friendly Fire didn&#8217;t invent a new kind of bug. Symlink confusion and prompt-following agents are both old problems wearing new clothes. What&#8217;s new is the shape of the failure: a human did everything right, clicked approve on exactly the box they were shown, and the system still executed something they&#8217;d never have authorized if they&#8217;d seen it clearly.</p><p>Last edition&#8217;s thesis was that generation got nearly free and judgment didn&#8217;t. This week&#8217;s addendum: judgment isn&#8217;t worth much either, human or automated, if the party being judged also controls the evidence. A skeptical evaluator with no memory of the generator&#8217;s framing. A policy layer that can&#8217;t read the agent&#8217;s justification. An approval box built from a resolved fact instead of an agent&#8217;s claim. Same principle, three places you have to build it in.</p><p>The wall only works if the agent can&#8217;t write to it.</p><p>See you next week.</p><p><strong>Charu Mitra Dubey</strong><br>Co-Editor-in-chief, Build with AI</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Build With AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Build with AI #18: The Agentic Engineering Playbook – Part 4]]></title><description><![CDATA[Everyone talks about what AI can do. Nobody's asking what that leaves for the human engineer.]]></description><link>https://packtbuildwithai.substack.com/p/build-with-ai-16-the-agentic-engineering-2f1</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/build-with-ai-16-the-agentic-engineering-2f1</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 16 Jul 2026 14:02:13 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ab3025d9-fbfa-4c60-97d2-6d43c9d87558_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <a href="https://x.com/packtwebdevpro?s=21">X </a></p></div><div class="callout-block" data-callout="true"><h1 style="text-align: center;"><strong><a href="https://www.eventbrite.com/e/ai-agents-for-cloud-native-web-applications-tickets-1993668868253?aff=oddtdtcreator&amp;keep_tld=true">AI Agents for Cloud-Native Web Applications: F</a></strong><a href="https://www.eventbrite.com/e/ai-agents-for-cloud-native-web-applications-tickets-1993668868253?aff=oddtdtcreator&amp;keep_tld=true">ree workshop, 31 Jul</a>y</h1><p>AI is reshaping far more than code generation. Discover how AI agents can accelerate cloud-native development while strengthening monitoring, troubleshooting, and incident response in production. </p><p>Join Microsoft Cross Solutions Architect Albert Tannure for a practical workshop on Spec-Driven Development, AI-assisted operations, and building resilient applications from design to deployment. Whether you&#8217;re a developer, platform engineer, or cloud architect, you&#8217;ll leave with actionable ideas you can apply immediately.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.com/e/ai-agents-for-cloud-native-web-applications-tickets-1993668868253?aff=oddtdtcreator&amp;keep_tld=true&quot;,&quot;text&quot;:&quot;Reserve a Spot&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.com/e/ai-agents-for-cloud-native-web-applications-tickets-1993668868253?aff=oddtdtcreator&amp;keep_tld=true"><span>Reserve a Spot</span></a></p></div><p>In <em><a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-16-the-agentic-engineering">Part 3</a></em>, we walked through the 70% Problem, and we closed on a case study that probably felt a little too familiar: a payment service that looked done, then quietly started double-charging customers because nobody told the AI what &#8220;done&#8221; actually meant.</p><p>That case study raises a question that&#8217;s been sitting under this whole series: if the AI is doing more and more of the actual typing, what&#8217;s left for the engineer to do?</p><p>That&#8217;s exactly what we&#8217;re digging into in Part 4.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Deas!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 424w, /__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 848w, /__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Deas!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3682136-6950-48ba-80de-98a686db08d1_1614x975.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1102222,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/206965230?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 424w, /__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 848w, /__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Deas!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3682136-6950-48ba-80de-98a686db08d1_1614x975.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If your gut answer is &#8220;review the code more carefully,&#8221; you&#8217;re only half right, and this part is about the other half.</p><p>Along the way, we&#8217;ll unpack:</p><p>&#10145;&#65039; Why syntax knowledge, the thing we used to hire and promote for, has stopped being the scarce resource<br>&#10145;&#65039; The difference between agency and autonomy, and why confusing the two is where things go wrong<br>&#10145;&#65039; The four decisions that separate a vibe coder from an agentic engineer, and why none of them involve writing code<br>&#10145;&#65039; Why the developer&#8217;s role is shifting from player to coach, and what that actually looks like day to day</p><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>The knowledge shift &#8211; Why judgment is the scarce resource</h1><blockquote><p>If the mechanical typing of syntax is no longer the primary bottleneck of software development, what happens to the developer?</p></blockquote><p>Historically, we valued engineers who had deep knowledge of language quirks, standard libraries, and the small, detailed rules of syntax. The developer who could instantly recall the exact argument order for obscure bash commands, or who could write flawless, highly optimized C++ without consulting documentation, was prized as a strong senior engineer. That fluency was real and hard-won.</p><p>Today, the cost to generate a syntactically valid loop, a Dockerfile, or a Kubernetes manifest is functionally zero. The LLM is an infinite, instantaneous documentation parser and syntax generator.</p><p>&#128161; <em>Because syntax knowledge has been commoditized, the nature of value within engineering has shifted. The new scarce resource, and the skill that will define the senior engineers of the next decade, is engineering judgment.</em></p><h2>Engineering judgment and agency</h2><p><em>Judgment</em> is the most visible facet of a larger change. The deeper shift is one of agency: the engineer moves from producing code to owning the decisions. </p><p>When you no longer type the implementation, your work is to decide what gets built, how it is shaped, and how tightly or loosely the agent is allowed to run. A simple, well-specified task can run on a long leash, the agent working many steps before you check it. A risky change to a payment path runs on a short one, where you inspect every step. </p><p><strong>Agency</strong> is the capacity to act and effect outcomes. For the engineer, it rests on three things: competence (the skill to do the work), authority (the standing to make the call), and information (knowing enough to choose well), plus the willingness to act under risk. The same three describe the agent. Its competence is the tools it can call, its authority is the access and permissions it has been granted, and its information is its memory and context. One concept, both sides of the work, human and agentic.</p><p><strong>Autonomy</strong> is something else, and putting it next to agency surfaces the failure mode that matters. Autonomy is what the engineer actually does on their own, the independent action they take without someone stepping in. The dangerous case is high autonomy paired with low agency: acting independently without the competence, authority, or information to act well. It breaks the same way on both sides. An engineer who acts beyond their competence or authority ships the wrong thing, and an agent granted autonomy without the tools, permissions, or context to succeed does exactly that too.</p><p><strong>Judgment</strong> is the capability that sits above the codebase. It is the coach&#8217;s work of directing play, not the player&#8217;s work of executing it.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out?  We&#8217;d love to hear about it&#8212;just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h2>Engineering judgment in agentic workflow</h2><p>But what exactly is engineering judgment in an agentic workflow? </p><p>This is the vibe-versus-agentic distinction taken one level deeper. The vibe coder pushes a request and accepts whatever comes back; the agentic engineer makes a series of decisions the vibe coder never reaches. Four of those decisions matter most:</p><ul><li><p><strong>Deciding what not to build</strong>: The fastest way to take on technical debt is to build a feature you don&#8217;t need simply because the AI makes it easy. Judgment is understanding the product requirements deeply enough to reject unnecessary complexity.</p></li><li><p><strong>Architectural boundary setting</strong>: A junior developer can ask an AI to build a logging service. A senior developer with judgment knows exactly how to decouple that logging service from the core business logic via interfaces, ensuring the AI cannot accidentally tightly couple the domains.</p></li><li><p><strong>Designing the system that reviews the work</strong>: When an AI generates a 500-line pull request in thirty seconds, reading every line by hand does not scale. The engineer builds the system that reviews instead: tests, type checks, security scans, and a second agent that screens the diff against the project&#8217;s constraints. Those gates catch the mechanical faults but never the whole picture, so judgment is deciding what they must catch and which high-risk paths an agent can pass clean while still being wrong, where a human still has to look.</p></li><li><p><strong>Curating context</strong>: Giving an AI model the exact right amount of information. Too little context, and the AI hallucinates. Too much context, and the AI becomes distracted and loses the core instruction. Judgment is knowing exactly which schemas, ADRs, and file paths are required for a specific task.</p></li></ul><p>As this shift takes hold, our identity changes with it. We stop being players who execute every move on the field and become the coach who directs the play. The coach sets up the system the players run inside, calls the shots they cannot see from the field, and owns the result when the whistle blows. The agents do the running; the engineer decides where they run and why.</p><p>This kind of role elevation has happened before. When higher-level languages arrived, developers who hand-wrote assembly lamented the loss of their craft, insisting that real coding meant managing registers yourself. The new tools pushed developers up to write higher-level logic, and the work got more valuable, not less. Agents continue that arc, though the comparison only goes so far: a compiler is deterministic and gives the same output every time, while an agent predicts and varies, which is exactly why it needs the discipline this book builds. The joy of programming isn&#8217;t dead. It has moved up the stack, to deciding what gets built and how it is shaped.</p><p>This shift in how we work also changes where real gains in productivity come from. That&#8217;s exactly what we explore in <em>Part 5</em> of <em>The Agentic Engineering Playbook</em>.</p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out?  We&#8217;d love to hear about it&#8212;just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Further reading</h1><p>&#10145;&#65039; <strong><a href="https://cacm.acm.org/news/the-end-of-the-coder/">What replaces code as the scarce resource</a></strong><br>Communications of the ACM&#8217;s piece on the shrinking role of pure coding quotes Carnegie Mellon&#8217;s Bill Nichols making the point plainly: the value shifts from being a scarce source of code to being a scarce source of well-formed decisions. It also digs into what happens to junior hiring when entry-level implementation work gets automated away.</p><p>&#10145;&#65039; <strong><a href="https://arxiv.org/pdf/2605.12105">Where agency and autonomy actually diverge</a></strong><br>This paper treats agency and autonomy as two separate dials rather than one sliding scale: an agent can have broad tools but tight supervision, or narrow tools with none at all. It&#8217;s a more rigorous version of the distinction we drew between the engineer&#8217;s agency and the agent&#8217;s, and worth reading if you want the formal framing behind it.</p><p>&#10145;&#65039; <strong><a href="https://www.cio.com/article/4169591/ai-coding-tools-are-changing-output-faster-than-they-are-changing-judgment.html">Why review is where judgment actually lives now</a></strong><br>This piece argues that basic review can be automated, but judgment review can&#8217;t be casually delegated, which is exactly the line we drew between the mechanical gates and the human call in &#8220;designing the system that reviews the work.&#8221; A good read if you want the case for why senior engineers get more valuable, not less, as review volume climbs.</p><p>&#10145;&#65039; <strong><a href="https://matthopkins.com/technology/cognitive-debt-the-hidden-cost-of-letting-ai-write-your-code/">The cost of losing the mental model</a></strong><br>This essay names a specific failure mode worth watching for: the code still compiles, but the team&#8217;s shared understanding of why it works quietly erodes. It&#8217;s a sharper way to think about what&#8217;s actually at stake when curating context and owning decisions slips.</p><p>&#10145;&#65039; <strong><a href="https://towardsdatascience.com/code-is-cheap-engineering-judgement-is-now-the-scarce-resource/">Why saying no is getting harder, not easier</a></strong><br>This piece makes the case that cheaper building doesn&#8217;t remove the need for judgment; it just moves the bottleneck to deciding what&#8217;s worth building at all. Directly relevant to the first of the four decisions we covered, deciding what not to build.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Co-Editor-in-Chief <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Loop Engineering is a Judgment Problem, Not a Scheduling Problem]]></title><description><![CDATA[The fourth layer above your harness &#8212; and the one detail most teams building loops are skipping.]]></description><link>https://packtbuildwithai.substack.com/p/loop-engineering-is-a-judgment-problem</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/loop-engineering-is-a-judgment-problem</guid><dc:creator><![CDATA[BuildWithAI]]></dc:creator><pubDate>Tue, 14 Jul 2026 18:46:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xyV4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><p>On June 7, Peter Steinberger posted two sentences that pulled millions of views: stop typing prompts into your coding agent, build the system that prompts it for you. Days earlier, Boris Cherny, who runs Claude Code at Anthropic, said the same thing from the inside: &#8220;I don&#8217;t prompt Claude anymore. I have loops running that prompt Claude. My job is to write loops.&#8221; Addy Osmani named the shift loop engineering &#8212; a fourth layer above prompt, context, and harness engineering. By month&#8217;s end, Anthropic and OpenAI had each shipped official loop docs, days apart.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Here&#8217;s what&#8217;s getting lost in the excitement: a loop is only as good as its ability to catch its own mistakes, and an agent asked to grade its own output tends to just praise it. The teams getting real leverage aren&#8217;t the ones with the cleverest scheduling. They&#8217;re the ones who built a skeptical evaluator, separate from the agent doing the work.</p><div class="callout-block" data-callout="true"><p>In this edition:</p><ul><li><p>Why loop engineering is a real fourth layer, not a rebrand</p></li><li><p>Why the same loop produces different outcomes in different hands</p></li><li><p>A working loop you can build this week, with generator/evaluator separation built in</p></li><li><p>What&#8217;s worth watching in AI this week</p></li></ul></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xyV4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 424w, /__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 848w, /__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xyV4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:235585,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/207054313?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 424w, /__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 848w, /__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xyV4!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2629239-ca0e-4545-8141-5fefa98604cc_2720x1360.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>The Mental Model - Why loops aren&#8217;t just bigger prompts</strong></h2><p>For two years, the workflow was: you write a prompt, you read the output, you write the next prompt. The agent is a tool you hold, one turn at a time. That&#8217;s prompt engineering, and no matter how good the context or the harness around it gets, you&#8217;re still the one closing the loop.</p><p>Loop engineering removes you from that position. You&#8217;re not writing better instructions &#8212; <strong>you&#8217;re building the system that decides what to prompt, when, and whether the result is good enough to move on</strong>. The agent still does the work. The loop decides whether to trust it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Every week, one architecture worth stealing and one habit worth breaking.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><br><strong>No hype. No fluff. Just the mental models and code that hold up in production.</strong></p><p><strong>The five moves every loop makes</strong></p><p><strong>1. Discovery</strong> - how the loop finds work to do, instead of waiting for you to hand it a task.</p><p><strong>2. Handoff</strong> - how work gets assigned to an agent, with enough context to act without you in the room.</p><p><strong>3. Verification</strong> - checking the result before it&#8217;s accepted. This is the one teams get wrong (more below).</p><p><strong>4. Persistence</strong> - recording what&#8217;s done, so the next tick doesn&#8217;t repeat work or lose state.</p><p><strong>5. Scheduling</strong> - deciding when the loop runs again, and when it stops.</p><p><strong>The part that breaks most loops</strong></p><p>Verification is usually built as the agent checking its own work. That fails predictably: an agent grading its own output tends to praise it, the same way a student grading their own exam finds fewer mistakes than a second reader would. The fix isn&#8217;t a better prompt for self-review &#8212; it&#8217;s a structurally separate evaluator: a different call, ideally a different model, whose only job is to be skeptical.</p><p>This is also why the same loop can produce opposite results in different hands. Two teams can build an identical five-move structure and get completely different reliability, because the judgment sitting inside the verification step is where the real engineering is. Generation is nearly free now. Judgment is the scarce resource.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">If you&#8217;re enjoying this article, <strong><a href="https://www.agentengineering.co/">Agentic Engineering</a></strong> might be another newsletter community you might love.</p><p style="text-align: center;">It&#8217;s a small but growing newsletter community for people building with agentic AI. Through Packt&#8217;s expert network, it shares failure fixes, live project builds, implementation playbooks and informed perspectives on the conversations shaping where the field goes next.</p><p style="text-align: center;">Agentic Engineering exists because agent coverage has become rather good at announcing things and less good at explaining whether they work. They&#8217;re handing out a free digital copy of <em>AI Agents in Practice</em> by Valentina Alto, plus 40% off the <em>Building Intelligent AI Agents with GraphRAG</em> workshop to new subscribers. Think of it as a unsubtle preview of how this newsletter operates: useful things first, fanfare optional.</p><p style="text-align: center;"><strong><a href="https://www.agentengineering.co/">Subscribe to get the useful stuff</a></strong></p></div><h2><strong>The Build - A loop with generator/evaluator separation built in</strong></h2><p>Most loop implementations skip straight to scheduling and treat verification as an afterthought &#8212; a quick &#8220;does this look right?&#8221; prompt to the same agent that did the work. That&#8217;s the mistake from the Mental Model section, and it&#8217;s the first thing this architecture fixes.</p><p>Stack: Node.js and the Anthropic SDK. Same as always &#8212; the pattern holds regardless of framework.</p><p><strong>The architecture before writing a line of code</strong></p><p>Five moves, four layers:</p><ul><li><p><strong>Discovery</strong> - finds the next unit of work without waiting for you to hand it over</p></li><li><p><strong>Handoff + Generation</strong> - the agent that does the actual task</p></li><li><p><strong>Verification</strong> - a separate model call, structurally isolated from the generator</p></li><li><p><strong>Persistence + Scheduling</strong> - state tracking, retry limits, and halt conditions</p></li></ul><p>Build verification as its own function, calling its own model instance, before you build anything else. Everything else can be simple. This can&#8217;t be an afterthought.</p><p><strong>Layer 1 - Discovery</strong></p><p>Discovery isn&#8217;t always &#8220;check a queue.&#8221; In practice it&#8217;s usually pulling from several sources &#8212; a ticket tracker, a failed CI run, a scheduled check &#8212; and normalizing them into a common task shape before anything downstream touches them.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;1238ed3f-2e74-4887-bd2e-3c8e20a94d02&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">async function discoverWork(sources) {
  const tasks = [];

  for (const source of sources) {
    const items = await source.fetchPending();
    tasks.push(...items.map(item =&gt; ({
      id: item.id,
      description: item.description,
      context: item.context,
      status: 'pending',
      source: source.name
    })));
  }

  // Prioritize: oldest first, but let a source mark urgency
  return tasks.sort((a, b) =&gt; (b.urgent === a.urgent) ? 0 : b.urgent ? 1 : -1);
}</code></pre></div><p>The point of splitting this out: discovery is where most loops silently break in production. If a source goes down or starts returning malformed items, you want that failure visible here, not three layers deep inside a generation call.</p><p><strong>Layer 2 - Handoff + Generation</strong></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;f875c243-5079-42c3-831c-d4d60830ca82&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">import Anthropic from '@anthropic-ai/sdk';

const anthropic = new Anthropic();

const MODELS = {
  GENERATOR: 'claude-sonnet-5',
  EVALUATOR: 'claude-sonnet-5'   // can be a different model entirely &#8212; the separation matters more than which model
};

async function generate(task) {
  const response = await anthropic.messages.create({
    model: MODELS.GENERATOR,
    max_tokens: 2000,
    system: `You are completing a discrete task. Do the work. Do not
evaluate your own output &#8212; a separate reviewer will do that.`,
    messages: [
      { role: 'user', content: `Task: ${task.description}\nContext: ${task.context}` }
    ]
  });
  return response.content[0].text;
}</code></pre></div><p>Note the system prompt explicitly tells the generator not to self-assess. Without that line, models tend to append their own confidence claims to the output &#8212; &#8220;this looks correct&#8221; &#8212; and those claims leak into downstream logic if you&#8217;re not careful to strip them out.</p><p><strong>Layer 3 - Verification</strong></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;41a2302a-591a-4bc7-a933-6bd139bcabc1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">async function evaluate(task, output) {
  const response = await anthropic.messages.create({
    model: MODELS.EVALUATOR,
    max_tokens: 500,
    system: `You are a skeptical reviewer. You did not do this work.
Find problems. Do not default to approval. Respond with APPROVE or
REJECT, followed by one sentence of reasoning.`,
    messages: [
      { role: 'user', content: `Task: ${task.description}\n\nOutput to review:\n${output}` }
    ]
  });
  const verdict = response.content[0].text;
  return { approved: verdict.startsWith('APPROVE'), verdict };
}</code></pre></div><p>Two design choices matter more than they look. First, the evaluator gets the task description fresh &#8212; it&#8217;s not inheriting the generator&#8217;s conversation history, so it isn&#8217;t anchored by the generator&#8217;s framing of its own success. Second, &#8220;do not default to approval&#8221; is doing real work in that system prompt. Left unstated, evaluators drift toward approval over long runs, the same self-praise bias the Mental Model section covered, just one hop removed.</p><p><strong>Layer 4 - Persistence + Scheduling</strong></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;9cd30b84-f8e4-4046-95af-42bf403ef2d9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">class LoopState {
  constructor(maxAttempts = 3, dollarCap = 5.00) {
    this.completed = [];
    this.attempts = {};
    this.maxAttempts = maxAttempts;
    this.dollarCap = dollarCap;
    this.spend = 0;
  }

  shouldRetry(taskId) {
    return (this.attempts[taskId] || 0) &lt; this.maxAttempts &amp;&amp; this.spend &lt; this.dollarCap;
  }

  recordAttempt(taskId, costEstimate) {
    this.attempts[taskId] = (this.attempts[taskId] || 0) + 1;
    this.spend += costEstimate;
  }

  recordCompletion(taskId, output) {
    this.completed.push({ taskId, output, at: Date.now() });
  }

  status() {
    return { completed: this.completed.length, spend: this.spend };
  }
}

async function runLoop(sources, state) {
  const queue = await discoverWork(sources);

  for (const task of queue) {
    if (!state.shouldRetry(task.id)) {
      console.log(`Task ${task.id} exceeded budget or attempts. Escalating to human.`);
      task.status = 'needs_human';
      continue;
    }

    state.recordAttempt(task.id, 0.03); // replace with a real per-call estimate
    const output = await generate(task);
    const review = await evaluate(task, output);

    if (review.approved) {
      task.status = 'done';
      state.recordCompletion(task.id, output);
    } else {
      console.log(`Rejected: ${review.verdict}`);
      // stays pending, retried next tick
    }
  }

  return state.status();
}

// Run on a schedule &#8212; cron, a queue worker, whatever fits your infra
// setInterval(() =&gt; runLoop(sources, state), 5 * 60 * 1000);</code></pre></div><p><strong>What this architecture gives you</strong></p><ul><li><p><strong>Discovery</strong> that fails loudly instead of silently starving the loop</p></li><li><p><strong>Generation</strong> that isn&#8217;t quietly grading its own homework</p></li><li><p><strong>Verification</strong> that&#8217;s structurally incapable of the self-praise bias, because it&#8217;s a different call with no memory of being the one who did the work</p></li><li><p><strong>Persistence and scheduling</strong> that stop a stuck task from becoming an unbounded bill &#8212; the dollar cap matters as much as the attempt cap</p></li></ul><p>This is close to a hundred lines for something that took most teams a pile of unmaintained bash a year ago. The generator/evaluator split adds one extra API call per task. That&#8217;s the cost of the whole idea &#8212; and it&#8217;s cheap relative to what an unchecked loop can quietly approve.</p><div><hr></div><h2><strong>The Radar</strong></h2><h3>01. ITBench-AA just proved that longer agent trajectories don&#8217;t mean better judgment &#8212; they often mean worse.</h3><p>IBM Research and Artificial Analysis launched ITBench-AA in late May, the first benchmark testing agents on real enterprise IT work: 59 Kubernetes incident-diagnosis tasks built from live SRE scenarios. Every frontier model scored below 50%, with Claude Opus 4.7 leading at 47% and GPT-5.5 close behind at 46%. The detail worth sitting with: models that took more investigation turns didn&#8217;t score higher &#8212; Gemini 3.1 Pro Preview averaged 83 turns and still landed below terser models. Open-weight models held their own on cost, with Gemma 4 31B scoring competitively at a fraction of proprietary pricing per task. That&#8217;s the loop engineering thesis in benchmark form &#8212; a system that just keeps iterating isn&#8217;t the same as a system that knows when it&#8217;s actually done, and it&#8217;s exactly why verification has to be a separate, skeptical step rather than &#8220;let it keep trying.&#8221; </p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://huggingface.co/blog/ibm-research/itbench-aa">Hugging Face &#8212; ITBench-AA</a></p></div><h3>02. Sonnet 5&#8217;s breaking changes are still catching teams mid-migration.</h3><p>Claude Sonnet 5 became the default across Claude plans on June 30, and Anthropic priced it aggressively at $2/$10 per million tokens through August 31 &#8212; but three changes are quietly breaking agent loops in production. Any code that sets temperature or top_p will now error out, one of the changes most likely to silently break agent loops. The new tokenizer also produces 1.0 to 1.35 times more tokens for the same text, which means budget enforcement calibrated on the old tokenizer will underestimate real spend. If you built a token-budgeting layer using last edition&#8217;s architecture, this is the week to re-check your assumptions against it.</p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://www.buildfastwithai.com/blogs/ai-news-today-july-1-2026">Build Fast with AI &#8212; AI News Today</a></p></div><h3>03. China&#8217;s AI companion law takes effect tomorrow, and Doubao and Qwen aren&#8217;t waiting to find out if they comply.</h3><p>China&#8217;s Interim Measures for AI Anthropomorphic Interactive Services takes effect July 15. ByteDance&#8217;s Doubao and Alibaba&#8217;s Qwen are both shutting down their humanlike and user-created agent features ahead of the deadline rather than retrofit compliance, a regulation jointly issued by China&#8217;s Cyberspace Administration and four partner agencies requiring anti-addiction systems, mandatory usage notifications, and instant-exit mechanisms for services that simulate human personality. Worth watching if you build anything with persistent agent personas &#8212; it&#8217;s an early signal of where personality-simulating AI regulation is headed globally, not just in China.</p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://www.buildfastwithai.com/blogs/ai-news-today-july-6-2026">Build Fast with AI &#8212; AI News Today</a></p></div><div><hr></div><h2>AI Tools of the Week</h2><h3><a href="https://claude.com/claude-code">Claude Code </a><code>/loop</code>: The native scheduling primitive</h3><p>The fastest way to try loop engineering without building infrastructure. A single slash command &#8212; <code>/loop babysit all my PRs. Auto-fix build issues, and when comments come in, use a worktree agent to fix them</code> &#8212; turns Claude Code into a scheduled, self-prompting system. It ladders from turn-based to goal-based to time-based to fully proactive loops, so you can start conservative and expand autonomy as you trust the verification step. No separate install, no new billing &#8212; it runs on whatever Claude Code plan you already have. The tradeoff: it&#8217;s Claude Code specific, so if your team runs a mixed-model stack, you&#8217;ll want a tool-agnostic layer alongside it.</p><h3><a href="https://hermes-agent.org">Hermes Agent</a>: Self-hosted, model-agnostic, learns as it runs</h3><p>Open source, MIT-licensed, and built by Nous Research &#8212; the closest thing to a persistent loop that lives outside your IDE entirely. Hermes runs on a $5 VPS or your own server, remembers projects and conventions across sessions via a layered memory system, and writes reusable skill documents after it solves something non-trivial, so it gets better at recurring work over time. Built-in cron scheduling means daily reports, nightly audits, and morning briefings run unattended, delivered to Telegram, Slack, Discord, or CLI. Model-agnostic by design &#8212; point it at Claude, GPT, or a local model. Free forever, no telemetry, all data stays on your machine. The honest caveat: self-improving skills only help when new work resembles old work; it&#8217;s not a shortcut around building a real verification step.</p><h3><a href="https://www.requesty.ai">Requesty</a>: Cost tracking and failover, built for loops specifically</h3><p>If last edition&#8217;s token budgeting layer felt like something you&#8217;d rather not maintain by hand, Requesty routes your loop&#8217;s calls through a layer that adds cost tracking and automatic failover from day one. Particularly relevant given this edition&#8217;s Build section: a scheduled loop with a verifier running after every turn can burn tokens fast and unevenly depending on cadence and how many sub-agents spawn, so having routing and spend visibility in front of the loop &#8212; not bolted on after a surprising invoice &#8212; is the more defensible starting position.</p><h3><a href="https://loopengineering.app/">loopengineering.app</a>: A minimal reference implementation of the five-move pattern</h3><p>Less a product, more a worked example &#8212; a small, readable project demonstrating discovery, handoff, verification, persistence, and scheduling as five distinct, inspectable steps, with the verification step explicitly built as a separate reviewer that rejects shortcuts like deleted tests or weakened checks. Worth bookmarking less for what it automates and more as a sanity check: if your own loop&#8217;s verification step can&#8217;t point to something this concrete, it&#8217;s probably not separate enough from the generator yet.</p><div><hr></div><p>If you made it this far, you&#8217;ve got a loop with a real verifier in it &#8212; which puts you ahead of most teams currently pointing an agent at itself and calling that review.</p><p>Steinberger and Cherny didn&#8217;t discover something new in June. They named the moment the leverage moved: from typing the prompt to designing the system that decides what gets prompted, when, and whether the result is good enough to trust. That&#8217;s the whole shift. Generation got nearly free. Judgment didn&#8217;t.</p><p>The five moves &#8212; discovery, handoff, verification, persistence, scheduling &#8212; aren&#8217;t hard to list. The hard part, and the part ITBench-AA just put a number on, is that more turns and more autonomy don&#8217;t buy you more judgment. A separate, skeptical evaluator does. Build that piece first. Let everything else be simple.</p><p>Four layers. One principle: the loop is only as trustworthy as the thing checking its work, and that thing can&#8217;t be the same agent that did it.</p><p>See you next week.</p><p><strong>Charu Mitra Dubey</strong><br>Co-Editor-in-chief, Build with AI</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Build with AI #16: The Agentic Engineering Playbook – Part 3]]></title><description><![CDATA[Why AI gets you 70% of the way there instantly, and what it takes to close the rest.]]></description><link>https://packtbuildwithai.substack.com/p/build-with-ai-16-the-agentic-engineering</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/build-with-ai-16-the-agentic-engineering</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 09 Jul 2026 14:01:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0e8cbed1-a32c-4217-8cdb-0e75e4b71fc3_2912x1942.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <a href="https://x.com/packtwebdevpro?s=21">X </a></p></div><p>In <a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-14-the-agentic-engineering?r=55ncj4">Part 2</a>, we drew a hard line between vibe coding and agentic engineering, and we closed on a question a lot of you already know in your gut: <em><strong>if AI can get you most of the way there in seconds, why does the last stretch feel like it takes forever?</strong></em></p><p>That&#8217;s exactly what we&#8217;re digging into in Part 3.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9ILX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 424w, /__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 848w, /__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9ILX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png" width="1456" height="969" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:969,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1221278,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/206010920?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 424w, /__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 848w, /__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9ILX!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb96f07f-05e0-46e1-b4a4-0d62c3aa926a_1537x1023.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That gap has a name, and once you see it, you can&#8217;t unsee it. We call it the 70% Problem, and it explains almost every AI horror story you&#8217;ve heard: the demo that looked flawless, the production incident that followed a week later.</p><p>Along the way, we&#8217;ll unpack:</p><p>&#10145;&#65039; Why the first 70% of a feature comes almost for free</p><p>&#10145;&#65039; Why the remaining 30% is where AI starts to struggle, and why that happens</p><p>&#10145;&#65039; A case study of what goes wrong when a team hits the wall without a plan for it</p><p>&#10145;&#65039; How agentic engineering turns the wall into a solvable problem instead of a dead end</p><p>&#10145;&#65039; Five reads to go deeper, from Stripe&#8217;s own idempotency writeup to real data on how AI-generated code stacks up against human-written code</p><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>The 70% problem &#8211; Where AI stops and engineering begins</h1><p>Every developer using modern AI tools has experienced a specific euphoria. You ask a complex question, and within seconds, the IDE generates a near-perfect scaffold of the feature you wanted. It feels like magic. The AI has instantly transported you 70% of the way to the finish line.</p><p><span>This creates a dangerous illusion of linearity. The human brain assumes that because the first 70% took three seconds, the remaining 30% will take only a few moments longer. This is known in the industry as the 70% Problem.</span></p><p>The reality implies a drastically different shape. The curve of AI productivity starts vertical and then immediately flattens into a punishing asymptote, approaching a limit it never quite reaches.</p><p><span>The first 70% of a feature consists of highly generic, frequently modeled patterns. Setting up an Express.js server, creating a React component hierarchy, or writing standard CRUD SQL queries are foundational tasks perfectly represented in the massive training data corpora of an LLM. The AI doesn&#8217;t need to reason to build this; it simply repeats the deeply grooved statistical patterns it has seen millions of times before.</span></p><p>But the final 30% is where your application deviates from the generic average. The final 30% involves:</p><ul><li><p><strong>Intricate business logic unique to your specific organization</strong>. These are rules shaped by your domain, not something the model has broadly seen before.</p></li><li><p><strong>Complex race conditions in asynchronous states</strong>. These occur when timing issues between concurrent processes create subtle and hard-to-reproduce bugs.</p></li><li><p><strong>Systemic security layers and authorization flows</strong>. These require precise enforcement of permissions and policies across multiple system boundaries.</p></li><li><p><strong>Deep integration with legacy, undocumented internal APIs</strong>. These systems often behave inconsistently and lack clear specifications, making them difficult to reason about.</p></li><li><p><strong>Memory leak debugging and highly specific performance tuning</strong>. These problems require careful observation of runtime behavior and fine-grained optimization.</p></li></ul><p>These edge cases are not well represented in training data. At this point, the AI cannot rely on pattern matching. It has to reason across very specific context. This is where LLMs begin to struggle.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!33JE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 424w, /__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 848w, /__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 1272w, /__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!33JE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png" width="660" height="440" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:440,&quot;width&quot;:660,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 424w, /__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 848w, /__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 1272w, /__u/substackcdn.com/image/fetch/$s_!33JE!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9763d8-b1f8-4914-9a01-2fc3592451a5_660x440.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>Figure 1 - The 70% problem, contrasting rapid early progress with a flattening completion curve</em></p><p><span>When a developer operating under the vibe coding paradigm hits the 70% wall, their instinct is to prompt harder. They tweak the wording, paste in more logs, and yell at the chat interface. Often, this makes the problem worse. The AI begins hallucinating fixes, tearing out components that worked, and introducing regressions elsewhere. The developer spends hours caught in an agonizing loop of prompt-and-pray, ultimately rewriting the code manually anyway.</span></p><p><span>Agentic engineering recognizes the 70% wall and structurally prepares for it.</span></p><p>To break through the final 30%, you cannot rely on generic instructions. You rely on deterministic engineering instead. This means transitioning the agent from writing entirely new code to passing highly specific verification tests. When a system is heavily typed, continuously integrated, and tested at both the unit and boundary levels, the human engineer can isolate the failing 30% and constrain the agent explicitly within those boundaries.</p><p>The prompt shifts from: </p><p><code>&#10140; Build this feature.</code> </p><p>to </p><p><code>&#10140; Here is the failing integration test, here is the exact interface specification, and here is a slice of the state machine. Fix the execution to satisfy the test</code><span>.</span></p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out?  We&#8217;d love to hear about it&#8212;just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h2><span>Case study &#8211; The 70% wall in practice</span></h2><p><span>To fully illustrate the danger of the 70% problem, we can examine a common enterprise scenario. Consider a mid-level engineer tasked with migrating an existing monolithic payment processor into a new serverless microservice. The engineer opens an AI coding assistant and provides a highly descriptive prompt: </span></p><p><code>Create a Node.js AWS Lambda function that accepts a JSON payload, parses the customer ID, queries our Redis cache for session validity, and then forwards the payload to the Stripe API. Make sure to handle rate limits and generic HTTP errors.</code></p><p><span>Within fifteen seconds, the agent produces a two-hundred-line file. It imports the correct AWS SDKs, sets up a beautiful try-catch block, uses modern async/await syntax, and correctly formats the Stripe payload. The developer pastes the code, hits save, and it deploys successfully. They test it with a generic payload, and it returns a 200 OK status.</span></p><p><span>The developer has just experienced the 70% magic. They feel like a 10x multiplier engineer. They have completed a day&#8217;s worth of scaffolding in fifteen seconds.</span></p><p><span>However, the final 30% breaks their spirit. In production, this microservice immediately begins dropping transactions under heavy load. The developer opens the chat interface again and pastes the CloudWatch error logs: </span><code>Redis connection timeout and Stripe Idempotency Key Missing</code><span>.</span></p><p><span>The AI agent, eager to please, suggests adding a massive retry loop and randomly generating an idempotency key on every single retry. The developer, operating purely on vibes rather than engineering judgment, copies the new code. Now, the service doesn&#8217;t just crash; it accidentally double-charges three premium customers because the randomly generated idempotency key circumvents Stripe&#8217;s exact protection mechanisms against duplicate network requests.</span></p><p><span>The developer has hit the wall. The AI model successfully predicted the statistical shape of a payment function, but it failed to synthesize the bespoke business requirement of deterministic idempotency across distributed network boundaries.</span></p><h3><span>How does agentic engineering solve this? </span></h3><p><span>Agentic engineering solves this by acknowledging the wall before the project even begins.</span></p><p>If this engineer had applied Dave Farley&#8217;s foundations, the workflow would look entirely different. Rather than asking the AI to build the Stripe function, the engineer would have first stated the intent precisely: that the service must decline a payment whose request ID matches one it has already seen inside the idempotency window. Note the move. The engineer expresses the requirement specifically, and the environment they set up turns it into a failing integration test the implementation has to satisfy. </p><p>The agent then runs against that constraint. It attempts the loose retry block, but the local test suite immediately fails and bounces the code back automatically. The agent, enclosed in this tight feedback loop, realizes its generic logic failed the structural constraint. It iterates, removing the random key generation and binding the idempotency check directly to the incoming request header.</p><p><span>By the time the human engineer looks at the code, the final 30% has been conquered not by a smarter prompt, but by a deterministic engineering boundary. The magic of the 70% leap is very real, but crossing the finish line requires acknowledging where the model stops being a guru and starts being a mechanism that needs strict logical framing.</span></p><p>Once we accept this limit, the focus shifts from generating code to exercising judgment over how that code is shaped and constrained.</p><div class="callout-block" data-callout="true"><h2><strong>Conclusion</strong></h2><p>That&#8217;s the real handoff this series has been building toward. The first two parts made the case for discipline. This one showed you where the model&#8217;s limits actually sit, and how a tight verification loop turns the final 30% from a wall into just another problem you can engineer your way through.</p><p>But conquering the wall raises the next question: if the model handles more and more of the typing, what&#8217;s actually left for the engineer to do?</p><p>In <em>Part 4</em> of <em>The Agentic Engineering Playbook</em>, we&#8217;ll dig into that shift directly, from typing code to owning decisions. We&#8217;ll look at why judgment, not typing speed, is becoming the scarce resource, what separates agency from autonomy, and why the engineer&#8217;s job increasingly looks less like playing the game and more like coaching it.</p></div><div class="pullquote"><p><em><strong><span>&#128221; Spotted something worth calling out?  We&#8217;d love to hear about it&#8212;just drop us a quick note through </span><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a><span>. We&#8217;re all ears!</span></strong></em></p></div><h1>Further reading</h1><p>&#10145;&#65039; <strong><a href="/__u/addyo.substack.com/p/the-70-problem-hard-truths-about">The piece that coined the term</a></strong><br>Addy Osmani&#8217;s original essay laid out the 70% problem in the context of AI app builders like Lovable and Bolt: fast, confident progress on the generic part of the work, then a wall around the last stretch that actually makes software production-ready. If you want the framing in its original form before reading anyone&#8217;s commentary on it, start here.</p><p>&#10145;&#65039; <strong><a href="/__u/addyo.substack.com/p/the-80-problem-in-agentic-coding">What changed a year later</a></strong><br>Osmani revisited the idea in a follow-up piece, noting that the completion percentage has crept up as tools improved, but argued the shape of the problem hasn&#8217;t changed, just where the wall sits. He also points to Armin Ronacher&#8217;s survey of thousands of developers showing how much manual typing has already disappeared from daily work. A good corrective if you assume better models eventually erase the wall entirely.</p><p>&#10145;&#65039; <strong><a href="https://dev.to/themachinepulse/ai-writes-70-of-your-code-now-heres-what-actually-breaks-34n3">What the data says about AI-generated code</a></strong><br>CodeRabbit&#8217;s comparison of AI-authored and human-authored pull requests found meaningfully more bugs and performance issues in the AI-authored set, concentrated in exactly the kind of subtle, load-dependent problems that don&#8217;t show up in a quick manual test. It&#8217;s a useful counterweight to the &#8220;it deployed and returned 200 OK&#8221; feeling from our case study.</p><p>&#10145;&#65039; <strong><a href="https://blog.reccehq.com/before-you-let-agents-touch-your-codebase-build-these-gates">A verification gate in practice</a></strong><br>This account from an engineer who rebuilt production AWS infrastructure with AI agents describes the actual gates his team uses, from pre-commit hooks to a pre-push suite to a combined agent-and-human PR review. It&#8217;s a concrete, unglamorous look at what &#8220;deterministic engineering boundary&#8221; means day to day, including what broke before the gates existed.</p><p>&#10145;&#65039; <strong><a href="https://www.loadsys.com/blog/ai-generated-code-production-ready/">How exposed AI-generated code really is</a></strong><br>Veracode&#8217;s 2025 GenAI Code Security Report tested a large set of language models across many coding tasks and found a substantial share of the generated code contained security vulnerabilities, with real variation by language. Useful context for why the final 30% so often includes a security layer nobody thought to ask for explicitly.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Co-Editor-in-Chief <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Token Costs Are an Architecture Problem, Not a Billing Problem (or whichever you chose)]]></title><description><![CDATA[The four cost drivers compounding your agent spend &#8212; and the architecture that fixes them.]]></description><link>https://packtbuildwithai.substack.com/p/token-costs-are-an-architecture-problem</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/token-costs-are-an-architecture-problem</guid><dc:creator><![CDATA[Charu Mitra Dubey]]></dc:creator><pubDate>Tue, 07 Jul 2026 19:32:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kWId!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><p>In April 2026, Uber&#8217;s CTO announced the company had burned through its entire annual AI coding budget in four months. Microsoft cut most of its internal Claude Code licenses shortly after. Meta started ranking engineers on a token usage leaderboard it calls &#8220;Claudeonomics.&#8221; These aren&#8217;t cautionary tales about AI adoption going wrong. They&#8217;re what happens when teams optimise for capability and ship without a cost architecture.</p><p>The assumption most teams made &#8212; that falling inference prices would eventually make token spend a non-issue &#8212; turned out to be wrong in a specific way. Agents don&#8217;t make one API call. They loop, plan, call tools, retry, hand off, and resend their full context at every step. By turn 20 of an agentic session, you&#8217;ve paid for the same history 20 times. A single unconstrained agent working through a complex software engineering task can cost $5 to $8 per run. Multiply that across a team, across a week, across a production system with no hard caps, and the bill takes a shape nobody planned for.</p><p>Token cost is not a billing problem. It&#8217;s an architecture problem. And it has an architecture solution.</p><div class="callout-block" data-callout="true"><p>In this edition:</p><ul><li><p>Why agent costs don&#8217;t behave like API costs &#8212; the four drivers that compound spend in ways per-token pricing tables don&#8217;t capture</p></li><li><p>A cost-aware agent architecture: prompt caching, model routing, session hygiene, and token budgeting as first-class engineering decisions</p></li><li><p>What&#8217;s worth paying attention to in AI this week</p></li></ul></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kWId!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 424w, /__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 848w, /__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kWId!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png" width="1406" height="688" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:688,&quot;width&quot;:1406,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:118094,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/205917536?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 424w, /__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 848w, /__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kWId!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066f78f0-de3a-48b2-83bc-99e4ef04e67b_1406x688.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>The Mental Model - Why agent costs don&#8217;t behave like API costs</h3><p>The mental model most developers bring to token costs comes from building with single-turn APIs. You send a prompt, you get a response, you pay for the tokens in both. The cost is predictable, linear, and easy to reason about.</p><p>Agentic systems break that model completely.</p><p>An agent isn&#8217;t making one call. It&#8217;s running a loop &#8212; planning, acting, observing, replanning &#8212; and at every step of that loop, it sends the full conversation history back to the model. Not a summary. Not a diff. The entire accumulated context, from turn one, every time. A 2026 Concordia University study put a number on this: a 2-to-1 input-to-output ratio across agentic sessions, with code review alone consuming 59% of every token spent. The technical term is the communication tax. The practical implication is that your token bill doesn&#8217;t scale linearly with the number of tasks your agent completes. It scales with session length, context depth, and failure rate &#8212; three things most teams aren&#8217;t measuring.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KplR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 424w, /__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 848w, /__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KplR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png" width="950" height="846" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:846,&quot;width&quot;:950,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:107429,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/205917536?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 424w, /__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 848w, /__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KplR!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd1e42f3-1872-4fc7-95c5-507f5a3875bc_950x846.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>The four cost drivers that actually matter</strong></h3><h4><strong>1. Session length</strong></h4><p>This is the biggest lever. A 20-turn session costs a fraction of an 80-turn session &#8212; but not proportionally. Later turns are more expensive per turn because context has accumulated. A session that runs twice as many turns can cost three to four times as much. The relationship is convex, not linear, and it punishes open-ended tasks where the agent iterates through failures without a stopping condition.</p><h4><strong>2. Context accumulation</strong></h4><p>Every tool call result, every intermediate reasoning step, every failed attempt gets appended to the context and resent on the next turn. An agent that calls ten tools before completing a task is carrying the output of all ten tools in its context window for every subsequent call &#8212; including the final ones where most of that history is irrelevant. Context doesn&#8217;t just cost money. It degrades performance. The model&#8217;s attention is diluted across information it no longer needs, which increases the likelihood of the kinds of errors that trigger retries, which adds more context, which compounds the problem.</p><h4><strong>3. Model selection</strong></h4><p>Most teams default to their most capable model for everything. This is the single most expensive mistake in agent architecture. Routing a simple classification task &#8212; &#8220;is this request in scope?&#8221; &#8212; through your most powerful model is like using a freight truck to deliver a letter. The capability gap between frontier models and smaller models has narrowed dramatically in 2026. For deterministic, well-scoped subtasks &#8212; intent classification, output validation, schema extraction &#8212; a smaller model is faster, cheaper, and often equally accurate. The cost difference is not marginal. Frontier model pricing runs 10 to 20 times higher per token than capable mid-tier models.</p><h4><strong>4. Retry loops</strong></h4><p>A demo that works 80% of the time is impressive. A production system that fails 20% of the time is useless &#8212; and expensive. Every retry is a full context resend. An agent stuck in a loop trying the same failed action can consume 50 times the tokens of a single linear pass. Unconstrained retry logic is the fastest path from a reasonable cost estimate to an invoice that requires a meeting to explain. This is the Unreliability Tax: the additional compute, latency, and spend required to mitigate the risk of failure in a system that wasn&#8217;t designed to fail gracefully.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rVq5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 424w, /__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 848w, /__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rVq5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png" width="1068" height="884" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:884,&quot;width&quot;:1068,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:174280,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/205917536?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 424w, /__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 848w, /__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rVq5!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F898abf54-cd30-4a81-a200-121c2b63693c_1068x884.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>The decision framework &#8212; where is your cost coming from?</strong></h3><p>Before reaching for an optimisation, you need to know which driver is responsible. The four drivers have different signatures:</p><ul><li><p><strong>Bill spikes on specific tasks</strong> &#8594; model selection problem. A single task type is over-provisioned.</p></li><li><p><strong>Bill grows faster than task volume</strong> &#8594; session length or context accumulation problem. Your agents are getting more expensive per task as usage scales.</p></li><li><p><strong>Bill spikes unpredictably</strong> &#8594; retry loop problem. An agent is hitting a failure mode and spinning.</p></li><li><p><strong>Bill is high but flat</strong> &#8594; prompt caching opportunity. You&#8217;re reprocessing static context you could be caching.</p></li></ul><p>Most teams have all four problems simultaneously. The Build section addresses each one with a concrete fix &#8212; but identifying which driver dominates your spend tells you where to start.</p><h4><strong>The shift that changes everything</strong></h4><p>Falling inference prices are real. GPT-4-class capability that cost $20 per million tokens in 2022 runs under a dollar today. Gartner expects inference to cost 90% less by 2030. The temptation is to treat this as a reason to defer the cost problem.</p><p>It isn&#8217;t. Agents consume tokens at a fundamentally different rate than single-turn applications &#8212; and as your agent count grows, the surface area for runaway spend grows with it. The teams that build cost architecture now ship more reliably, debug more easily, and don&#8217;t end up in the position Uber&#8217;s engineering org found itself in: a year&#8217;s budget, gone in four months, with no visibility into where it went.</p><p>In The Build, we implement the four-layer cost architecture: prompt caching, model routing, session hygiene, and token budgeting &#8212; the fixes that address each driver directly.</p><div><hr></div><h3>The Build &#8212; A cost-aware agent architecture</h3><p>Most cost problems in production agents aren&#8217;t model problems. They&#8217;re architecture problems &#8212; missing layers that let spend accumulate unchecked until the bill makes them visible. This section adds those layers explicitly, one at a time.</p><p>Stack: Node.js and the Anthropic SDK. The patterns are framework-agnostic &#8212; the same logic applies whether you&#8217;re using LangChain, LlamaIndex, or a bare API client.</p><p><strong>The architecture before writing a line of code</strong></p><p>Four layers, each targeting a different cost driver:</p><ul><li><p><strong>Prompt caching</strong> &#8212; stop reprocessing static context on every turn</p></li><li><p><strong>Model routing</strong> &#8212; right-size the model to the task</p></li><li><p><strong>Session hygiene</strong> &#8212; know when to start fresh</p></li><li><p><strong>Token budgeting</strong> &#8212; hard caps, anomaly detection, loop prevention</p></li></ul><p>Each layer is independent. You can add them in any order and see immediate impact on each corresponding cost driver. In practice, implement them in the order above &#8212; caching and routing give the fastest return, hygiene and budgeting prevent the catastrophic cases.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Every week, two concepts that make you a better AI engineer. No hype. No fluff. Just the mental models and code that matter in production.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong>Layer 1 &#8212; Prompt caching</strong></h3><p>Every agent has static content that doesn&#8217;t change between turns: the system prompt, tool definitions, knowledge base chunks, persona instructions. Without caching, this content is retokenised and billed on every single API call. With caching, you pay once and reference the cache for every subsequent call &#8212; reducing input costs by approximately 90% and latency by approximately 75% for the cached portion.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;a7eda59a-bd65-42d9-b408-60ab21bf0955&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">import Anthropic from '@anthropic-ai/sdk';

const anthropic = new Anthropic();

// Static content &#8212; defined once, cached across all turns
const SYSTEM_PROMPT = `You are a code review agent. Your job is to 
analyse pull requests for security vulnerabilities, performance issues, 
and adherence to engineering standards. Be specific. Be concise.`;

const KNOWLEDGE_BASE = `
[Your static knowledge base content here &#8212; coding standards, 
security checklists, architectural patterns. The longer this is, 
the more caching saves you.]
`;

async function reviewWithCaching(prContent, conversationHistory) {
  const response = await anthropic.messages.create({
    model: 'claude-sonnet-4-6',
    max_tokens: 1000,
    system: [
      {
        type: 'text',
        text: SYSTEM_PROMPT,
        cache_control: { type: 'ephemeral' } // cache the system prompt
      },
      {
        type: 'text',
        text: KNOWLEDGE_BASE,
        cache_control: { type: 'ephemeral' } // cache the knowledge base
      }
    ],
    messages: [
      ...conversationHistory,
      { role: 'user', content: prContent }
    ]
  });

  // Track cache performance
  const usage = response.usage;
  console.log({
    input_tokens: usage.input_tokens,
    cache_read_tokens: usage.cache_read_input_tokens,
    cache_write_tokens: usage.cache_creation_input_tokens,
    output_tokens: usage.output_tokens,
    cache_hit_rate: usage.cache_read_input_tokens / usage.input_tokens
  });

  return response;
}</code></pre></div><p>Cache read tokens cost approximately 10% of standard input token prices. On a system prompt and knowledge base that together total 10,000 tokens, called 1,000 times a day, caching alone saves roughly $45 per day at standard Sonnet pricing. At scale, it&#8217;s the cheapest optimisation available.</p><h3><strong>Layer 2 &#8212; Model routing</strong></h3><p>Not every task needs your most capable model. Intent classification, output validation, schema extraction, and simple summarisation are all tasks where a smaller, faster, cheaper model performs equivalently. The routing logic is deterministic &#8212; you&#8217;re not asking an LLM to decide which LLM to use. You&#8217;re classifying the task type and mapping it to the right model tier.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;2ab5c0e4-63f8-466d-b0db-ba84d4bddbfd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">// Model tiers &#8212; adjust based on your provider and pricing
const MODEL_TIERS = {
  FAST: 'claude-haiku-4-5-20251001',   // classification, validation, routing
  STANDARD: 'claude-sonnet-4-6',        // most agent tasks
  POWERFUL: 'claude-opus-4-6'           // complex reasoning, architecture decisions
};

// Task classifier &#8212; deterministic, no LLM call needed
function classifyTask(task) {
  const taskLower = task.toLowerCase();

  const fastTasks = [
    'classify', 'validate', 'extract', 'format',
    'summarize', 'route', 'check', 'verify', 'parse'
  ];

  const powerfulTasks = [
    'architect', 'design', 'reason', 'analyze complex',
    'debug intricate', 'review security', 'plan system'
  ];

  if (fastTasks.some(keyword =&gt; taskLower.includes(keyword))) {
    return MODEL_TIERS.FAST;
  }

  if (powerfulTasks.some(keyword =&gt; taskLower.includes(keyword))) {
    return MODEL_TIERS.POWERFUL;
  }

  return MODEL_TIERS.STANDARD;
}

async function routedCompletion(task, messages) {
  const model = classifyTask(task);

  console.log(`Routing task to: ${model}`);

  const response = await anthropic.messages.create({
    model,
    max_tokens: 1000,
    messages
  });

  return { response, model };
}</code></pre></div><p>The cost difference between Haiku and Opus on the same task is roughly 20x. A pipeline where 60% of tasks are classification or validation &#8212; a reasonable distribution for most production agents &#8212; can cut model spend by more than half with routing alone.</p><h3><strong>Layer 3 &#8212; Session hygiene</strong></h3><p>Context accumulation is silent and convex &#8212; it gets more expensive faster than it gets more useful. The fix is knowing when a session has reached diminishing returns and starting fresh rather than continuing to accumulate. Two signals matter: turn count and context token depth.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;ad3088dd-6501-4ccb-b0e1-eceff23ff676&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">const SESSION_LIMITS = {
  MAX_TURNS: 20,           // start fresh after 20 turns
  MAX_CONTEXT_TOKENS: 50000, // start fresh above 50k context tokens
  SUMMARY_THRESHOLD: 15    // summarise at turn 15, before hitting the limit
};

class SessionManager {
  constructor() {
    this.history = [];
    this.turnCount = 0;
    this.totalContextTokens = 0;
  }

  shouldRefresh() {
    return (
      this.turnCount &gt;= SESSION_LIMITS.MAX_TURNS ||
      this.totalContextTokens &gt;= SESSION_LIMITS.MAX_CONTEXT_TOKENS
    );
  }

  shouldSummarise() {
    return this.turnCount === SESSION_LIMITS.SUMMARY_THRESHOLD;
  }

  async summariseHistory() {
    // Compress history into a single context-preserving summary
    const summary = await anthropic.messages.create({
      model: MODEL_TIERS.FAST, // use the cheap model for summarisation
      max_tokens: 500,
      messages: [
        {
          role: 'user',
          content: `Summarise this conversation history concisely, 
preserving all decisions made, actions taken, and current task state.
Keep it under 300 words.

History: ${JSON.stringify(this.history)}`
        }
      ]
    });

    // Replace full history with compressed summary
    this.history = [
      {
        role: 'user',
        content: `Previous session summary: ${summary.content[0].text}`
      }
    ];

    console.log(`History compressed at turn ${this.turnCount}`);
  }

  async addTurn(role, content, tokenCount) {
    if (this.shouldSummarise()) {
      await this.summariseHistory();
    }

    this.history.push({ role, content });
    this.turnCount++;
    this.totalContextTokens += tokenCount;

    return this.shouldRefresh();
  }

  getStatus() {
    return {
      turns: this.turnCount,
      contextTokens: this.totalContextTokens,
      needsRefresh: this.shouldRefresh()
    };
  }
}</code></pre></div><h3><strong>What this architecture gives you</strong></h3><div class="callout-block" data-callout="true"><p>Four layers, four guarantees:</p><ul><li><p><strong>Prompt caching</strong> &#8212; you never pay full price for static context twice</p></li><li><p><strong>Model routing</strong> &#8212; every task runs on the cheapest model that can handle it reliably</p></li><li><p><strong>Session hygiene</strong> &#8212; context never accumulates past the point of diminishing returns</p></li><li><p><strong>Token budgeting</strong> &#8212; no task runs away, no retry loop compounds unchecked</p></li></ul></div><p>The full cost-aware stack is around 150 lines. None of it is complex. All of it is skipped by most teams until the bill arrives. The teams that build it upfront ship more predictably, debug more easily, and don&#8217;t end up explaining to a CTO why four months of budget is already gone.</p><div><hr></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/token-costs-are-an-architecture-problem?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Found this useful? Share it with a developer who's shipping agents without a cost layer.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/token-costs-are-an-architecture-problem?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/p/token-costs-are-an-architecture-problem?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h3>The Radar</h3><h4>01. Uber burned through its entire annual AI coding budget in four months &#8212; and it&#8217;s not an edge case.</h4><p>In April 2026, Uber&#8217;s CTO Praveen Neppalli Naga confirmed to The Information that the company had exhausted its full-year AI coding budget before the end of Q1. The cause wasn&#8217;t a rogue experiment or a misconfigured billing account. Claude Code spread across roughly 5,000 engineers faster than finance models had anticipated, with monthly costs per engineer ranging from $150 to $2,000 depending on workflow intensity. Naga himself reportedly spent $1,200 in a single two-hour session. The company subsequently capped per-employee AI spend at $1,500 per month. Microsoft followed by cutting most of its direct Claude Code licenses, moving engineers toward GitHub Copilot CLI instead. Neither company had a token budgeting layer in place before adoption scaled. The architecture from this edition&#8217;s Build section is what both needed before rollout, not after.</p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/">Fortune &#8212; Uber burned through its entire 2026 AI budget in four months</a> </p></div><h4>02. Meta&#8217;s Claudeonomics &#8212; what a token leaderboard teaches us about measuring AI productivity.</h4><p>An employee at Meta independently built an internal leaderboard called &#8220;Claudeonomics&#8221; that ranked the company&#8217;s 85,000+ employees by token consumption. In a single 30-day window, total usage exceeded 60 trillion tokens, with the top user logging 281 billion. The leaderboard awarded titles like &#8220;Token Legend&#8221; and &#8220;Cache Wizard&#8221; and was taken down two days after The Information reported on it. The story went public in April 2026 &#8212; and then got more interesting in June, when an internal memo revealed token consumption had hit 73.7 trillion in a subsequent month, and Meta warned costs were approaching billions of dollars annually. The company is now building &#8220;AI Gateway,&#8221; a centralised spend monitoring platform with anomaly detection and structured token budgets. Meta CTO Andrew Bosworth&#8217;s framing is worth sitting with: his best engineer was spending the equivalent of their salary in tokens and was 5 to 10 times more productive as a result. Token spend without output context is noise. Token spend mapped to measurable output is a signal. The distinction is exactly what good budget architecture surfaces.</p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://fortune.com/2026/04/09/meta-killed-employee-ai-token-dashboard/">Fortune &#8212; Meta killed its AI token dashboard</a> </p></div><h4>03. 60% of LLM production errors are rate limit failures &#8212; and most teams are misdiagnosing them.</h4><p>Datadog&#8217;s State of AI Engineering 2026 report analysed LLM call failures across production traces from thousands of organisations. The finding: 5% of all LLM call spans reported an error, and nearly 60% of those errors were caused by exceeded rate limits &#8212; not model failures, not hallucinations, not tool misuse. In March 2026 alone, that translated to approximately 8.4 million rate limit errors. Most teams see a 429 and reach for a retry. The correct response is to reach for a model router. Rate limit failures at scale are a signal that call volume is concentrated on a single high-demand model when smaller, faster alternatives could handle a significant portion of that load. The fix maps directly to Layer 2 of this edition&#8217;s Build section. Rate limits are not a provider problem. They are an architecture problem with an architecture solution.</p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://www.datadoghq.com/state-of-ai-engineering/">Datadog &#8212; State of AI Engineering 2026</a> </p></div><div><hr></div><h2>AI Tools of the Week</h2><h3><a href="https://www.vantage.sh">Vantage</a>: FinOps for AI token spend</h3><p>The tool that finally makes token costs visible in the way infrastructure costs have been for years. Vantage applies FinOps principles to AI spend &#8212; LLM token allocation, virtual tagging, budget alerts, and anomaly detection across providers including Anthropic, OpenAI, and Azure. Where most observability tools show you what your agents are doing, Vantage shows you what they&#8217;re costing and why, mapped to the business context that raw billing data lacks. The most useful feature for engineering teams: connecting token spend to engineering output so you can answer the question Uber&#8217;s COO couldn&#8217;t &#8212; is this spend producing anything measurable? Free tier available; paid plans from $20/month.</p><h3><a href="https://www.helicone.ai">Helicone</a>: Proxy-based LLM cost and usage visibility</h3><p>The fastest path to production cost visibility for teams that don&#8217;t want to redesign their agent instrumentation. Helicone sits in the request path as a proxy &#8212; one line of code change &#8212; and immediately starts logging requests, responses, latency, token usage, and spend across providers. You get a dashboard showing cost per model, cost per user, cost per session, and cache hit rates within minutes of setup. It&#8217;s less suited to explaining the internal execution path of a complex multi-agent system, but for teams that need cost visibility fast without an observability overhaul, it&#8217;s the right starting point. The rate limit error data from Datadog&#8217;s report? Helicone surfaces exactly that, per provider, in real time. Free tier up to 10,000 requests per month.</p><h4><a href="https://opentelemetry.io/docs/specs/semconv/gen-ai/">Prometheus + OpenTelemetry GenAI conventions</a>: The open standard for agent telemetry</h4><p>Not a product &#8212; a standard worth knowing and implementing now. OpenTelemetry&#8217;s GenAI semantic conventions define how to instrument AI agents with structured spans covering model calls, tool executions, token usage, latency, and errors. As of v1.41, the spec defines agent, workflow, tool, and model spans with required latency and token usage metrics. Nearly all major observability platforms &#8212; Datadog, Grafana, Honeycomb &#8212; are converging on this standard. Instrumenting your agents with OTel GenAI conventions now means your telemetry is portable across tools and you&#8217;re not locked into any single vendor&#8217;s schema. The attributes are still in Development stability, meaning names can change without a major version bump, so pin your instrumentation to a specific version and review on upgrades.</p><h3><a href="https://docs.anthropic.com/en/docs/build-with-claude/token-counting">Anthropic Token Counter</a>: Know your costs before the call</h3><p>A small tool that deserves more attention than it gets. Anthropic&#8217;s token counting API lets you calculate the exact token cost of a prompt &#8212; including tools and system prompts &#8212; before making the actual API call. For teams building the budgeting layer from this edition, it&#8217;s the right companion: count tokens before each turn, compare against your budget ceiling, and gate the call if it would breach the limit. No approximations, no surprises. Works synchronously and doesn&#8217;t consume output tokens. Free to use within your existing API access.</p><div><hr></div><p>If you made it this far, your agent architecture now has a cost layer &#8212; and that puts you ahead of most teams shipping in production today.</p><p>Uber had 5,000 engineers and a CTO who didn&#8217;t see the bill coming. Meta had 85,000 employees and a leaderboard that rewarded the wrong metric. Neither organisation was careless. They were building fast in a discipline that didn&#8217;t have established cost architecture yet. That architecture exists now. You just built it.</p><p>Four layers. One principle: token spend is an architecture decision, not a billing outcome. Prompt caching means you never pay full price for static context twice. Model routing means every task runs on the cheapest model that can handle it reliably. Session hygiene means context never accumulates past the point of diminishing returns. Token budgeting means no agent runs away, no retry loop compounds unchecked, no invoice arrives as a surprise.</p><p>The teams that build this upfront don&#8217;t just spend less. They debug faster, scale more predictably, and have the visibility to answer the question Uber&#8217;s COO couldn&#8217;t answer out loud: is this spend producing anything?</p><p>That&#8217;s the question worth building infrastructure to answer.</p><p>See you next week.</p><p><strong>Charu Mitra Dubey</strong><br>Co-Editor-in-chief, Build with AI</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share Build With AI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share Build With AI</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Build with AI #14: The Agentic Engineering Playbook – Part 2]]></title><description><![CDATA[Your AI wrote it in ten minutes. Here's why you'll be rewriting it for ten weeks.]]></description><link>https://packtbuildwithai.substack.com/p/build-with-ai-14-the-agentic-engineering</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/build-with-ai-14-the-agentic-engineering</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Thu, 02 Jul 2026 14:01:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nYva!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><div class="pullquote"><p>Enjoying Build with AI? Join us on social media for more AI news, practical tips, and updates between issues.</p><p>Follow us on: <a href="https://www.linkedin.com/company/packt-build-with-ai/">LinkedIn</a> | <a href="https://www.instagram.com/buildwithai_pro?igsh=MXJoN2l1bDgzbHpmaQ==">Instagram</a> | <a href="https://x.com/packtwebdevpro?s=21">X </a></p></div><p>Last week, we kicked off <strong>The Agentic Engineering Playbook</strong> with <em><a href="/__u/packtbuildwithai.substack.com/p/build-with-ai-12-the-agentic-engineering?r=55ncj4">Part 1</a></em>, where we explored a simple but powerful idea: AI hasn&#8217;t made software engineering less important. If anything, it has made good engineering practices more valuable than ever. The teams that will get the most out of AI won&#8217;t necessarily be the ones generating the most code. They&#8217;ll be the ones with the strongest systems for managing complexity, validating changes, and learning quickly.</p><p>But that naturally raises another question: if AI can generate working code in seconds, why do so many AI-assisted projects become difficult to maintain?</p><p>That&#8217;s exactly what we&#8217;ll explore in Part 2.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nYva!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nYva!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1594215,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/204386166?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nYva!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1529ea5-035b-4959-94a2-ab7ad43c7d79_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Along the way, we&#8217;ll unpack:</p><p>&#10145;&#65039; What developers really mean when they talk about <em>vibe coding</em></p><p>&#10145;&#65039; Why vibe coding feels incredibly productive, especially at the start of a project</p><p>&#10145;&#65039; The three structural reasons it begins to break down in production</p><p>&#10145;&#65039; How agentic engineering uses specifications, context, and verification to turn AI into a dependable engineering partner</p><p>&#10145;&#65039; Seven reads to go deeper, from real cost comparisons to a working engineer&#8217;s day-to-day workflow</p><p>Let&#8217;s dive in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1><span>Vibe coding versus agentic engineering &#8211; Why the distinction matters</span></h1><p>To proceed, we need to draw a sharp line between what has become known as vibe coding and the professional discipline of agentic engineering.</p><div class="pullquote"><p><span>The term vibe coding captures a very real, very widespread phenomenon. Developers sit down with an advanced tool like Claude Code or GitHub Copilot, open a chat window, and begin throwing natural language instructions at the model: Build me a React dashboard with a dark mode toggle that pulls data from this endpoint. The model churns out a massive block of boilerplate. The developer pastes it into their IDE, runs it, and hits an error. They copy the stack trace, paste it back into the chat window, and say, Fix this.</span></p></div><p><span>In this workflow, the developer is coding entirely by vibe. </span>They are relying on the LLM&#8217;s probabilistic predictions<span> to stumble towards a working solution. They are not defining strict architectures or writing tests. They are merely providing a directional push and letting the model hallucinate the details.</span></p><p><span>For the first forty-eight hours of a greenfield project (a prototype, a hackathon, a disposable script), vibe coding is arguably the most efficient way to build. The lack of constraints allows the model to draw on its training data freely, rapidly bridging the blank-page problem.</span></p><p><span>The trap, however, is that vibe coding doesn&#8217;t scale. It fails structurally and spectacularly when applied to enterprise systems.</span></p><p>Let&#8217;s examine why vibe coding fails in production.</p><h2>The reason vibe coding fails in production</h2><p>Three failures sit at the same structural level, and each one affects a different part of the workflow: output variance, context collapse, and the absence of a verification gate before deployment:</p><ul><li><p><strong>Output variance</strong>: Ask for the same feature on Tuesday and Thursday and the model can hand you two different architectures, because it samples each next token with randomness baked in. One run gives you a clean repository pattern, the next inlines the same queries across three controllers. Neither is wrong on its own, but the work no longer converges on a stable shape you can build on.</p></li><li><p><strong>Context collapse</strong>: As the codebase grows, the developer using the vibe approach has to paste more and more context into the chat window. Eventually, the signal-to-noise ratio degrades. The AI becomes confused, overwriting unrelated components or losing track of the initial objective. The developer makes this worse without noticing, because LLMs are sycophantic by design and exist to satisfy the prompt. Suggest a flawed approach and a vibe-coding agent will happily generate thousands of lines executing it, never pushing back.</p></li></ul><div class="callout-block" data-callout="true"><p><strong>&#10067; Have you hit context collapse in a real project? <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">Tell us in the survey</a>; we might feature it in Part 3.</strong></p></div><ul><li><p><strong>No verification gate</strong>: Vibe coding runs no automated check between generation and acceptance. The developer eyeballs the output, runs it once, and moves on. Nothing catches the variance or the drift before the code reaches production, so the first two failures compound silently. </p></li></ul><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/build-with-ai-14-the-agentic-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/build-with-ai-14-the-agentic-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/p/build-with-ai-14-the-agentic-engineering?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h2>How agentic engineering solves the problem </h2><p>Agentic engineering rejects the vibe approach. Instead of loose chats, it relies on strict, explicit specifications.</p><p>An agentic engineer does not write code, but they do not blindly chat with an AI either. They engineer the system that guides the AI. Rather than hand-writing the test suite, they set up an environment that makes the agent produce tests before it writes a line of implementation, so the agent is bound by verifiable constraints it cannot route around. They curate highly specific context files containing <strong><a href="https://adr.github.io/">Architectural Decision Records</a></strong> (<strong><a href="https://adr.github.io/">ADRs</a></strong>) and <strong><a href="https://www.ibm.com/think/topics/database-schema">data schemas</a></strong>, ensuring the agent is aware of the exact boundaries it must operate within.</p><div class="callout-block" data-callout="true"><p><strong>&#128269; What are ADRs and data schemas?</strong></p><p><strong>Architectural Decision Records (ADRs):</strong> Short documents that capture important architectural decisions, why they were made, and the trade-offs behind them. They help both developers and AI agents understand the intended design of a system and make changes that align with those decisions.</p><p><strong>Data schemas:</strong> Structured definitions of how data is organized, including the fields, data types, relationships, and constraints that govern it. They provide AI agents with a clear understanding of the shape and rules of the data they are working with.</p></div><p><span>The distinction matters because vibe coding leads to systems you cannot maintain, while agentic engineering creates systems that scale gracefully.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-POG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 424w, /__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 848w, /__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-POG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png" width="1456" height="1199" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1199,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:370652,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/204386166?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 424w, /__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 848w, /__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-POG!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382f43ff-f8e4-4308-a9b0-ffdf564cde50_2720x2240.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>Figure 1 - The three ways vibe coding breaks down, and what agentic engineering does instead</em></p><p><span>Imagine hiring fifty junior developers. Would you put them in a room, give them root access to production, and tell them to &#8220;build the vibes&#8221;? Of course not. You would establish style guides, implement code review processes, enforce CI/CD pipelines, and provide strict API contracts. Autonomous coding agents require the exact same infrastructure, scaled for computational speed.</span></p><p>With that distinction in place, we are ready to examine where AI delivers value and where it begins to break down.</p><p>In <em><strong>Part 3</strong></em><strong> </strong>of <em><strong>The Agentic Engineering Playbook</strong></em>, we'll explore one of the most common experiences every developer using AI coding tools eventually encounters: <strong>the 70% Problem</strong>. We'll look at why AI can take you from zero to a working prototype in seconds, why progress suddenly grinds to a halt, and how agentic engineering uses deterministic feedback loops to bridge the final 30% from "it works" to production-ready software.</p><div class="pullquote"><p><em><strong><span>&#128221; Spotted something worth calling out?  We&#8217;d love to hear about it&#8212;just drop us a quick note through </span><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a><span>. We&#8217;re all ears!</span></strong></em></p></div><h1>Further reading</h1><p><strong>&#10145;&#65039; <a href="https://tripleminds.co/blogs/technology/agentic-engineering-vs-vibe-coding/">Where the term &#8220;vibe coding&#8221; actually broke down at scale</a></strong><br>A blunt, cost-focused comparison from a dev shop that spends its days cleaning up vibe-coded production messes. Includes real numbers on what each approach costs in time and money, useful if you want the business case rather than just the philosophical one.</p><p><strong>&#10145;&#65039; <a href="https://wendelladriel.com/blog/my-take-on-vibe-coding-vs-agentic-engineering">A working engineer&#8217;s version of the workflow we described</a></strong><br>Wendell Adriel lays out his actual day-to-day process for staying in the driver&#8217;s seat while agents handle the tedious work, less theory, more &#8220;here&#8217;s what this looks like on a Tuesday.&#8221; A good companion if you want to see the ADR and spec approach applied to a real codebase.</p><p><strong>&#10145;&#65039; <a href="https://simonwillison.net/2026/May/6/vibe-coding-and-agentic-engineering/">The counterargument worth sitting with</a></strong><br>Not everyone agrees the line is as sharp as we made it sound. Simon Willison argues the two approaches are converging faster than he&#8217;d like, as tools built for one increasingly borrow from the other. Worth reading if you want the pushback before you fully commit to the framing.</p><p><strong>&#10145;&#65039; <a href="https://motherduck.com/blog/vibe-coding-dangerous-agentic-engineering-wes-mckinney/">What pandas&#8217; creator does instead of vibe coding</a></strong><br>Wes McKinney talks through his spec-driven workflow and continuous AI code review setup, and lands on a line worth stealing: when code is free, saying no is your last line of defense. Directly relevant to the verification gate we covered above.</p><p><strong>&#10145;&#65039; <a href="https://thebcms.com/blog/spec-driven-development">The ADRs and specs section, expanded</a></strong><br>We touched on ADRs and data schemas briefly above. This guide goes much deeper into spec-driven development as a category, including how GitHub Spec Kit, Kiro, and Claude Code each implement the idea differently. Useful if you&#8217;re evaluating tools for your own team.</p><p><strong>&#10145;&#65039; <a href="https://www.augmentcode.com/guides/what-is-spec-driven-development">For teams that need governance, not just guardrails</a></strong><br>If your org needs audit trails or multi-team coordination on top of the spec-first approach, this is the more technical version, covering the spectrum from specs that constrain code to specs that effectively become the source of truth.</p><p><strong>&#10145;&#65039; <a href="https://medium.com/@dave-patten/the-state-of-ai-coding-agents-2026-from-pair-programming-to-autonomous-ai-teams-b11f2b39232a">Why context engineering is the real skill now</a></strong><br>This piece makes a claim that lines up closely with what we said about context collapse: the hard part isn&#8217;t prompting anymore, it&#8217;s keeping specs and architecture notes persistently available to the agent as it works.</p><p><strong>&#10145;&#65039; <a href="https://dev.to/shrsv/why-agentic-engineering-must-replace-vibe-coding-339f">A concrete example of a verification gate</a></strong><br>We described the &#8220;no verification gate&#8221; failure mode in the abstract. This is what closing that gap looks like in practice: a lightweight tool that reviews every commit before it lands, rather than trusting a human to eyeball the output once.</p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><p><strong><span>Cheers!<br>Adrija Mitra<br>Co-Editor-in-Chief, <br></span>Build with AI</strong><span> </span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Harness Engineering is the New Prompt Engineering - Your Model Isn't the Bottleneck Anymore]]></title><description><![CDATA[Reliability in production agents comes from the runtime layer, not the model. Here's how to build the system around it.]]></description><link>https://packtbuildwithai.substack.com/p/harness-engineering-is-the-new-prompt</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/harness-engineering-is-the-new-prompt</guid><dc:creator><![CDATA[Charu Mitra Dubey]]></dc:creator><pubDate>Tue, 30 Jun 2026 14:05:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iLPq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Welcome back to Build with AI.</p><p>Over the last two editions we built the protocol layer (MCP, A2A, ACP) and the memory layer (vector stores, retrieval, the four types of memory your agent actually needs). Both editions assumed something: that once the model has the right tools and the right context, it&#8217;ll behave reliably.</p><p>It won&#8217;t. Not by default.</p><p>This week we cover the layer that sits around the model and decides whether all that careful protocol and memory work actually survives contact with production: the harness.</p><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong><a href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w"><span>Social engineering is about manipulating people&#8217;s emotions. Identify the susceptibilities that hackers use to exploit people.</span></a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 424w, /__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 848w, /__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_zdj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png" width="519" height="346" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:200,&quot;width&quot;:300,&quot;resizeWidth&quot;:519,&quot;bytes&quot;:9987,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtnetpro.substack.com/i/204084771?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 424w, /__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 848w, /__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_zdj!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F348ac97a-030d-4de1-9f07-4b22c5a7cac1_300x200.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong>This <a href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w"><span>NINJIO Insights Report</span> </a>dives into the key emotional susceptibilities that make social engineering work and offers concrete steps that your security team can take to equip your workforce to resist cyberattacks.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w&quot;,&quot;text&quot;:&quot;Download the Guide&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w"><span>Download the Guide</span></a></p></div><div class="callout-block" data-callout="true"><p>In this edition:</p><ul><li><p>Why harness design, not model capability, is what separates agents that work in a demo from agents that work at 2am when nobody&#8217;s watching</p></li><li><p>How to build a cross-language, cross-framework multi-agent handoff using A2A, with explicit state checkpoints and fail-safe routing to a human</p></li><li><p>What&#8217;s worth paying attention to in AI this week</p></li></ul></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iLPq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 424w, /__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 848w, /__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iLPq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png" width="1346" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/be2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1346,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:132964,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/204249784?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 424w, /__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 848w, /__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iLPq!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe2a69fb-d9b0-4e7a-a71a-621d61b14cb2_1346x742.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Mental Model - Stop tuning the model. Start engineering the harness.</h2><p>For the past two years, the default move when an agent misbehaved was to swap models, rewrite the prompt, or wait for the next release. That instinct made sense when models were the limiting factor. It doesn&#8217;t anymore.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9Hya!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 424w, /__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 848w, /__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9Hya!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png" width="1338" height="638" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:638,&quot;width&quot;:1338,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127842,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/204249784?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 424w, /__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 848w, /__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9Hya!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1232c4c-a3e1-4c57-b573-ced46e0c23c9_1338x638.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The pattern shows up across nearly every team running agents in production right now: a more capable model doesn&#8217;t fix a flaky agent. A better harness does. There&#8217;s a useful framing going around in the agent engineering community this month: classify every agent failure into one of four buckets, and you&#8217;ll find the fix is almost never &#8220;use a smarter model.&#8221;</p><p><strong>The four failure types</strong></p><ul><li><p><strong>Context failures</strong> - the agent didn&#8217;t have the information it needed, or had too much of the wrong information. This is a retrieval and assembly problem, not a reasoning problem.</p></li><li><p><strong>Constraint failures</strong> - the agent did something it technically could do but shouldn&#8217;t have. Wrong permission scope, wrong tool for the job, no guardrail stopping it.</p></li><li><p><strong>Verification failures</strong> - nobody checked the output before it shipped. The agent was confident and wrong, and confidence was the only signal anyone was watching.</p></li><li><p><strong>Planning failures</strong> - the agent took a reasonable first step toward an unreasonable overall plan, and nothing caught the drift until five steps in.</p></li></ul><p>Each of these maps to a harness component, not a model parameter. Context failures get fixed by your retrieval layer. Constraint failures get fixed by your permission model. Verification failures get fixed by adding a checker step. Planning failures get fixed by checkpointing and bounded execution, not by hoping the model plans better next time.</p><p><strong>Four things every agent harness needs</strong></p><p>Strip away the framework-specific jargon and a production harness is really just four components:</p><ul><li><p><strong>An agent loop</strong> - the control flow that decides what happens next: call a tool, ask for clarification, hand off, or stop.</p></li><li><p><strong>A tool interface</strong> - how the agent reaches the outside world. This is your MCP layer from two editions ago.</p></li><li><p><strong>Context management</strong> - what the model sees on each turn. This is your memory layer from last edition.</p></li><li><p><strong>Control mechanisms</strong> - retries, timeouts, approval gates, rollback. The part most teams skip until something breaks in production.</p></li></ul><p>Most teams have built the first three. Almost nobody has built the fourth properly, and it&#8217;s the one that determines whether your agent is a toy or infrastructure.</p><p><strong>Why this matters right now</strong></p><p>Two things converged this month that make this less theoretical. First, GPT-5.4 and Claude&#8217;s frontier models both shipped meaningful capability jumps in agentic reasoning. Teams that haven&#8217;t separated their harness logic from their model logic are now locked into upgrading both together, every time, with no way to isolate what actually improved. Second, OWASP&#8217;s &#8220;Excessive Agency&#8221; risk category has moved from a niche security concern to something showing up in actual vendor security reviews &#8212; over-provisioned permissions and missing approval gates are now things procurement teams ask about.</p><blockquote><p>The takeaway: build your harness so it&#8217;s the thing you&#8217;d keep even if you swapped the model underneath it tomorrow. If a model upgrade requires you to rewrite your retry logic or your permission scopes, that logic was never separated from the model in the first place.</p></blockquote><p>In The Build, we put this into practice with the harness pattern most teams need next: a multi-agent handoff with explicit state checkpoints, a deterministic validator that doesn&#8217;t trust the LLM&#8217;s output blindly, and a fail-safe route to a human when something doesn&#8217;t check out.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>The Build - A cross-language agent handoff with checkpoints and a fail-safe</h2><p>Most multi-agent demos show two LLM agents talking to each other through A2A and call it done. That&#8217;s not a harness, it&#8217;s a happy path. This is the version with the parts that make it production-safe: a deterministic validator that doesn&#8217;t reason, just checks; a checkpoint that survives a crash mid-task; and an explicit route to a human when the validator isn&#8217;t confident.</p><blockquote><p>Stack: a Python extraction agent (the one doing LLM reasoning), a Go validation service (deterministic, no LLM, just rules), connected via A2A, with shared state checkpoints in Postgres.</p></blockquote><p><strong>The architecture before writing a line of code</strong></p><p>Two services, two languages, one job split cleanly between them.</p><ul><li><p><strong>Extraction agent (Python)</strong> - receives a document, uses an LLM to extract structured fields, and sends the result to the validator via A2A. It does not decide whether its own output is correct. That&#8217;s not its job.</p></li><li><p><strong>Validation service (Go)</strong> - deterministic rule checks against the extracted fields. No LLM. Fast, cheap, and either passes, fails, or flags for human review. It writes its decision to the same checkpoint table the extraction agent reads.</p></li></ul><p>The reason this split matters: LLMs are good at extraction and bad at grading their own work. A deterministic validator written in a few hundred lines of Go will catch malformed dates, out-of-range values, and missing required fields more reliably and more cheaply than asking the model to check itself.</p><h4><strong>Step 1 - the checkpoint table</strong></h4><p>Every task gets a row. State moves forward, never silently. If the process crashes mid-task, the next run picks up exactly where it left off instead of starting over or losing the task.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;sql&quot;,&quot;nodeId&quot;:&quot;cd2ea434-c42b-4c2f-9373-039c77b6aec1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-sql">create table agent_tasks (
  id            uuid primary key default gen_random_uuid(),
  document_id   text not null,
  status        text not null default 'extracting',
  -- extracting -&gt; validating -&gt; approved | needs_review | failed
  extracted     jsonb,
  validation    jsonb,
  attempt_count int not null default 0,
  created_at    timestamptz default now(),
  updated_at    timestamptz default now()
);</code></pre></div><h4><strong>Step 2  - the extraction agent</strong></h4><p>This is the only place an LLM touches the task. It writes its result and moves the status forward. It does not call the validator directly &#8212; it hands off via A2A and lets the validator pull the task when ready.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;af013809-0307-4737-bebd-766c76b61023&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import json
import anthropic
from db import get_connection

client = anthropic.Anthropic()

def extract(document_id: str, document_text: str):
    conn = get_connection()
    task = conn.execute(
        "insert into agent_tasks (document_id, status) "
        "values (%s, 'extracting') returning id",
        (document_id,)
    ).fetchone()

    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=500,
        system=(
            "Extract structured fields from this document. "
            "Return only valid JSON with keys: invoice_date, "
            "amount, currency, vendor_name. Do not guess "
            "missing fields, return null instead."
        ),
        messages=[{"role": "user", "content": document_text}]
    )

    extracted = json.loads(response.content[0].text)

    conn.execute(
        "update agent_tasks set extracted = %s, "
        "status = 'validating', updated_at = now() "
        "where id = %s",
        (json.dumps(extracted), task["id"])
    )

    notify_validator(task["id"])
    return task["id"]


def notify_validator(task_id: str):
    import requests
    requests.post(
        "http://validator-service:8080/tasks/send",
        json={
            "id": str(task_id),
            "message": {
                "role": "user",
                "parts": [{"type": "text", "text": str(task_id)}]
            }
        },
        timeout=5
    )</code></pre></div><h4><strong>Step 3 - the validator (Go, no LLM)</strong></h4><p>This is the harness control mechanism doing its job. Deterministic checks only. Three outcomes: approved, needs_review, or failed. Nothing here calls a model.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;go&quot;,&quot;nodeId&quot;:&quot;e2975778-9396-4730-9436-f9a011f04e2b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-go">package main

import (
    "encoding/json"
    "net/http"
    "time"
)

type Extracted struct {
    InvoiceDate string  `json:"invoice_date"`
    Amount      float64 `json:"amount"`
    Currency    string  `json:"currency"`
    VendorName  string  `json:"vendor_name"`
}

func validate(e Extracted) (status string, reasons []string) {
    if e.InvoiceDate == "" {
        reasons = append(reasons, "missing invoice_date")
    } else if _, err := time.Parse("2006-01-02", e.InvoiceDate); err != nil {
        reasons = append(reasons, "invoice_date not ISO 8601")
    }

    if e.Amount &lt;= 0 {
        reasons = append(reasons, "amount must be positive")
    }
    if e.Amount &gt; 1000000 {
        // outlier &#8212; not necessarily wrong, but worth a human look
        reasons = append(reasons, "amount exceeds review threshold")
    }

    if e.VendorName == "" {
        reasons = append(reasons, "missing vendor_name")
    }

    switch {
    case len(reasons) == 0:
        return "approved", nil
    case containsHardFailure(reasons):
        return "failed", reasons
    default:
        return "needs_review", reasons
    }
}

func containsHardFailure(reasons []string) bool {
    for _, r := range reasons {
        if r == "missing invoice_date" || r == "missing vendor_name" {
            return true
        }
    }
    return false
}

func taskHandler(w http.ResponseWriter, r *http.Request) {
    var payload struct {
        ID string `json:"id"`
    }
    json.NewDecoder(r.Body).Decode(&amp;payload)

    extracted := fetchExtracted(payload.ID)
    status, reasons := validate(extracted)
    writeValidationResult(payload.ID, status, reasons)

    w.Header().Set("Content-Type", "application/json")
    json.NewEncoder(w).Encode(map[string]any{
        "id":      payload.ID,
        "status":  status,
        "reasons": reasons,
    })
}

func main() {
    http.HandleFunc("/tasks/send", taskHandler)
    http.ListenAndServe(":8080", nil)
}</code></pre></div><h4><strong>Step 4 - the fail-safe route</strong></h4><p><code>needs_review</code> is the status that makes this a harness and not a happy-path demo. Tasks that land there don&#8217;t retry automatically and don&#8217;t get marked done. They go to a queue a human actually looks at.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7bb0c921-bf8e-435d-a522-a310cf391b49&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def check_for_review_queue():
    conn = get_connection()
    pending = conn.execute(
        "select id, document_id, extracted, validation "
        "from agent_tasks where status = 'needs_review' "
        "order by created_at asc"
    ).fetchall()

    for task in pending:
        send_to_review_queue(task)
        # status stays needs_review until a human resolves it &#8212;
        # never auto-transitions to approved</code></pre></div><p><strong>What this gives you that the two-agent demo doesn&#8217;t</strong></p><ul><li><p><strong>A validator that can&#8217;t be talked into agreeing with a wrong extraction.</strong> It&#8217;s deterministic. It doesn&#8217;t have a context window to be persuaded in.</p></li><li><p><strong>A checkpoint that survives a crash.</strong> If the extraction service dies after writing <code>extracted</code> but before notifying the validator, the next sweep picks the task back up from <code>validating</code> instead of losing it.</p></li><li><p><strong>An explicit failure mode that doesn&#8217;t autoresolve.</strong> <code>needs_review</code> is a dead end until a human moves it forward. No silent retries that quietly approve something nobody checked.</p></li></ul><p>This is maybe 200 lines across two languages. None of it is hard. The part that took thought is deciding where the LLM&#8217;s judgment ends and where deterministic rules and human review begin &#8212; and that boundary is the actual harness design decision, not the code.</p><div><hr></div><h2>The Radar</h2><h3><strong>01. Google published an open spec for how agents find each other at runtime, the discovery layer your A2A handoffs have been missing.</strong> </h3><p>Agentic Resource Discovery (ARD), backed by eleven companies including Microsoft, GitHub, Salesforce, and Hugging Face, standardizes how organizations publish a machine-readable catalog of their MCP servers, A2A agents, and APIs at a well-known URL, so agents can search for the right capability at runtime instead of having every integration pre-wired by hand. It&#8217;s an early draft, adoption is close to zero so far, but the spec is real, Apache 2.0 licensed, and GitHub already shipped a reference implementation (agent finder) the same day. If your harness currently hardcodes every tool and agent endpoint, this is the standard to watch for the next layer up. </p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://developers.googleblog.com/announcing-the-agentic-resource-discovery-specification/">Announcing the Agentic Resource Discovery specification &#8212; Google Developers Blog</a></p></div><h3><strong>02. New survey data: 88.4% of organizations running AI agents had a security breach in the past year, and visibility is getting worse, not better.</strong> </h3><p>AvePoint&#8217;s third annual State of AI report found the share of organizations unable to detect unsanctioned agent activity nearly tripled year over year, even as 46.9% of employees now use AI agents weekly or daily. The uncomfortable detail: confidence is rising while actual monitoring coverage isn&#8217;t, organizations are getting more comfortable with a risk they haven&#8217;t reduced. If your harness&#8217;s control mechanisms layer doesn&#8217;t include agent-specific logging and access auditing yet, this is the data point to bring to whoever owns that budget conversation. </p><p>Read it here: <a href="https://www.globenewswire.com/news-release/2026/06/29/3318982/0/en/AvePoint-Research-Reveals-AI-Visibility-Gaps-Have-Nearly-Tripled-as-AI-Agents-Scale-and-Almost-Half-of-Enterprise-Employees-Now-Rely-on-Agents-Daily-or-Weekly.html">AvePoint Research Reveals AI Visibility Gaps Have Nearly Tripled as AI Agents Scale &#8212; GlobeNewswire</a></p><h3><strong>03. Only about 7% of companies run fully autonomous agents in production, a useful reality check against the hype.</strong> </h3><p>Despite the spending and headline numbers, a fresh practitioner analysis this week found true end-to-end automation remains rare, and that many agentic projects are at risk of cancellation specifically because teams bolt agents onto existing workflows instead of redesigning around them, without clear baselines for time, error rate, or human effort. It&#8217;s a useful gut check before your next harness investment: the gap between hype and live deployment is exactly where the discipline in this edition pays off, since the teams that survive are the ones measuring before they scale. </p><div class="callout-block" data-callout="true"><p>Read it here: <a href="https://aiagentstore.ai/ai-agent-news/this-week">Daily AI Agent News </a></p></div><div><hr></div><h2>AI Tools of the Week</h2><p><strong><a href="https://laminar.sh">Laminar</a></strong> - open-source, OpenTelemetry-native observability built specifically for long-running agents rather than retrofitted from single-call LLM logging. The standout feature is its agent rollout debugger: when a 30-minute browser agent fails on tool call 1,847 of 2,000, it shows you which span actually caused it instead of forcing a manual scroll. Self-hostable, free tier at 1GB.</p><p><strong><a href="https://github.com/leondz/garak">garak</a></strong> - NVIDIA&#8217;s LLM vulnerability scanner. Probes your agent for prompt injection susceptibility, jailbreaks, and data leakage before someone else finds them in production. Most teams run security scans on their infrastructure and never think to scan the model layer itself; this fills that gap and it&#8217;s free.</p><p><strong><a href="https://logfire.pydantic.dev">Logfire</a></strong> - observability from the Pydantic team, built Python-native rather than bolted on through a generic SDK. If your harness uses Pydantic-AI or you&#8217;re already structuring outputs with Pydantic models (which you should be, for the validation layer in last edition&#8217;s Build), this gives you tracing that understands your types instead of just logging opaque JSON blobs.</p><p><strong><a href="https://github.com/traceloop/openllmetry">OpenLLMetry</a></strong> - an OpenTelemetry instrumentation library for LLM apps, not a platform in itself. The value is vendor neutrality: instrument once, then route traces into Grafana, Datadog, or whatever your team already runs, instead of marrying your harness to one observability vendor&#8217;s proprietary SDK. Easy to overlook because it doesn&#8217;t have a dashboard to demo.</p><p><strong><a href="https://github.com/agentnotary">agentnotary</a></strong> - a small, sharp tool that does one thing most harness checklists skip: cryptographically notarizes and audits agent runs, with built-in EU AI Act documentation generation and an adversarial fuzzer for testing how your agent handles malformed or hostile inputs. Worth a look given this edition&#8217;s CISA item on accountability gaps; it&#8217;s the kind of control-mechanism tooling that exists specifically because regulators started asking for audit trails.</p><div><hr></div><p>If you made it this far, you&#8217;ve stopped asking &#8220;which model should I use&#8221; and started asking &#8220;what does my harness do when the model is wrong.&#8221; That&#8217;s the more useful question, and most teams shipping agents right now still haven&#8217;t gotten there.</p><p>You have the failure taxonomy, the four-component model, and a working pattern for the part everyone skips: the deterministic checker and the fail-safe that doesn&#8217;t autoresolve.</p><p>If this edition changed how you think about your agent stack, forward it to whoever on your team is still debugging flaky agents by swapping models. The fix probably isn&#8217;t the model.</p><p>See you next week.</p><p><strong>Charu Mitra Dubey,</strong> </p><p>Editor-in-chief, </p><p>Build with AI</p>]]></content:encoded></item><item><title><![CDATA[Build with AI #12: The Agentic Engineering Playbook – Part 1]]></title><description><![CDATA[Why speed without structure is a trap, and how AI amplifies your engineering fundamentals]]></description><link>https://packtbuildwithai.substack.com/p/build-with-ai-12-the-agentic-engineering</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/build-with-ai-12-the-agentic-engineering</guid><dc:creator><![CDATA[Adrija Mitra]]></dc:creator><pubDate>Fri, 26 Jun 2026 14:02:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/46dae638-9047-460a-86d6-9b22bb5c8578_1456x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><p>Let&#8217;s be honest: AI coding tools are a little intoxicating right now. You type a prompt, hit Enter, and suddenly there&#8217;s a React component, a REST API, or half a feature sitting in front of you as if it appeared by magic. It&#8217;s fast. It&#8217;s fun. It&#8217;s also exactly why we need to get more serious about how we build with AI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JVkU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 424w, /__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 848w, /__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JVkU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png" width="1456" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:816,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2302861,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/203515932?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 424w, /__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 848w, /__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JVkU!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49edc775-8ae1-45f3-90d5-d956aae4b72e_2912x1632.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Over the next five weeks in <strong>Build with AI</strong>, we&#8217;re bringing you <strong>The Agentic Engineering Playbook</strong>, a five-part series by <em><strong><a href="https://www.linkedin.com/in/alexmerced/">Alex Merced</a></strong></em><strong> </strong>and<strong> </strong><em><strong><a href="https://www.linkedin.com/in/benedikt-stemmildt/">Benedikt Stemmildt</a></strong></em> that explores what really changes when AI moves from &#8220;helpful autocomplete&#8221; to an active participant in the software delivery process. Along the way, we&#8217;ll unpack:</p><p>&#10145;&#65039; Where AI genuinely shines<br>&#10145;&#65039; Where it consistently breaks down<br>&#10145;&#65039; And what it takes to build workflows that turn raw model speed into reliable, production-grade software</p><p>By the end of the series, you&#8217;ll have a practical framework for using AI as an engineering multiplier, not just a faster code generator.</p><p>We&#8217;re kicking things off with the biggest mindset shift of all. Part 1 is about why AI doesn&#8217;t make engineering discipline less important. It makes it more important. If software development&#8217;s old bottleneck was typing, AI has blown that bottleneck wide open. The question now isn&#8217;t how fast AI can write code. It&#8217;s what happens after the code is generated.</p><p>Let&#8217;s find out.</p><div class="pullquote"><p>AI has commoditized syntax, not engineering judgment. The teams that benefit most from AI will not be the ones generating the most code, but the ones with the strongest systems for managing complexity, validating changes, and learning fast.</p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>Beyond Vibe Coding</h1><p>Software development is undergoing its most radical transition in decades. For most of that history, the fundamental bottleneck in engineering was the human translation of intent into syntactic execution. We possessed grand architectures in our heads, but typing them out, debugging semicolons, satisfying borrow checkers, and memorizing standard libraries was a painstaking, manual craft.</p><p>Today, that bottleneck has largely evaporated.</p><p>LLMs and autonomous agents have commoditized syntactic generation. A developer can now describe a complex state machine or a fully functional REST API in natural language, and a system will generate it in seconds. Unlike a compiler, which translates one precise notation into another and produces the same output every time, these systems predict. They vary. </p><div class="pullquote"><p>The same request can produce different code twice. That difference is the whole reason engineering discipline matters more here, not less.</p></div><p>However, this newfound velocity has introduced a severe vulnerability into our workflows. Early adopters of this technology have embraced a paradigm often referred to as <strong>vibe coding</strong>. Coined by <em><strong>Andrej Karpathy</strong></em> in early 2025 to describe an intent-first, syntax-second approach, vibe coding champions a fast, loose, iterative generation style. </p><blockquote><p>The developer acts as a Head Chef, barking orders at algorithmic sous-chefs and letting the AI handle the mechanical slicing and dicing. It sounds liberating, and for small scripts or weekend side projects, it absolutely is. But when the goal is production-grade enterprise software, the same looseness becomes dangerous.</p></blockquote><p>When you apply unstructured vibe coding to massive legacy codebases with strict security, performance, and scaling constraints, the magic quickly turns into a nightmare. AI agents are non-deterministic, bounded reasoners. Left to their own devices, they will write code that appears functionally correct but is structurally disastrous. They will tightly couple components, hallucinate dependencies, and introduce security vulnerabilities.</p><p>This leads to a fundamental thesis that runs through this entire series: <strong>agentic software engineering is the pinnacle of traditional engineering rigor, not a break from it.</strong> The disciplines that make an agent productive are the disciplines you already know: testing, version control, modular design, continuous integration, and tight feedback loops. None of this is new. What changes is who does the typing and how much leverage each decision now carries.</p><p>We need to move beyond vibe coding and integrate the rapid generative powers of AI with the structural discipline of classic software engineering. By doing so, we shift our focus from typing code to designing the systems that produce it, managing complexity, and optimizing for learning.</p><p>The central argument of this series explains <em>why</em> agentic engineering works. A second question is where the industry currently stands in adopting it. We can understand that progression through four eras:</p><ul><li><p><strong>Craft</strong>: AI is autocomplete. It speeds up the individual at the keyboard, and nothing else changes. The large majority of companies still work here.</p></li><li><p><strong>Manufactory</strong>: The engineer steers the agent step by step and validates each step by hand. The leverage is real, but a human sits in the loop for every move. A growing minority of teams operate this way.</p></li><li><p><strong>Factory</strong>: The agent runs inside the loop, with deterministic orchestration between its steps, and the engineer&#8217;s job becomes designing the factory floor rather than walking it. Only a small fraction of teams have reached this. This is the era we build toward across the series.</p></li><li><p><strong>Compound Factory</strong>: The system improves itself each cycle, with knowledge accumulating so that every pass starts smarter than the last. The first concrete piece of this is post-merge self-healing, where the factory learns from its own production incidents.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mys_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 424w, /__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 848w, /__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mys_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png" width="1392" height="1130" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1130,&quot;width&quot;:1392,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1003531,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/203515932?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 424w, /__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 848w, /__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mys_!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd209bec-8c3c-4eb0-badb-e85b425a7e92_1392x1130.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li></ul><p style="text-align: center;"><em><strong>Figure 1 - The four eras of agentic engineering</strong></em></p><p>These eras describe how widely the practice has spread, not whether it is sound. <em><strong>The Pinnacle Thesis</strong></em>, that is, the idea that agentic software engineering is the pinnacle of traditional engineering rigor, not a break from it, answers the second question.</p><div class="pullquote"><p><em><strong><span>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it&#8212;just drop us a quick note through </span><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a><span>. We&#8217;re all ears!</span></strong></em></p></div><h2>The Pinnacle Thesis &#8211; From modern software engineering to agentic engineering</h2><p><a href="https://www.linkedin.com/in/dave-farley-a67927/">Dave Farley</a>, a pioneer of <em>Continuous Delivery (CD)</em>, defined the core of the discipline in his book <em>Modern Software Engineering</em>. Farley stripped away the noise of agile ceremonies, specific languages, and fads, reducing engineering down to two non-negotiable foundations:</p><ul><li><p><em>Optimizing for learning</em></p></li><li><p><em>Managing complexity</em></p></li></ul><p>Any practice that supports these foundations is good engineering; any practice that hinders them is bad engineering.</p><p>Before we explore how these apply to AI, we should understand them in a purely human context.</p><blockquote><p><em><strong>Optimizing for learning</strong></em></p></blockquote><p>Optimizing for learning acknowledges that software development is fundamentally an exercise in discovery. We rarely know the exact solution before we start coding. Because we operate in an environment of high uncertainty, we need systems that let us learn quickly and safely. This involves adopting an empirical mindset, making hypotheses, running experiments, and measuring results. We implement tight feedback loops, leaning heavily on automated testing and Continuous Integration (CI) to tell us immediately if our last change broke the system. We favor incremental, evolutionary design over massive upfront planning because we learn more from a running system than a theoretical diagram.</p><blockquote><p><em><strong>Managing complexity</strong></em></p></blockquote><p>Managing complexity acknowledges that scale is the enemy of understanding. As systems grow, they exceed the limits of human working memory. To survive, we partition problems into manageable chunks. We use techniques like modularity, high cohesion, low coupling, and clear abstractions. We establish strict boundaries and interfaces so that a developer can reason about a single component without needing to understand the entire application. We isolate state and side effects. By managing complexity, we ensure the system remains safe to change.</p><div class="callout-block" data-callout="true"><p><strong><span>Core design principles</span></strong></p><p><span>These terms describe foundational techniques used to manage complexity in software systems:</span></p><ul><li><p><em><span>Modularity</span></em><span>: Breaking a system into smaller, independent components so each part can be developed, tested, and changed in isolation.</span></p></li><li><p><em><span>High cohesion</span></em><span>: Ensuring each component has a single, well-defined responsibility. A cohesive module does one thing and does it well.</span></p></li><li><p><em><span>Low coupling</span></em><span>: Minimizing dependencies between components so changes in one area do not ripple unpredictably through the system.</span></p></li><li><p><em><span>Clear abstractions</span></em><span>: Defining simple, stable interfaces that hide internal details, allowing developers (and AI agents) to work with a component without needing to understand its implementation.</span></p></li></ul><p><span>Together, these principles make systems easier to understand, safer to modify, and more resilient as they scale.</span></p></div><p><span>When generative AI entered the mainstream, a dangerous false dichotomy emerged. Many assumed that since the AI could generate code instantaneously, the slow, methodical practices of engineering were obsolete. Why worry about decoupling when the AI can just refactor the whole file? Why write tests when the AI can read the entire codebase and ensure correctness?</span></p><blockquote><p>This assumption is catastrophic in practice</p></blockquote><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/build-with-ai-12-the-agentic-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/build-with-ai-12-the-agentic-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/p/build-with-ai-12-the-agentic-engineering?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h3>AI models do not replace Farley&#8217;s foundations</h3><p>The Pinnacle Thesis says that AI models do not replace Farley&#8217;s foundations; they multiply them. Let us examine why.</p><p><span>AI agents are phenomenal at generating raw volume, but they suffer from bounded rationality. An LLM has a finite context window and a limited ability to track deeply nested, cross-file dependencies. If you ask an agent to modify a highly coupled, monolithic ball of mud, it will almost certainly fail. The AI will lose the plot, hallucinate connections that do not exist, and introduce regressions precisely because the complexity is unmanaged.</span></p><p>However, if you point an AI agent at a system with aggressively managed complexity (one with high cohesion, strict modularity, and well-defined interfaces), the agent becomes a superpower. The AI only needs to hold a small, bounded context in its memory to operate successfully. Traditional engineering practices (like micro-architectures and interface segregation, keeping interfaces narrow so a client depends only on the methods it uses) were originally designed to protect human cognitive limits. Ironically, they are the exact same practices required to protect LLM context windows.</p><p>Optimizing for learning matters even more when your code generator is non-deterministic. When a human writes code, they have an internal feedback loop. They know why they made a change. An AI agent does not. It is predicting the most likely next token based on a prompt, not reasoning through your business logic.</p><p>That means your external feedback loops have to do the heavy lifting. They need to be tight and reliable. If a human introduces a bug, you might spend hours tracking it down. An AI can introduce thousands of lines of broken logic in seconds. Without strong, automated checks in place, that speed works against you. You are not moving faster. You are accumulating unverified, hard-to-change code more quickly.</p><p><em><strong>This is the stake of the whole series: AI does not fix a weak engineering organization; it multiplies whatever is already there.</strong></em> </p><div class="callout-block" data-callout="true"><p>Picture two teams handed the same agent. Team A has tests, clear interfaces, and tight feedback loops; the agent runs along those rails and the team ships verified work several times faster. </p><p>Team B has flaky tests, tangled coupling, and tribal knowledge no document captures; the agent produces mediocre code faster, and every shortcut compounds. Same tool, opposite outcomes. </p><p>The gap between Team A and Team B does not stay fixed either. It widens with every cycle, because the strong team&#8217;s structure keeps paying off while the weak team&#8217;s debt keeps charging interest. </p></div><h2>Traditional engineering versus agentic engineering</h2><p>The following diagram shows how the role of the engineer changes when agents become part of the development process. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KbUR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 424w, /__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 848w, /__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KbUR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png" width="660" height="440" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97e8df86-80de-42e8-847a-04bedab7012a_660x440.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:440,&quot;width&quot;:660,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 424w, /__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 848w, /__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KbUR!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e8df86-80de-42e8-847a-04bedab7012a_660x440.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><strong>Figure 2 - The agentic engineering paradigm, showing AI bounded by traditional structural constraints</strong></em></p><p>In traditional engineering, the developer takes an idea and turns it directly into production code. In agentic engineering, the engineer spends more time defining the goal, setting the right context, and checking the result, while the agent handles more of the code generation work. The work does not become fully automated, but the human role moves closer to direction, judgment, and quality control.</p><p><strong>Agentic software engineering</strong> is the practice of wrapping the raw, generative power of AI models in the strict, deterministic systems of traditional engineering: the tests, type checks, and pipelines that either pass or fail. We do not throw away the foundational rules of the craft. We apply them to constrain and guide the agent.</p><p><span>Understanding this definition requires drawing a clear line between agentic engineering and the more ad hoc approach often called vibe coding.</span></p><p><em>In Part 2 of <strong>The Agentic Engineering Playbook</strong>, we&#8217;ll explore why vibe coding feels almost magical for prototypes but falls apart in production, and how agentic engineering uses structure, context, and verification to turn AI into a reliable engineering partner.</em></p><div class="pullquote"><p><em><strong>&#128221; Spotted something worth calling out? We&#8217;d love to hear about it&#8212;just drop us a quick note through <a href="https://forms.cloud.microsoft/e/1wFeWTwewK">this 1-minute survey</a>. We&#8217;re all ears!</strong></em></p></div><h1>Further reading</h1><p><strong>&#10145;&#65039; <a href="https://www.amazon.com/Modern-Software-Engineering-Better-Faster-ebook/dp/B09GG6XKS4">Modern Software Engineering</a></strong> by Dave Farley is the direct source for the Pinnacle Thesis framing in this post. Farley distills the discipline into two core exercises: learning and exploration, and managing complexity, and defines principles that improve everything from your mindset to the quality of your code. Essential background if the ideas here resonated. </p><p><strong>&#10145;&#65039; <a href="https://x.com/karpathy/status/1886192184808149383?lang=en">Andrej Karpathy&#8217;s original &#8220;vibe coding&#8221; post</a></strong> (X, February 2, 2025) is the artifact that started the conversation. In the post, Karpathy described a new kind of coding where you &#8220;fully give in to the vibes, embrace exponentials, and forget that the code even exists,&#8221; made possible by how capable LLMs had become. Worth reading in its original form before reading any commentary about it. </p><p><strong>&#10145;&#65039; <a href="https://karpathy.bearblog.dev/sequoia-ascent-2026/">Karpathy&#8217;s Sequoia Ascent 2026 summary</a></strong> is Karpathy&#8217;s own cleaned-up write-up of his April 2026 fireside chat with Sequoia Capital, where he formally introduced agentic engineering as what comes after vibe coding. He describes agentic engineering as the professional discipline of coordinating fallible agents while preserving correctness, security, taste, and maintainability and draws a sharp line between that and casual prototyping. The best single primary source on the concept this series builds on. </p><p><strong>&#10145;&#65039; <a href="https://www.coderabbit.ai/blog/a-semantic-history-how-the-term-vibe-coding-went-from-a-tweet-to-prod">A Semantic History of Vibe Coding</a></strong> (CodeRabbit, March 2026) traces how the term evolved from a throwaway tweet into an industry flashpoint. The piece argues that as the phrase broadened into shorthand for any prompt-driven development, it became a signal that many engineering teams had added AI to their stacks and were struggling with the quality of its output and its downstream effects. Good context for why the distinction between vibe coding and agentic engineering matters in practice. </p><h1 style="text-align: center;"><strong>And that&#8217;s a wrap &#127916;</strong></h1><p><span>We&#8217;re glad you joined us for this edition of Build with AI!</span><br><br><span>If you have any thoughts, questions, or feedback on this edition, or if you&#8217;d like to share what you&#8217;d love to see next, feel free to take our </span><strong><a href="https://forms.cloud.microsoft/e/1wFeWTwewK">1-minute survey</a></strong><span>. We&#8217;d love to hear from you.</span><br><br><span>Thanks for following along! Until next time, keep learning and keep building.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[MCP, A2A, ACP — Three Protocols, Three Jobs, One Architecture Decision You Can't Ignore]]></title><description><![CDATA[The agent ecosystem now has three protocols. Each solves a different problem at a different layer of your stack. Confusing them isn't a technical debt. It's a design mistake.]]></description><link>https://packtbuildwithai.substack.com/p/mcp-a2a-acp-three-protocols-three</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/mcp-a2a-acp-three-protocols-three</guid><dc:creator><![CDATA[Charu Mitra Dubey]]></dc:creator><pubDate>Wed, 17 Jun 2026 11:45:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!z_fe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Welcome back to Build with AI.</strong></p><p>In our last edition we covered context windows and memory architecture &#8212; why stuffing transcripts into context is a design mistake, and how to build a proper retrieval layer instead. The response made one thing clear: developers are hungry for the mental model fix, not just the implementation.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This week, we fix another one.</p><p>If you&#8217;ve been building with AI agents in 2026, you&#8217;ve seen three acronyms everywhere: MCP, A2A, ACP. You&#8217;ve probably read that they&#8217;re &#8220;complementary.&#8221; You&#8217;ve probably also spent an afternoon in spec docs trying to figure out what that actually means for your architecture.</p><p>Here&#8217;s what most of that content doesn&#8217;t say: <strong>they solve completely different problems at completely different layers of your stack.</strong> Treating them as alternatives &#8212; or worse, picking one and ignoring the others &#8212; is how teams end up with agents that can use tools but can&#8217;t talk to each other, or agents that coordinate beautifully in isolation and fall apart the moment a third-party system is involved.</p><p>In this edition:</p><ul><li><p>The three-layer protocol stack &#8212; what MCP, A2A, and ACP each do, where they live, and why the &#8220;which one should I use&#8221; question is usually the wrong question</p></li><li><p>How to implement the two-layer stack most production teams actually need right now &#8212; MCP for tools, A2A for agent coordination</p></li><li><p>What&#8217;s worth paying attention to in AI this week</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!z_fe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 424w, /__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 848w, /__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 1272w, /__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!z_fe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png" width="1418" height="714" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:714,&quot;width&quot;:1418,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:125478,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/202393802?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 424w, /__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 848w, /__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 1272w, /__u/substackcdn.com/image/fetch/$s_!z_fe!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa7bd3a8-f213-45c7-90f1-eba5d84b1eab_1418x714.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>The Mental Model &#8212; Your agent stack now has a protocol layer. Here&#8217;s how it&#8217;s structured.</h3><p>Six months ago, MCP was the only protocol most AI engineers needed to know. That&#8217;s no longer true.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fec7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 424w, /__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 848w, /__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fec7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png" width="1134" height="964" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:964,&quot;width&quot;:1134,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:130669,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/202393802?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 424w, /__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 848w, /__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fec7!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c9363e-5b4c-4ba4-a560-ddb4e2d49814_1134x964.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The agent ecosystem has matured fast enough that three distinct protocols now exist &#8212; each solving a different problem at a different layer of your stack. The confusion isn&#8217;t that developers don&#8217;t understand the protocols. It&#8217;s that they&#8217;re comparing things that aren&#8217;t in competition.</p><p>Here&#8217;s the one-line version of each:</p><ul><li><p><strong>MCP</strong> &#8212; how your agent talks to tools</p></li><li><p><strong>A2A</strong> &#8212; how agents talk to each other</p></li><li><p><strong>ACP</strong> &#8212; how agents message each other when you want REST and nothing else</p></li></ul><p>Conflating any two of these is like asking whether HTTP or SQL is better. They operate at different layers. You don&#8217;t choose between them &#8212; you choose which combination your architecture needs.</p><p><strong>MCP: the tool layer</strong></p><blockquote><p>Model Context Protocol, launched by Anthropic in late 2024 and now governed by the Linux Foundation, is the foundational layer. It gives agents structured, authenticated access to external tools &#8212; databases, APIs, file systems, third-party services. Think of it as the USB-C standard for AI: one protocol, every tool.</p></blockquote><p>MCP is why your agent can query a database, send a Slack message, and read a PDF without you writing bespoke integration code for each. It&#8217;s the layer every team needs first, and for most single-agent applications, it&#8217;s the only layer they&#8217;ll ever need.</p><p>The numbers reflect this. MCP hit 97 million monthly SDK downloads and 10,000+ deployed servers by its first anniversary. OpenAI, Google, Microsoft, and AWS all adopted it. It is, at this point, infrastructure.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/mcp-a2a-acp-three-protocols-three?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/mcp-a2a-acp-three-protocols-three?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/p/mcp-a2a-acp-three-protocols-three?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p><strong>A2A: the coordination layer</strong></p><blockquote><p>Agent-to-Agent protocol, created by Google and donated to the Linux Foundation in June 2025, solves a different problem entirely. Not how an agent uses tools &#8212; but how agents hand off tasks to each other across systems, vendors, and organisational boundaries.</p></blockquote><p>A2A hit v1.0 in April 2026. That matters because it&#8217;s no longer a directional bet &#8212; it&#8217;s a real implementation target. The protocol defines how agents discover each other, negotiate capabilities, exchange messages, and track task state across a handoff. Think of it as HTTP for agent coordination: a universal communication layer that doesn&#8217;t care what framework or vendor built the agent on either end.</p><p>When do you need it? When your architecture has multiple autonomous agents that need to delegate to each other &#8212; especially across vendor or service boundaries. A single agent connected to ten MCP servers doesn&#8217;t need A2A. An orchestrator agent handing off a compliance check to a specialist agent built by a different team on a different framework does.</p><p><strong>ACP: the lightweight alternative</strong></p><blockquote><p>Agent Communication Protocol, developed by IBM and now under the Linux Foundation, covers similar ground to A2A but with a different design philosophy. Where A2A uses JSON-RPC 2.0 over HTTP and SSE, ACP is REST-native &#8212; it uses patterns most backend developers already know. No new transport to learn, no SSE to manage.</p></blockquote><p>ACP is the right choice when your team wants inter-agent messaging without the overhead of adopting a new protocol stack. It&#8217;s particularly well-suited to controlled environments where all agents are within your own infrastructure. For cross-vendor or cross-organisation coordination, A2A&#8217;s richer capability negotiation is worth the added complexity.</p><p><strong>The decision framework</strong></p><p>You don&#8217;t need all three. Most teams need one or two, depending on where they are in their agent architecture:</p><ul><li><p><strong>Single agent, multiple tools</strong> &#8594; MCP only. This is the majority of teams right now.</p></li><li><p><strong>Multiple agents, same infrastructure</strong> &#8594; MCP + ACP. Lightweight coordination without protocol overhead.</p></li><li><p><strong>Multiple agents, cross-vendor or cross-org</strong> &#8594; MCP + A2A. The two-layer stack that enterprise teams are converging on.</p></li><li><p><strong>All three</strong> &#8594; Only when you have internal agent messaging (ACP), cross-vendor coordination (A2A), and tool access (MCP) as distinct requirements. Rare, and usually premature.</p></li></ul><p>The shift isn&#8217;t complicated. But it requires treating the protocol layer as an architectural decision you make deliberately &#8212; not something you bolt on when your single-agent system starts failing at scale.</p><div class="callout-block" data-callout="true"><p>In The Build, we implement the two-layer stack most production teams actually need: MCP for tool access, A2A for agent coordination, with a concrete handoff pattern you can drop into your own architecture.</p></div><h3>The Build &#8212; Implementing the two-layer stack: MCP for tools, A2A for agent handoff</h3><p>Most teams don&#8217;t need all three protocols. They need two: MCP to give agents access to tools, and A2A to let agents hand off tasks to each other. This is the stack we&#8217;re building &#8212; an orchestrator agent that uses MCP to query a database, and delegates a specialist task to a second agent via A2A.</p><p>Stack: Node.js, the MCP SDK, and the A2A client library. The pattern is framework-agnostic &#8212; the same architecture works with LangChain, LlamaIndex, or a bare OpenAI client.</p><p><strong>The architecture before writing a line of code</strong></p><p>Two agents. One job each.</p><ul><li><p><strong>Orchestrator agent</strong> &#8212; receives the user&#8217;s request, uses MCP to access tools (database, APIs, file system), decides when a task is outside its own capability, and delegates via A2A.</p></li><li><p><strong>Specialist agent</strong> &#8212; exposes an A2A-compliant endpoint, receives delegated tasks, executes them, and returns structured results to the orchestrator.</p></li></ul><p>The handoff is the critical piece. The orchestrator doesn&#8217;t call the specialist directly &#8212; it sends an A2A task, waits for a response, and incorporates the result into its own context. The specialist has no knowledge of the orchestrator. They&#8217;re decoupled by design.</p><h3><strong>Step 1 &#8212; set up MCP tool access for the orchestrator</strong></h3><p>Install the MCP SDK and define your tool server. This example connects the orchestrator to a database tool.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;cc78052f-3898-4e7b-99f9-ec17f0b52d1e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';

async function createMCPClient() {
  const transport = new StdioClientTransport({
    command: 'node',
    args: ['./tools/database-server.js']
  });

  const client = new Client(
    { name: 'orchestrator', version: '1.0.0' },
    { capabilities: { tools: {} } }
  );

  await client.connect(transport);
  return client;
}

async function queryDatabase(mcpClient, query) {
  const result = await mcpClient.callTool({
    name: 'query_database',
    arguments: { sql: query }
  });
  return result.content[0].text;
}</code></pre></div><p>Your MCP tool server defines what the orchestrator can access. Keep each server focused &#8212; one server per domain (database, file system, external APIs). Don&#8217;t build a monolithic MCP server that does everything. The same principle applies here as with microservices: small, bounded, replaceable.</p><h3><strong>Step 2 &#8212; build the specialist agent&#8217;s A2A endpoint</strong></h3><p>The specialist exposes two endpoints: an agent card (capability discovery) and a task handler. A2A requires both.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;dab44e00-6fbf-4385-84d8-4f6abfcf5bb4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">import express from 'express';
import Anthropic from '@anthropic-ai/sdk';

const app = express();
app.use(express.json());

const anthropic = new Anthropic();

// Agent card &#8212; tells the orchestrator what this agent can do
app.get('/.well-known/agent.json', (req, res) =&gt; {
  res.json({
    name: 'compliance-specialist',
    version: '1.0.0',
    description: 'Runs compliance checks on structured data',
    capabilities: {
      streaming: false,
      pushNotifications: false
    },
    skills: [
      {
        id: 'compliance_check',
        name: 'Compliance check',
        description: 'Validates data against regulatory requirements',
        inputModes: ['application/json'],
        outputModes: ['application/json']
      }
    ]
  });
});

// Task handler &#8212; receives delegated tasks from the orchestrator
app.post('/tasks/send', async (req, res) =&gt; {
  const { id, message } = req.body;

  const userMessage = message.parts
    .filter(p =&gt; p.type === 'text')
    .map(p =&gt; p.text)
    .join('\n');

  const response = await anthropic.messages.create({
    model: 'claude-sonnet-4-6',
    max_tokens: 1000,
    system: `You are a compliance specialist. 
Analyse the provided data and return a structured 
compliance report. Be specific about violations 
and recommendations.`,
    messages: [{ role: 'user', content: userMessage }]
  });

  res.json({
    id,
    status: { state: 'completed' },
    artifacts: [{
      name: 'compliance-report',
      parts: [{
        type: 'text',
        text: response.content[0].text
      }]
    }]
  });
});

app.listen(3001, () =&gt; {
  console.log('Specialist agent running on port 3001');
});</code></pre></div><h3><strong>Step 3 &#8212; build the orchestrator&#8217;s A2A client</strong></h3><p>The orchestrator decides when to delegate. The decision logic is straightforward: if the task requires specialist capability, send it via A2A. Everything else, handle locally with MCP tools.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;b8aefc66-a55b-4734-b523-46c23c37f53e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">import { v4 as uuidv4 } from 'uuid';

async function delegateToSpecialist(specialistUrl, taskData) {
  const taskId = uuidv4();

  const response = await fetch(`${specialistUrl}/tasks/send`, {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      id: taskId,
      message: {
        role: 'user',
        parts: [{ type: 'text', text: JSON.stringify(taskData) }]
      }
    })
  });

  const result = await response.json();

  if (result.status.state !== 'completed') {
    throw new Error(`Task ${taskId} failed: ${result.status.state}`);
  }

  return result.artifacts[0].parts[0].text;
}

async function orchestrate(userRequest, mcpClient) {
  // Step 1: use MCP to fetch the data needed
  const data = await queryDatabase(
    mcpClient,
    `SELECT * FROM transactions WHERE date &gt; NOW() - INTERVAL '30 days'`
  );

  // Step 2: decide whether to delegate
  const needsComplianceCheck = userRequest
    .toLowerCase()
    .includes('compliance');

  if (needsComplianceCheck) {
    // Step 3: delegate via A2A
    const report = await delegateToSpecialist(
      'http://localhost:3001',
      { request: userRequest, data: JSON.parse(data) }
    );
    return report;
  }

  // Handle locally if no specialist needed
  return data;
}</code></pre></div><p>Step 4 &#8212; wire it together</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;2ba4aa80-aee4-4f1f-844f-be3440cd21db&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">async function main() {
  const mcpClient = await createMCPClient();

  const userRequest =
    'Run a compliance check on last month\'s transactions';

  const result = await orchestrate(userRequest, mcpClient);

  console.log('Result:', result);

  await mcpClient.close();
}

main().catch(console.error);</code></pre></div><p><strong>What this architecture gives you</strong></p><p>Three things that matter in production:</p><ul><li><p><strong>Decoupled specialists.</strong> The compliance agent doesn&#8217;t know who called it. You can swap it out, version it independently, or replace it with a third-party A2A-compliant service without touching the orchestrator.</p></li><li><p><strong>Clean capability boundaries.</strong> MCP handles tool access (what the agent can <em>read and write</em>). A2A handles task delegation (what the agent can <em>ask another agent to do</em>). Mixing these responsibilities into one layer is how architectures become unmaintainable.</p></li><li><p><strong>Protocol-level interoperability.</strong> Because A2A is an open standard, your specialist agent can receive tasks from any A2A-compliant orchestrator &#8212; not just yours. That&#8217;s the difference between integration and infrastructure.</p></li></ul><h3><strong>The decision your architecture needs to make explicitly</strong></h3><p>One question that comes up when building this: how does the orchestrator decide when to delegate versus handle locally?</p><p>Don&#8217;t use an LLM call for this decision &#8212; it&#8217;s slow and expensive for a routing problem. Use deterministic rules first: keyword matching, task type classification, input schema validation. Only reach for an LLM-based router if your routing logic is genuinely complex and rules-based approaches fail. Most teams over-engineer the routing layer and under-engineer the handoff contract.</p><p>The full two-layer stack &#8212; MCP tool access, A2A task delegation, deterministic routing &#8212; is around 120 lines of working code. The architecture decisions above are harder than the code. Get the boundaries right and the implementation follows.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The Radar</h2><h3>01. A CVSS 9.8 vulnerability just dropped in an MCP integration &#8212; and it changes how you think about tool access.</h3><p>MCP is infrastructure now. And infrastructure gets attacked. A critical vulnerability was disclosed in May 2026 in an MCP integration &#8212; the kind of severity score that means remote code execution is on the table. The practical implication for teams building on MCP: your tool servers are now an attack surface, and the default assumption that MCP integrations are internal and therefore safe is wrong. Two things to do now: audit every MCP server you&#8217;re running for authentication gaps, and treat your tool server permissions the same way you&#8217;d treat database credentials &#8212; least privilege, scoped access, rotated regularly. The protocol is solid. The implementations are where the risk lives.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://www.cve.org">MCP Security Vulnerability Disclosure &#8212; May 2026</a></p></div><h3>02. A2A hit v1.0 in April &#8212; which means agent-to-agent handoffs are now a real implementation target, not a directional bet.</h3><p>For the past year, A2A has been something teams watched but didn&#8217;t build on. The 0.3 preview spec was unstable enough that production commitments felt premature. That changed in April 2026. V1.0 means a stable spec, committed SDK support, and &#8212; most importantly &#8212; enterprise vendors treating it as infrastructure. SAP has integrated A2A into its Joule AI assistant, which runs ERP systems for the majority of the Global Fortune 500. When SAP commits to a protocol, it becomes the de facto standard for enterprise agent coordination across finance, supply chain, and procurement. If you&#8217;ve been waiting for A2A to stabilise before building on it, the wait is over.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://a2a-protocol.org">A2A Protocol v1.0 Release &#8212; Linux Foundation</a></p></div><h3>03. Half of all enterprise agents still can&#8217;t talk to each other &#8212; and that&#8217;s now a competitive gap, not just a technical debt.</h3><p>The average enterprise runs 12 AI agents. Half of them operate in complete isolation. That number comes from Salesforce&#8217;s 2026 Connectivity Benchmark Report, and it reframes the protocol conversation entirely. The teams asking &#8220;should we implement A2A?&#8221; are already behind. The right question is &#8220;how quickly can we retrofit coordination into our existing agent architecture?&#8221; The answer, for most teams, is: faster than you think, if you start with the two-layer stack from this edition&#8217;s Build section and add A2A to the agents that actually need to delegate tasks. Don&#8217;t boil the ocean. Pick one handoff that matters and make it protocol-compliant first.</p><div class="callout-block" data-callout="true"><p style="text-align: center;">Read it here: <a href="https://salesforce.com">Salesforce 2026 Connectivity Benchmark Report</a></p></div><h2>AI Tools of the Week</h2><h3><a href="https://github.com/modelcontextprotocol/inspector">MCP Inspector</a>: Debug your MCP servers visually</h3><p>The tool you wish existed the first time an MCP server silently failed. MCP Inspector is an open-source visual debugger that lets you connect to any MCP server, browse its available tools, fire test calls, and inspect the raw request/response cycle &#8212; all without writing a line of code. Think of it as Postman for MCP. If you&#8217;re building the two-layer stack from this edition, this is the first thing to install before you write a single tool server. Catching a malformed tool schema in Inspector takes thirty seconds. Catching it in production takes thirty minutes and a lot of logs.</p><h3><a href="https://agentboard.dev">Agentboard</a>: Observability for multi-agent systems</h3><p>The missing monitoring layer for teams running more than one agent in production. Agentboard traces task flows across agents, visualises A2A handoffs, and surfaces latency and failure points in multi-step pipelines. The dashboard shows you which agents are doing the most work, where tasks are stalling, and which handoffs are failing silently &#8212; the exact failure modes that are invisible when you&#8217;re logging individual agents but not the coordination layer between them. Free tier covers up to 5 agents and 10,000 traces per month. Worth connecting before your agent count grows past the point where mental models stop working.</p><h3><a href="https://www.getzep.com">Zep</a>: Temporal knowledge graph for agent memory</h3><p>Worth calling out again in this context: as your agent architecture grows from one agent to many, memory becomes a coordination problem, not just a storage problem. Which agent remembered what? Which facts are still valid? Zep&#8217;s Graphiti engine stores fact validity windows rather than timestamped snapshots &#8212; it knows that a fact <em>was</em> true until a certain point, not just that it <em>was stated</em> at a certain point. In a multi-agent system where different agents update the same knowledge base at different times, that distinction prevents an entire class of stale-context bugs. If the memory edition two weeks ago convinced you to build a retrieval layer, Zep is the right tool for the coordination layer that comes next.</p><h3><a href="https://betterclaw.io">Betterclaw</a>: Managed agent infrastructure with MCP and A2A built in</h3><p>The managed option for teams that want the two-layer stack without running the infrastructure themselves. Betterclaw handles MCP server provisioning, A2A endpoint routing, agent identity, and capability discovery out of the box &#8212; the four things that take the most setup time when you build the stack from scratch. Verified skills, encrypted secrets, and a free-forever tier make it a reasonable starting point for smaller teams validating the architecture before committing to a self-hosted setup. If the Build section felt like a lot of moving parts, start here and work backwards to understand what you eventually want to own.</p><div><hr></div><p>If you made it this far, your agent architecture now has a protocol layer &#8212; and you understand why it exists.</p><p>Most teams shipping agents in 2026 are still treating MCP as the whole story. They&#8217;ll hit the coordination ceiling eventually: agents that can use tools but can&#8217;t talk to each other, architectures that scale to three agents and then become unmaintainable, integrations that break every time a third-party vendor changes their API. You just opted out of that trajectory.</p><p>You have the mental model &#8212; three protocols, three layers, three different jobs. You have the implementation &#8212; MCP for tool access, A2A for task delegation, deterministic routing between them. And you have the decision framework for knowing which combination your architecture actually needs, today, without over-engineering.</p><p>If this edition changed how you think about agent interoperability &#8212; forward it to the engineer on your team who&#8217;s still asking &#8220;MCP or A2A?&#8221; They&#8217;re asking the wrong question. Now you can tell them why.</p><p>See you next week.</p><p><strong>Charu Mitra Dubey</strong><br>Editor-in-chief, Build with AI</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Your AI server just asked the AI for help. Here's why that's not as strange as it sounds.]]></title><description><![CDATA[MCP sampling is one of the protocol's most underexplained features &#8212; and one of the most useful. It's about delegation: when a server hands part of its job to the client's LLM, and why that changes wh]]></description><link>https://packtbuildwithai.substack.com/p/your-ai-server-just-asked-the-ai</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/your-ai-server-just-asked-the-ai</guid><dc:creator><![CDATA[BuildWithAI]]></dc:creator><pubDate>Wed, 10 Jun 2026 08:59:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0z30!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most people think of an MCP server as something that <em>responds</em>. A client calls a tool, the server does the work, the server sends back a result. Clean, linear, predictable.</p><p>Sampling breaks that model &#8212; and in doing so, opens up a class of use cases that the standard request-response flow simply can&#8217;t handle.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Here&#8217;s the core idea: sometimes a server needs to complete a task that it isn&#8217;t equipped to do on its own. Generating a compelling product description. Summarizing a blog post draft. Giving voice to a game character. These are generative tasks &#8212; and generative tasks need an LLM. The server doesn&#8217;t have one. But the client does.</p><p>So the server asks. That ask is called a sampling request.</p><div class="callout-block" data-callout="true"><p><strong>In this post:</strong></p><ol><li><p>What sampling is &#8212; and what problem it actually solves</p></li><li><p>The four participants in a sampling flow</p></li><li><p>Three scenarios where sampling makes the work better</p></li><li><p>What&#8217;s actually inside a sampling request</p></li><li><p>Building it: server side and client side</p></li><li><p>The human-in-the-loop piece that most explanations skip</p></li></ol></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.packtpub.com/en-in/product/learn-model-context-protocol-with-python-9781806103232" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 424w, /__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 848w, /__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0z30!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1265594,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.packtpub.com/en-in/product/learn-model-context-protocol-with-python-9781806103232&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/201414207?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 424w, /__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 848w, /__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0z30!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e70f1f2-297d-467d-81e6-3a2674cafa53_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>What sampling is &#8212; and what problem it actually solves</h2><p>The word comes from the dictionary definition: &#8220;the action or process of taking samples of something for analysis.&#8221; In MCP, the server sends a sample &#8212; a prompt, a piece of content, a task &#8212; to the client for analysis. The client&#8217;s LLM processes it and sends back a response. The server then uses that response to complete whatever it was working on.</p><p>This is delegation, not dependency. The server stays in control of the overall task. It just hands off the parts that require generative intelligence &#8212; the parts it genuinely can&#8217;t do as well on its own.</p><p>The pattern is more common than it might first appear. Any time a system needs to produce language &#8212; descriptions, summaries, tags, dialogue &#8212; and that language needs to be good, there&#8217;s a case for sampling.</p><div class="callout-block" data-callout="true"><p></p></div><h2>The four participants</h2><p>A sampling interaction involves four distinct roles, and it&#8217;s worth keeping them clear:</p><p><strong>The user</strong> is involved at two points: as the person who triggered the original action, and as the human in the loop who can review or modify the sampling request before it goes to the LLM. This second touchpoint is important &#8212; it&#8217;s what keeps sampling human-supervised rather than fully automated.</p><p><strong>The server</strong> sends the sampling request. This happens from within a server feature &#8212; typically a tool call, a resource read, or a prompt template. The server is the one that recognizes it needs LLM help and initiates the delegation.</p><p><strong>The client</strong> receives the sampling request and passes it to the user for review. The server can suggest a specific model, token limit, and system prompt &#8212; but these are recommendations, not instructions. The client and the user have the final say.</p><p><strong>The LLM</strong> on the client side does the actual work. It takes the prompt from the server, runs its generative process, and returns a response that the client sends back.</p><div class="callout-block" data-callout="true"><p><strong>The flow is:</strong> user action &#8594; server processes &#8594; server needs LLM help &#8594; sampling request &#8594; client &#8594; user reviews &#8594; LLM generates &#8594; response back to server &#8594; server completes the task.</p></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XV0E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 424w, /__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 848w, /__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XV0E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png" width="352" height="325.7838328792007" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1019,&quot;width&quot;:1101,&quot;resizeWidth&quot;:352,&quot;bytes&quot;:1731659,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/201414207?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c9a5e93-31bc-4b5b-aec4-7963ecd8b424_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 424w, /__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 848w, /__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XV0E!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2b2245f-3c34-4054-ba44-fa11ab7fb39e_1101x1019.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;">Author: Christoffer Noring</p><h2>Three scenarios where this matters</h2><p><strong>Writing a blog post.</strong> A user submits a draft to the server. The server stores it, but realizes it needs tags and a summary abstract &#8212; the kind of work an LLM does better than a rules-based function. A sampling request goes out with the draft as context. The client&#8217;s LLM analyzes it, generates the metadata, and sends it back. The server updates the post and returns a complete result.</p><p><strong>Back office e-commerce.</strong> An admin adds a new product: name, keywords, a few properties. Writing the description is time-consuming and easy to get wrong. The server uses those keywords as context in a sampling request and asks the client to generate something compelling. The result gets attached to the product automatically. No copy-paste, no separate tool.</p><p><strong>NPC dialogue in a game.</strong> Non-player characters in most games are limited by pre-scripted responses. With sampling, a server can retrieve a character&#8217;s backstory, personality, and motivations, and send all of that as a sampling request. The client&#8217;s LLM produces a contextually appropriate response &#8212; one that feels alive rather than canned. The same character can have a different conversation every time.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>What&#8217;s actually inside a sampling request</h2><p>When the server sends a sampling request, it isn&#8217;t just passing a prompt into the void. The request is structured, and the structure gives you real control. Here&#8217;s what the key fields look like:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;efc36aeb-c0a9-453f-92c5-601ceec08961&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">{

  &#8220;method&#8221;: &#8220;sampling/createMessage&#8221;,

  &#8220;params&#8221;: {

    &#8220;messages&#8221;: [

      {

        &#8220;role&#8221;: &#8220;user&#8221;,

        &#8220;content&#8221;: {

          &#8220;type&#8221;: &#8220;text&#8221;,

          &#8220;text&#8221;: &#8220;Write a compelling description of this product: tomato, here&#8217;s some keywords: red, vegetable, fresh&#8221;

        }

      }

    ],

    &#8220;modelPreferences&#8221;: {

      &#8220;hints&#8221;: [{ &#8220;name&#8221;: &#8220;claude-3-sonnet&#8221; }],

      &#8220;intelligencePriority&#8221;: 0.8,

      &#8220;speedPriority&#8221;: 0.5

    },

    &#8220;systemPrompt&#8221;: &#8220;You&#8217;re a professional writing assistant who writes descriptions in a poetic way&#8221;,

    &#8220;maxTokens&#8221;: 100

  }

}</code></pre></div><p>A few things worth noting here. The <code>modelPreferences</code> field lets the server <em>suggest</em> a model and express priorities around intelligence versus speed &#8212; but the client doesn&#8217;t have to honor it. The <code>systemPrompt</code> shapes how the LLM responds, which means the server can give the LLM a persona or a constraint specific to the task. And <code>maxTokens</code> keeps output scoped to what the task actually requires.</p><h2>Building it: server side and client side</h2><p><strong>On the server</strong>, the pattern is straightforward. Inside a tool, after the initial processing, you call <code>ctx.session.create_message()</code> with a <code>SamplingMessage</code> containing your prompt. The call is <code>await</code> &#8212; the server pauses and waits for the client to come back with a response before continuing.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;802b2117-ce58-4359-8e53-0d0b42f6cb47&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">result = await ctx.session.create_message(
    messages=[
        SamplingMessage(
            role="user",
            content=TextContent(type="text", text=prompt),
        )
    ],
    max_tokens=100,
)
product["description"] = result.content.text</code></pre></div><p>The server doesn&#8217;t know or care which LLM the client uses. It just sends the request and handles the result.</p><p><strong>On the client</strong>, two things need to happen. First, you declare sampling support when initializing the session:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;2c5b2dad-b097-4b85-b46e-730dd229f697&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">{ "capabilities": { "sampling": {} } }</code></pre></div><p>Second, you register a callback that handles incoming sampling requests:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;e00c4b19-dfdc-4522-b4ce-c2526a38c4e3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">async with ClientSession(read, write,
    sampling_callback=handle_sampling_message) as session:</code></pre></div><p>The callback receives the sampling request, extracts the prompt, calls the LLM, and returns a structured response. The client controls which LLM is used &#8212; it could be GPT-4o, Claude, a local model, anything. The server has no say in the implementation detail, only in the task description.</p><p>Here&#8217;s what the handler looks like in practice:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c42c7192-c3cd-4973-95be-2ac5fa480495&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">async def handle_sampling_message(context, params):
    message = params.messages[0].content.text
    response = await call_llm(message, "You're a helpful assistant. Create a compelling product description.")
    return types.CreateMessageResult(
        role="assistant",
        content=types.TextContent(type="text", text=response),
        model="gpt-4o-mini",
        stopReason="endTurn",
    )</code></pre></div><p>When this runs end-to-end &#8212; a user prompts &#8220;create product called tomato with keywords red and vegetable and delicious&#8221; &#8212; the server calls the tool, issues the sampling request, and comes back with something like:</p><blockquote><p><em>&#8220;Introducing our Red Garden Medley &#8212; a vibrant selection of the freshest, most delicious red vegetables nature has to offer! Each hand-picked assortment features juicy tomatoes, crisp red bell peppers, and sweet red radishes, bursting with flavor and color.&#8221;</em></p></blockquote><p>That&#8217;s not a template. That&#8217;s a generative response, produced on demand, shaped by the keywords the admin entered.</p><h2>The human-in-the-loop piece that most explanations skip</h2><p>Sampling isn&#8217;t fire-and-forget. The design of the protocol explicitly includes a point where the user can inspect the sampling request before it runs &#8212; reviewing the prompt, adjusting the model preference, modifying token limits. This is intentional.</p><p>The server&#8217;s sampling request is a <em>recommendation</em>, not a command. The client and the user are always in control of what actually gets sent to the LLM. That framing matters for anyone thinking about production use cases: sampling keeps humans in the decision loop by design, not as an afterthought.</p><p>This is the same principle that governs tool approval in hosts like VS Code. Every step where an AI system is about to do something consequential gets a checkpoint. Sampling applies that same logic to generative sub-tasks.</p><p>The deeper pattern here is something that shows up across well-designed AI systems: knowing where to delegate. A server that tries to do everything itself will either be limited or brittle. A server that understands what it&#8217;s good at &#8212; structured data, persistence, tool orchestration &#8212; and knows when to hand off to generative intelligence will be both more capable and more maintainable.</p><p>Sampling is how MCP formalizes that handoff.</p><div class="callout-block" data-callout="true"><p><em>This post is adapted from <strong>Learn Model Context Protocol with Python: Build agentic systems in Python with the new standard for AI capabilities</strong> by Christoffer Noring &#8212; engineer at Microsoft, Oxford tutor, and the person who somehow makes protocol internals feel approachable.</em></p><p><em>The full book covers everything from building STDIO and SSE servers from scratch, to elicitation, security, and production deployment &#8212; with working Python code and assignments throughout.</em></p><p>Get the book &#8594; <a href="https://www.amazon.com/Learn-Model-Context-Protocol-Python/dp/1806103230/ref=tmm_pap_swatch_0?_encoding=UTF8&amp;dib_tag=se&amp;dib=eyJ2IjoiMSJ9.ohfyUjfDPBcFVAwpQzJKTI47EmNe2u6L3H1BE0oEJvZMpG249pf-vbjkkO-hub1V_9iCPNhFg9ZdmDz6tHUkRkW7jm3OJSI19V_p9gZUkKVStFzTTfrZ1kKiL2DiROq-EshH6EwoAm-thJ0441SyTZo3U4VwoSaUyZxFRQIvK7IMe9TA7Tl8nsy0JT_JFQYE0VYbzbVqw2oGx6Z79QSplw.Lkkic50yphgzik69snbr4CuQvjejien9iI0L3fFh0hg&amp;qid=1781074837&amp;sr=1-1">Amazon</a> | <a href="https://www.packtpub.com/en-in/product/learn-model-context-protocol-with-python-9781806103232">Packt</a></p></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Dev Workflow is Broken. Vibe Coding Fixes the Right Parts.]]></title><description><![CDATA[What actually changes &#8212; and what doesn&#8217;t &#8212; when AI enters every stage of your development process]]></description><link>https://packtbuildwithai.substack.com/p/the-dev-workflow-is-broken-vibe-coding</link><guid isPermaLink="false">https://packtbuildwithai.substack.com/p/the-dev-workflow-is-broken-vibe-coding</guid><dc:creator><![CDATA[Charu Mitra Dubey]]></dc:creator><pubDate>Wed, 03 Jun 2026 09:37:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Gs8a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most developers treat vibe coding as a replacement for disciplined software development. It isn&#8217;t. It&#8217;s a remix of it. And if you don&#8217;t understand the original before you remix it, you&#8217;re going to ship something fragile &#8212; just faster than before.</p><div class="callout-block" data-callout="true"><p><em>This edition was adapted from</em> <strong><a href="https://www.packtpub.com/en-us/product/vibe-coding-with-cursor-windsurf-and-lovable-9781807301637">Vibe Coding with Cursor, Windsurf, and Lovable</a> </strong><em>by <strong>Greg Lim</strong>. </em></p></div><div class="callout-block" data-callout="true"><p><strong>In this post</strong></p><ol><li><p>Why the SDLC still matters in an AI-assisted workflow</p></li><li><p>Planning: the stage AI can&#8217;t replace</p></li><li><p>Design: the decisions that still belong to you</p></li><li><p>Implementation: where AI actually earns its keep</p></li><li><p>Testing: the part most vibe coders skip</p></li><li><p>Deployment and maintenance: what changes, what doesn&#8217;t</p></li><li><p>How to pick a tech stack when AI lowers every barrier</p></li></ol></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Gs8a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_webp, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Gs8a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1292649,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://packtbuildwithai.substack.com/i/200410813?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_424, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_848, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_1272, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gs8a!, /__u/packtbuildwithai.substack.com/w_1456, /__u/packtbuildwithai.substack.com/c_limit, /__u/packtbuildwithai.substack.com/f_auto, /__u/packtbuildwithai.substack.com/q_auto:good, /__u/packtbuildwithai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849a5d63-4851-4d9f-bd9a-8ce3d9374dfa_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>Why the SDLC still matters in an AI-assisted workflow</h2><p>The software development lifecycle has six stages. Planning, design, implementation, testing, deployment, maintenance. You know them.</p><p>For decades, the bottleneck was implementation. Writing the code was the slow, expensive part. Everything else &#8212; planning, design, testing &#8212; existed to protect the code from going wrong.</p><p>Now that AI can write code faster than most engineers, the bottleneck has shifted. But the stages haven&#8217;t disappeared. What&#8217;s changed is <em>where the leverage is</em> &#8212; and where the new failure modes are hiding.</p><p>Here&#8217;s what that looks like across each stage.</p><h2>Planning: the stage AI can&#8217;t replace</h2><p>This one doesn&#8217;t change much. And that&#8217;s the point.</p><p>Before AI, bad planning led to slow, expensive rewrites. With AI, bad planning leads to <em>fast, cheap, wrong</em> builds that still need expensive rewrites.</p><p>The speed of AI implementation doesn&#8217;t make upfront planning less valuable. It makes it more valuable. You can now spiral into the wrong thing at 10x the velocity.</p><blockquote><p>What good vibe coding looks like here: use an LLM as a consultation partner <em>before</em> you write a single prompt to your IDE. Ask it to poke holes in your idea. Ask it what you haven&#8217;t considered. Treat it like a senior engineer in the room &#8212; not a code dispenser.</p></blockquote><p>The planning stage is where you go from idea to requirements. Who is this for? What do they need? What does the project actually involve, scoped into concrete tasks? Don&#8217;t underestimate it. It&#8217;s the foundation for everything that follows.</p><h2>Design: the decisions that still belong to you</h2><p>Architecture decisions still belong to you.</p><p>Which tech stack? How does the data model look? What APIs are you relying on? Where do you draw the service boundaries? What does the UI flow look like?</p><p>AI will happily make all of these decisions for you if you let it. And it&#8217;ll make them <em>confidently</em> &#8212; and incorrectly for your specific context.</p><blockquote><p>The vibe coding trap here is mistaking speed for correctness. A poorly architected app built in 30 minutes is still a poorly architected app. The AI didn&#8217;t know about your team&#8217;s constraints, your infra costs, or the edge cases in your domain. You do.</p></blockquote><p>The practical move: bring your architecture questions to an LLM before you open your AI IDE. Use it to stress-test decisions, not make them. Your codebase will thank you six weeks from now.</p><h2>Implementation: where AI actually earns its keep</h2><p>This is where AI earns its keep. Seriously.</p><p>Code generation has gotten genuinely good. If you have a clear specification and a well-scoped task, a tool like Cursor or Windsurf can implement entire features in the time it would take you to write the function signature.</p><p>Two things still matter though.</p><blockquote><p>One: the specification has to exist first. Zero-shot prompts (&#8221;build me a Kanban board&#8221;) produce zero-shot quality. The AI isn&#8217;t bad. The instruction was bad. Invest time upfront in a proper spec document &#8212; what the app should look like, how it should behave, what technologies to use. Then break it into a phased TODO list and work incrementally. One feature at a time.</p></blockquote><blockquote><p>Two: you still have to read the code. AI is non-deterministic. The same prompt produces different outputs on different runs. Treating generated code as ground truth without reviewing it is how bugs get shipped quietly &#8212; and in an AI-assisted codebase, the blast radius of a broken change is larger because the AI has touched every file.</p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><em>Subscribe to BuildWithAI for weekly breakdowns of AI tools and workflows for engineers.</em></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Testing: the part most vibe coders skip</h2><p>Don&#8217;t skip it.</p><p>Automated tests aren&#8217;t bureaucracy. They&#8217;re the thing that lets you keep moving fast <em>after</em> the first 20% of the build. Every feature you add without a corresponding test is debt you&#8217;ll pay in debugging time later.</p><blockquote><p>The good news: AI is useful for writing tests too. Ask your tool of choice to write end-to-end tests for every feature you complete. Make it a rule, not an afterthought.</p></blockquote><p>Some tools can be configured to run tests automatically after every change so you never have to remember to ask. In Cursor, for example, you can set a project rule &#8212; &#8220;after making changes to the codebase, run all tests to ensure they pass&#8221; &#8212; and it&#8217;ll happen without you prompting it each time. That&#8217;s not a nice-to-have. That&#8217;s how you stay confident as the codebase grows.</p><p>The types worth knowing: unit tests cover individual functions, integration tests check how components work together, and end-to-end tests simulate real user interactions. An AI-assisted project should have all three. If you&#8217;re building fast, start with E2E tests for your critical paths.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/the-dev-workflow-is-broken-vibe-coding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/p/the-dev-workflow-is-broken-vibe-coding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/packtbuildwithai.substack.com/p/the-dev-workflow-is-broken-vibe-coding?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h2>Deployment and maintenance: what changes, what doesn&#8217;t</h2><p><strong>Deployment</strong></p><p>The pipeline hasn&#8217;t changed structurally. Staging environments, version control, production releases &#8212; the process is the same.</p><blockquote><p>What&#8217;s changed: AI can now walk you through the entire deployment process conversationally, including catching build errors and suggesting fixes in real time. The friction that used to live here &#8212; figuring out how to configure Vercel, what the build flags need to be, what the CI error means &#8212; is largely gone. It&#8217;s now a conversation, not a documentation spelunk.</p></blockquote><p><strong>Maintenance and iteration</strong></p><p>This is the stage that breaks most AI-assisted projects.</p><p>Bug reports come in. Features need updating. Dependencies get flagged. And now you&#8217;re back in the codebase &#8212; except you didn&#8217;t write most of it.</p><p>The engineers who do this well are the ones who maintained a specification document throughout the build. They kept their TODO list updated. They reviewed every line the AI generated. They committed frequently to version control so rolling back to a working state is one prompt away.</p><blockquote><p>Vibe coding done right is still disciplined coding. The discipline just lives in different places now &#8212; in the spec, the commit history, the test suite, the rules file. Move those things and the maintenance phase becomes painful regardless of how fast the build went.</p></blockquote><h2>How to pick a tech stack when AI lowers every barrier</h2><p>Pick the stack you know, not the stack that&#8217;s trending.</p><p>This sounds obvious. It isn&#8217;t practiced.</p><p>AI tools are seductive here. They&#8217;ll happily scaffold a project in any language or framework you name. That doesn&#8217;t mean you should reach for a new stack just because the barrier to entry is lower now.</p><p>The engineers who ship fastest with AI-assisted tools are the ones using stacks they already understand. They can catch the AI&#8217;s mistakes. They know what &#8220;correct&#8221; looks like. They can debug when the AI confidently produces something broken.</p><blockquote><p>If you&#8217;re a JavaScript/TypeScript person, stay there. Full-stack with React on the frontend and Node on the backend is still one of the most productive setups for AI-assisted development &#8212; one language end-to-end, massive community, tons of tooling. If you&#8217;re a Python person, your backend instincts plus AI assistance is a strong combination, and Python&#8217;s dominance in AI/ML means the ecosystem keeps getting better.</p></blockquote><p>New stack, new language, AI-assisted build, tight deadline &#8212; that&#8217;s four variables you don&#8217;t want active at the same time.</p><h3>The one-line summary</h3><p>Vibe coding shifts the bottleneck from <em>writing</em> code to <em>thinking clearly</em> about what to build.</p><p>The SDLC isn&#8217;t a relic. It&#8217;s the foundation that makes AI-assisted development actually work. Skip the stages you think AI has replaced and you&#8217;ll find out exactly why they existed.</p><div class="callout-block" data-callout="true"><p><em>This edition was adapted from</em> <strong>Vibe Coding with Cursor, Windsurf, and Lovable </strong><em>by <strong>Greg Lim</strong>. </em></p><p><em>The full book covers every stage hands-on &#8212; building a Math Practice app with Cursor, a Kanban board with Lovable, and working on existing codebases with Windsurf.</em></p><p><em>Get the book &#8594; <a href="https://www.packtpub.com/en-us/product/vibe-coding-with-cursor-windsurf-and-lovable-9781807301637">Packt</a> | <a href="https://www.amazon.com/Vibe-Coding-Cursor-Windsurf-Lovable/dp/180730163X/ref=tmm_pap_swatch_0">Amazon</a></em></p></div><p></p><p style="text-align: center;"><strong>Get the book. On us.</strong></p><p style="text-align: center;"><strong>We&#8217;re giving away a free copy of </strong><em><strong>Vibe Coding with Cursor, Windsurf, and Lovable</strong></em><strong> by Greg Lim to every new subscriber of BuildWithAI.</strong></p><p style="text-align: center;"><strong>Subscribe below and we&#8217;ll send it your way.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://packtbuildwithai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>