<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Enterprise Context Management]]></title><description><![CDATA[Enterprise Context Management is an emerging technology that helps enterprises extract and apply context, turning everyone’s AI into your AI.]]></description><link>https://enterprisecontextmanagement.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!sf2k!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2553112d-caae-469a-8fb4-58e47a82db75_1126x1126.png</url><title>Enterprise Context Management</title><link>https://enterprisecontextmanagement.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 05:01:33 GMT</lastBuildDate><atom:link href="/__u/enterprisecontextmanagement.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[AI One]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[enterprisecontextmanagement@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[enterprisecontextmanagement@substack.com]]></itunes:email><itunes:name><![CDATA[AI One]]></itunes:name></itunes:owner><itunes:author><![CDATA[AI One]]></itunes:author><googleplay:owner><![CDATA[enterprisecontextmanagement@substack.com]]></googleplay:owner><googleplay:email><![CDATA[enterprisecontextmanagement@substack.com]]></googleplay:email><googleplay:author><![CDATA[AI One]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Your AI Agent Is Slowly Turning Into a DAG]]></title><description><![CDATA[When your prompt starts looking like application logic, it probably is.]]></description><link>https://enterprisecontextmanagement.substack.com/p/your-ai-agent-is-slowly-turning-into</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/your-ai-agent-is-slowly-turning-into</guid><dc:creator><![CDATA[Emanuele Melis]]></dc:creator><pubDate>Wed, 26 Aug 2026 18:10:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8mZS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Agent-first development is a fast way to discover a workflow. As the workflow becomes predictable, stable execution should move into code, while the model remains available for the cases where fixed rules and conventional software are no longer enough. But how do you tell when that&#8217;s happening?</strong><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8mZS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8mZS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ChatGPT Image Aug 5, 2026 at 11_28_27 PM (1).png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="ChatGPT Image Aug 5, 2026 at 11_28_27 PM (1).png" title="ChatGPT Image Aug 5, 2026 at 11_28_27 PM (1).png" srcset="/__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8mZS!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e687ef5-499a-479d-82a2-af7d58b2b0dc_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We have all seen this movie while building an agent.<br><br>You begin with a question, refine the prompt, add another instruction, and try again. The answers improve until the system appears capable of completing useful work. It may still be rough, but it has moved beyond a demo and started to resemble an agent.<br>Then you let it run.<br><br>You change the prompt here, add a tool there, and introduce a skill somewhere in the middle. After a few iterations, the model is selecting tools, deciding how to sequence them, passing results between steps, and working out what to do when the expected output does not arrive.<br><br>I started calling this the <strong>LLM-as-orchestrator</strong> pattern.<br><br>It is often a very effective way to start the build because at the beginning of a project, you rarely understand the workflow well enough to define every branch, dependency, and failure mode. Letting the model attempt the task helps expose what the process actually requires.<br><br>The agent reveals missing tools, recurring decisions, awkward handoffs, and inputs that are less structured than everyone originally claimed, so you learn the workflow by watching an agent try to execute it.<br><br>The architecture, however, needs to evolve as that learning accumulates. Often, the agent begins following the same sequence across most runs, and much of its work becomes predictable execution. Keeping the LLM in control of every step then starts to add token cost, latency, and behavioral variance to a path the team already understands.<br><br>The useful production question is therefore fairly simple:</p><blockquote><p>Which parts of the agent&#8217;s workflow have become ordinary software, and where does the code still run out of road?</p></blockquote><p>And when that happens, what are the common signs, or smells, to look out for?<br></p><h3>When the prompt starts looking suspiciously like code</h3><p>You can usually see the transition by reading the prompt.<br><br>Early prompts describe an objective and provide enough context for the model to attempt the task. Later prompts contain increasingly precise instructions about which tool should run first, what fields should be extracted, how the output should be transformed, which conditions should trigger another action, and what should happen when a value falls outside an expected range.<br><br>The prompt gradually assumes responsibility for application logic.<br><br>At some point, it is worth asking whether much of that agent behavior would be clearer and safer in Python. Other languages remain available, apparently, although Python appears to have won the argument by refusing to leave.<br><br>This is where the phrase &#8220;prompt engineering&#8221; occasionally gives us more comfort than it deserves.<br><br>A prompt can contain substantial application logic while providing very few of the controls normally associated with software engineering: interfaces remain implicit and testing often focuses on the final answer rather than the route taken to produce it. A change in one instruction can affect behavior somewhere else, and the regression may remain hidden because the output still <em>sounds</em> plausible.<br><br>Plausibility is one of the less charming failure modes of language models. The system can produce a coherent answer after skipping a required check, choosing the wrong tool, or passing incomplete information downstream.<br><br>As the workflow becomes clearer, however, repeated behavior should start moving out of the prompt. That is when known transformations can become functions, stable payloads can become typed interfaces, and validation can move into schemas and assertions. At this point repeated sequences can be handled through deterministic orchestration and the model should remain where fixed rules, schemas, and ordinary control flow cannot resolve the case cleanly.<br></p><h3>The accidental DAG</h3><p>This transition normally happens one piece at a time.<br><br>A repeated instruction becomes a tool, a model-generated object becomes a structured payload, and one LLM call disappears because a rule can now produce the same result reliably.<br><br>Run it long enough, and eventually the traces show that the agent has called the same seven tools in the same order for the last twenty successful runs.<br><br>If you squint hard enough, this can resemble a self-improving loop.<br><br>It is, in a sense. You still have to inspect the traces, identify the pattern, extract it into code, test it, deploy it, and keep it alive while the system contributes mainly by generating the evidence and the invoice. Alternatively, you automate this with a &#8220;grade and trajectory analysis&#8221;, but in either case the agent is likely to converge into making ten tool calls back to back, passing the result of one into the next, before using a final model call to translate, classify, or summarize the output.<br><br>After that, you need to decide when the agent runs, so you add scheduling.<br>When it wakes up and there is nothing to process, you pay for the privilege of discovering that nothing happened.<br><br>When something important occurs between scheduled runs, processing is delayed.<br><br>When the workflow fails halfway through, you need state, retries, idempotency, and a way to recover from partial execution.<br><br>By this stage, the agent has known dependencies, defined stages, an execution schedule, and predictable failure modes.<br><br>You have accidentally converged on a directed acyclic graph.<br><br>Data and software engineers have spent decades solving this class of problem. Once the path is understood, conventional orchestration can execute it more cheaply and reliably, while giving the team clearer control over state, retries, monitoring, and deployment.<br><br>This is useful progress. The system has revealed enough of its structure to become ordinary software.<br><br>The only disappointment is that &#8220;scheduled workflow with a semantic transformation step&#8221; sounds considerably less magical than &#8220;autonomous agent&#8221;. Production systems have an unfortunate habit of losing some whimsy as they improve.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HqGp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HqGp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ChatGPT Image Aug 5, 2026 at 11_27_51 PM (2).png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="ChatGPT Image Aug 5, 2026 at 11_27_51 PM (2).png" title="ChatGPT Image Aug 5, 2026 at 11_27_51 PM (2).png" srcset="/__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HqGp!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88c360d0-32f1-4e00-b1e8-3b9cc685ce4e_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>You can arrive from the other direction</h3><p>You can also begin designing the workflow from the tools:<br>One deterministic function performs a defined operation. Another follows it.<br><br>The sequence is explicit, testable, and scheduled through conventional infrastructure.<br>You continue connecting steps until the path between two tools becomes difficult to express in code: perhaps the input is unstructured, or perhaps the next action depends on language that can reasonably be interpreted in several ways. Perhaps mapping one system into another would require a forest of brittle rules and regular expressions that nobody will want to maintain six months later (if you know Perl, I hear you may disagree).<br><br>That breakpoint is a natural place for an LLM.<br><br>The LLM receives the context, resolves the case, produces a structured result, and returns control to the deterministic workflow. It is being used because the software cannot cover the situation economically or reliably through fixed rules.<br><br>This route often produces a lighter agent because the LLM is invoked only where the normal runtime is insufficient while scheduling, retries, state, and observability are already handled by infrastructure designed for those responsibilities.<br><br>Agent-first and tool-first development therefore represent two different ways of discovering the same eventual boundary.<br><br>Agent-first development starts with broad flexibility and gradually extracts stable behavior. Tool-first development starts with a defined path and introduces the model at the points where the code is not enough.<br><br>Both can converge on the same production architecture.<br></p><h2>The emerging architecture</h2><p>The architecture I keep seeing emerge has three layers:<br></p><ol><li><p><strong>A deterministic backbone</strong></p></li><li><p><strong>An agentic exception path</strong></p></li><li><p><strong>A human escalation boundary</strong></p></li></ol><p><br>The deterministic backbone handles the normal path by owning stable sequencing, scheduling, validation, state management, retries, permissions, and observability.<br>The agentic exception path activates when the normal workflow reaches a case that the known rules cannot handle confidently. Here, &#8220;exception&#8221; covers more than a runtime error and includes any situation where the software lacks enough structure, context, or coverage to continue.</p><p>Consider, for example, all the cases where an input fails to match an existing schema, a document is incomplete, or two systems return conflicting information. And we could go on and on... a request may contain ambiguity that materially affects the next action, and a result may fall outside the range covered by the original rules.<br><br>In those &#8220;exception&#8221; cases, the model can inspect the context, call tools, gather additional evidence, and select a recovery path dynamically.<br><br>That is still agentic orchestration; it simply happens at the edge of the known workflow instead of supervising every routine step, and when the LLM resolves the issue, control returns to the deterministic path.<br><br>This approach paves the way for higher escalation. When the information remains insufficient or the action carries too much risk, the workflow escalates to a human in the loop.<br><br>The resulting three-step architecture gives the team a clear operating model where the known path is predictable, the exception path is adaptive and bounded, and the escalation boundary has a clear owner.<br><br>I quite like the approach as it is more useful than debating whether the system is sufficiently &#8220;agentic&#8221;: that term now covers everything from a loop around an LLM call to software with access to production systems and corporate credit cards. Some precision feels overdue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!IMx2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!IMx2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ChatGPT Image Aug 5, 2026 at 11_28_27 PM (2).png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="ChatGPT Image Aug 5, 2026 at 11_28_27 PM (2).png" title="ChatGPT Image Aug 5, 2026 at 11_28_27 PM (2).png" srcset="/__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IMx2!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb23d029e-862a-423a-b01b-1ad3e1d9932e_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Agent-first and tool-first solve different learning problems</h3><p>Agent-first development works well when the workflow remains uncertain.<br>The model can attempt the task, expose missing tools, reveal hidden dependencies, and show which decisions actually recur. This reduces the cost of learning because the full process does not need to be formalized before anyone knows whether it works.<br>Its main risk is inertia: once the prototype begins producing useful results, delivery pressure rewards adding another instruction. The next stakeholder demonstration is approaching, someone needs the feature by Thursday, and restructuring the workflow feels slower than making the prompt slightly longer.<br><br>Over time, the prompt becomes the permanent runtime and the token cost, latency, and variance of the experimentation phase become production characteristics because nobody owns the industrialization.<br><br>Tool-first development works well when the execution path is already understood, or when reliability, auditability, and control matter from the beginning. The team implements the known workflow and introduces model calls where the software cannot cover the case cleanly.<br><br>That route has its own risk: teams can formalize the process too early and spend weeks implementing rules for a workflow they barely understand while several agent-driven attempts may have exposed the real structure much faster.<br><br>The starting point should reflect how much the team already knows and the final architecture should reflect how much uncertainty remains after the workflow has been exercised.<br></p><h3>Four operating gates for every model call</h3><p>When reviewing an agent workflow, I find these four questions useful.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!sQ7s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!sQ7s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df794514-3857-4ec8-ae8e-700acf875797_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ChatGPT Image Aug 5, 2026 at 11_27_51 PM (1).png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="ChatGPT Image Aug 5, 2026 at 11_27_51 PM (1).png" title="ChatGPT Image Aug 5, 2026 at 11_27_51 PM (1).png" srcset="/__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sQ7s!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf794514-3857-4ec8-ae8e-700acf875797_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Is the execution path predictable?</h3><p>Look at the traces.<br><br>When similar inputs consistently produce the same tools in the same order, the agent is executing a known path and that sequence is ready to be considered for deterministic orchestration.<br><br>The length of the prompt is only a clue, while the agent&#8217;s behavior is stronger evidence. The useful questions are how often the path repeats, where it branches, and which steps continue to vary meaningfully.<br></p><h3>Can the result be checked mechanically?</h3><p>When an output must match a schema, satisfy a rule, reconcile against a known value, or pass an assertion, software should perform that check directly.<br><br>Asking a second model whether the first model followed a fixed structural rule adds another probabilistic step to a question with a deterministic answer, although it remains a surprisingly popular design, perhaps because one uncertain opinion feels lonely.<br><br>Model-based evaluation can still help when quality depends on meaning, completeness, or contextual relevance. Nonetheless, mechanical correctness should remain mechanical.<br></p><h3>Is code enough to resolve the case?</h3><p>Some steps are easy to express through fixed rules, typed interfaces, and normal control flow while others become expensive or brittle when forced into deterministic logic.<br><br>A model call earns its place when the software cannot cover the case cleanly, such as when it involves unstructured input, incomplete context, conflicting evidence, unfamiliar formats, or a large number of possible variations.<br><br>Moving data between known systems, applying fixed transformations, checking explicit conditions, and following a stable sequence rarely need a model. Token cost is only one part of the calculation, as even an inexpensive call introduces latency, variance, another failure mode, and another component that needs to be observed.<br>The LLM call should contribute something the runtime cannot provide more effectively.<br></p><h3>Is the agentic path bounded?</h3><p>Once the model can choose tools dynamically, the boundaries need to be explicit.<br>Which tools may it call? What data can it access? Which actions can it take without approval? How is recovery validated? When does control return to the deterministic path? Which conditions require human review?<br><br>These boundaries make agentic orchestration operable, and very importantly, make the system easier to explain to security, compliance, operations, and customers because the known path is predictable, the adaptive path is constrained, and the escalation point has an owner.<br><br>Without those controls, &#8220;agentic exception handling&#8221; can become a sophisticated phrase for granting broad permissions and hoping the traces are educational.<br></p><h3>The organizational failure usually happens gradually</h3><p>Most teams do not make a deliberate decision to keep an LLM responsible for predictable orchestration forever; they arrive there through a series of reasonable local decisions.<br><br>The prototype works &#8594; A customer asks for another case &#8594; A tool is added &#8594; The prompt receives another instruction &#8594; The deadline moves closer, and nobody owns the moment when stable behavior should be extracted into code!<br><br>At the speed of prototyping, each change makes sense on its own. Put together and left unchecked, however, those changes can create a production architecture whose cost and reliability remain tied to the way the team experimented.<br><br>Engineering leaders need an explicit review point for this transition. That review should examine path stability, token cost per successful run, latency, validation coverage, exception frequency, recovery behavior, and the proportion of model calls handling cases that software cannot yet cover.<br><br>The purpose is simple: as the team understands more of the workflow, the architecture should reflect that understanding.<br><br>Without this review, the agent remains the orchestrator by default. Stable behavior stays buried in prompts, and the organization continues paying for flexibility it no longer uses.<br></p><h2>Let the architecture learn with you</h2><p>Broad agentic orchestration can be a very effective way to explore a workflow early.<br>As the system runs, repeated patterns emerge. Known sequences can move into code. Common decisions can become explicit rules. Stable outputs can gain mechanical validation. Scheduling, state, retries, and observability can move into the runtime designed to handle them.<br><br>The agent remains available for the cases where code is no longer enough. It can handle unfamiliar inputs, reconcile conflicting information, coordinate recovery, and return the workflow to the known path.<br><br>That gives you the flexibility of an agent without charging the agent tax on every ordinary execution.<br><br>Use the model to discover the workflow. Move the parts you understand into code. Preserve agentic orchestration for the points where fixed rules and conventional software can no longer resolve the case efficiently.<br><br>Or, in terms my cloud bill can understand:</p><blockquote><p><strong>Thou shalt not spend tokens on software-engineering problems.</strong></p></blockquote><p></p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Agentic development did not fix delivery. It exposed it]]></title><description><![CDATA[How we built FieldOne to turn compressed build time into trustworthy production delivery]]></description><link>https://enterprisecontextmanagement.substack.com/p/agentic-development-did-not-fix-delivery</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/agentic-development-did-not-fix-delivery</guid><dc:creator><![CDATA[Emanuele Melis]]></dc:creator><pubDate>Thu, 23 Jul 2026 14:30:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Z2AC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Enterprise delivery has always moved slower than development. A feature can be built quickly, but getting it into production still depends on access, security reviews, environments, permissions, data ownership, integration, UAT, operational readiness, and customer adoption.</p><p>For years, we could manage that gap through spreadsheets because build effort still represented a meaningful share of the delivery timeline. Agentic development changed that balance dramatically and exposed how wide the gap could become.</p><p>With agentic development, an integration that once took weeks can now be built in a day, while its path to production remains largely unchanged. The result is a widening gap between what a team can produce and what an enterprise can safely adopt.</p><p>This creates a new leadership problem: work can look complete while remaining blocked by the conditions required to operate it. Projects can appear active without moving materially closer to production, and engineers can finish the build while the delivery state remains unresolved.</p><p>We quickly found that traditional measures of progress had become less reliable. They were designed for a world in which development effort and delivery progress moved together closely enough for one to serve as a proxy for the other. We needed a way to make delivery readiness visible, evidenced, and reusable across projects so that each project could make the next one easier.</p><p>FieldOne is the system we built to do that.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Z2AC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 424w, /__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 848w, /__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Z2AC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png" width="526" height="495.65384615384613" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1372,&quot;width&quot;:1456,&quot;resizeWidth&quot;:526,&quot;bytes&quot;:373651,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/207315925?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 424w, /__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 848w, /__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Z2AC!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F383121f1-c29b-4c1e-975e-86f671ecdfa1_2400x2262.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>The spreadsheet worked until build time stopped being the constraint</strong></h3><p>FieldOne started as a spreadsheet that tracked live projects, effort, owners, dates, blockers, and slippage, which made the traditional delivery model visible and gave us a single place to plan the work.</p><p>The spreadsheet could show that an engineer was allocated to a project, but it could not capture the widening gap between compressed build time and the realities of enterprise deployment. Most importantly, a spreadsheet could not show the significant opportunity that the new way of working was presenting: engineers could now complete a build, switch context, support another project, or feed work back into product while an integration sat waiting on access, validation, or a customer environment.</p><p>Agentic development changed what we needed to see: the amount of code being produced no longer told us how much delivery progress had actually been made.</p><p>For a growing field engineering team operating across more concurrent projects, the spreadsheet could no longer represent the work. We needed a shared state that showed what was active, what was ready, what was blocked, and where agentic speed was creating leverage or risk. If the original spreadsheet tracked hours, the next system needed to track delivery truth.</p><h3>The spreadsheet showed us what the operating model was missing</h3><p>The second iteration of FieldOne was still a spreadsheet, but it gave us a more useful view of delivery by plotting milestones over time and layering in the skills, owners, and dependencies required at each stage. What looked like a planning tool soon exposed the operating shape of the problem: projects were no longer moving in step with engineering effort, but through periods of build, waiting, validation, handoff, and escalation, often with responsibility shifting between field, product, the customer, and external system owners along the way.</p><p>That view made readiness easier to locate. We could see where a project had stopped needing active engineering but was still unable to advance, where the next dependency sat outside the person formally assigned to the work, and where delays were accumulating without appearing in the capacity plan. Access, environments, permissions, data ownership, validation, security, UAT, and operational readiness were shared delivery conditions, and when they were not surfaced early, they returned later as rework, delay, and commercial risk.</p><p>As the model became more useful, the spreadsheet became less able to contain it. The relationships between projects, milestones, owners, dependencies, evidence, and risk were becoming too complex for rows and formulas, while the reporting and workflow logic increasingly belonged in a system rather than a shared file. So we moved FieldOne into a dedicated application with its own database, backend, user experience, reporting layer, and workflow model, all built on our own product layer, ContextOne.</p><p>That transition turned FieldOne from a planning aid into the system of record for delivery state, with structured objects and lifecycles for projects, phases, ownership, evidence, risk, and readiness. ContextOne supplied the integration, automation, and agent capabilities around it. Delivery conditions became visible, repeatable, and governable without requiring agents or operators to reconstruct the truth from scattered updates.</p><p>Once delivery existed as a structured state rather than rows in a spreadsheet, we could begin designing the operating model around what the data was showing us. That model settled into <strong>four practices: classify work by delivery archetype, promote repeated work into certified capability, make delivery health evidence-based, and make the reliable path the default.</strong></p><h4><strong>1) <span>Classify work by delivery archetype</span></strong></h4><p>The first pattern FieldOne made clear was that projects differed less by size than by the shape of the work required to move them forward. No two projects were the same, but they rhymed. Most fell into three broad archetypes, each with a different mix of skills, dependencies, and operating risk.</p><p><strong>Agent-heavy projects</strong> reused a significant amount of capability from ContextOne and required relatively light last-mile configuration or integration. Their progress depended less on new infrastructure and more on subject-matter expertise, workflow validation, and clear acceptance criteria. The technical work often involved configuring an agent loop around a known set of tools, behaviors, and validation steps, while the harder part was confirming that the workflow reflected how the customer actually operated.</p><p><strong>Integration-heavy projects</strong> connected fragmented data and processes across multiple systems. They required more field customization and more coordination across system owners, permission models, data structures, and reconciliation logic. They also generated some of the most valuable product learning, because connectors and workflows that began as project-specific work often revealed patterns that could be hardened and reused elsewhere.</p><p><strong>Build-heavy projects</strong> extended the boundaries of the product itself. These were first-of-a-kind deployments that required new capability, such as a fully air-gapped on-premise environment, a custom operating system kernel, a 1,000-user concurrency certification, or a customer-led UAT process supported by dashboards that operational teams could use without us in the room. They needed deeper involvement from product, engineering, and infrastructure because the delivery work was also creating part of the future product.</p><p>The archetypes helped us understand why projects with similar timelines could require completely different interventions. Agent-heavy work needed more subject-matter expert (SME) access and UAT coordination. Integration-heavy work needed stronger project management and cross-system ownership. Build-heavy work needed deeper product and architectural support.</p><p>Once those differences were explicit, we could design delivery pods around the actual shape of the work rather than staffing every project as though it followed the same path. The same model improved capacity planning, clarified where risk was likely to appear, and helped us distinguish between work that should remain customer-specific and work that warranted product investment.</p><h4><strong><span>2) Promote repeated work into certified capability</span></strong></h4><p>The next question was what should happen when the same integration, workflow, connector, or agent behavior appeared across several projects.</p><p>A working implementation was not enough to make something reusable. Production constraints such as environment parity, permission models, scale, rate limits, payload size, timeouts, monitoring, rollback, and operational ownership could still be untested, even when the feature worked correctly in development. Reusing that work without understanding those limits simply moved the uncertainty into the next project.</p><p>We created a promotion path that turned repeated project work into certified capability. To move into the product path, the work needed a named owner, supporting evidence, defined prerequisites, known operating limits, reuse criteria, and production-readiness checks. This gave the next team something they could depend on rather than a previous implementation they first had to reverse-engineer.</p><p>The purpose was to make learning portable. A project should leave behind more than code that once worked in a particular environment. The next project should inherit progress, not repay the zero-to-one cost.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c9Ws!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 424w, /__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 848w, /__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c9Ws!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png" width="526" height="512.2355140186916" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1042,&quot;width&quot;:1070,&quot;resizeWidth&quot;:526,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 424w, /__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 848w, /__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c9Ws!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7cf3bea-82ed-48f3-b94f-997fb205193b_1070x1042.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Defining what made a capability reusable also forced us to define its production conditions earlier. Many of those conditions were knowable before kickoff, which changed how we approached pre-sales. We introduced a small set of questions covering the systems involved, the applicable permission model, the available environments, known technical limits, customer ownership, and the evidence that would count as production-ready in that operating context.</p><p>These questions did not add a new approval layer. They made assumptions visible before they became embedded in the plan, giving both sides a clearer view of the conditions required for delivery to proceed.</p><h4><strong><span>3) </span></strong><span>Make delivery health evidence-based</span></h4><p>As the portfolio grew, project health became increasingly difficult to interpret because terms such as &#8220;on track,&#8221; &#8220;nearly complete,&#8221; and &#8220;green&#8221; often described different realities depending on who was speaking. One person might be referring to the state of the build, another to customer validation, and another simply to the absence of a newly reported blocker. Reviews therefore spent too much time reconciling competing versions of the project before the team could decide what required action.</p><p>We wanted those conversations to begin from a common set of facts: the dates and scope that had been committed, the current delivery phase, the evidence available, the age and ownership of open blockers, and the risks being carried into the next checkpoint. These questions became the basis of the Theta Dashboard, named after the options concept of time decay. In FieldOne, theta represented the delivery cost that accumulated as the committed plan and the current reality moved further apart.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!V95o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 424w, /__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 848w, /__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 1272w, /__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!V95o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png" width="526" height="495.65384615384613" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1372,&quot;width&quot;:1456,&quot;resizeWidth&quot;:526,&quot;bytes&quot;:584074,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/207315925?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 424w, /__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 848w, /__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 1272w, /__u/substackcdn.com/image/fetch/$s_!V95o!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb9c286-c57c-4526-9758-747703043b93_2400x2262.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>FieldOne calculated carry, slippage, and trajectory for each phase by comparing planned dates and scope with the current forecast, evidence freshness, blocker age, active exceptions, ownership, and phase-gate status. The underlying fields remained deliberately simple, but together they made the delta visible early enough to act on. A project described as on track now needed an aligned forecast, current evidence, clear ownership, and an acceptable blocker and exception state, turning health from a summary judgment into a claim that could be inspected.</p><p>To support that model, we defined phase gates and checklists around the production conditions that mattered in practice. A gate was complete only when the required evidence existed or an explicit exception had been recorded. &#8220;Access verified,&#8221; for example, required an approval artifact or a successful test query linked to the milestone and a named owner, while &#8220;Environment parity verified&#8221; required evidence that the deployed configuration matched the production target. By tying sign-off to evidence, unresolved conditions became visible before they could pass silently into the next phase.</p><h4><strong><span>4) Make the reliable path the default</span></strong></h4><p>Evidence made risk visible, but it did not ensure that teams followed the reliable path once delivery came under pressure. Optional process disappears the moment urgency shows up, so the reliable path had to become the default path, built into the way delivery operated.</p><p>Deployment patterns, agent workflow templates, integration checklists, validation routines, rollback plans, and production-readiness criteria became first-class delivery assets rather than supporting documents scattered across folders and chat threads. Each pattern followed a consistent structure covering prerequisites, environment and permission checks, implementation steps, validation, rollback, and the criteria required for production use.</p><p>That structure turned individual experience into something another engineer could apply without first reconstructing the reasoning behind the previous deployment. A new field engineer could begin with a known pattern, confirm which prerequisites were already satisfied, and focus on the parts of the environment that were genuinely different. Customers and delivery teams benefited from the same clarity because both sides could see the established path, the decisions still outstanding, and the places where the deployment departed from the norm.</p><p>The model still allowed teams to move forward when reality required a deviation, but every bypass had to be recorded as an explicit exception with a rationale, owner, expiry date, available evidence, mitigation, and follow-up action. That preserved flexibility while turning silent risk into visible risk.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nJwM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nJwM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg" width="526" height="391.97115384615387" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1085,&quot;width&quot;:1456,&quot;resizeWidth&quot;:526,&quot;bytes&quot;:228952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/207315925?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!nJwM!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f9a16f-3331-4e2d-a29f-e4cd2fd3dca7_1840x1371.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Over time, the exceptions themselves became a source of operational and product learning. When several projects bypassed the same gate for the same reason, the pattern pointed to a weak template, a missing readiness check, an unrealistic process, or a capability ContextOne needed to provide. Delivery friction could then be converted into a stronger default rather than rediscovered by the next team.</p><p>By this stage, FieldOne was no longer recording delivery after the fact. It was governing how work moved across field, product, agents, customers, and the systems they depended on.</p><h3><strong><span>What we learned</span></strong></h3><p>The most important lesson from building FieldOne was that managing delivery in an agentic organization requires more than adding automation to an existing process. Once agents began compressing build time, shifting engineering capacity between projects, and turning delivery work into reusable product capability, we needed a system that could represent the new operating model directly.</p><p>That meant owning the delivery object model first. Projects, phases, owners, evidence, risks, decisions, dependencies, staffing, product signals, and health all needed to exist as structured state with clear relationships and lifecycles. Without that foundation, automation would have accelerated the production of updates without improving our understanding of what was actually ready, blocked, or at risk.</p><p>The separation between delivery state and the automation around it made FieldOne both useful and governable. FieldOne remained responsible for the record of what had happened, what had been approved, what was waiting, and what needed to happen next, while ContextOne provided the agents, integrations, and workflows that could act on that state. Agents could prepare meeting briefs, surface aging risk, draft customer updates, or recommend next actions, but they were operating within an explicit delivery model rather than reconstructing one from scattered context.</p><p>We also learned that signals beat status. A weekly review should not begin with each owner explaining whether a project feels on track, but with changes to the forecast, aging blockers, missing evidence, active exceptions, phase-gate status, and ownership of the next action. This shifted reviews away from reconciling competing narratives and toward deciding where intervention was required.</p><p>The same principle applied to process. Optional practices tend to disappear under pressure, so gates, templates, readiness checks, and evidence requirements had to become part of the default path. Teams could still bypass a gate when delivery required it, but the exception had to be named, owned, justified, time-bounded, and tied to mitigation. Flexibility remained possible without allowing risk to become invisible.</p><p>Over time, those exceptions, escalations, and postmortems created a compounding loop. A repeated problem could become a stronger gate, a clearer readiness check, a certified capability, or a reusable delivery pattern. Each project could therefore improve not only its own outcome, but the system used to deliver the next one.</p><p>That is the larger role FieldOne came to play: managing delivery in a world where agents had changed the speed, allocation, and economics of the work.</p><p>As AI reduces the effort required to build, the scarce resource becomes trustworthy progress: work that survives contact with a real customer environment, can be inspected and governed, and leaves behind reusable capability rather than another isolated implementation.</p><p>The goal is not to make enterprise delivery simple. It is to ensure every project makes the next one easier.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Determinism Problem]]></title><description><![CDATA[Why Agent Systems Need a Deterministic Execution Layer]]></description><link>https://enterprisecontextmanagement.substack.com/p/the-determinism-problem</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/the-determinism-problem</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Mon, 06 Jul 2026 16:13:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!QVu6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Current generation agent systems usually reach production in a state that is difficult to describe honestly. They work often enough to justify the investment, fail rarely enough to resist simple diagnosis, and vary just enough between runs to make every incident expensive.</span></p><p><span>More often than not, the failure is not particularly dramatic. Two sessions receive the same request, use the same model, and appear to call the same tools. One completes. The other takes a slightly different path, nothing dramatic, retrieves the same records in another order, retries a tool after a timeout and arrives at a completely different answer. Both traces look plausible. Neither explains why they diverged.</span></p><p>This is the point at which an agent platform stops behaving like traditional software and starts behaving like several nearby versions of the same software, selected by timing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!QVu6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 424w, /__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 848w, /__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!QVu6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif" width="1120" height="629" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:629,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:282952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204483845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 424w, /__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 848w, /__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!QVu6!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e492a44-98ad-4b0f-aa57-dfd8337aa94f_1120x629.gif 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>It&#8217;s not the model</h4><p><span>The &#8216;traditional&#8217; response to this is to investigate the model. Temperature is reduced, prompts are tightened, perhaps a stronger model is substituted. These changes can improve answer quality, but they do little for execution consistency because the model is rarely the only source of variation.</span></p><p><span>A production run is shaped by the order in which messages arrive, which tool result returns first, whether a request hits a cache, how a retry is scheduled, which retrieval candidate receives a marginally higher score, and what state was left behind by the previous turn. Even timestamps and generated identifiers can alter later behavior if they enter the context or are used as lookup keys.</span></p><p><span>None of these variations is large enough to look like a platform fault on its own, but together they destroy reproducibility.</span></p><p><span>So does that matter? Well, yes. Reproducibility is the boundary between an operational system and a convincing demonstration. A demonstration needs to work, perhaps just to justify a budget.  An operational system needs to explain what it did, resume after failure, and produce evidence that a proposed fix addresses the same execution that failed.</span></p><p><span>In short - </span><strong><span>a log or a trace is not enough.</span></strong></p><p><span>Logs are observations emitted by a running process. They are frequently incomplete, asynchronously written, and detached from the state that gave an event its meaning. A tool log may show that a request was submitted twice without revealing whether the second submission reused the first execution, started a duplicate job, or observed different upstream data. A model trace may contain the final prompt while omitting the retrieval ordering, evicted tool output, or cache decision that produced it.</span></p><h4>The Missing Layer: Deterministic Execution</h4><p><span>The missing component to all this is a </span><strong><span>deterministic execution layer</span></strong><span> between the model and the distributed systems around it.</span></p><p><span>That layer does not necessarily attempt to make language generation mathematically deterministic. What it does do is make the surrounding computation deterministic enough that model variation is visible rather than mixed with orchestration noise.</span></p><h4>Strict Event Ordering</h4><p><span>The first requirement is </span><strong><span>strict event ordering</span></strong><span>.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kIXE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 424w, /__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 848w, /__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kIXE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif" width="1120" height="629" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:629,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:452003,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204483845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 424w, /__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 848w, /__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!kIXE!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1348d877-2f50-498c-b97f-20bf144cfc9d_1120x629.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Agent runtimes are naturally exposed to concurrency. User messages arrive while tools are running, and multiple tools may be dispatched in parallel. Provider streams emit incremental output. Recovery policies introduce retries, route changes, and context compaction. If these events mutate session state as they arrive, wall-clock timing will naturally become part of the program.</span></p><p><span>A hundred real world variations can cause this to happen. A tool that completes 20 milliseconds earlier may enter context first. A retry callback may race with a cancellation. A session update may be accepted between two tool results. The resulting trace will still look reasonable, but the state transition order is no longer stable.</span></p><p><span>A more reliable runtime treats inbound activity as an ordered event stream. Events may be produced concurrently, but state changes are applied synchronously by one authority. The core consumes one ordered inbox and advances the session through explicit transitions. Model requests, tool completions, retry decisions, recovery actions, and external inputs all acquire a defined position in that sequence.</span></p><p><span>This design can sound conservative to teams accustomed to extracting parallelism from every service, but crucially, the restriction applies to state mutation, not to all work. Tool calls can still execute concurrently. Retrieval can still fan out. Model traffic can still be streamed. What cannot remain ambiguous is the order in which their results become part of the agent&#8217;s state.</span></p><p><span>A single-threaded deterministic core is often the simplest implementation. It removes an entire class of locking, race, and causal-ordering problems. More elaborate state machines can provide the same property, but they still need a single agreed sequence for state-changing events and a replayable transition function.</span></p><p><span>Timestamps and identifiers belong inside this boundary as well. If they are generated from ambient wall time or random UUIDs during replay, the replay is already different. A deterministic clock and reproducible identifier generation allow restored execution to recreate the same internal references and event metadata.</span></p><h4><strong>Materialized Execution Snapshots</strong></h4><p><span>This ordered core becomes the foundation for the next requirement: </span><strong><span>materialized execution snapshots.</span></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ys68!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 424w, /__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 848w, /__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ys68!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif" width="1120" height="629" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:629,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:364678,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204483845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 424w, /__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 848w, /__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!ys68!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d61e6e6-502f-482c-984e-ae5ef96b82be_1120x629.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Agent sessions are commonly persisted as conversation transcripts. That captures what the user and model said, but not the execution environment in which the conversation occurred. The transcript may refer to files that have since changed, tool results that were truncated, retrieval outputs that were never stored, or subagent state that existed only in memory.</span></p><p><span>A production checkpoint needs to include the state the agent could actually act upon.</span></p><p><span>That means sealing the core session state together with the execution trajectory and the session&#8217;s working environment. If the agent can create files, transform data, spill large tool outputs to storage, or continue a child agent later, those artefacts are part of the session. Persisting the messages without them is comparable to storing a database transaction log while discarding the database pages.</span></p><p><span>A practical checkpoint can be taken at the end of each turn, when the runtime has reached a durable boundary. The snapshot should contain the model-neutral conversation representation, active policies, tool and recovery state, session labels, durable child conversations, the virtual file system, and a complete trajectory of decisions and external results.</span></p><p><span>The archive can then be sealed for integrity and written to object storage as one recoverable unit. A compatible runtime should be able to load it and continue from the last committed turn without reconstructing the session from scattered service logs. Before resumption, the target runtime should verify that the required tools, model capabilities, and configuration contract are still available. Why? Because silent restoration into a materially different environment is another form of nondeterminism.</span></p><p><span>Once snapshots exist, context construction must be derived from them rather than rebuilt opportunistically from live systems, and this is where many otherwise careful platforms lose determinism. A replay loads the same transcript, then reruns retrieval against a changed index. A tool catalogue is searched again and returns a slightly different ranking. Context compaction chooses a different boundary because token estimates changed. An old tool result is re-fetched from a system whose data has moved on.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9Lo0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 424w, /__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 848w, /__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9Lo0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif" width="1120" height="664" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:664,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:412211,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204483845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 424w, /__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 848w, /__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!9Lo0!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff737d2da-9c03-41c8-9a7f-7943b11f8e6f_1120x664.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Instead of a replay event asking &#8220;What context would the system build now?&#8221;, a better question is  &#8220;What context did this state produce then?&#8221;</span></p><p><span>Context policy can still be dynamic of course. Large results may be moved from the prompt into a session file system. Older turns might be compacted. Tool responses may be re-encoded into a more token-efficient representation. The important property is that these actions are explicit state transitions recorded in the trajectory, and that these transitions are accurately reflected in the snapshot.</span></p><p><span>The same discipline applies to retrieval. Search systems routinely produce ordering variance because of index updates, floating-point scoring, replica differences, or ties resolved by execution timing. A production agent should either materialize the selected retrieval set as an event or query against a versioned, time-bounded view that can answer as the data existed at the time of the decision.</span></p><p><span>Without all of this, &#8220;replay&#8221; means submitting a similar request to a different world.</span></p><h4><span>Tools as State-Bound Events</span></h4><p><strong><span>Tool execution</span></strong><span> is the most consequential boundary because tools create side effects.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hDoM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 424w, /__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 848w, /__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hDoM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif" width="1120" height="664" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:664,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:281276,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204483845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 424w, /__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 848w, /__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!hDoM!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe69fa70a-7343-4538-9c62-9611953d28f1_1120x664.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Many orchestration frameworks treat a tool call as an ordinary function invocation. The runtime sends arguments, waits, and inserts the result into the transcript. This abstraction conceals the distributed transaction underneath. The request may be retried by the client, the service mesh, the gateway, or the tool provider. A timeout may mean that the operation failed, or that it succeeded and the response was lost.</span></p><p><span>A deterministic runtime must therefore treat a tool invocation as a state-bound event with a stable identity. The call must carry an idempotency key derived from the session, turn, and tool event. A retry must resolve to the original execution wherever the tool contract permits it. It must not create a new side effect merely because transport delivery was uncertain.</span></p><p><span>The gateway can remain asynchronous internally while presenting a synchronous result to the agent. It may submit a job, poll its status, and return only when the job succeeds, fails, or times out. The agent does not need to reason about the gateway&#8217;s scheduling. The runtime does need to preserve the job identity, dispatch decision, arguments, result, and retry relationship.</span></p><p><span>The rule is simple even when the implementation is not: </span><strong><span>uncertainty about the response must not become uncertainty about whether the side effect occurred</span></strong><span>.</span></p><p><span>Mutable storage needs the same treatment. Agent file writes should be versioned and protected by optimistic concurrency. A caller that edits a file must identify the version it observed. If the head has changed, the write should fail and force a re-read rather than silently overwriting a concurrent update. Each successful mutation should create an immutable version linked to its predecessor and attributed to the source session and tool event.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bQgp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 424w, /__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 848w, /__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bQgp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif" width="1120" height="664" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:664,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:334792,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204483845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 424w, /__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 848w, /__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!bQgp!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81a115c-dda8-47ee-b6bd-808b1ba9028a_1120x664.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>This will produce more than simple auditability. It will ensure state drift is visible at the point where it occurs.</span></p><h4><span>An Incremental Implementation Path</span></h4><p><span>For in-house builds, the implementation path is usually incremental. The ordered event core should come first because every later guarantee depends on it. Existing tools can initially be wrapped as events even if their downstream idempotency remains imperfect. Turn-boundary snapshots can then capture conversation state, event history, and tool outputs. The session file system and other working state can be added once the restore contract is stable. Context policies can subsequently be moved into deterministic transitions, followed by stronger tool deduplication and versioned external state.</span></p><p><span>Trying to begin with a complete reference architecture tends to create a large program with little early operational value. The useful milestone is narrower: take one failed production session, restore it from a committed boundary, and reproduce the same state transitions without contacting systems that were not contacted in the original run.</span></p><p><span>The architecture we describe here evolved over a focused period of real world implementations, culminating in the &#8220;ContextOne&#8221; product. We use an ordered, single-threaded core, deterministic clocks and identifiers, turn-boundary archives containing sealed session and file-system snapshots, durable trajectories, idempotent tool identities, version-aware writes, and recovery policies owned by the core rather than hidden inside transport layers. For the curious, single threaded does not mean slow either - in a 30 minute long agent turn, our own overhead is less than 10 milliseconds.</span></p><p><span>These choices impose cost. More state is persisted. Tool results must be retained. Event schemas need versioning. Compatibility checks become part of deployment. Engineers must distinguish between parallel execution and ordered state application. Product teams lose some freedom to patch behavior through invisible middleware.</span></p><p><strong><span>The return is not simply fewer failures. It is control over failures.</span></strong></p><p><span>When a model produces a different answer from the same preserved state, that difference can be measured as model behavior. When a tool returns different data, the changed external observation is visible as an event. When a retry occurs, the system can show whether it reused or duplicated execution. When a deployment changes prompts, tools, or recovery policy, recorded sessions can be replayed against the candidate configuration and checked for regressions.</span></p><p><span>That is the operational value of determinism in agent systems. </span><strong><span>It turns an opaque sequence of plausible actions into a state machine that can be recovered, audited and compared.</span></strong></p><p><span>Senior engineering leaders do not need agent systems to behave identically forever. They need to know which differences were intentional, which were produced by the model, which came from changing data, and which were introduced by the runtime itself. Without a deterministic layer, those causes remain entangled. Every incident becomes a reconstruction exercise, and every apparent fix is tested on a run that may not be the one that failed.</span></p><p><span>Production trust begins when the system can preserve its own history closely enough to disagree with it.</span></p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[False Summits: Why Enterprise AI Keeps Looking Finished ]]></title><description><![CDATA[A practitioner's map of the six milestones that feel like the finish line, why the best teams still stall, and the three disciplines that actually reach production.]]></description><link>https://enterprisecontextmanagement.substack.com/p/false-summits-why-enterprise-ai-keeps</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/false-summits-why-enterprise-ai-keeps</guid><dc:creator><![CDATA[Conor Twomey]]></dc:creator><pubDate>Tue, 30 Jun 2026 20:47:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!P1kK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Every enterprise AI program I walk into believes it is almost finished. Most are halfway up the mountain, standing on a ridge that looks exactly like the top.</span></p><p><span>Climbers have a name for that: the false summit. You crest what is plainly the peak, and the real one appears behind it, higher, colder, and further off than anything you have climbed. Enterprise AI is a mountain made almost entirely of these, and the cost of mistaking one for the top is measured in quarters, not days. A decade ago there were a few hundred AI companies; today there are tens of thousands, and every one of them is standing on one of these ridges, telling you it is the top.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!P1kK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!P1kK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1975668,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/204337009?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P1kK!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfe301ad-2cce-4aca-beea-6c0a9b98946e_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Two ridges come before the real climb even starts, and nearly everyone clears them and calls it progress.</span></p><p><span>The first is the demo. An agent answers a hard question against a clean, hand-picked set of sources, in a sandbox, in front of the executive committee. It is convincing precisely because everything that makes the enterprise hard has been removed: the data is tidy, the questions are the ones the demo was built for, and nobody is asking what happens on Monday when a thousand people use it. A demo proves something is possible. It does not prove you can deliver it.</span></p><p><span>The second is enterprise chat. You switch on an assistant for everyone, wire it to email and documents, and the adoption charts climb. The gains are real, but they are small, and they show up in the employee&#8217;s day, not on the income statement. I have started calling this corporate wellness: a decent thing to offer, and people feel supported, but it was never going to change the financials. Set a usage target like &#8220;use the assistant four times a day&#8221;, and people will ask about the weather to hit the number.</span></p><p><span>Both ridges make the same mistake, and it is the one hiding under every false summit higher up: they confuse the how for the why. Across the 477 of these conversations my team has sat in since January, everyone wants the same thing, and it is not AI. They want operational leverage, the ability to do more with what they have. AI is the how; operational leverage is the why.</span></p><p><span>It helps to be precise about the work AI should actually do. The sweet spot is high complexity, low judgment tasks. High complexity, because if the task were simple you would have automated it years ago with if-this-then-that rules. Low judgment, because you do not want a model deciding which drug to prescribe or whether to approve a mortgage. You want it to do the gathering and cross-checking, the swivel-chair work, and hand a complete picture to the expert who makes the call. Keep that test in mind and the real false summits explain themselves.</span></p><p><span>Above the tree line, the climb gets serious. These are the four that cost real years.</span></p><h4><span>Summit one: ten times the code</span></h4><p><span>The first real false summit is AI-assisted software development, and it is seductive because the gain is visible: the tools work, and engineers produce dramatically more code. The mistake is reading code velocity as product velocity.</span></p><p><span>Across the teams I see, 10x the code is not turning into 10x the product, or even a noticeably faster roadmap. One CIO told me an AI agent had rebuilt the same component 76 times in a month, a component that already existed, leaving him with 76 versions to maintain and a strong wish to put the genie back in the bottle. He has since taken the tools off his developers, not because they do not work, but because the old way of specifying, reviewing, and shipping software cannot absorb that volume. Writing the code was never the bottleneck. The real constraints were deciding what to build, checking it, and integrating it safely, and those are the parts the tools do not yet speed up.</span></p><p><em><span>Code was never the constraint. Until the process around it is redesigned, more code just means more to review.</span></em></p><h4><span>Summit two: centralizing the data</span></h4><p><span>Once teams build agents that act rather than answer, they hit the real state of enterprise data: scattered across dozens of systems, structured and unstructured, some of it true, some stale, much of it contradictory. The instinct is the one the industry reaches for every cycle: centralize. Get everything into one place so the agent has one surface instead of forty. Whole programs are commissioned on that theory, often with a multi-year price tag.</span></p><p><span>Then you finish the migration and find that location was never the hard part. The agent can reach everything and still does not know what anything means. Roughly a third of the meaning is clear. Another third is ambiguous: is the &#8220;customer&#8221; in one system the same as the &#8220;customer&#8221; in another? The final third was never written down at all; it lives in the heads of people who have done the job for fifteen years. Moving data recovers none of that. The first third is software, the second needs experts resolving cases one at a time, and the third can only be learned by watching how the work is actually done.</span></p><p><span>There is a second cost that rarely gets priced. When you centralize, you hand your hardest-won knowledge to whichever platform you migrated into. For a one-off report, fine. For the understanding that makes your business run, it is worth asking whether you want it living inside a single vendor.</span></p><p><em><span>The battle is over meaning, not location. Pay for understanding before you pay for migration.</span></em></p><h4><span>Summit three: the first dozen agents</span></h4><p><span>This is the summit that fools the most capable teams, because by the time they reach it they have done everything right and have the scars to prove it. I had dinner recently with senior AI leads from some of the largest financial institutions in the world. The pattern around the table was remarkably consistent: a dozen, maybe two dozen, meaningful agentic experiences in production. Built by dedicated teams that started around fifteen people and grew to thirty-five over the twelve months it took.</span></p><p><span>Now do the arithmetic out loud, which almost nobody does. A win is not a dozen agents; inside a large enterprise the real surface area is thousands, maybe tens of thousands, of these workflows. If the first dozen took a year and a team of thirty-five, the current approach gets you to full scale somewhere around 2035. Every leader I put that number to goes quiet, because they have done the same math privately and kept it to themselves.</span></p><p><span>The reason sits in how those first agents were built. Getting an answer at all is the hard 30%, the context work, wired by hand for each agent. Earning the right to trust the answer is the other 70%: the evaluations and checks, the machinery that catches a mistake before a customer does. Almost all of it is hand-built, per agent. The thirteenth agent costs roughly what the first one cost. Nothing accumulates. That is the defining feature of a false summit: the effort does not turn into height.</span></p><p><em><span>Track the cost of your next agent, not the count of your last ten. A flat cost curve at agent twelve is your next five years.</span></em></p><h4><span>Summit four: the bill</span></h4><p><span>The last summit arrives as a reward for success. One bank built an internal assistant for its developers and tested it on five small systems, where it was accurate and cheap. Then they pointed it at three dozen real ones. The answers held up; the cost did not. Queries that mattered now ran past two dollars each across tens of thousands of engineers. The options were all bad: move everything into a warehouse, pay the token bill and hope, or rebuild the data a third time as a graph. None of those is a strategy.</span></p><p><span>This is the part nobody prices in during the pilot. Metered tokens look harmless at prototype volumes. At enterprise scale, with agents making dozens or hundreds of model calls per task, the bill starts to behave like a cloud bill: opaque until you instrument it, volatile until you govern it, and politically radioactive once it shows up in operating expense. CFOs have seen this film before, when predictable capital spend turned into unpredictable cloud spend, and they remember the ending. Two things make it worse. The same task that costs ten dollars on a frontier model is often a fifty-cent task on a smaller one, but only if your setup can route between models at all. And the &#8220;best&#8221; model keeps changing, so anything hardwired to a single provider is a depreciating asset.</span></p><p><em><span>What matters is cost per finished piece of work at a set quality, and the ability to swap the model underneath it in weeks, not quarters.</span></em></p><h4><span>The actual mountain</span></h4><p><span>So what does the real summit look like? It is a system that actually holds up in real-world conditions: answers it can trust and stand behind, on the messy data it already has, fast and affordable across the whole enterprise. Every false summit falls short of that in a different way. The coding wave produced far more code, but not more finished products. Centralizing the data put everything in one place, but the agent still could not tell what any of it meant. The first hand-built agents worked, but each new one cost about as much as the last: the cost never fell and nothing carried over, so the approach could never scale. And the moment real usage arrived, the token bill made it impossible to sustain.</span></p><p><span>Underneath all of them is one shift most enterprises have not absorbed: this is the first major wave of technology built on probabilistic foundations. Every previous system gave the same answer to the same question every time. A language model does not. Ask it the same question five times and you can get five different answers, each delivered with full confidence. Everything an enterprise means by &#8220;production&#8221; was defined in a deterministic world: auditability, repeatability, accountability. Crossing that gap is the whole climb.</span></p><p><span>What the crossing demands turns out to be specific, and it is not more clever agents. It is three disciplines.</span></p><p><span>The first is a living data foundation, built where the data already sits. Not another migration into a warehouse, but software that reads every source in place, whatever the system or data model, structured and unstructured alike. It works out which entities are which, fills in the missing context, and keeps re-resolving as schemas, joins, and meaning shift underneath it. A static catalog is stale the moment it is finished, because the business does not hold still. A foundation that re-resolves against live sources and real workflows is the thing centralizing was reaching for and never caught.</span></p><p><span>The second is context and guardrails built by software, not by hand. The 30% that surfaces an answer and the 70% that earns the right to trust it: the evaluations, the judges, the checks, generated and enforced across a whole population of agents at once, with evaluation kept permanently in the loop. Crafting one agent by hand is tractable. Governing thousands of them, as models and prompts drift beneath you, is not. Not by people, not ever. It is the only thing that makes effort accumulate into altitude instead of starting again from zero at agent thirteen.</span></p><p><span>The third is the move from exploration to dependable execution. Let an agent reason a task out once, in full; grade that path; and when it is proven, promote it to plain code that runs the same way every time, with the model verifying the result rather than rediscovering it. The same question stops returning different answers. The cost stops rising every time you ask. Proven work hardens into infrastructure, which is exactly what a regulated business means by production.</span></p><p><span>Between them, those three clear the false summits that cost companies years: the meaning, the hand-built dozen, and the bill. None of it looks like building impressive individual agents. It looks like building the thing that builds them: the shared understanding of the business, the machinery that checks the work, and the path from probabilistic exploration to deterministic execution. The agents are the easy part. The discipline underneath is the climb, and it is where the advantage that compounds for a decade actually lives.</span></p><p><span>Two principles hold it together. Start with the AI working in the background while a person signs off on every result, and let it earn more autonomy only on evidence, not enthusiasm. And design accountability in from the first day, because when an AI-assisted decision goes wrong, &#8220;the system told me&#8221; is not an answer.</span></p><p><span>None of this is a counsel of despair. The mountain is climbable, and the teams that climb it will hold an advantage that compounds for years. But the route matters more than the enthusiasm. So one question to carry into your next AI review: are we higher than we were last quarter, or are we standing on another ridge? The teams that ask it early save themselves years.</span></p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Work Is Inside the Parentheses]]></title><description><![CDATA[The next frontier is systems that know what they are checking for.]]></description><link>https://enterprisecontextmanagement.substack.com/p/the-work-is-inside-the-parentheses</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/the-work-is-inside-the-parentheses</guid><dc:creator><![CDATA[Chuck]]></dc:creator><pubDate>Tue, 23 Jun 2026 14:47:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BueM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A line from Boris Cherny at Anthropic has been making the rounds: &#8220;I don&#8217;t prompt Claude anymore. My job is to write loops.&#8221; It is a good line because it captures a real transition in how people are starting to work with powerful models. The early era of LLMs was dominated by prompting: write the instruction, read the answer, adjust the instruction, try again. The next era is more procedural. Instead of asking the model to produce a single output, we build a system that can evaluate a state, take an action, observe the result, and continue until some condition has been satisfied.</p><p>I agree with the sentiment. The job is increasingly to write loops. But the more important thing is what a loop actually is.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BueM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 424w, /__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 848w, /__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BueM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:943040,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/202587982?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 424w, /__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 848w, /__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BueM!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357e0956-0672-426a-ba32-0f3673c6b941_2400x1350.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A loop is not simply &#8220;let the model run for a while.&#8221; In programming, a loop has structure. It says: while this condition is not true, do this thing. Or: continue until this condition is met. There is always an evaluation followed by an execution. The system checks the state of the world, decides whether the condition has been satisfied, and either exits or performs another action. That is what makes it a loop rather than a long prompt.</p><p>So the real work is not merely creating an agent that can keep acting. It is defining the evaluation that tells the agent whether the last action mattered. What did we put inside the parentheses? What counts as done? What counts as better? What counts as a regression? What state needs to be carried forward so the next iteration is more informed than the last one?</p><p>That is why I think &#8220;loop engineering&#8221; is directionally right but incomplete. The leverage is in the evaluation. We are not just writing loops; we are creating systems that can evaluate chunks of work, decide whether they satisfy a condition, and then execute the next step accordingly.</p><p>Without that evaluation, a loop is just automated continuation. A powerful model can continue for a very long time. It can generate plans, search through tools, inspect files, make edits, summarize its progress, and try again. If the original goal is vague enough, it may eventually produce something plausible. But plausibility is not the same as correctness. If we say &#8220;build me App X,&#8221; and provide little else, then the model has to infer what App X means. It has to infer the users, the data model, the workflows, the edge cases, the integrations, the permissions, the interface, and the definition of success. A better model and a larger token budget will improve that process, but they do not eliminate the basic problem. The loop is spending intelligence to compensate for the absence of evaluation.</p><p>Software engineering gives us the cleanest version of this idea because software already has a culture of evaluation. A code agent can write code, run tests, inspect failures, patch the implementation, and repeat. In that environment, the loop has something concrete to check. Did the test pass? Did the build complete? Did the type checker succeed? Did the output match the expected fixture? The agent does not need to guess forever whether it is making progress. It can be told.</p><p>This is also how we think about our own agent infrastructure. One of the most important things we built was not a more elaborate prompt, but a deterministic way to evaluate the harness around the model. By using a mock LLM interface, we can test tool calling behavior, component interaction, handoffs, and context transitions without relying on a live model response every time. That lets us test the system around the non-deterministic component. Once that structure exists, the evaluation surface expands. We can run throughput tests, send millions of messages through the system, watch context compact and rehydrate, and observe agents operating as long-running processes rather than isolated chat sessions.</p><p>The lesson is not that everything should be deterministic. The lesson is that the non-deterministic part needs a deterministic frame around it. The model may decide how to interpret an ambiguous situation, but the system should know what state it is in, what tools are available, what result is expected, what action was taken, and what should be evaluated next. That is the difference between an agent that is merely active and an agent that is participating in a controlled process.</p><p>This becomes more important when we move from software engineering into enterprise workflows. Codebases come with files, tests, logs, types, git history, and build systems. Enterprises have SaaS tools, spreadsheets, process documents, ticket histories, Slack/Teams threads, data warehouses, browser-only workflows, tribal knowledge, and business rules that live in people&#8217;s heads. The process may be obvious to the team that runs it every day, but invisible to an agent unless it has been represented somewhere.</p><p>Take a simple business task: review a situation, decide what should happen next, and route it through the right process. On the surface, that sounds small. In reality, it depends on business context that is rarely contained in the prompt. The agent needs to know what the relevant entities are, where the data lives, which systems are authoritative, what the organization means by status or risk or priority, what actions are allowed, and what outcome would count as correct. A human employee often carries that context implicitly. They know which dashboard matters, which field is stale, which exception is normal, and which process doc is out of date. An agent does not know any of that unless the organization has represented it somewhere.</p><p>A high-end model can help discover these things. Given enough access and enough budget, it can inspect systems, read documentation, query endpoints, observe patterns, ask questions, and assemble a plausible ontology of the business. That is incredibly valuable. But if every future loop has to rediscover that ontology from scratch, then we have not built a system. We have built an expensive ritual.</p><p>The better pattern is to spend intelligence on discovery once, then preserve the result as managed context. This is how we think about the Agent Ontology Service, or AOS. A capable model can be given read-only access to explore the relevant systems, validate endpoints, map entities, identify relationships, and build an ontology layer around a business process. It can connect the language people use in the business to the systems where that language becomes data.</p><p>Once that ontology exists, the loop changes. The agent no longer has to begin with &#8220;go figure out what this company means.&#8221; It can begin with a more precise instruction: execute this process against this known context, under these constraints, and evaluate these outcomes. That shift matters enormously. It changes the cost profile because the loop is not spending tokens rediscovering the same background reality. It changes the reliability profile because the loop is now operating against a shared representation of the business. And it changes the improvement curve because the evaluation layer can become more complete over time.</p><p>This is the enterprise version of evaluation-driven development. Every business process contains a mix of deterministic and non-deterministic work. The judgment calls will remain judgment calls. The model may need to classify a situation, infer intent, summarize evidence, recommend an action, or decide which precedent applies. But once that ambiguous step has been resolved, the surrounding process can often become deterministic. If a risk is identified, route it to the right owner, attach the evidence, update the relevant system, create the follow-up task, and log the rationale. If required data is missing, request it from the source of truth. If an exception violates policy, escalate it through the correct path. If the workflow completes, write back the outcome and update the context for the next loop.</p><p>That is where context management becomes the control plane for agentic work. Context is not just memory. It is the durable state that allows loops to evaluate correctly across time. It tells the agent what exists, what matters, what has already been tried, what constraints apply, and what success looks like. It turns discovery into an asset rather than an expense. It allows expensive, high-autonomy work to be distilled into lower-cost, higher-determinism execution.</p><p>So yes, the job is to write loops. But in the enterprise, the job is more specifically to design the evaluations that make those loops worth running. The loop is the structure: evaluate, execute, repeat. The evaluation is the leverage. Context is what lets that learning persist.</p><p>The teams that win with agents will be the teams that know how to define success, capture business context, preserve ontology, and convert repeated judgment into durable process. The next frontier is not agents that can keep going. It is systems that know what they are checking for when they do.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[What Woke the Agent?]]></title><description><![CDATA[Use tokens for managing business events, not checking for silence.]]></description><link>https://enterprisecontextmanagement.substack.com/p/what-woke-the-agent</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/what-woke-the-agent</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Thu, 04 Jun 2026 18:15:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rloR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At 08:00, the renewal-risk agent wakes because it is 08:00. It checks Salesforce, scans the support queue, reads a Teams channel, asks for product-usage changes, and pulls the last account note, which was already stale when somebody wrote it. Nothing has happened, so it returns to sleep with a clean run log and a small bill.</p><p>By lunch, it has done this a dozen times. The team may call that proactive AI. The finance system will call it compute. The platform team will call it load. The account manager will ignore the notifications because most of them say nothing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rloR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rloR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2233681,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/200521882?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rloR!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e295d7-6afc-4f96-b5ba-28361a63c9d7_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Then, at 14:37, the thing that matters happens. A support ticket from a renewal account is raised to severity one. Usage has already dropped. The executive sponsor changed jobs two weeks ago. Legal has an unresolved contract question. The event is sitting there in the business systems, crisp and timestamped.</p><p>The agent still waits for the next scheduled run.</p><p>When it finally wakes, it has to rediscover why the account matters. It loads too much context, misses the contract issue, pings the wrong owner, and drafts a generic account note that the customer-success lead quietly rewrites. Essentially, the save motion has started too late. The first internal escalation is noisy. The customer receives a thin response precisely when they need reassurance that someone truly understands their needs. Nobody built a broken model here, they built the wrong waiting architecture.</p><p>At <a href="https://openai.com/business/intelligence-at-work/">OpenAI&#8217;s Intelligence at Work livestream</a>, Sam Altman pointed to &#8220;constantly running proactive AI&#8221; as the next major phase after chat models and agent-style systems such as Codex. His larger point was that companies will have to change how they implement AI, starting now.</p><p>That implementation change is where things get interesting. If enterprises hear &#8220;constantly running&#8221; and respond by scheduling more model calls, they will recreate the oldest mistake in distributed systems with the most expensive component in the stack.</p><p>The architecture question has moved: what event gives the agent a reason to think?</p><p>That is the first question I would ask of any proactive AI design. If the answer is &#8220;a timer,&#8221; the system may still be useful. It can produce a Friday report, check a known queue, or send a recurring reminder. A timer proves that time passed. It says nothing about whether a business event occurred.</p><p>A serious proactive system starts with the event. A complaint crosses a severity threshold. Inventory falls below a buffer. A payment fails. A fraud score jumps. A supplier misses a milestone. A policy changes. A VIP customer goes quiet after a bad onboarding call. These are deterministic facts before they are reasoning problems.</p><p>Those signals rarely arrive ready for an agent. The runtime has to turn the raw change into a wake decision: resolve the business object, enrich the event just enough to route it, check the agent registry, and decide whether any hibernating agent should resume. Sometimes the answer should be no. Sometimes the event should be stored against an agent&#8217;s state and wait for a second signal. Sometimes it should wake the agent immediately because the next useful move requires judgment, language, or action.</p><p>That is the practical meaning of proactive AI at enterprise scale. The agent does not wander around the company looking for something to care about. The company tells the agent when something happened.</p><p>This distinction sounds small until it touches volume. Scheduled agents spend money on absence. Event-driven agents spend money on change. Once there are hundreds of agents watching thousands of queues, accounts, contracts, tickets, suppliers, documents, and policies, that difference stops being elegant architecture and becomes the operating cost of the AI program.</p><p>The agent runtime also needs a real lifecycle. Wake is only the first transition. A production agent should be able to acquire a work item, load its last checkpoint, take a lease on the relevant state, reason, call tools, write back a new checkpoint, and release the lease. If the work remains unresolved, it may stay active. If it is waiting on a human, a downstream system, a cooling-off period, or a future business event, it should hibernate with an audited state record.</p><p>That is the technical line I would draw. Perpetual loops are certainly not inherently wrong. Unbounded loops are. The problem is an agent that keeps thinking because nobody gave the runtime a safe way to stop, persist, and resume. A good design lets an agent stay awake while the decision is alive, and forces it to sleep once the next meaningful transition belongs somewhere else.</p><p>I think there are three somewhat unglamorous pieces that matter.</p><ul><li><p>The first is a trigger layer. Business systems have to publish state changes in a form other systems can trust. If a high-value account becomes a renewal risk, the signal should not live only as a human-readable note in a Teams channel. It should become a routable event.</p></li></ul><ul><li><p>The second is an agent registry. Every serious agent needs a declared surface area: events it cares about, entities it owns, state it maintains, tools it may use, actions it can propose, actions it can take, and conditions under which it must hibernate. Without that registry, orchestration becomes vibes at machine speed.</p></li></ul><ul><li><p>The third is audited state persistence. A hibernating agent has to remember where it left off, and the enterprise has to know what state it preserved. It should not rebuild its world from prompt fragments every time a scheduler wakes it. The account risk posture, last commitment, escalation history, open decision, pending dependency, and allowed next actions belong in durable state with a reason for the next wake-up.</p></li></ul><p>This is where Enterprise Context Management becomes more than a governance phrase. Context is the condition that lets the agent resume intelligently after the business changes. It includes the event that woke it, the state it carried into hibernation, the business meaning attached to the customer or inventory item or policy, and the record of what the agent did next. At AI One, this is the architectural layer we think enterprises will have to get right as agents move from chat into operational workflows.</p><p>Governance must also be in the room, but I would not let it run the meeting. The more urgent implementation error is simpler: we are seeing teams use reasoning where infrastructure should do the waiting. That choice will show up as cost, latency, duplicate work, noisy escalations, and agents that seem busy while missing the moment they were supposed to catch.</p><p>It also changes vendor evaluation. Connector lists are table stakes. Scheduled runs are table stakes. Teams integration is table stakes. In a serious architecture review, I would ask for the wake trace: the source event, the routing decision, the agents left asleep, the checkpoint loaded into the agent that woke, the hibernation policy, the action boundary, and the record after execution. That trace tells you whether the product has an architecture for proactive AI or a scheduling feature with a model behind it.</p><p>OpenAI and Anthropic&#8217;s current workspace-agent direction is already pointing enterprises toward shared agents that can run workflows across tools, schedules, and applications like Email and Teams. That is a useful bridge because it gets companies thinking beyond chat sessions. The next step is richer than a schedule. Proactive AI needs a business event fabric underneath it and a runtime that knows when to stop thinking.</p><p>The cost pressure will help force this discussion. <a href="https://www.axios.com/2026/06/02/altman-openai-top-token-user">Axios reported Altman saying OpenAI&#8217;s top token user was using about 100 billion tokens per month</a>, and that cost had become a huge issue amongst clients. Whether or not your enterprise is anywhere near that scale, the pattern is clear: once AI leaves the pilot budget, waste becomes visible. A single scheduled agent is easy to ignore. A fleet of agents polling every account, queue, channel, supplier, policy, and exception all day is a bill with an architecture diagram attached.</p><p>The right preparation is a wake-map exercise. Put one workflow on the wall. For renewal risk, write down the few events that deserve intelligence: severity-one ticket, usage drop, champion departure, procurement delay, legal blocker, failed implementation milestone. Draw the agents beside those events. Mark which agent owns account state, which one can act, which one should only recommend, what checkpoint it must write before sleeping, and which events wake a human immediately. If nobody in the room can say where the agent sleeps between those moments, the design is still a scheduled automation wearing an agent badge.</p><p>Altman also said that if there were one thing to prepare for over the next year, he would pick proactive AI. I agree. The preparation starts in the design review, with one awkward question on the table: what wakes this agent, what state does it carry back to sleep, and why is the model still thinking after the business event has moved somewhere else?</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[From Forward Deployed Engineers to Forward Deployed Software]]></title><description><![CDATA[Forward deployed engineers can earn the first use case. Forward deployed software is what turns the second, third, and tenth use case into repeatable ROI.]]></description><link>https://enterprisecontextmanagement.substack.com/p/from-forward-deployed-engineers-to</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/from-forward-deployed-engineers-to</guid><dc:creator><![CDATA[Fergus Keenan]]></dc:creator><pubDate>Thu, 14 May 2026 17:14:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YeKu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The market has discovered forward deployed engineers. That does not mean customers actually want them.</em></p><p>In the last two weeks, the three leading frontier labs have each reached for the same instrument. <a href="https://openai.com/index/openai-launches-the-deployment-company/">OpenAI launched the OpenAI Deployment Company</a> with more than $4 billion in backing at a $10 billion valuation, anchored by TPG, Bain Capital, Brookfield, Goldman Sachs, McKinsey and Capgemini. <a href="https://www.theinformation.com/briefings/google-hire-hundreds-engineers-help-customers-adopt-ai?rc=jjsw78">Google Cloud CEO Thomas Kurian announced</a> plans to hire hundreds of forward deployed engineers to form a new team inside Google Cloud. Anthropic partnered with Blackstone, Hellman &amp; Friedman and Goldman Sachs on <a href="https://www.cnbc.com/2026/05/04/anthropic-goldman-blackstone-ai-venture.html">a $1.5 billion venture</a> to embed engineers inside private equity portfolio companies and redesign workflows around Claude. The framing in each case is roughly the same: <strong>enterprise AI is hard, customers need help, send the engineers.</strong> <strong>This is a symptom of where we are in the adoption curve, not the answer to it.</strong> So what does the answer actually look like, and what are customers actually buying?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This rise of the forward deployed engineer tells us that the market has found a real problem: AI does not become valuable just because a model is available, a workflow has been demoed, or an integration exists in theory. <strong>The hard part begins when the system meets the customer&#8217;s actual operating environment</strong>, where the real process is not quite the documented process, handoffs are messy, ownership is unclear, approvals depend on context, and exceptions are often known by everyone but written down by no one.</p><p>In that environment, it makes sense that companies are sending engineers closer to the customer. Someone has to understand how the work actually happens before software can do anything useful with it. <strong>The mistake is to assume that because forward deployment is often necessary at the beginning, it is what the customer ultimately wants to buy.</strong></p><p>The customer does not wake up wanting a forward deployed engineer. They do not want another technical team embedded in their business, another delivery model to manage, or another standing implementation function that quietly becomes part of the operating model. What they want is simpler and much harder: <strong>they want the work done, inside the environment where the work already happens, with enough reliability and control that they can trust the outcome.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YeKu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 424w, /__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 848w, /__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YeKu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39530,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/197544423?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 424w, /__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 848w, /__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YeKu!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdfa0a86-d47b-4779-8d8d-2e5e51b58a15_2400x1350.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>The first use case is allowed to be messy</strong></h2><p>There is nothing wrong with using people to get the first use case right. In enterprise AI, it is often the only honest way to begin.</p><p><strong>The first deployment is where the real workflow is discovered.</strong> It is where you learn which system is trusted, which process is mostly performative, which exception matters, where the risk sits, and what evidence the customer needs before they will let the software act with confidence. These things are not always available in a requirements document because the requirements document is often a partial fiction. It captures what the process is supposed to be, not necessarily how value, judgment, and accountability actually move through the organization.</p><p>That is the useful role of the FDE. They get close enough to the customer to expose the context that the product does not yet know how to capture for itself. They translate operational reality into something software can understand, and they find the gap between the demo and the deployment.</p><p>But that should be the point of the motion: to <strong>teach the software.</strong> If every new deployment requires the same amount of human discovery, bespoke configuration, and engineering effort to make the customer successful, then the company has not found a scalable AI model. It has found <strong>a consulting model with a better interface.</strong></p><h2><strong>Consulting starts where product learning stops</strong></h2><p>The line between forward deployment and consulting is not the title on the team, <strong>it is whether the work compounds.</strong></p><p>If an engineer sits with the customer, understands the workflow, builds around the exceptions, and leaves behind a system that makes the next use case faster, more reliable, and less dependent on human interpretation, <strong>that is product learning.</strong> If the same engineer, or another one just like them, has to come back and repeat the same exercise again and again, <strong>that is consulting.</strong></p><p>This distinction matters because AI companies can easily hide a services business inside a software story. They can call it deployment velocity, customer obsession, outcome orientation, or any other phrase that makes the human effort feel strategic. Some of it is strategic. The danger is not being close to the customer; the danger is becoming <strong>permanently dependent</strong> on that closeness.</p><p>Customers will accept human help when it accelerates value, but they will not confuse it for the outcome. A long-running FDE engagement may feel reassuring at first because it means smart people are paying attention. Over time, though, it starts to <strong>look like exactly what customers were trying to avoid:</strong> another project, another dependency, another team whose knowledge lives partly in software and partly in people&#8217;s heads.</p><p>The promise of AI in the enterprise is not that customers get a more technical consulting team, it is that more of the work can move into software without losing the context, control, and judgment that made the human process work in the first place.</p><h2><strong>Forward deployed software</strong></h2><p>That is why the more interesting idea is <strong>forward deployed software</strong>.</p><p>Not software that waits for the customer to adapt to it, and not software that sits outside the workflow asking people to come to yet another place to get value. Forward deployed software enters the customer&#8217;s operating environment, <strong>learns how work is actually done,</strong> and gradually takes on more of the deployment burden that would otherwise sit with an FDE.</p><p><strong>This does not mean pretending every workflow can be fully automated.</strong> That is usually where AI products become either brittle or irresponsible. In many enterprise processes, judgment still matters. Subject matter experts still matter. Review, escalation, auditability, and <strong>deterministic controls still matter</strong>. The point is not to remove humans from the loop everywhere; it is to stop using humans for work that software should be able to handle once the first deployment has taught it what matters.</p><p>Forward deployed software should learn the workflow, remember the exceptions, and understand where probabilistic reasoning is useful versus where deterministic logic is required. It should know when a task can be completed automatically, when it needs human review, and what evidence must be preserved so the customer can trust the result. In other words, software should start taking on the tasks that made forward deployment necessary in the first place.</p><p><strong>That is what changes the economics for the customer.</strong> The first use case may require a heavier lift, but the second should be easier and the third should be faster. By the tenth, the customer should not feel like they are starting again. They should feel the return compounding.</p><h2><strong>The customer buys the outcome</strong></h2><p>Context graphs, integrations, and workflow memory all matter. But they are not what the customer is buying. They are the machinery behind the outcome. The customer buys the confidence that the work will happen correctly, consistently, and with less human effort over time.</p><p>That is the useful way to think about the infrastructure layer. It matters because enterprise AI cannot own real work if it does not understand the environment it is operating inside. It needs access to the systems, rules, exceptions, permissions, approvals, and prior decisions that shape how work actually gets done. But those things only matter if they make the outcome faster, safer, cheaper, or more repeatable.</p><p><strong>No customer is buying a context graph because they want a graph.</strong> They are buying fewer stalled processes, faster approvals, cleaner handoffs, less operational drag, and more confidence that the same result can be delivered again without assembling another project team around it.</p><p>That is the customer-centric test for AI deployment. Did the work get done? Did it happen in the environment where the business already operates? Did it require less human effort over time? Did the next use case benefit from the last one?</p><p>If the answer is yes, then forward deployment has done its job. <strong>If the answer is no, then the FDE has become the product,</strong> and the customer is back in the world of consulting, even if everyone is using more modern language.</p><h2><strong>The real prize</strong></h2><p>The real prize is not more forward deployed engineers. It is <strong>software that can do more of what forward deployed engineers are currently being asked to do:</strong> understand the customer&#8217;s environment, translate messy workflows into executable systems, preserve the right controls, involve experts where judgment matters, and turn the first deployment into a faster path for the next one.</p><p><strong>That is the difference between a services motion and a compounding software-based outcome.</strong> It is also the difference between a customer buying another project and a customer buying repeatable ROI.</p><p>Forward deployed engineers can be the bridge, but if you need them for too long, the bridge has become the destination. Customers do not want the bridge. They want to get to the other side.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Stop Hiring Your Agents]]></title><description><![CDATA[The enterprise doesn't need digital employees. It needs to rethink how work gets done.]]></description><link>https://enterprisecontextmanagement.substack.com/p/stop-hiring-your-agents</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/stop-hiring-your-agents</guid><dc:creator><![CDATA[Fergus Keenan]]></dc:creator><pubDate>Wed, 22 Apr 2026 16:18:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wAo3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Walk into the majority of enterprise AI pitches right now and you&#8217;ll hear a version of the same story: the agents are here, they come with job titles, and you manage them the way you manage people. Workday has announced an &#8220;Agent System of Record,&#8221; an HR platform for your digital employees where you hire, onboard, assign responsibility, and manage outcomes &#8220;the same way businesses manage people.&#8221; Salesforce is shipping pre-built role templates (HR Agent, Finance Agent, Banker Agent) alongside its new Headless 360 announcement that exposes the entire platform as APIs and MCP tools. Microsoft is talking about &#8220;agent bosses&#8221; and &#8220;human-agent ratios&#8221; as if we&#8217;re optimizing a staffing model.</p><p>The common thread isn&#8217;t really about pricing or packaging, it&#8217;s about rerunning the same playbook. These incumbents are taking the org chart we built for scarce human intelligence and laying it down as the template for abundant machine intelligence, and the framing is intuitive enough that most buyers aren&#8217;t questioning it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!wAo3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!wAo3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1891065,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/195050819?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wAo3!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445f6b52-1475-44a0-b302-c45a30bbf335_4368x3144.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Anthropomorphism was a useful bridge but it&#8217;s the wrong destination.</strong></p><p>When ChatGPT launched, making AI feel human (conversational, personable, a little quirky) was a smart move because it lowered the barrier. People talked to it like a colleague and it worked. The anthropomorphism served a purpose: it made an alien technology approachable.</p><p>The problem is what happened next. A UX choice that worked at the chat interface got promoted into an operating model, and the metaphor that helped a single user feel comfortable talking to a chatbot is now being used to describe how entire enterprises should deploy, govern, and manage AI at scale. That&#8217;s a much bigger claim, and the metaphor doesn&#8217;t stretch that far. Making something feel familiar to a first-time user is not the same thing as designing how an organization should run.</p><p><strong>Roleification places unnecessary limits.</strong></p><p>&#8220;Paul the Paralegal&#8221; and &#8220;Sally the Salesperson&#8221; are agents defined by the specialization constraints of existing knowledge workers: an HR Agent that can only do what an HR person does, a Finance Agent bound by the walls of the finance department.</p><p>When you &#8220;onboard&#8221; an agent into a role, you are telling it to think within those walls. Of course, the agent does need to respect the access controls, data boundaries, and policies that define the task, and that part genuinely matters. But beyond those guardrails, the whole promise of generalized intelligence is that the thinking doesn&#8217;t have to be bound by the same walls that the job description is. Within what it&#8217;s permitted to see, the agent should be free to reason across whatever the work requires. Roleification optimizes for familiarity, making AI look like something a manager already understands, and familiarity is not the same as capability. The gap between the two is where all the value leaks out.</p><p><strong>The work, not the worker.</strong></p><p>The reframe that matters is this: stop thinking about what agents <em>are</em> and start thinking about what work needs to get done.</p><p>Take employee onboarding, which doesn&#8217;t really live in HR. It touches recruiting, IT provisioning, facilities, payroll, compliance, and the hiring manager&#8217;s calendar, and it&#8217;s slow today precisely because it&#8217;s distributed across six departments with their own systems, handoffs, and queue times. The HR Agent framing doesn&#8217;t fix this, it just automates HR&#8217;s slice and leaves the handoff friction untouched.</p><p>The more useful approach isn&#8217;t deciding which department an agent belongs to, but what outcome the business actually wants, what information is needed to get there, and what the fastest path to completion looks like. That path will almost always cut across the silos you&#8217;ve built, which is the whole point. The organizations starting to see real returns from AI aren&#8217;t the ones that bolted agents onto existing processes, they&#8217;re the ones willing to rethink the processes underneath.</p><p><strong>This is process reengineering, not headcount planning.</strong></p><p>Vendors don&#8217;t lead with this message, because process reengineering is hard, unsexy, and impossible to sell as a per-seat license, but it&#8217;s where the value actually sits.</p><p>Some of the incumbents are starting to acknowledge, at least in their marketing, that the game is moving. Salesforce going headless is the most visible example: the company that invented per-seat SaaS is repackaging itself as an API surface for agents, which is a recognition that the agent is the interface and the data layer is where the value will be captured, while also being a move to make sure they remain the ones capturing it. Our view has always been that the systems of record were going to have to become data companies, and exposing the API surface is the price of admission rather than a gift to the ecosystem.</p><p>When intelligence becomes abundant, and any node in your enterprise can reason, synthesize, and act, the binding constraint changes. It&#8217;s no longer whether you have enough smart people in the right roles, it&#8217;s whether your information architecture is set up to let intelligence flow to where the work actually is. That&#8217;s a fundamentally different question, and it demands a fundamentally different investment: not in agent personas, but in data flows, in breaking down the information silos that were built for a world where intelligence was scarce and expensive and had to be rationed into departments.</p><p>Your company&#8217;s competitive advantage was never the org chart anyway. It was the institutional knowledge, the customer relationships, and the proprietary processes, the things that actually differentiate you. Agents don&#8217;t protect that advantage by mimicking your current structure, they unlock it by making that knowledge actionable across every surface of the business.</p><p><strong>The real infrastructure investment.</strong></p><p>Getting this right starts with the layer underneath the agents. Before you deploy anything that calls itself an agent, you need a runtime that can govern what it sees, what it does, and what it&#8217;s allowed to conclude. Your proprietary data, business rules, and strategic priorities need to be extracted, structured, and made available as living context, rather than trapped inside department-specific tools or siloed in applications that don&#8217;t talk to each other. The question isn&#8217;t which agent you deploy, it&#8217;s whether any intelligence, human or machine, can access the right context, act on it safely, and be held to account for the outcome.</p><p>That means investing in how your organizational knowledge is modeled, how context is delivered to the point of work, how memory accumulates across interactions, and how every action an agent takes can be governed and audited. The cross-cutting layer where all of that lives cannot be owned by any single vendor whose data it&#8217;s governing access to. It has to sit on your side of the line, portable across models and across systems, or the independence you think you&#8217;re buying is an illusion.</p><p>When that foundation is in place, you don&#8217;t need to &#8220;hire&#8221; an HR agent or a Finance agent. You compose capabilities (reason, research, draft, decide, escalate) dynamically around whatever the work demands. You give your agents a goal and guardrails.</p><p><strong>Where we think things will go.</strong></p><p>The companies that will look back on this era and wonder what they were thinking are the ones building agent org charts today, meticulously onboarding digital employees into the same departmental silos that have been slowing them down for decades. Right behind them will be the ones who accepted a vendor&#8217;s headless pivot as the whole answer, and only discovered later that the runtime governing their agents was never really theirs.</p><p>The ones that will win are doing something less photogenic but far more consequential: rewiring how their business thinks. Not adding AI to the machine but rebuilding the machine around what AI makes possible.</p><p>Your agents don&#8217;t need job titles. Your enterprise needs intelligent plumbing.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Single-Provider Trap Is Coming to Enterprise AI ]]></title><description><![CDATA[The most expensive AI decision many companies will make in the next few years will feel like the safest one.]]></description><link>https://enterprisecontextmanagement.substack.com/p/the-single-provider-trap-is-coming</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/the-single-provider-trap-is-coming</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Thu, 16 Apr 2026 19:21:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!na9T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The most expensive AI decision many companies will make in the next few years will feel like the safest one. It will look just like many previous cloud decisions did: pick a vendor, move fast, and assume portability can be cleaned up later. In AI, that means choosing a single frontier model provider, adopting its hosted agent stack, and assuming any architectural consequences can be unwound down the road. That works for a pilot because everything is convenient at the same time - one SDK, one hosted memory layer, one eval surface, one commercial relationship, one vendor telling a coherent story. Then the workload shifts from chat to agents, the spend curve stops being linear, the best model for one class of work is no longer the best model for another, and the framework that made it easy to get started begins to make leaving feel less like a migration and more like open-heart surgery.</p><p>That is why I increasingly think LLM agnosticism is becoming a strategic requirement for enterprise AI, not a technical preference or a purity argument. The recent <a href="https://www.theinformation.com/articles/anthropic-changes-pricing-bill-firms-based-ai-use-amid-compute-crunch">Information Report</a> on Anthropic&#8217;s changes to enterprise pricing for heavy business usage is significant because it shows how quickly the economics can shift once workloads become serious. Providers will price for their own compute constraints, capacity limits, and margin objectives, and they will work hard to pull customers deeper into their own hosted agent ecosystems. That is entirely rational from their side. It is simply not a stable basis on which to build the rest of your company&#8217;s operating mode in a market this volatile, this politically exposed, and this fast-moving.</p><p>The durable asset is not the model endpoint. It is the agentic system that governs how models interact with your data, memory, context, tools, and policy. If that layer belongs to you, providers compete for your workloads, the best models can be chosen task by task, and switching remains an engineering exercise rather than a strategic crisis. If that layer belongs to them, your economics, privacy posture, and exit costs are all downstream of somebody else&#8217;s platform strategy. Avoiding that trap requires more than a gateway. It will require owning the runtime.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!na9T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 424w, /__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 848w, /__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 1272w, /__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!na9T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2691714,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/194440499?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 424w, /__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 848w, /__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 1272w, /__u/substackcdn.com/image/fetch/$s_!na9T!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ea983e6-00a6-43d8-96e9-f9079a10d2ff_4608x3072.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Why This Is Becoming Urgent</h3><p>The Anthropic pricing change is a useful signal because it exposes a deeper truth about agentic AI economics. Once usage shifts from seat-based productivity into long-running coding agents, scheduled work, tool use, retries, and autonomous workflows, pricing stops being a clean per-user abstraction and starts to reflect raw consumption. The Information piece suggests that some heavy Claude Enterprise users could see costs double or even triple under the new structure, and it notes similar usage-sensitive moves elsewhere in enterprise AI. This is probably not an outlier, but what happens when providers discover where the expensive workloads are.</p><p>And price is only one axis of dependency. The frontier model market is moving too quickly, its economics remain too unsettled, and its policy environment is too uncertain for any serious company to commit its architecture to one provider&#8217;s roadmap. You do not know which firms will still be structurally advantaged in three years, whose gross margins will hold, which models will end up constrained by regional regulation or procurement rules, or where governments will decide strategic control points need to sit. Frontier models are already close enough to strategic infrastructure that export controls, access restrictions, or safety interventions can no longer be treated as remote possibilities.</p><p>That is too much uncertainty to absorb passively. The practical response is to make model choice reversible.</p><h3>The Best Model Depends on the Job</h3><p>This matters because different model families genuinely excel at different kinds of work, and the best answer is rarely &#8220;pick one and standardize everything around it.&#8221;</p><p>In many enterprise teams, the picture has already shifted noticeably over the last three months. For coding workflows in particular, many teams have begun moving from Claude toward Codex, not because Claude has stopped being useful, but because OpenAI is often better at following precise instructions, operating inside tighter agentic control loops, and using tools in a more disciplined way over longer execution chains. Claude still has a distinct strength of its own: it is often better at finding a user&#8217;s underlying intent inside messy, ambiguous conversation, which matters in workflows where the hardest problem is understanding what was actually meant before any action is taken.</p><p>Neither choice should dictate the architecture of the whole company.</p><p>There is another advantage to model plurality that is still underused: one model can check another&#8217;s work. A model that proposes an action is often a poor judge of its own blind spots, especially if the same family is also being asked to validate the output, grade the result, and decide whether to continue. Once you let a second model from a different provider review the work, or better still attack it, you reduce correlated failure. One model writes the code change, another looks for the edge case; one proposes a customer action, another searches for the compliance or policy problem; one drafts the plan, another tries to break it. What matters is not merely a second opinion, but a second failure distribution.</p><p>That diversity also gives you a more rational cost structure. Not every step justifies the most expensive model, and not every task should be solved with the model you happened to start with six months earlier. The ability to route by task, validate across families, and change those choices without a replatform is what low switching cost actually looks like in practice.</p><h3>The Real Lock-In Is Where Enterprise Meaning Lives</h3><p>The strongest form of lock-in is not the API call. It is where your memories, context, ontology, and workflow semantics end up living.</p><p>Hosted agent frameworks are attractive because they compress months of work into days. They ship with managed memory, built-in tracing, default planners, integrated evals, native tools, and smooth developer ergonomics. That convenience is real, but so is the functional limit. Most are optimized for their own abstractions, their own storage assumptions, their own tool semantics, and their own view of governance. They tend to treat context as a provider-native payload rather than a governed architectural layer, which is exactly why they feel so productive at the start and so constraining later.</p><p>This is where sovereignty stops sounding abstract and becomes operational. Data sovereignty in this sense does not just mean &#8220;keep the files private.&#8221; It means the enterprise owns its memories, context layers, business ontology, approved tools, versioned queries, access rules, and audit trail. It means a model can invoke a reviewed and approved query by name rather than inventing live query logic in production. It means credentials remain at the gateway or data-access layer rather than inside the agent loop. It means the map of enterprise meaning, how customers, products, policies, mandates, exceptions, and actions are actually defined, stays on your side of the boundary.</p><p>That is the layer that compounds with use, and it is the layer you do not want to hand away.</p><p>We have seen this pattern before in the cloud. Single-provider choices did not usually hurt in year one; they became painful in year three, when data gravity, egress, proprietary services, and rewrite costs converted earlier speed into later dependency. AI providers are now trying, quite sensibly, to move up the same stack. They do not want to provide only inference; they want the hosted agent framework, the memory layer, the eval system, the tracing surface, and the developer workflow that makes your application harder to move. The easiest architecture to start with is becoming, again, the hardest one to leave.</p><h3>Agnostic Does Not Mean Generic</h3><p>There is, however, one important caveat that gets missed in simplistic &#8220;multi-model&#8221; discussions: an LLM-agnostic system is not just a gateway.</p><p>The transport layer is the easy part. The hard part is behavior. Different model families want different prompt structures, tool descriptions, context packing strategies, result validators, retry policies, stop conditions, and cache behavior. Real agnosticism does not mean pretending Claude, OpenAI, and a local open model are interchangeable. It means keeping the invariant layers, your data fabric, memory, ontology, policy controls, audit, and tool execution, on your side, while versioning prompt packs and task variants by model family and even by model version so each one is optimized for the work it is doing.</p><p>If you want low switching cost without dropping into lowest-common-denominator behavior, you need a canonical internal representation of prompts and tasks, with provider-specific translation at invocation time, and you need different prompt versions retained for different families because the same task often performs differently across them. In other words, low switching cost is not achieved by denying model differences; it is achieved by containing them.</p><p>This is one of the reasons we designed our ContextOne architecture in the way we did. The runtime is LLM-agnostic, but it does not assume identical behavior across model families. It preserves the enterprise&#8217;s control over memory, ontology, audit, and policy while allowing prompts and tasks to be optimized for specific providers and versions. That is the appropriate kind of complexity, because it buys flexibility where it matters and specialization where it pays.</p><h3>Local Open Models Belong Beside the Frontier Loop</h3><p>Once you own orchestration, smaller local open models stop being ideological and start being practical.</p><p>They can run in side-chains alongside the main agentic loop, handling tasks that are narrow, repetitive, privacy-sensitive, or simply not worth frontier-model pricing. A local model can classify inbound documents, normalize entity names across messy systems, rerank retrieved passages, or perform first-pass policy and code checks before the frontier model is called.</p><p>These side chains can run in parallel with the main loop, which means the gain is not only lower cost but also lower latency and reduced exposure. The frontier model spends its budget on reasoning rather than housekeeping, while the most sensitive or repetitive steps remain inside the enterprise boundary. For regulated businesses, that hybrid pattern is usually much better than the false choice between &#8220;everything hosted&#8221; and &#8220;everything local.&#8221;</p><h3>Board-Level Takeaway</h3><p>The strategic question is not which model vendor looks strongest this quarter. It is whether the company owns the runtime that governs data, memory, ontology, tools, and policy, because that layer determines pricing leverage, privacy posture, resilience, and exit cost. It is the layer that compounds with use. A company that owns it can let model vendors compete inside its system; a company that does not is effectively underwriting the platform risk of the vendor it chose first.</p><h3>Conclusion</h3><p>For most serious enterprises, the disadvantages of running their own agentic system in-house are front-loaded and manageable, while the benefits compound over time. You take on more architecture, more governance work, and more operational discipline up front; in return, you get lower structural risk, better cost control, stronger privacy, clearer ownership, reduced lock-in, and the freedom to use the best model for each task rather than the model you happened to standardize on early.</p><p>None of this means provider ecosystems are useless. They will remain valuable learning surfaces, and for many teams they will still be the fastest way to get started. But as soon as AI touches core workflows, proprietary data, and real budgets, convenience stops being the right optimization target. Risk, cost, lock-in, ownership, and privacy move to the front of the queue.</p><p>That is why the more durable pattern is to keep the context and control architecture on the enterprise side of the boundary, let frontier and local models work together inside it, and make model choice reversible. We build around that principle: LLM-agnostic at the runtime layer, optimized by model family where it matters, and structured so the enterprise keeps the memories, ontology, and controls that actually become more valuable with use.</p><p>The companies that benefit most from this cycle will not be the ones that guessed the eventual winning model provider. They will be the ones that refused to bet the rest of their architecture on that guess.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Evaluation Is How Agentic AI Earns the Right to Act. Here's What We're Finding.]]></title><description><![CDATA[If memory is how an agent learns, evaluation is how it earns the right to act.]]></description><link>https://enterprisecontextmanagement.substack.com/p/evaluation-is-how-agentic-ai-earns</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/evaluation-is-how-agentic-ai-earns</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Wed, 15 Apr 2026 14:10:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iZS_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>An agent books a seat before checking availability. Another approves a loan that fits the customer request but violates the firm&#8217;s risk posture. A third says it has updated a reservation, but the downstream system never changed at all. All three are failures, but of very different types, and none should be handled by the same mechanism.</p><p>This is where we see some of the current conversations about &#8220;evals&#8221; breaking down. The term has become a catch-all for runtime guardrails, policy review, task verification, benchmark scoring, and sometimes deterministic controls that are not really evaluations in the first place. The lack of precision matters. If you collapse all of those into one category, you end up using the wrong instrument for the wrong job. The industry still tends to treat agent failure as a model-quality problem, however in our experience it is increasingly a runtime-design problem.</p><p>At AI One, we have been working on this problem inside ContextOne, our enterprise agent runtime. What has become clear is that evaluation only becomes reliable when it is treated as part of the architecture. Not a score at the end. Not a judge prompt bolted onto a swollen context window. A governed system that decides what the agent may do next, what it must prove before it continues, and what the platform should learn from the result.</p><p>These are early findings, not final answers. But they suggest that evaluation is less a single technique than a stack of distinct mechanisms, each matched to a different class of failure.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iZS_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 424w, /__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 848w, /__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iZS_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png" width="1456" height="966" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:966,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3222143,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/194297295?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 424w, /__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 848w, /__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iZS_!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57867aa1-f829-4627-9825-04aaa616ff15_4029x2674.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Why &#8220;Evals&#8221; Is Too Broad a Word</h3><p>In traditional software, verification is easier because the system is deterministic. A function returns the right output or it does not. Agentic systems behave differently. The same model can take different paths to the same objective. It can look reasonable while violating sequence, policy, or downstream state. Once an agent can read enterprise data, call tools, and write to live systems, &#8220;does the response look plausible?&#8221; stops being a serious standard.</p><p>Manufacturing figured this out a long time ago. You do not use end-of-line inspection to decide whether a bolt should have been tightened earlier in the assembly process, and you do not use a torque sensor to decide whether the finished engine actually runs. Different failure modes require different checks at different points in the process. Agentic systems are no different.</p><p>Consider an investment-bank operations agent running a pre-open allocation workflow. If it allocates a block trade before checking inventory, that is a sequencing failure. If it proposes an allocation that fits the mechanics but breaches client mandate or desk risk posture, that is a judgment failure. If it reports the booking as complete but the order management system still shows the old state, that is an outcome failure. The stakes are not academic: broken trades, manual repair, capital and compliance exposure, and the familiar question of why was a proof of concept allowed anywhere near production?</p><p>In practice, four different questions hide inside the word &#8220;evaluation&#8221;:</p><ol><li><p>Can the agent take the next step?</p></li><li><p>Is the proposed action aligned with policy and intent?</p></li><li><p>Did the claimed outcome actually occur in the underlying system?</p></li><li><p>Does the architecture materially improve autonomous task completion across realistic workflows?</p></li></ol><p>A resilient system answers each with a different mechanism.</p><p>If those questions are collapsed into one category, system design gets confused. Teams ask an LLM judge to do work that deterministic logic should have handled. They use benchmark scores as a proxy for runtime safety. Every control gets called an eval and sight is lost of what each component is there to solve.</p><h3>Runtime Evaluation Has Three Jobs</h3><p>Our research indicated we needed to settle on a three tier architecture. Instant Eval asks whether the next move is <strong>allowed</strong>. The LLM Judge asks whether the move is <strong>aligned</strong> with higher-level intent. The Agentic Grader asks whether the job was actually <strong>done</strong>.</p><h4>Instant Eval: Can the agent take the next step?</h4><p>Instant Eval enforces deterministic, real-time constraints on sequencing and preconditions of tool calls. Before an agent can execute a tool, the runtime checks whether the required prior steps have occurred. If the check fails, the agent receives actionable feedback rather than a silent rejection.</p><p>In an airline workflow, if the agent attempts to reserve a seat before it has checked availability, the call should be blocked and the agent should be told exactly which prerequisite is missing. That sounds simple. It is also the difference between a self-correcting system and a corrupted record in production. A surprising share of enterprise failure is structural, not conceptual.</p><h4>LLM Judge: Is the decision consistent with policy and intent?</h4><p>Some decisions cannot be reduced to sequence. They depend on judgment, tradeoffs, and principles that are too broad or too fluid to encode as a strict set of preconditions. This is where the LLM Judge belongs. It sits at designated checkpoints and evaluates whether a proposed action aligns with higher-level objectives and constraints, using only the slice of context relevant to that assessment.</p><p>Consider a loan approval workflow. An agent proposes a $50,000 loan to a customer with a credit score of 620. The sequence may be correct. The decision can still violate the firm&#8217;s risk posture. That is a judgment problem, not a sequencing problem. This is where an LLM Judge is useful, and also where it should stop. It should not be asked to do work that deterministic logic can do better.</p><h4>Agentic Grader: Did the work actually happen?</h4><p>The final runtime question appears when an agent says it is done. In many deployments, this is still the weakest link. An agent can produce a convincing summary of a task it did not actually complete. It can claim success after making only part of the required change. It can update the wrong record and still sound confident.</p><p>The Agentic Grader verifies the outcome against the underlying system. For tasks that can be checked deterministically, it can call Python scripts, database queries, or gateway tools directly. For tasks that require a more interpretive check, it can dispatch dedicated sub-agents with introspection tools and feed their output into a separate judge. If an agent says it updated a customer&#8217;s seating assignment and the reservations system does not show the change, the task is not complete. The agent goes back and fixes the discrepancy.</p><p>That separation matters because agents fail before an action, at the point of action, and after the work is supposedly complete.</p><h3>Some Controls Should Not Be Left to Evals</h3><p>One of the most common confusions in agentic AI is treating every control surface as an eval. Some mechanisms exist to judge behavior or outcome. Others exist to make certain classes of failure impossible. Both matter. They do different jobs.</p><p>Formal constraint enforcement is the clearest example. Some business rules are too important to leave to probabilistic judgment. In ContextOne, those rules can be translated into logical constraints and checked with a formal constraint solver before execution. If a proposed action violates the rule, the action is blocked. A loan approval that conflicts with a formal credit policy should not be &#8220;graded down&#8221; later. It should fail closed.</p><p>In regulated environments, query generation belongs in the same category. For example, in ContextOne, we use the concept of  Named Queries to let the agent invoke fixed, versioned query structures by name, supplying parameters rather than generating live query logic at runtime. Credentials stay out of the Agent Harness, access is enforced at runtime, and activity is recorded in a cryptographically signed audit ledger. These are not alternative flavors of eval. They are the architectural conditions under which evaluation becomes meaningful.</p><h3>Why This Belongs in the Runtime</h3><p>A useful evaluation framework cannot sit outside the agent as a wrapper around one giant prompt. It has to live where the work is being done. In ContextOne, that is the Agent Harness.</p><p>The architectural split is simple but important. The loop reasons and the harness governs. That separation creates a stable place to apply runtime checks, manage permissions, validate tool calls, gate results back into context, and halt or escalate when thresholds are reached. In ContextOne&#8217;s governed loop, we made that explicit: prompt assembly, context management, model invocation, result validation, tool execution, result collection and context gating, and control condition evaluation all happen inside a managed execution pipeline. Large tool outputs are offloaded and referenced rather than blindly stuffed back into context. Cost and iteration limits are checked before an agent is allowed to continue.</p><p>The economics matter too. ContextOne is stateless and event-driven. Agents materialize when work is required, yield resources between iterations, and resume from shared state in single-digit milliseconds. That makes governed, step-by-step execution practical at enterprise concurrency.</p><p>The practical takeaway is simple: a model that looks intelligent in a proof of concept is not production-ready until you can explain what stops it, what judges it, what verifies its work, and what it is deterministically prevented from doing.</p><h3>Benchmark Evaluation Answers a Different Question</h3><p>Runtime evaluation tells you whether a live task should proceed. Benchmark evaluation answers a different question: does the architecture materially improve autonomous task completion across realistic workflows?</p><p>That is the role benchmark evaluations play. On Tau Bench 2, ContextOne achieved 95% task completion in the telecoms domain against a baseline of roughly 30%. In the airline domain, it achieved 78% on a manually verified solvable subset against a roughly 40% baseline, and 52% on the full task set against a roughly 25% baseline. The same prompts were used for both the baseline and ContextOne runs. The gains came from architecture rather than prompt optimization.</p><p>But benchmark evaluation should not be confused with runtime safety. A Tau score does not tell you whether a particular loan approval, trade allocation, or customer update should be blocked right now. It tells you whether the architecture has moved the system from constant-supervision territory toward genuine delegation. At around 30% task completion, an agent still demands babysitting. At 90%+, you are starting to talk about real business process execution, with human oversight reserved for genuine exceptions.</p><h3>The Board-Level Test</h3><p>Before any agent moves from proof of concept to production, management should be able to answer four plain-language questions: what blocks it, what judges it, what verifies that the work actually happened, and what it is deterministically prevented from doing. &#8220;The model looked good in testing&#8221; is not a control framework.</p><h3>What Comes Next</h3><p>What comes next is not a search for one universal eval. It is better architecture for separating questions that are currently collapsed into that one word.</p><p>Which actions should be blocked deterministically? Which decisions should be reviewed against principles? Which outcomes must be verified against live system state? Which failures belong in a benchmark, and which belong in the runtime? Which signals should feed memory and improve the next execution? These are the design questions that matter.</p><p>And this connects directly to memory. In our earlier <a href="/__u/enterprisecontextmanagement.substack.com/p/memory-is-the-next-frontier-for-ai">SubStack article on memory</a>, we argued that enterprise agents need governed ways to retain what happened, how tools should be used, and which contextual facts matter. Evaluation is what makes those memories trustworthy. Memory without evaluation compounds both skill and bad habits. Evaluation without memory catches the same failure forever. Joined together, they create a feedback loop in which successful trajectories become reusable patterns, failed actions become correctable guidance, and autonomy can expand only when the system has earned it.</p><p>Our current findings point in one direction. In enterprise AI, evaluation architecture may be one of the most under-leveraged variables in system performance. Models matter, and prompts matter. But in agentic workflows that touch real systems, the more consequential lever may be the architecture that determines what an agent is allowed to do, what it must prove, what it is prevented from doing altogether, and what the platform learns from the result.</p><p>We do not claim to have the final answer. The evidence so far suggests that the systems that earn production trust will be the ones that know when to block, when to judge, when to verify, and when to learn. They will treat evaluation not as a single score, but as a governed part of the runtime itself. That is the standard proof-of-concept projects should be held to, before anyone calls them production-ready.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Memory Is the Next Frontier for AI. Here’s What We’re Finding. ]]></title><description><![CDATA[Our approach to enterprise memory architecture, and what the early results show.]]></description><link>https://enterprisecontextmanagement.substack.com/p/memory-is-the-next-frontier-for-ai</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/memory-is-the-next-frontier-for-ai</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Thu, 26 Mar 2026 19:37:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/41f912ee-6fba-439f-992f-59919e3eba27_4522x2566.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Memory is one of the next major frontiers for AI. Models are increasingly capable, but they don&#8217;t retain anything between invocations. An agent can&#8217;t learn from yesterday&#8217;s execution to improve today&#8217;s. It can&#8217;t accumulate skill with a tool that&#8217;s used a hundred times. The industry knows this. What&#8217;s less clear is how to solve it.</p><p>At AI One, we&#8217;ve been tackling this &#8220;memory problem&#8221; for large enterprises. We&#8217;ve created a Memory Engine with three asynchronous pipelines that create a feedback loop between agent execution and agent learning. Each pipeline captures a different class of knowledge from each run and injects it into the next. These are early findings, not final answers. But they suggest structured memory may be a more consequential lever than model sophistication or prompt engineering.</p><p>This article shares our approach.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4ujv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4ujv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9497342,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/192231693?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4ujv!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27986ed0-8281-406c-9d41-06e147f1f74b_4368x3144.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>Why Memory Is the Frontier</strong></h3><p>Every serious engineering system has a feedback loop. Gas pipelines have PID controllers (Proportional, Integral, and Derivative). In machine learning, gradient descent makes training observable, debuggable, and repeatable. Software teams have CI/CD. The dominant agentic frameworks have none. Every invocation is independent. Context is assembled fresh each time: conversational history, tool definitions, and retrieved documents are appended into a single payload that grows until it hits the token limit. Nothing carries forward.</p><p>The consequences are many. An agent queries a technology catalog where the same vendor is listed as &#8220;IBM&#8221; in one system and &#8220;International Business Machines Corp.&#8221; in another. The agent fails on entity resolution. It failed on the same resolution yesterday. It will fail again tomorrow. An agent calls an API that returns 15,000 tokens, consuming budget that should go to reasoning. Another agent violates a business rule. Nothing stops it from repeating the same mistakes over and over.</p><p>The LLM isn&#8217;t the problem. The problem is the absence of a feedback loop to retain important learnings, discard what&#8217;s irrelevant, or compound skills over time.</p><h3><strong>Three Pipelines, Not One Memory</strong></h3><p>Our approach consists of three asynchronous pipelines. Each handles a different class of learned knowledge: episodic, procedural, and semantic. Their design is guided by the idea that agents need to remember different kinds of things; each requires distinct storage, retrieval, and update mechanisms. All three run in parallel with the main <a href="/__u/enterprisecontextmanagement.substack.com/p/stop-building-agent-chains-start">agent loop</a>.</p><h4>Episodic Memory: Trajectory Distillation</h4><p>The first pipeline captures episodic memory: what happened.</p><p>Every agent execution produces a trace: reasoning steps, tool calls, and outcomes. Episodic memory pipelines review these traces, cluster similar activities, and identify &#8220;golden trajectories&#8221; &#8212; execution paths that represent best practices for a given task class.</p><p>This isn&#8217;t log storage. The distillation process compresses full execution histories into reusable patterns. When an agent encounters a similar task later, it starts from the best-known path, not from zero. It retains the freedom to deviate when specifics demand it. Agents get faster and more accurate on repeated task classes. No manual prompt tuning required.</p><h4>Procedural Memory: Tool Learning</h4><p>Enterprise agents interact with dozens or hundreds of tools: APIs, databases, and internal services. Each has its own quirks, failure modes, and optimal invocation patterns. The procedural memory pipeline captures how to use a tool by analyzing tool outcomes and updates prompts that guide future tool use.</p><p>For instance, imagine an agent that queries a technology catalog where the same vendor appears under three names across three systems. First attempt: entity resolution fails. The procedural pipeline captures that failure, analyzes the tool&#8217;s response, and updates the guidance. Next invocation: the agent accounts for the naming inconsistency. We update our procedural memory again. And on it goes. Over time, the agent&#8217;s skill with each tool compounds.</p><p>The system doesn&#8217;t just remember that a tool exists. It remembers how to use it well.</p><h4>Semantic Memory: Contextual Facts</h4><p>The semantic memory pipeline captures facts. It extracts contextual information from agent interactions: business rules, user preferences, organizational hierarchies, and domain terminology. It builds a persistent, governed model of the enterprise&#8217;s operational reality.</p><p>Semantic memory is what allows an agent to know, without being told each time, that &#8220;net revenue&#8221; means something different in the EMEA division than in North America. Or that a particular approval workflow requires two VP sign-offs, not one. These facts accumulate. Because they&#8217;re stored in a structured layer rather than buried in conversational history, they can be shared across agents, audited, and corrected.</p><h3><strong>Context Management Memory Strategies</strong></h3><p>The three pipelines feed a catalog of context management strategies. These strategies are invoked during each agent iteration. They include techniques such as token-efficient reformatting, result offloading, interaction summarization, noise injection, and prompt cache management. Let&#8217;s examine two of these strategies in detail to illustrate problems that only surface in production and that no amount of prompt engineering can fix.</p><h4>Noise Injection: Breaking Degenerate Loops</h4><p>After lengthy sequences of repetitive tool calls, LLMs can become trapped in degenerate loops. The model produces the same sequence of actions, receives the same results, and repeats. Each individual step is technically correct. Deterministic controls can&#8217;t catch it because no single action violates a rule. The failure is emergent: the model has been, in effect, lulled by the regularity of its own output.</p><p>This failure mode was first documented by the Manus team. We address this by injecting noise: small, semantically meaningless perturbations introduced into the context flow at controlled intervals. These perturbations disrupt the repetitive pattern without altering the task&#8217;s semantics. The agent breaks out of the loop and resumes productive reasoning. It&#8217;s a pragmatic, empirically validated countermeasure to a problem that exists at the boundary between deterministic system design and probabilistic model behavior.</p><h4>Prompt Cache Management: Co-Designing Context and Cost</h4><p>Context management systems reformat, offload, summarize, and mutate context between iterations. Each mutation can invalidate cached prompt prefixes. This matters because prompt caching is one of the most effective cost and latency optimizations offered by modern LLM providers. Naively applying context management without considering cache implications results in a perverse outcome: reduced token counts per call but higher effective costs due to constant cache invalidation.</p><p>One way to address this problem is a &#8220;harness-managed prompt caching&#8221; strategy that&#8217;s aware of the full context management pipeline configuration. The harness knows which regions of the context are stable (system prompts, ontology definitions) and which are volatile (tool results, summarized history). It structures the prompt to maximize prefix reuse even as downstream content changes. This optimization is only possible when context management and cache management are co-designed rather than treated as independent concerns. The platform provides dedicated views into context stability and cache performance, so operators can tune strategies not just for accuracy but for economic efficiency.</p><p>The broad design choice of context management strategies is that they happen at the architectural level, not the prompt level. Each content class (tool results, conversational history, ontological facts, memory-derived guidance) is optimized independently. This keeps context windows small and clean as task complexity grows, instead of the monotonically ballooning payloads that characterize monolithic approaches.</p><h3><strong>Memory and Human-in-the-Loop Governance</strong></h3><p>Reducing business rule violations doesn&#8217;t come from memory alone. It requires deterministic controls, including an SMT solver that enforces formal logical constraints on agent behavior. Memory is what makes those controls practical.</p><p>Without memory, every invocation is a cold start. The agent has no record of which actions trigger compliance issues, which tool sequences produce reliable results, or which edge cases require human escalation. A memory-free workflow catches violations reactively; with memory, the agent is less likely to violate a rule in the first place. The result? Fewer interruptions. Fewer escalations. A more efficient human-in-the-loop workflow.</p><p>This is &#8220;tunable autonomy.&#8221; Enterprises need to adjust the tradeoff between speed and accuracy, cost and confidence. Human in and out of the loop. Memory helps make those tradeoffs. An agent with rich episodic and procedural memory can operate with higher autonomy because it has demonstrated competence. An agent encountering a new domain should operate with tighter controls. Memory is the mechanism that distinguishes the two.</p><h3><strong>The Compounding Effect</strong></h3><p>The most important property of the Memory Engine, based on what we&#8217;ve observed so far, is that its knowledge compounds continuously. Each execution improves the next. Skills learned from one agent&#8217;s interactions become available to others through the shared ontology. Tool guidance refined through procedural memory applies across every agent that uses that tool. Contextual facts persist and accumulate.</p><p>This separates a learning system from a stateless one. A stateless agent&#8217;s performance is bounded by prompt quality and context window capacity. A learning agent&#8217;s performance is bounded by the breadth of its accumulated experience. Over time, the gap widens. The stateless agent stays flat. The learning agent improves.</p><p>Early manufacturers who adopted electricity didn&#8217;t see transformative gains until they redesigned their factories around the technology, replacing centralized steam engines with distributed electric workstations placed where the work happened. The same principle applies to LLMs and memory. Bolting an LLM onto existing workflows produces marginal improvements. Redesigning the architecture so agents accumulate knowledge, refine skills, and operate within governed memory structures produces a different class of system. One that gets better with use.</p><h3><strong>What Comes Next</strong></h3><p>What comes next is not a search for one universal memory architecture. It is a better way to decide what should be remembered, what should be forgotten, and what should be treated as canonical. Episodic, procedural, and semantic memory have different update patterns, retrieval requirements, and failure modes, so any serious enterprise memory architecture has to remain extensible as new domains and new operational realities appear.</p><p>The harder problem is deciding which memories matter the most for the task in hand. Some memories carry more weight because they are recent. Others should dominate because they encode stable business rules or durable tool-use patterns. Some because they were decisive in a prior successful trajectory, even if they were not the most obvious part of the trace. That attribution problem is where our focus is now: ranking memories, using temporal lineage, and ascertaining which prior experiences materially improve the next execution.</p><p>Those answers will not come from benchmarks alone. They will come from production deployments, under real constraints of governance, latency, and cost, where compounding effects become measurable over weeks and months: skill transfer across workflows, fewer repeated failures, lower human intervention, and improved token economics.</p><p>Our early findings point in one direction. In enterprise AI, memory architecture may be one of the most under-leveraged variables in system performance. Models matter. Prompts matter. But in repeated, governed workflows, the more consequential lever may be the memory architecture that determines what an agent keeps, reuses, and learns from. We do not claim to have the final answer. The evidence so far suggests that the next major gains will come from agents that do not merely reason at inference time, but actively accumulate governed experience over time.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Beyond Logs: A Lineage-First Approach to Production Observability for AI Agents]]></title><description><![CDATA[The deployment of AI agents in production environments has revealed a fundamental mismatch in existing observability paradigms.]]></description><link>https://enterprisecontextmanagement.substack.com/p/beyond-logs-a-lineage-first-approach</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/beyond-logs-a-lineage-first-approach</guid><dc:creator><![CDATA[Shaun Laurens]]></dc:creator><pubDate>Wed, 11 Mar 2026 16:01:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xJF6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The deployment of AI agents in production environments has revealed a fundamental mismatch in existing observability paradigms. The established pillars of modern observability (logs, metrics, and distributed tracing) were designed for deterministic software systems in which the same input reliably produces the same output. AI agents, by their very nature, violate this assumption. Their behavior is shaped not only by code and configuration but by the dynamic content of their context window, the state of external systems at the moment of access, the stochastic nature of LLM reasoning, and the emergent interactions within multi-agent architectures.</p><p>Production-grade observability for AI agents requires a different architectural principle: one that elevates data lineage to a first-class observable, captures the complete causal chain from data source to agent decision, and maintains this chain across agent boundaries, tool invocations, and temporal gaps.</p><p>First, let&#8217;s explore why observability in the context of AI agents is different.</p><h3><strong>Why AI Agents are Harder to Observe</strong></h3><p>The operational monitoring of software systems has matured considerably, converging on structured logging, OpenTelemetry-based distributed tracing, and Prometheus-style metrics. For traditional distributed applications, this stack provides a recoverable causal chain: a trace shows which services were called, logs show what each service did, and metrics reveal the resource conditions.</p><p>AI agents break this model. An agent processes a request according to its code, its prompt, the current contents of its context window, the results of any tools it has called, and the probabilistic reasoning of an LLM. Two identical requests, submitted seconds apart, can produce entirely different outcomes&#8212;not because of a system malfunction, but because the data retrieved by a tool call changed, the context window was managed differently due to token pressure, or the LLM&#8217;s sampling produced a different reasoning path.</p><p>This non-determinism is challenging enough for a single agent. It becomes dramatically harder to reason about in multi-agent systems, where agents consume one another&#8217;s outputs as inputs. In these environments, a log entry stating &#8220;Agent B processed the data&#8221; tells an operator almost nothing about why the output contains what it does, where the underlying data came from, or whether it was transformed in transit. The existing observability stack can tell you <em>what happened</em>. It often fails to explain <em>why</em> the agent decided what it decided, and <em>where</em> the data that informed that decision came from. For enterprise deployments, where audit, compliance, and explainability are non-negotiable, this is a critical gap.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xJF6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 424w, /__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 848w, /__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xJF6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png" width="1456" height="809" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:809,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:13535206,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/190506483?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 424w, /__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 848w, /__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xJF6!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa573c1e8-c090-4f37-9b9e-b5df47bc68c4_4340x2410.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>The Three Dimensions of Agent Observability</strong></h3><p>Meaningful observability for AI agents must operate across three dimensions that traditional tooling addresses unevenly at best.</p><h4>1. Operational Observability: What Happened</h4><p>This is the dimension best understood by existing tools. Distributed traces show the sequence of service calls. Metrics capture latency, throughput, and error rates. Logs record discrete events. The industry has a robust ecosystem, largely built around the <a href="https://opentelemetry.io/">OpenTelemetry standard</a>, that provides this foundational layer of visibility. This dimension answers the first, most basic question an operator asks, but for AI agents, it is rarely the last.</p><h4>2. Decisional Observability: Why It Decided What It Decided</h4><p>The second dimension requires capturing the full reasoning context that led to an agent&#8217;s decisions. This is more than just a log of the final output; it&#8217;s a snapshot of the agent&#8217;s complete execution context: the system prompt, the tool results present in the context window, the conversational history, and the specific content of the LLM&#8217;s response.</p><p>A new class of LLM-native observability tools, such as <a href="https://www.langchain.com/langsmith/observability">LangSmith</a> and the open-source <a href="https://langfuse.com/">Langfuse</a>, has emerged to address this. They capture detailed traces or trajectories of an agent&#8217;s execution, providing a step-by-step view of its internal operations, including the full content of the context window at each decision point. This is a significant and necessary step beyond traditional APM. The <a href="https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/">OpenTelemetry project&#8217;s GenAI semantic conventions</a> are formalizing a standard vocabulary for these traces, defining attributes for <strong>gen_ai.input.messages, gen_ai.system_instructions</strong>, and tool call results. However, treating decisional observability as a separate, specialized form of tracing risks isolates the <em>why</em> from the <em>where</em>. A trace that shows an agent made a decision based on a piece of data is useful; it becomes operationally essential when that trace is natively linked to the full, verified lineage of that data.</p><p>This requires more than capturing additional data in a trace; it requires structuring that trace as a queryable record of the agent&#8217;s full decisional state, where every element in the context window is not just a string, but a reference to a verifiable source.</p><h4>3. Provenance Observability: Where the Data Came From</h4><p>The third dimension is data lineage. When an agent acts on a piece of information, the enterprise must be able to trace that information back to its source: the specific query, table, or API call that produced it. When that information flows between agents, the lineage must be maintained without interruption.</p><p>Open standards like <a href="https://openlineage.io/">OpenLineage</a> provide a crucial foundation, offering a vendor-neutral framework for capturing lineage metadata from data pipeline components at the point of access. This is essential for dataset-level lineage. However, agent observability requires a finer granularity. It&#8217;s not enough to know that Agent A used a dataset that came from a specific database. We need to know that a <em>specific cell</em> in a spreadsheet generated by Agent B is a result of a transformation of a <em>specific record</em> retrieved by Agent A at a precise moment in time. The <a href="https://www.w3.org/TR/prov-overview/">W3C PROV standard</a> and its recent extension, <a href="https://arxiv.org/abs/2508.02866">PROV-AGENT</a>, demonstrate that the academic community has been developing the conceptual models for this level of granularity. The production challenge is implementing these models at the architectural layer of an agent runtime, not as an instrumentation library.</p><p>This requires lineage to be more than a record reconstructed after the fact; it must be an intrinsic, inseparable property of the data itself. Lineage should be stamped at the point of access and propagated automatically through every transformation and every agent handoff. Critically, access controls must propagate with lineage. If data carries a sensitivity attribute&#8212;&#8220;US Staff Only,&#8221; &#8220;MNPI&#8221;&#8212;that attribute must not be lost when the data is transformed or repackaged by a downstream agent. This is not something that can be easily bolted on; it must be enforced by the underlying architecture, making data sensitivity a structural property of the system.</p><h3><strong>Multi-Agent Observability: Maintaining Lineage Across Boundaries</strong></h3><p>The challenges of these three dimensions compound in multi-agent architectures. When a single agent fails, a combination of operational and decisional observability might be sufficient. When multiple agents collaborate, a new class of failure emerges: failures of interaction that are invisible at the individual agent level.</p><p>Consider a scenario where Agent A retrieves customer data, Agent B applies a risk model, and Agent C generates a recommendation. If the recommendation is incorrect, the root cause could lie in any of the three agents, in the data itself (which may have changed between retrieval and consumption), or in a subtle reasoning error by Agent C operating on correct inputs. Without a continuous, cross-agent lineage chain, diagnosing this requires a manual, error-prone reconstruction of the data flow.</p><p>True multi-agent observability requires a single, connected graph of provenance that allows an operator to trace any output back to its original sources through every intermediate transformation. This is not a correlation of independent logs from different agents; it is a unified view of the complete data flow across all participating agents, their tool calls, and the external systems they accessed.</p><h3><strong>Cryptographic Integrity and the Future of Audit</strong></h3><p>For regulated enterprises, observability data is not merely operational&#8212;it is evidence. The integrity of this evidence is paramount. While not yet a mainstream practice, the use of cryptographic techniques to ensure the immutability of audit trails is a logical and necessary evolution. <a href="https://www.w3.org/TR/did-core/">W3C Decentralized Identifiers (DIDs)</a> provide a path to giving each agent instance a verifiable, persistent identity, enabling cryptographic signing of traces, trajectories, and lineage records. A hash-chain-backed log, as demonstrated in the <a href="https://www.mdpi.com/2079-9292/15/1/56">AuditableLLM framework</a>, can ensure that the observability record is tamper-evident and independently verifiable&#8212;a requirement that will become increasingly common in finance, healthcare, and defense. The key property is that verification must be possible without access to the platform that generated the records; the cryptographic proof must be self-contained.</p><h3><strong>Conclusion: From Patchwork to Platform</strong></h3><p>The observability tools the software industry has built over the past decade are necessary but insufficient for the unique challenges of AI agents in production. They answer <em>what happened,</em> but struggle with <em>why the agent made the decision it did</em> and <em>where the data came from</em>. As enterprises move from single-agent experiments to multi-agent production deployments, this gap will become the primary obstacle to operational confidence, regulatory compliance, and organizational trust.</p><p>The path forward is a lineage-first approach to agent observability. This does not mean abandoning the principles of logs, metrics, and traces. It means re-centering them around a new architectural principle: a unified graph of data provenance. The principles are straightforward: capture lineage at the point of access, propagate it automatically through every transformation, enforce access controls as an inseparable property of the data, and provide cryptographic guarantees that the record is trustworthy.</p><p>The challenge is not conceptual; it is architectural. The solution lies not in better tools applied after the fact, but in an architecture that makes observability an inherent, structural property of the system. This means treating the lineage graph as the primary data structure of the agent runtime, not as a secondary artifact produced by instrumentation. Logs, traces, and metrics remain essential; they become views into the graph rather than independent streams. Access policies become edge attributes on the graph rather than gateway rules. Cryptographic signatures become properties of graph nodes rather than optional add-ons. When observability is structural, it cannot be omitted, misconfigured, or bypassed. That is the architecture the enterprise AI industry needs to build toward.</p><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Why We Deleted Our Visual Workflow Builder ]]></title><description><![CDATA[Six months ago, we deleted our visual workflow builder. This is why.]]></description><link>https://enterprisecontextmanagement.substack.com/p/why-we-deleted-our-visual-workflow</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/why-we-deleted-our-visual-workflow</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Wed, 04 Mar 2026 19:59:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Qumf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Six months ago, we deleted our visual workflow builder.</p><p>I don&#8217;t mean we hid it behind a feature flag or stopped talking about it. We removed it from the product, ripped out the code paths, and stopped designing around it. It wasn&#8217;t a rage decision. It was the end of a slow, unmistakable realization: the builder was training us to solve the wrong problem.</p><p>When we first shipped it, it felt like progress. Engineers could point to a canvas and say, &#8220;This is what happens.&#8221; Boxes, arrows, a clean sequence from input to output. It made the system easy to grasp, especially for people seeing it for the first time. But easy to grasp isn&#8217;t the same as correct.</p><h4>Not all &#8220;agents&#8221; are the same</h4><p>At some point we had to admit we were mixing up different categories of &#8220;agentic&#8221; products. From a distance, it can look like everything is an &#8220;agent&#8221; now. But under the hood, it&#8217;s really a few distinct paradigms wearing the same label.</p><p>The first is the classic <strong>visual workflow builder</strong>. But it&#8217;s not really an agent. It&#8217;s a workflow: a structured, predefined sequence where the user specifies the path up front - usually with drag-and-drop blocks. It&#8217;s great when the operation is repeatable and the world stays stable. But it doesn&#8217;t reason. It doesn&#8217;t adapt. It mostly just advances.</p><p>The second is what I&#8217;d call the <strong>framework-first approach</strong> - the ecosystems where you wire up tool calls, loops, and multi-step behaviors with prompts and orchestration code. Those tools made &#8220;agent development&#8221; accessible. But they also push the hard problems into the least reliable place: the model&#8217;s behavior. The more you rely on prompt compliance for governance, the more your security and policy guarantees become <em>aspirations</em> instead of <em>structure</em>.</p><p>The third is the <strong>power-tool</strong> category: things like coding agents and lightweight task runners. They can be shockingly capable in a constrained environment. But they usually don&#8217;t give you the things organizations need when the work matters: centralized policy enforcement, reproducibility, consistent audit logs, and a way to explain why it made the decisions it did.</p><p>We realized what we were building is closer to a fourth category: an <strong>enterprise-aware agent system</strong>. The experience should feel as immediate as the best power-tools - type a goal, give it tools, watch it work - but with the operational guarantees that make it deployable in the real world: enforceable boundaries, evidence and lineage, and auditability.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KhQ3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 424w, /__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 848w, /__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KhQ3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png" width="1456" height="769" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:769,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:15571815,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/189791651?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 424w, /__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 848w, /__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KhQ3!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2ced155-1462-4bc5-94fe-388376fcd6ed_6144x3243.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>The First Shift: From Steps to Outcomes</h4><p>Back to our visual builder. The first cracks didn&#8217;t show up in the demo environment. They showed up in the unglamorous places: the runs where the data didn&#8217;t match the happy path, where a tool returned something plausible but incomplete, where a downstream system changed a response shape without telling anyone. The builder did what it was built to do: it moved forward. It advanced to the next box.</p><p>And that&#8217;s exactly when we started losing the plot.</p><p>Because the system we were trying to build wasn&#8217;t supposed to &#8220;advance.&#8221; It was supposed to reach a goal reliably, even when the path wasn&#8217;t the one we expected.</p><p>That was the first shift: we stopped thinking in steps and started thinking in outcomes. In a builder, the path is the artifact. In an agentic system, the path is disposable, a means to an end.</p><p>Once you see that difference, you can&#8217;t unsee it.</p><p>We tried to reconcile the two worlds for a while. We added more structure. More branches. Retries. Validation blocks. Error handlers. We made the canvas smarter the way everyone makes canvases smarter: more boxes. The diagrams got impressive&#8212;dense, &#8220;enterprise,&#8221; serious.</p><p>They also got fragile.</p><p>Every new block was a commitment to predict the future: which failure modes matter, which recoveries are safe, which validations are sufficient, which conditions are worth branching on. And as soon as production taught us a new lesson, we encoded it as another node. That created a familiar pattern: the canvas wasn&#8217;t representing reality; it was chasing it.</p><p>Here&#8217;s the distinction we couldn&#8217;t escape: workflow graphs are execution plans. You pre-author a path. Agentic systems are control policies&#8212;they observe, decide, act, and verify against the goal. Adding boxes just hard-codes yesterday&#8217;s lessons; it doesn&#8217;t buy you adaptation.</p><h4>The Second Shift: From Calls to Actors</h4><p>The second shift made the decision inevitable: we stopped treating the model like a single call and started treating it like an actor in a system.</p><p>The best runs shared a simple shape:</p><ul><li><p>give the model a goal</p></li><li><p>give it tools and skills</p></li><li><p>constrain it with guardrails</p></li><li><p>let it cook</p></li></ul><p>The surprise was how often the right path differed. The model changed tool order based on what it found, recovered from failures, asked clarifying questions when needed, and validated intermediates because the goal demanded it.</p><p>A concrete example made this obvious.</p><p>Take onboarding. Not &#8220;fill in a form&#8221; onboarding, real onboarding where inputs arrive as emails, PDFs, spreadsheets, and half-structured notes. One run is clean; the next has mismatched identifiers, missing evidence, or a policy threshold that changes what checks are required. The system has to ask a question, route to a reviewer, pull a source of truth, re-check assumptions, and then prove the outcome is supported.</p><p>In a canvas world, that&#8217;s not one workflow. It becomes a library of workflows: branches for every policy fork, variants for every product line, region, customer type, evidence format, and exception path. Over time it turns into hundreds of slight variations, most of them identical except for one step, one validation, one integration edge case.</p><p>An agentic system treats those variations differently. The goal is stable. The constraints are explicit. The tools are bounded. The agent decides the path at runtime based on what it sees, and it can explain why it took the route it did.</p><p>That&#8217;s what &#8220;agentic&#8221; meant in practice. It is not a buzzword, but a behavior: understanding the environment, adapting to change and error, and staying oriented to the goal.</p><p>And that&#8217;s where the canvas became more than inconvenient. It became actively counterproductive.</p><p>A visual workflow builder rewards certainty. It rewards pre-definition. It encourages you to treat the journey as sacred: draw it carefully, lock it in, and the system will follow it faithfully.</p><p>But effective automation isn&#8217;t faithfulness to a plan. It&#8217;s faithfulness to an outcome.</p><p>The builder was making the journey more important than the destination.</p><p>We felt this most sharply in debugging. With a canvas, a run &#8220;fails&#8221; when a node errors. But many of the worst failures in automation don&#8217;t throw errors, they produce outputs that look reasonable. A workflow graph can happily deliver something that passes through every box and still be wrong, because the world changed and the flow didn&#8217;t notice. The diagram stays clean. The result quietly degrades.</p><p>To solve that, you don&#8217;t add more arrows. You add a control plane: explicit checks, verifiable tool boundaries, consistency tests, and guardrails that stop the system from &#8220;hallucinating to completion.&#8221; You make the system able to say: I can&#8217;t support this conclusion with evidence. Or: I need clarification. Or: This output fails validation; I&#8217;m taking a different approach.</p><p>Those are not canvas nodes. They&#8217;re governance primitives.</p><h4>The Third Shift: Operations Change Everything</h4><p>The third shift was operational: supporting deployments over time means living in continuous change. Data formats drift. Vendors update. Policies and thresholds move. Exceptions emerge. Tools develop new failure modes. Teams change how approvals and escalations work.</p><p>Once you account for that, the math is brutal: every diagram is an artifact you now have to keep in sync with a moving world. It starts as &#8220;a few workflows,&#8221; becomes a library, then a forest of near-duplicates.</p><p>Onboarding makes the cost obvious. If step X exists in 500 variants, a small upstream change isn&#8217;t one fix, it&#8217;s 500 edits, 500 tests, and 500 chances to miss a corner.</p><p>In an agentic system, the unit of change is smaller and more powerful. You improve a tool. You refine a prompt. You add memory. You tighten a guardrail. One change, applied everywhere. The system adapts at runtime instead of demanding that engineers pre-adapt it in diagrams.</p><h4>The Decision to Delete</h4><p>Once we accepted that, keeping a workflow builder around wasn&#8217;t harmless. It was pulling the product toward a worldview we no longer believed: that the primary job of automation is to predefine the path.</p><p>So we deleted it.</p><p>Not as a rejection of visual builders in general, just an acknowledgment of what we were actually building. In an agentic-native system, the interface isn&#8217;t a canvas. The interface is principles, intent, constraints, tools, and a trace of decisions that you can inspect, audit, and replay.</p><p>The goal stays fixed. The route is allowed to change because the world does.</p><p>And after we committed to that, a workflow builder wasn&#8217;t a feature anymore. It was a contradiction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Qumf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 424w, /__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 848w, /__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Qumf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png" width="1456" height="795" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:18909586,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/189791651?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 424w, /__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 848w, /__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Qumf!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ba68c49-facd-4832-8b9e-b977c3e819d8_5919x3230.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Ship More, Type Less: A Leader's Guide to AI-Assisted Development]]></title><description><![CDATA[How AI-enablement is changing the economics of software, shifting the pain of development, and redefining what makes a great engineer]]></description><link>https://enterprisecontextmanagement.substack.com/p/ship-more-type-less-a-leaders-guide</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/ship-more-type-less-a-leaders-guide</guid><dc:creator><![CDATA[Shaun Laurens]]></dc:creator><pubDate>Thu, 26 Feb 2026 15:10:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xZxu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The economics of software development have fundamentally shifted. Things that were previously too expensive to build in-house may no longer be. While I don&#8217;t expect most large enterprises to abandon their typical posture of focusing on core strengths and outsourcing the rest, past decisions grounded in cost and risk deserve a fresh look. The calculus has changed, and organizations that don&#8217;t revisit those assumptions will overpay for commodity solutions and cede innovation to vendors when they could build, own, and control better solutions themselves.</p><p>So how do you actually translate this shift into shipped systems?</p><h4><strong>Start With the Right People</strong></h4><p>Not every developer will thrive in this new model. Some have always coded for the beauty of code itself rather than for the goal of producing working business systems. AI agents will displease those who love coding for its own sake, and empower those who love to ship real outcomes.</p><p>Find a handful of those people and experiment with them first. I don&#8217;t believe there&#8217;s a strong level of experience correlation here; what matters is mindset. Specifically, look for developers who:</p><ul><li><p>Are curious and willing to learn. Working with AI agents is a new discipline. The people who will succeed are those who approach it with genuine curiosity and a willingness to experiment.</p></li><li><p>Are driven to ship business value, not just code. The agent handles much of the raw coding. What remains is the harder, more critical work of translating business problems into working systems.</p></li><li><p>Understand the business itself. People who grasp what the organization is trying to accomplish will be far more effective at steering agents toward the right outcome. This matters more than ever, because the bottleneck is no longer typing speed, it&#8217;s judgment.</p></li><li><p>Have used AI coding tools and can talk about what they&#8217;ve done. There&#8217;s simply no excuse for not having used a number of tools at this point. If you meet someone who hasn&#8217;t committed to that growth mindset and at least done personal work with it, start to ask a lot of questions.</p></li></ul><h4><strong>Pick the Right Problems</strong></h4><p>Once you have one or more small teams, the challenge is finding the right kind of problem to experiment with. AI agents work best on greenfield projects, though they absolutely help with legacy systems too.</p><p>Regarding scope, the obvious answer is often the right one - you should start small with a single end-to-end feature. However, more importantly, you must start with something verifiable. AI agents shine when they can validate their own work. Focus on problems with clear, measurable results - a well-defined API, a calculation that can be checked, a workflow with observable outputs. Avoid anything ambiguous where success is subjective and hard to confirm.</p><h4><strong>Recognize Where the Pain Shifts</strong></h4><p>This is a point that catches many organizations off guard. The development itself is no longer the long pole. The pain moves upstream and downstream: to specification and to review.</p><p>Before an agent can build the right thing, someone needs to clearly articulate what the right thing is. This loops back to having developers with a solid handle on the business problem being solved. They should also have a strong understanding of architecture, organizational standards, and deployment constraints so that the agent doesn&#8217;t produce something technically impressive that can never actually ship.</p><p>On the other end, review becomes the critical bottleneck. AI agents can generate substantial volumes of working code, but someone still needs to verify that it does the right thing, in the right way, within the right guardrails. Plan for this. Build review capacity into your workflow from the start - and that doesn&#8217;t have to be manual.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xZxu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xZxu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xZxu!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101c273d-326e-446b-bf2d-e4b425438e63_1024x559.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>To make this a bit more concrete: inside the product engineering team at AI One, we have shifted design reviews to the left: the story initiator writes the initial design document (AI-assisted), one or more other team members review it and provide feedback, and then we assign a single developer to take it forward from that point to production. During development, we use AI tooling (alongside heavily automated, high coverage end-to-end testing) to perform reviews across a number of different dimensions: correctness, functionality, and security. Then, once the developer has proven the story is in a state of readiness, they submit the PR for another team member to review. The team member then reviews the evidence and the high risk areas of the PR before approving it (or requesting changes). This sits around level 3 to level 4 of <a href="https://www.danshapiro.com/blog/2026/01/the-five-levels-from-spicy-autocomplete-to-the-software-factory/">Dan Shapiro&#8217;s Five Levels of AI coding</a>. As a SOC II compliant organization, we require at least some human oversight of the code so this approach fits well into a highly regulated context.</p><h4><strong>Invest in Infrastructure and Test Environments</strong></h4><p>This follows naturally from the review constraint. AI agents perform dramatically better when they can interact with real systems in controlled environments. If you&#8217;re building something that connects to a custom internal API, give the agent access to a live instance of that API in a sandbox. If this cannot be achieved, build as close a digital twin as you can. Then, let it validate that API calls actually did what it expected.</p><p>Agents are remarkably good at feeling their way around systems and figuring out how best to use them once given real access. Static documentation alone is a poor substitute.</p><h4><strong>A Practical Example: Building Integration Gateways</strong></h4><p>We recently needed to build gateways for both Google Workspace (authentication and access to Drive, Sheets, Calendars, and more) and Microsoft 365 (authentication, Teams, SharePoint, and related services). Rather than handing agents documentation and hoping for the best, we stood up full sandbox environments for them to work against.</p><p>We populated these environments with real data - files in SharePoint and Drive, channels in Teams, calendar entries - and gave the agents well-defined tasks: authenticate, retrieve a file from SharePoint, list items in a Drive folder, post to a Teams channel. The agents had full access to these sandbox environments, access to the internet for API documentation, and crucially, the ability to verify every step of their work against real system responses.</p><p>The result was high-quality integration gateways that were proven from the start to work correctly. They weren&#8217;t built in theory and then tested, they were built through testing, with the agent iterating against live systems until everything behaved as expected. We still reviewed the output for security, architectural fit, and edge cases, but the heavy lifting - the tedious, iterative work of getting OAuth flows, API pagination, error handling, and data mapping right - was all done by the agents.</p><p>This is the pattern that works: give agents real environments, clear objectives, and the ability to validate their own results. The quality of the output is dramatically higher than what you get from agents working blind against documentation alone.</p><p>Then, critically, ensure the agent builds out the tests necessary to enable quick and safe changes in the future. The first build is just the beginning - the test suite is what makes it sustainable.</p><h4><strong>Don&#8217;t Overthink Costs - Invest to Experiment</strong></h4><p>It&#8217;s surprisingly uncommon for developers to be given truly high-quality tools - whether that&#8217;s powerful laptops, top-tier IDEs, or $200/month subscriptions to AI agents. Use your small teams to experiment and build the business case - it&#8217;s a very small investment for a potentially massive return.</p><p>As a closing piece of advice, I&#8217;d also strongly encourage organizations not to be too restrictive about letting curious developers use their subscriptions at home for personal projects. Building a small side project, automating something in their own life, experimenting with a new framework - this is all part of learning to work effectively with the tools. That fluency will flow directly back into their professional output. Treat it as professional development, because that&#8217;s exactly what it is.</p><p></p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Who Keeps the Surplus? The Coming Fight Over AI's Productivity Dividend]]></title><description><![CDATA[Early AI-driven cost cuts are the opening move in a much larger economic renegotiation]]></description><link>https://enterprisecontextmanagement.substack.com/p/who-keeps-the-surplus-the-coming</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/who-keeps-the-surplus-the-coming</guid><dc:creator><![CDATA[Fergus Keenan]]></dc:creator><pubDate>Fri, 20 Feb 2026 21:44:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5kYs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>An underreported item of corporate news may signal an early shift in how AI-era services are priced: <a href="https://www.ft.com/content/c891c47c-b21f-4e0f-84b3-b80c794eff3d">KPMG pushed its own auditor</a>, Grant Thornton UK, for a fee reduction on the logic that AI should make the audit cheaper. The fee disclosed in its UK filings fell 14 percent year over year (from $416,000 to $357,000, reported in dollar terms). For the same audit.</p><p>The number itself is less important than the precedent. The implicit demand is clear: <strong>if AI made you faster, I want the discount.</strong></p><p>That demand exposes a long-standing economic tension. The buyer&#8217;s instinct is understandable, but it rests on an assumption that has always been fragile: that hours billed are equivalent to value delivered. The billable hour has historically functioned as a proxy for value in professional services, not the value itself. But in reality, <strong>clients do not purchase time; they purchase judgment, accumulated context, risk transfer, and accountable outcomes.</strong> When AI compresses the hours required to produce an audit, a diligence report, or a strategic memo, it does not automatically compress the expertise or liability embedded in the final deliverable.</p><p>What the KPMG example illustrates is not just fee pressure in auditing. It signals that &#8220;AI pass-through&#8221; is becoming a baseline expectation. Wherever a buyer can point to a visible process and ask, &#8220;How many hours is this still taking, and why?&#8221;, pricing models anchored to labor time will come under scrutiny.</p><p>The mechanism will vary by industry, but the pressure is similar. In software and professional services, AI reduces labor minutes. In operational settings&#8212;factories, logistics networks, supply chains&#8212;it reduces variance: fewer defects, fewer returns, less downtime, tighter inventory. In both cases, measurable efficiency gains create procurement leverage.</p><p>Service providers will respond by arguing that <strong>AI does not merely make the same output cheaper; it shifts the quality frontier.</strong> A faster audit that surfaces anomalies with greater precision, a compliance workflow with stronger traceability, or a supply chain with materially lower variance is not strictly the same product at a lower cost. It is a product with different performance characteristics. From the vendor&#8217;s perspective, the relevant question is not &#8220;How many hours did this take?&#8221; but <strong>&#8220;How much risk was removed, how much upside was unlocked, and how much better is the outcome?&#8221;</strong></p><p>The problem is that <strong>procurement teams rarely price on abstract improvements in quality when visible cost savings exist.</strong> The adjustment is unlikely to be linear. The adoption path more closely resembles a J-curve followed by an S-curve.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5kYs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 424w, /__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 848w, /__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5kYs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8734923,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/187773012?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 424w, /__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 848w, /__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5kYs!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93d298d6-48ad-421e-8174-9a01a32ee102_13200x7296.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The J-curve is the awkward phase in which efficiency gains exist but are difficult to monetize cleanly: duplicated workflows, model risk controls, data integration overhead, security reviews, and persistent edge cases that pull manual work back into the loop. During this period, prices may not fall meaningfully, and in some cases may even rise as vendors absorb transition costs.</p><p>The S-curve follows once tooling standardizes, model workflows are audited, and AI is embedded in operating procedures. At that point, procurement gains confidence in the stability of the new process and begins pricing against unit economics rather than prototypes. <strong>Explicit AI discounting becomes harder to resist.</strong> Once a few vendors concede, peers risk appearing overpriced overnight.</p><p>This creates acute pressure for firms that lag in implementation. Vendors unable to demonstrate credible efficiency gains&#8212;or unable to absorb pass-through pricing&#8212;will struggle to compete.</p><p><strong>Yet the long-run dynamic is more complex than simple deflation.</strong> When technology meaningfully increases the productive capacity of expertise, <a href="https://www.npr.org/sections/planet-money/2025/02/04/g-s1-46018/ai-deepseek-economics-jevons-paradox">demand often expands</a>. As AI raises the ceiling on what a professional can analyze, oversee, or optimize, the binding constraint shifts from labor time to judgment and accountability. Pricing anchored purely to hours becomes increasingly incoherent, but total value creation may grow.</p><p>Whether buyers <strong>pay for that expanded capability or simply capture the efficiency</strong> gains depends largely on market structure. In highly competitive markets, savings are more likely to pass through, producing pockets of disinflation even as quality improves. In concentrated markets, vendors may retain a larger share of the productivity dividend, expanding margins while performance rises.</p><p>For firms early in their AI implementation cycle, the practical constraint is not model performance but proof. Efficiency gains that cannot be measured, governed, and allocated are difficult to defend in pricing negotiations. <strong>Buyers demanding discounts will increasingly demand evidence</strong>: auditable trails that separate real throughput improvements from shifted risk or hidden rework.</p><p>The evidentiary standard cannot stop at hours saved. If the debate is framed purely in terms of cycle time or headcount compression, the buyer&#8217;s logic dominates. The more durable strategy is to make surplus legible across multiple dimensions: throughput, error reduction, variance compression, regulatory defensibility, risk transfer, and outcome quality.</p><p>When a buyer can quantify not only cycle time and exception rates, but also control strength, auditability, and decision quality with the same rigor applied to financial metrics, the conversation changes. <strong>It shifts from &#8220;How much labor was removed?&#8221; to &#8220;What new level of assurance or performance is now possible?&#8221;</strong></p><p>The audit fee reduction is therefore less about one negotiation and more about a structural question that will echo across sectors: <strong>when AI increases productive capacity, who captures the surplus?</strong> The answer will depend on competitive intensity, evidentiary discipline, and how effectively vendors re-anchor pricing away from time and toward accountable outcomes.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Memory is not Authority]]></title><description><![CDATA[Enterprise agent systems have matured enough that the interesting problems are no longer about access or intelligence.]]></description><link>https://enterprisecontextmanagement.substack.com/p/memory-is-not-authority</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/memory-is-not-authority</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Wed, 11 Feb 2026 20:05:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vitc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Enterprise agent systems have matured enough that the interesting problems are no longer about access or intelligence. Retrieval works. Tools execute. Models reason well enough. What now differentiates serious deployments is something more prosaic and harder to fake: whether the system can be trusted to act inside an organization shaped by policy, risk, precedent, and drift.</p><p>This is why the concept of &#8220;context graphs&#8221; entered the conversation, and why they matter. The recent framing is precise. Context graphs are not knowledge graphs. They are decision trace graphs. They preserve what mattered at decision time so that justification, precedent, and accountability do not dissolve into chat logs.</p><p>That is a meaningful advance. It is crucial to include historical decisions, exceptions, and edge cases when making decisions about how to act. But over the last year, when actually implementing enterprise wide production systems, <em>we quickly found that it is not the end state</em>.</p><p>The mistake many teams make is subtle. They correctly observe that decision traces are consulted during agent reasoning, and then quietly slide from &#8220;informs action&#8221; to &#8220;authorizes action.&#8221; Once that line is crossed, the architecture starts to accumulate risk in ways that only appear months later.</p><p>What follows is the perspective that emerged when we treated decision traces as advisory reasoning surfaces, and deliberately placed authority elsewhere.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vitc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 424w, /__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 848w, /__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vitc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png" width="1456" height="795" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4120449,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/187644984?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 424w, /__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 848w, /__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vitc!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245f8c6-b6f9-43a2-8658-5f2c3d0be31b_3072x1677.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Why The Context Graph Concept Invites Overreach</strong></p><p>In practice, the strongest solutions in this area are active participants in agent reasoning. They are queryable during planning and execution in order to surface precedent, exceptions, constraints, and prior approvals. Used this way, decision traces materially improve reasoning quality. They ground agents in organizational memory, reduce hallucination, and make decisions legible after the fact. This is precisely what makes the graph concept so tempting to treat as a source of authority. Graph nodes look executable. They encode conditions, actions, and rationale. It is easy to slide from &#8220;this informs what we should do&#8221; to &#8220;this permits us to do it.&#8221;</p><p>That slide is where risk accumulates. A true context graph would preserve what mattered at decision time, but would not enforce what is valid now. When allowed to authorize action, three failure modes reliably appear. Precedent outlives policy because time is represented but not enforced. Rare exceptions become reusable templates as agents optimize for citation rather than validation. As graphs grow richer, systems become better at justifying actions than at checking whether they are still allowed. These are not flaws in graph design. <em>They are the natural outcome of treating memory as control.</em></p><p>Separately, in our practical experience, we found that <em>trajectories</em> solve a different and equally important problem. A trajectory is the time ordered trace of how work actually unfolds, including goals, actions, failures, escalations, overrides, and outcomes. Reasoning over trajectories shifts agents away from inventing plans and toward reproducing successful ones. Recovery improves because failure patterns are visible. Planning becomes empirical rather than speculative. But trajectories are descriptive, not normative. They explain how work gets done here, not what is permitted to run.</p><p><strong>Behavior Before Permission</strong></p><p>The missing primitive is executable authority. Rather than storing justification as narrative memory, we embedded it into time versioned units of execution - skills. A skill is not a suggestion. It is an enforcement point. Each skill defines the action it permits, the preconditions that must hold, the authoritative sources it may consult, and the subsequent deterministic validations that must be passed, including business rules, policy constraints, and risk controls. It also specifies any approvals or exceptions required, the evidence that must be emitted, and the window in which the skill is valid. Skills are evaluated synchronously at execution time and fail closed. If the rules no longer pass, the policy no longer applies, or the justification has expired, the action does not run. This is the moment where &#8220;why&#8221; stops being an explanation and becomes control.</p><p>Time bounding is the property that turns governance from an aspiration into a system behavior. Without it, organizations slowly accumulate what we came to think of as precedent rot. Temporary workarounds harden into defaults. Policy interpretations linger long after their intent has expired. Nothing breaks loudly, but authority quietly drifts. By contrast, a time bound skill expires in the open. It must be refreshed from sources, revalidated against policy, or deliberately replaced. Drift becomes visible and actionable instead of silent. Decision traces remain invaluable in this process because they preserve the history of how and why decisions were made. But they can only observe drift. Skills are what stop it.</p><p>When combined deliberately, the division of labour becomes clear. Different primitives serve different roles, and problems that once felt entangled separate cleanly. Trajectories ground agents in how work actually succeeds by exposing real sequences of action and recovery. Decision traces ground reasoning by preserving precedent, rationale, and explanation. Skills sit at execution time and decide what is allowed to run now, under current conditions. Agents consult all three continuously, but only one has the authority to say yes. This resolves a tension many teams feel but rarely articulate. Decision traces (or implementations of context graphs) should absolutely influence what an agent proposes to do. They simply should not be the thing that permits it.</p><p><strong>From Reasoning to Control</strong></p><p>The enterprise choice is straightforward. <em>Either reasoning is allowed to imply permission, or permission is enforced independently of reasoning. Only one of these scales safely. </em>Systems built on trajectories and skills are much easier to transition from prototype to production, and far harder to break once they are there.</p><p></p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The SaaS Selloff Is a Verdict on Platforms That Don’t Learn ]]></title><description><![CDATA[The recent a16z piece on The Palantirization of everything by Mark Andrusko is one of the clearer treatments of a problem many enterprise AI companies are already feeling: early traction is often driven by deeply embedded teams doing bespoke work, and without discipline, that path leads to a services business wearing a software costume.]]></description><link>https://enterprisecontextmanagement.substack.com/p/the-saas-selloff-is-a-verdict-on</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/the-saas-selloff-is-a-verdict-on</guid><dc:creator><![CDATA[Fergus Keenan]]></dc:creator><pubDate>Thu, 05 Feb 2026 18:58:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EhuK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The recent a16z piece on <a href="https://www.a16z.news/p/the-palantirization-of-everything">The Palantirization of everything</a> by <a href="/__u/substack.com/@marcandrusko">Mark Andrusko</a> is one of the clearer treatments of a problem many enterprise AI companies are already feeling: early traction is often driven by deeply embedded teams doing bespoke work, and without discipline, that path leads to a services business wearing a software costume. That diagnosis is largely correct - and worth taking seriously.</p><p>But the more interesting question is not whether forward-deployed work is dangerous. It&#8217;s <em>when</em> it becomes dangerous, and what actually separates a compounding platform from a high-end delivery shop in an era where software itself is getting cheaper and faster to build.</p><p>In practice, many of the most valuable enterprise AI problems still demand real proximity to the customer. The hardest work lives in messy, high-stakes domains where production outcomes matter: reconciling exceptions across fragmented systems, continuously monitoring disparate systems for early risk and next-best-action signals, and capturing the &#8220;in-between&#8221; context - decisions, SOPs, undocumented overrides - that determines whether automation is explainable, auditable, and trusted. These are not problems that vanish with better models alone.</p><p>The mistake is assuming that because this work is necessary, it should be permanent.</p><p>Forward deployment should be a <em>use-case scaling strategy</em>, not an operating model. The first phase - going from zero to one - benefits enormously from tightly integrated field and platform engineering. But that phase only creates leverage if it is explicitly designed to end. The purpose is not to deliver forever; it is to extract a repeatable pattern that can be scaled with far less human involvement.</p><p>This is where the article&#8217;s warning is most important, and where execution discipline matters more than rhetoric. Time-boxing initial deployments, constraining customization to the edges, and protecting a stable, upgradeable core are not nice-to-haves. They are the mechanisms that force compounding. Without them, every new customer quietly fragments the product, and the organization accumulates entropy rather than advantage.</p><p>Where the piece feels underweighted is on how moats actually form in the 2026 timeframe.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!EhuK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 424w, /__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 848w, /__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!EhuK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png" width="1456" height="795" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5031425,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/186983254?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 424w, /__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 848w, /__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EhuK!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f1db8b7-2e06-45be-9185-b920f282f8d0_3072x1677.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The traditional platform moat, large inventories of bespoke integrations and proprietary primitives, made sense when software was expensive and slow to build. That world is fading. AI-assisted engineering is collapsing development time and cost, which means static platforms are easier to replicate and harder to defend. The connector layer, by itself, is no longer durable.</p><p>What compounds now is not the surface area of the platform, but the <em>operational playbook</em> behind it. The ability to repeatedly drive real 0&#8594;1 outcomes, transition those outcomes to 1&#8594;N scale, and feed the resulting learnings back into the system. Most critically, it is the accumulation of context and memory: entity maps, workflows, decisions, state, and institutional knowledge that make each deployment smarter than the last.</p><p>That &#8220;context layer&#8221; becomes a durable advantage for customers, much like data did in prior eras, but only if the system remains modular and upgradeable. Lock-in through rigidity is brittle; advantage through adaptability is not.</p><p>Finally, the article implicitly treats SaaS as the default endpoint business model. That assumption is increasingly questionable (and probably requires its own post based on <a href="https://www.economist.com/business/2026/02/01/why-software-stocks-are-getting-pummelled">the stock market events of this week</a>!). As reliability improves and margins converge toward software economics, outcomes-based pricing becomes viable - first in hybrid form (platform plus outcomes), and increasingly outcomes-weighted over time.</p><p>The key distinction is that outcomes cannot simply be layered on top of bespoke delivery. The winners will be companies that can tie outcomes to platforms that genuinely compound, where each success lowers the cost and increases the reliability of the next one.</p><p>The real risk in enterprise AI isn&#8217;t forward deployment. It&#8217;s mistaking early momentum for durable leverage, and failing to design for compounding from day one.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Owning Your Enterprise AI Brain]]></title><description><![CDATA[Moving Past Pilots: The Strategic Imperative of Scalable Enterprise Context]]></description><link>https://enterprisecontextmanagement.substack.com/p/brain</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/brain</guid><dc:creator><![CDATA[Conor Twomey]]></dc:creator><pubDate>Wed, 17 Dec 2025 15:56:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XrWJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every company is building an AI brain. Most are doing it by accident, with random acts of AI prompting and data integration. The memories and knowledge they feed it are scattered like loose pages from a book.</p><p>A piece lives in your CRM&#8217;s new chatbot. Another is in the BI tool that turns questions into SQL. A third is in the agent framework your engineering team uses. These fragments don&#8217;t talk to one another. They can&#8217;t reason across your business. It&#8217;s a flawed foundation that won&#8217;t scale.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>There&#8217;s a better way: a single, coherent system of context. A brain you own and control. When you own the context you feed AI, you control how your AI reasons about every question your business asks. This requires a better architecture.</p><p>We&#8217;ve previously <a href="/__u/enterprisecontextmanagement.substack.com/p/stop-building-agent-chains-start">argued</a> that the Hybrid Loop is the gold standard for building AI systems that think. This essay is about how to train that brain. The strategic question isn&#8217;t <em>how</em> your AI reasons, but <em>what</em> it reasons with.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XrWJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 424w, /__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 848w, /__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XrWJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1523541,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/175492402?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 424w, /__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 848w, /__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XrWJ!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8b44e-efe1-4477-a7f9-1d2efff93e72_2160x2160.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><strong>Two Paths to AI Failure</strong></h4><p>Most leaders see two paths for building with AI. Both are traps that cause you to lose control of your enterprise context.</p><p><strong>Path 1: Build it yourself on a legacy framework.</strong> Developers designed toolkits like LangChain and CrewAI when models were weak and could only reason one step at a time. Their solution was to chain models together, letting them improvise in a loop. But this creates a contextual mess. Each step adds noise and replays history inefficiently. There is no central source of truth. The framework designed to help you instead creates the chaos you must overcome.</p><p><strong>Path 2: Rent it from a SaaS vendor.</strong> This path lets your vendors handle it. Your CRM gets a copilot. Your document store gets a summarizer. This approach is seductive and offers quick wins, but the long-term cost is control. Your customer data, product information, and internal processes are now managed inside someone else&#8217;s black box. You cannot standardize it. You cannot move it. You have outsourced your company&#8217;s memory. The result is a lobotomized organization. Your sales AI can&#8217;t learn from a support ticket, and your support AI can&#8217;t see a customer&#8217;s purchase history. Every interaction starts from zero.</p><p>Chaos or a lobotomy. Neither is a good choice</p><p>There&#8217;s a better way.</p><h4><strong>The Third Path: Own Your Enterprise Context</strong></h4><p>There is a third road: treat your enterprise&#8217;s context as your most valuable asset, managed by a dedicated system. This is the discipline of <a href="/__u/enterprisecontextmanagement.substack.com/p/intro">Enterprise Context Management, or ECM</a>. ECM is like a library and head librarian for your AI. It doesn&#8217;t just store the books; it knows which page to open for any given question. The library is organized, the librarian is precise, the building is secure, and you choose who gets a library card.</p><p>When you own your context, you own your company&#8217;s brain. The lessons you&#8217;ve learned and inferred principles live in your library, not in a vendor&#8217;s application or scattered in code. It enables the AI <a href="/__u/enterprisecontextmanagement.substack.com/p/stop-building-agent-chains-start">Hybrid Loop</a> architecture to function at scale based on intelligence that&#8217;s organized in multiple layers (like a <a href="/__u/enterprisecontextmanagement.substack.com/p/skyscraper">Context Skyscraper</a>).</p><p>What does this brain look like in practice? It has five key functions:</p><ul><li><p><strong>A Centralized and Curated Context Catalog</strong>. Instead of disconnected data or metadata silos, ECM is built on a Context Catalog. As explained in the<a href="/__u/enterprisecontextmanagement.substack.com/p/intro"> introduction to Enterprise Context Management</a>, this catalog doesn&#8217;t just store data or create semantic data views. It enriches context with operational data. It knows which customers are new, which contracts are expiring, and which internal processes are critical. It catalogs the history of the prompts that worked and didn&#8217;t. It enforces privacy. Your AI uses this context to reason, ensuring every decision is based on a complete picture of reality.</p></li></ul><ul><li><p><strong>Principle-Based, Not Rule-Based Context. </strong>Our human brains are pattern-matching engines. We see a new situation and instantly process it based on previous experience. AI makes judgements in much the same way, based on predictions and inference. Business context, therefore, must be designed for fuzzy, principle-based judgements and predictions. So ECM is built on principles that guide decisions based on context.</p></li></ul><ul><li><p><strong>Context Delivered with Precision</strong>. A catalog is only useful if the right information can be accessed at the right time. An ECM platform shapes and delivers the exact context needed for each step of a task. For a high-level planning step AI takes, it leans on strategic principles. For a detailed verification step, it might consider a single, critical data point. This precision allows the AI to operate on flexible principles derived from your business.</p></li></ul><ul><li><p><strong>A Secure Control Plane in Your Environment</strong>. Owning your brain means keeping it secure. ECM platforms run within your own environment, close to your data, respecting your data residency and privacy rules. It is managed by your team, not a third party. This is enforced through a control plane that implements a zero-trust security model. Every request from a user or an AI agent is authenticated and authorized, operating on a principle of least privilege. This ensures that as AI gains autonomy, it does so within the boundaries you defined.</p></li></ul><ul><li><p><strong>Control Over Your AI Models</strong>. An ECM platform separates the brain (your context) from the voice (the AI model). This gives you the freedom to choose, mix, and swap the foundational AI models that interact with your context. You are not locked into a single vendor&#8217;s ecosystem. When a better, faster, or cheaper model arrives, you can adopt it without having to rebuild your entire enterprise brain. You own the asset; the model is the tool you choose to apply to it.</p></li></ul><h4><strong>The Context Brain in Action</strong></h4><p>Your enterprise now has a brain, a central library of its own. But a library is only as good as the wisdom you can draw from it. The next step is to teach it to think. Forget what you know about rules. The core insight is this: agents need principles plus context, not data plus rules.</p><p>The old era of rules was about looking backward. You mined historical data to create rigid checklists. You told your system to halt a process if any box was unticked. The logic was brittle because it lacked situational awareness. It knew the rule, but it didn&#8217;t understand the customer or the goal.</p><p>The ECM era gives agents that awareness. It lets them apply your business principles, even when the data is ambiguous. That ability to navigate subjectivity is a feature, not a bug. It is judgment.</p><p>Imagine AI managing new customer onboarding. The old rule is simple: halt the process if any required document is missing. A high-value enterprise customer is ready to go, but a non-critical setup form is missing. The system stops. The customer waits. Momentum is lost. But AI with a context brain operates on a principle: prioritize onboarding steps based on customer value and compliance risk. It sees the missing form but also knows the customer is strategic and the document is not a legal requirement. It allows the technical integration to proceed while flagging the missing form for the account manager to handle personally. It distinguishes between an administrative snag and a genuine roadblock. This is not just automation. This is intelligence.</p><p>A principle-driven brain doesn&#8217;t just execute tasks; it understands priorities. It learns to apply the same nuanced thinking your best people use.</p><p>This is how your business already operates, through unwritten principles:</p><ul><li><p>In banking, you prioritize high-value clients, not just high-value transactions.</p></li><li><p>In pharma, you escalate patient reports of unusual side effects, even if they are not on an official list.</p></li><li><p>In your supply chain, you review uncommon shipping delays, not just any delay over 48 hours.</p></li><li><p>In insurance, you assess claims where multiple risk signals converge, not a dollar amount.</p></li><li><p>In marketing, you review AI-generated copy that sound robotic.</p></li><li><p>In HR, you investigate atypical patterns in employee expense reports, not by dollar amount.</p></li><li><p>In manufacturing, you flag unexplained dips in production quality, not every deviation from the norm.</p></li><li><p>In your legal team, you prioritize review of contracts with non-standard liability clauses.</p></li><li><p>And in customer service, you escalate support tickets with language that indicates frustration.</p></li></ul><p>AI can make these principle-based judgements. But each requires context and a process to apply it in the right way at the right moment. AI that runs on principles, powered by the unified context of your enterprise brain, can finally provide it. It can act not just on what the data says, but on what the situation demands.</p><h4><strong>The Strategic Choice</strong></h4><p>The choice before you is not about which AI tool to use. It is about what you want your AI to be and how your initial pilots can work at scale. Do you want AI to come from a third party? Serendipity? Or do you want to control your brain and harness the wisdom of your data, people, and processes?</p><p>AI from a third party can tell you who&#8217;s in your CRM. AI with wisdom can tell you which customers are active on Zendesk and determine they may be at risk of leaving and how to save the relationship. That leap from data to wisdom is impossible when your company&#8217;s memory is fragmented or outsourced.</p><p>ECM platforms are the engine of that transformation. They gather scattered context and turn it into knowledge. But once you stop feeding agents with data and rules and start feeding them with principles and context, you unlock enormous upside and expose all the messy, unsolved problems of principle design, debugging, and accountability. That&#8217;s the &#8220;next chapter&#8221; of enterprise AI that AI One has been treating as a personal mission.</p><p>Once agents are operating on principles instead of narrow rules, their effective agency increases. That creates a new class of enterprise challenges: how you represent principles in policy, how you audit decisions that were made under uncertainty, and how you explain &#8220;why the agent did that&#8221; to regulators and risk teams. This is the governance surface we&#8217;ve been obsessing over for the past two years, and it will be the topic of our next post.</p><p>This is how you create an AI that truly grows with your business. The Hybrid Loop may be the blueprint for an architecture that thinks, but the brain itself is made of context. When you own that context, every question your employees ask makes the AI smarter. Every decision becomes part of its permanent memory. Every relationship it maps deepens its understanding. This is not rented intelligence. This is the enduring mind of your enterprise.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rjRj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rjRj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg" width="852" height="852" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:852,&quot;width&quot;:852,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38729,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/175492402?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!rjRj!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2af9ad5-02dd-4ba9-8815-bc56fc9bbf1b_852x852.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">ECM helps turn data into wisdom when working with AI. <a href="https://www.gapingvoid.com/want-to-know-how-to-turn-change-into-a-movement/">Gapingvoid</a></figcaption></figure></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Stop Building Agent Chains. Start Building Hybrid Loops.]]></title><description><![CDATA[Why enterprises need architectures that think, not agents that drift.]]></description><link>https://enterprisecontextmanagement.substack.com/p/stop-building-agent-chains-start</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/stop-building-agent-chains-start</guid><dc:creator><![CDATA[AI One]]></dc:creator><pubDate>Tue, 02 Dec 2025 16:23:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Krq_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If you work with modern AI models, you will soon feel a strange tension. On one hand, you are told to give the model everything in one large prompt. This includes long context, rich instructions, a list of tools, and then you let the model figure it out. On the other hand, the problems you care about are messy. They are long running and tied to systems in the real world. They can involve thousands of tool calls, changing state, partial failures, and many chances to get things quietly wrong.</p><p>This essay describes the architecture that resolves this conflict: the hybrid loop. It is written for leaders who know LLMs are powerful, but who now must build durable, real-world systems that think. The core idea is simple:</p><ul><li><p>Let the model think deeply and globally at a few key points.</p></li><li><p>Let deterministic code own execution, state, and safety.</p></li><li><p>Let a hybrid agentic architecture guide the whole thing.</p></li></ul><p>Before we examine the hybrid loop, consider the architectures that came before it.</p><h4><strong>From Stepwise Agents to Single Shot Planning</strong></h4><p>In the first wave of LLM systems, the model was a short-term problem solver. It could not hold a long chain of reasoning. It had limited context. Its planning was weak and its tool use was fragile. So engineers built what they had to: stepwise agents.</p><p>The pattern was always the same. The model would:</p><ol><li><p>Look at the current context.</p></li><li><p>Propose a small next action.</p></li><li><p>Trigger a tool or fetch some data.</p></li><li><p>See the result.</p></li><li><p>Decide what to do next.</p></li></ol><p>Each decision was another call to the model. For anything complex, you quickly had long chains of calls. Every step replayed most of the past context. Every step added a little more noise. And every step was a chance for the model to drift. Frameworks emerged to manage these chains: strict REACT style, Planner-executors, Multi agent, self-refinement, and a raft of legacy orchestration frameworks such as LangGraph, CrewAI and AutoGen to construct and support them. But they all shared the same fundamental flaws:</p><ul><li><p>Latency and cost grew with the number of steps.</p></li><li><p>Long chains were brittle and hard to debug.</p></li><li><p>The same queries behaved differently when the context order changed.</p></li><li><p>There was no clean way to define success or set budgets.</p></li></ul><p>Then the models changed. Context windows grew. Tool use became more reliable. Planning improved. The best models started to feel less like chatbots and more like strong planners. They could take a clear goal and lay out a multi-phase approach in one shot.</p><p>This created a new temptation: Single Shot prompting: push everything into one big prompt, ask for the full plan and the final answer, and be done with it.</p><p>Sometimes that works. If the environment is static and the tools are simple, a single planning and execution pass can be enough. But most business problems are not like that. Markets move. Data changes. APIs fail. The Single Shot method is too rigid for the real world.</p><p>The chain of agents is too brittle. The Single Shot is too blind. A third way was needed.</p><h4><strong>The Hybrid Loop: Planning, Supervision, and Verification</strong></h4><p>A better way has emerged: the Hybrid Loop. It starts from a different division of labor. Instead of asking one model to &#8220;be the agent&#8221; and do everything, it separates the task into clear roles and underlying manager. This combines the strengths of today&#8217;s powerful models with the control of iterative execution.</p><p>The hybrid loop is an architecture of five actors, each with a distinct role. Four of them perform the work; a fifth, the Context Manager, directs the flow of information between them.</p><p>First, there is a <strong>planner model</strong>. This is a frontier reasoning LLM used for what it is now very good at: taking a high-level goal, a list of tools, requirements, and constraints, and producing a structured plan. Not a paragraph of text, but a multi-phase object that describes what to do, in what order, and with what success criteria.</p><p>Second, there is a <strong>supervisor</strong>. This is not a model. It is regular code. It holds the current plan and the execution state. It owns tool calls, asynchronous fanout, parallelism, budgets, and state transitions. It knows nothing about language. It just executes and enforces.</p><p>Third, there is a <strong>verifier</strong>. This combines deterministic code and LLMs, but has a narrow mandate: judge, never act. It is fed compact context snippets, like a generated SQL query, a JSON structure, an LLM output, or even execution traces. It produces<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> structured labels such as &#8220;valid,&#8221; &#8220;block,&#8221; or &#8220;unsafe<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>.&#8221; These labels then feed back into the Supervisor.</p><p>Fourth, there is the <strong>tool layer</strong>: your APIs, databases, and other services. That is where the actual work happens. Tool calls are asynchronous, often parallel, and frequently the dominant cost and latency in the system. They are made by a combination of code, and LLM instantiated calls.</p><p>Underpinning them all is the <strong>Context Manager</strong>, the operation&#8217;s central nervous system. It sends targeted pulses of information to each actor. Just what they need, exactly when they need it. Nothing more, nothing less.</p><p>Once you see the architecture in these terms, the question is no longer &#8220;how do I chain agent steps,&#8221; but &#8220;how do I build clean interactions between these roles.&#8221;</p><p>The loop is &#8220;agentic&#8221; because the system as a whole acts like an agent: it turns goals into actions. It is &#8220;hybrid&#8221; because the intelligence is shared between models and code in a deliberate way.</p><h4><strong>A Day in the Life of a Hybrid Loop</strong></h4><p>To make this concrete, imagine a standard business task: a user asks for a set of risk reports. This requires joining data across multiple systems, selecting the right tables, generating valid SQL queries, and writing a report.</p><p>Here is what happens inside a hybrid loop.</p><p>The planner model sees the request, a description of available tools, and a set of constraints. Instead of improvising step by step, it returns a plan object. That plan might say:</p><ul><li><p><strong>Phase 1:</strong> Discover relevant tables using a catalog search tool.</p></li><li><p><strong>Phase 2:</strong> Fetch detailed schema information only for candidate tables.</p></li><li><p><strong>Phase 3:</strong> Generate one or more SQL queries that answer the question.</p></li><li><p><strong>Phase 4:</strong> Execute the queries and check the results.</p></li><li><p><strong>Phase 5:</strong> Write a report and send it to the user.</p></li></ul><p>This plan is not created in a vacuum. The Planner constructs it using a curated set of information provided by the <a href="/__u/enterprisecontextmanagement.substack.com/p/skyscraper">Context Manager</a>, which draws from a Context Catalog of available tools, data sources, and constraints, organized in levels (for more on the architecture of context, read <a href="/__u/enterprisecontextmanagement.substack.com/p/skyscraper">The AI Context Skyscraper</a> on this Substack)</p><p>The resulting plan includes success criteria (e.g., &#8220;at least N candidate tables with a date column and a measure column,&#8221; hard constraints (&#8220;never use tables whose name starts with archive_ or tmp_&#8221;), budgets, (maximum tool calls or time per phase), and instructions on how to summarize tool outputs before they&#8217;re sent back to a model.</p><p>The supervisor picks up this plan and begins execution. During phase 1, it might call a table listing tool ten times, often in parallel, with different search terms. It collects and filters the results according to selection rules in the plan, like ignoring certain tables or requiring a time dimension. All of this is done in code, not by a model.</p><p>To generate the query in Phase 3, the supervisor does not send the model a raw data dump. It asks the Context Manager for a surgical strike. The Manager prepares a tailored package of information of the user&#8217;s goal and a lean summary of candidate schemas. The model can then reason without distraction.</p><p>If the stakes are high, the supervisor can call a verifier model with the SQL and schema summary. The verifier does not try to rewrite anything. It simply labels the query: for example &#8220;valid, warn&#8221; or &#8220;invalid, block, attempts to access unauthorized PII data<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>.&#8221;</p><p>If the label is acceptable, the supervisor proceeds. If not, it records the problem and decides whether to ask for a new plan or to stop.</p><p>All the while, the supervisor is monitoring whether the plan&#8217;s expectations are being met. If the plan is not working, it can prepare a summary of what has happened and send that back to the planner for a patch or a complete replan.</p><p>The hybrid loop is not a free-form conversation. It is a state machine that moves through planning, execution, verification, and replanning, with models and tools playing defined roles at each stage.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Krq_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Krq_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png" width="728" height="524" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:529272,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/180443571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Krq_!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda91cf1-d33f-4cdb-99af-4b6d4d7160c6_4368x3144.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Hybrid Agent Loop: Plan. Execute. Verify. Judge. It is guided by injecting the right context at the right time throughout this continuous loop, with replanning, plan modification, escalation, and ultimately, more accurate answers.</figcaption></figure></div><h4><strong>When the Plan is Wrong</strong></h4><p>In a toy example, the first plan might be perfect. In the real world, plans can be wrong. The hybrid architecture assumes this and makes being wrong a first-class concern.</p><p>When the plan does not match the world, there are two basic responses.</p><p>If the plan is broadly sound but a detail is off, the supervisor can ask the planner for a <strong>patch</strong>. It sends a small evidence pack, and the planner responds with changes. The supervisor applies those changes and continues.</p><p>If the plan&#8217;s assumptions are fundamentally broken, the supervisor can ask for a <strong>rewrite</strong>. This time it compiles a richer summary of the environment. The planner uses that to construct a new plan. The supervisor switches to this new plan and starts again.</p><p>There is a final, crucial guardrail for <strong>when the system itself loses confidence</strong>. The hybrid loop can be designed to know what it doesn&#8217;t know. If the planner is stuck, or if a result falls outside a defined boundary of certainty, the system stops and escalates. It can invoke the most reliable tool of all: a human in the loop.</p><p>In both cases, the combination of planner, supervisor, and verifier lets the system change course without becoming opaque. Plans, patches, and rewrites are all explicit objects that can be logged, compared, and audited.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!saPi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!saPi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:381760,&quot;alt&quot;:&quot;User Requests are sent to the Hybrid Loop for planning. The supervisor executes the plan, verifies intermediate results, and determines if the plan is complete; depending on its judgement, the plan is either tweaked, completely replanned, or escalated for human judgement. The loop continues until results are delivered&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/180443571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="User Requests are sent to the Hybrid Loop for planning. The supervisor executes the plan, verifies intermediate results, and determines if the plan is complete; depending on its judgement, the plan is either tweaked, completely replanned, or escalated for human judgement. The loop continues until results are delivered" title="User Requests are sent to the Hybrid Loop for planning. The supervisor executes the plan, verifies intermediate results, and determines if the plan is complete; depending on its judgement, the plan is either tweaked, completely replanned, or escalated for human judgement. The loop continues until results are delivered" srcset="/__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!saPi!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db1e803-d3f2-4304-a643-681c37b669c6_4368x3144.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">User Requests are sent to the Hybrid Loop for planning. The supervisor executes the plan, verifies intermediate results, and determines if the plan is complete; depending on its judgement, the plan is either tweaked, completely replanned, or escalated for human judgement. The loop continues until results are delivered.</figcaption></figure></div><h4><strong>Context is a Design Surface</strong></h4><p>Raw context is noise. The hybrid loop works because it treats context as a design surface, and the Context Manager is the designer. It derives intent, builds the complete context window, and optimizes each section using the most efficient techniques for that section.</p><p>The context manager keeps the world view the planner needs that&#8217;s rich enough to propose a sensible strategy, but not so cluttered that it gets lost in details: The verifier needs tightly shaped inputs that let it answer specific questions; Patch and replan prompts need the right amount of evidence (too little and the planner will guess. Too much and it will waste cycles rewriting what already works.); Even tool outputs need to be treated carefully.</p><p>The quality of the system depends as much on how you use the context manager to shape and compress these contexts as on which model you use. The same model with well-optimized context will often outperform a larger one wrapped in noise.</p><h4><strong>Determinism, Cost, and Speed</strong></h4><p>Hybrid loops are gaining ground because they hit a sweet spot between predictability and power.</p><p><strong>Determinism</strong> comes from the fact that the supervisor is real code and the models are constrained. There is no hidden control logic inside long prompts that change without anyone noticing.</p><p><strong>Cost</strong> is contained because the number of heavy model calls is small. You are not paying for the model to re-read its entire context on every micro-decision.</p><p><strong>Speed</strong> follows from the fact that the supervisor can aggressively fan out tool calls. There is no need for a single model to alternate between reading, thinking, acting, and reading again.</p><p>At the same time, you <strong>keep the human advantages of modern LLMs</strong>. They still do the hardest part: turning vague goals into workable plans.</p><p>The end result is a system that can do serious work, in real environments, at a cost and latency profile that leadership can live with, and with a level of traceability that risk and audit teams can understand.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AyNI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 424w, /__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 848w, /__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AyNI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png" width="4360" height="2686" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2686,&quot;width&quot;:4360,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:662730,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/180443571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb5dad3d-d754-414f-9efb-a35f0c36e7b6_4368x3144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 424w, /__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 848w, /__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AyNI!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa06d6476-c006-4452-83aa-a526a5e8b4fd_4360x2686.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">In summary: Stepwise Agents versus Single Shot Prompts versus Hybrid Agent Loops.</figcaption></figure></div><h4><strong>Architecture That Thinks</strong></h4><p>If you accept this picture, it has a simple consequence for how you think about AI platforms.</p><p>Your AI infrastructure should not be a model with prompts; it must be a reasoning architecture. The Context Manager is its heart, pumping precisely the right information to the planner, supervisor, and verifiers. The hybrid loop is an architecture that does not just execute; it reasons.</p><p>For technical leaders, the important point is not the label. It is the underlying separation of concerns. If your systems still think in terms of stepwise chats with an &#8220;agent,&#8221; you will find yourself fighting the tools as models continue to improve. If your systems treat models as planners and judges wrapped in a deterministic execution fabric, you will be able to take advantage of that improvement with far less friction.</p><p>Hybrid loops are, in that sense, less a feature and more a sign that your architecture has caught up with what the models can actually do.</p><div><hr></div><h4>For More About Enterprise Context Management</h4><p>Read our inaugural post: <strong><a href="/__u/enterprisecontextmanagement.substack.com/p/intro">Your AI is Blind: Enterprise Context Management Arrives</a> by </strong><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Conor Twomey&quot;,&quot;id&quot;:261184426,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e24bab1-6d79-40de-add0-7e2a451d349b_3344x3344.jpeg&quot;,&quot;uuid&quot;:&quot;3b53fbfb-9a9d-43cd-b6ea-49e6938a5d92&quot;}" data-component-name="MentionToDOM"></span> and <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Fergus Keenan&quot;,&quot;id&quot;:261638784,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f6bfdf5-ddcd-4c75-b0c1-558cd45ecee6_5464x5464.jpeg&quot;,&quot;uuid&quot;:&quot;6e6d2bc4-990e-435f-997e-72c343100b51&quot;}" data-component-name="MentionToDOM"></span>. </p><p>For more on an organizational architecture for context, check out <a href="/__u/enterprisecontextmanagement.substack.com/p/skyscraper">The AI Context Skyscraper</a> by <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Mark Sykes&quot;,&quot;id&quot;:62529052,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/009ddbf2-dc50-4d5d-bc66-c6c819125615_144x144.png&quot;,&quot;uuid&quot;:&quot;0d51b120-12e7-4f34-8891-92e748a46e2a&quot;}" data-component-name="MentionToDOM"></span>. </p><div><hr></div><h4>Subscribe to the ECM Substack</h4><p>Subscribe to the <a href="/__u/enterprisecontextmanagement.substack.com/">Enterprise Context Management Substack</a> to track the evolving world of context management.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/enterprisecontextmanagement.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h4>Connect With Us!</h4><p>Follow us. Because Context Matters :)</p><ul><li><p><a href="https://www.linkedin.com/company/ai-one-1/">AI One on LinkedIn</a></p></li><li><p>Co-founder <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Conor Twomey&quot;,&quot;id&quot;:261184426,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e24bab1-6d79-40de-add0-7e2a451d349b_3344x3344.jpeg&quot;,&quot;uuid&quot;:&quot;ce65e7fe-0ad7-4cea-bc42-f38f556485cc&quot;}" data-component-name="MentionToDOM"></span> on Substack and <a href="https://www.linkedin.com/in/conortwomey/">LinkedIn</a> </p></li><li><p>Co-founder <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Fergus Keenan&quot;,&quot;id&quot;:261638784,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f6bfdf5-ddcd-4c75-b0c1-558cd45ecee6_5464x5464.jpeg&quot;,&quot;uuid&quot;:&quot;057f5db5-13e5-4232-b473-4552f44a8bf0&quot;}" data-component-name="MentionToDOM"></span> on Substack and <a href="https://www.linkedin.com/in/fergus-keenan/">LinkedIn</a></p></li><li><p>The author of this article, <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Mark Sykes&quot;,&quot;id&quot;:1536418,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18ae698e-2d06-44ba-aa93-663790252b97_751x751.jpeg&quot;,&quot;uuid&quot;:&quot;05d4dd07-44d1-470b-8477-4b32f701b3de&quot;}" data-component-name="MentionToDOM"></span>, Chief AI Officer of AI One</p></li><li><p>Editor and contributing writer to this Substack, <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Mark Palmer&quot;,&quot;id&quot;:6384101,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/1ea6f5d5-7ddd-4484-9ef4-7d0260deccdb_2226x2226.jpeg&quot;,&quot;uuid&quot;:&quot;cdd8c11e-2a0f-4fe9-b689-8ea108352a5b&quot;}" data-component-name="MentionToDOM"></span> (and follow <a href="https://www.linkedin.com/in/markwpalmer/">Mark on LinkedIn</a> or his Substack, <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Techno Sapien&quot;,&quot;id&quot;:1234291,&quot;type&quot;:&quot;pub&quot;,&quot;url&quot;:&quot;https://open.substack.com/pub/technosapien&quot;,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/e54176ff-1833-40be-a04d-5fc6b356ddf4_500x500.png&quot;,&quot;uuid&quot;:&quot;aef73a8d-18a6-4465-84dd-55e864604f0c&quot;}" data-component-name="MentionToDOM"></span>.)</p></li></ul><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><h6>Using a mixture of deterministic routines and side-chain LLM classifiers.</h6></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><h6>The key to making this work is ensuring the <em>right</em> validator runs for the <em>right</em> tool. For example: invoke an external LLM judge for free-text outputs, a deterministic library for SQL validation, an external tool for PII risks, SMT solvers for formally-defined workflows. Choosing appropriately makes the difference between a dependable workflow and a brittle one.</h6></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><h6>Query labelling is a powerful supervisor task. It can annotate queries with any additional context, like semantic validity or formalized correctness techniques, depending on the nature of the task.</h6></div></div>]]></content:encoded></item><item><title><![CDATA[The AI Context Skyscraper]]></title><description><![CDATA[A guidebook to mastering the architecture of context]]></description><link>https://enterprisecontextmanagement.substack.com/p/skyscraper</link><guid isPermaLink="false">https://enterprisecontextmanagement.substack.com/p/skyscraper</guid><dc:creator><![CDATA[Mark Sykes]]></dc:creator><pubDate>Wed, 29 Oct 2025 14:05:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AjNj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For AI, great answers come from more than great questions. They come from great context. The right context includes the guardrails, historical data, environmental information, and guidance that turn raw outputs into real insights.</p><p>But while &#8220;context engineering&#8221; is an up-and-coming discipline in enterprises adopting AI, they&#8217;re often looking at the context the wrong way &#8211; as an engineering task &#8211; instead of a strategic business asset. Context is the very soul of the current AI age enterprise; the architecture of how you want your business to operate. Expecting engineers to engineer context alone is like a carpenter making cabinets without a blueprint for where they belong. Enterprise Context Management (ECM) is a rising enterprise software category that helps capture, curate, and deploy context at the right time and right place.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>But to use context effectively, a new mental model is needed: The AI Context Skyscraper.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AjNj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_webp, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AjNj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7776230,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://enterprisecontextmanagement.substack.com/i/177397069?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_424, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 424w, /__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_848, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 848w, /__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_1272, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AjNj!, /__u/enterprisecontextmanagement.substack.com/w_1456, /__u/enterprisecontextmanagement.substack.com/c_limit, /__u/enterprisecontextmanagement.substack.com/f_auto, /__u/enterprisecontextmanagement.substack.com/q_auto:good, /__u/enterprisecontextmanagement.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fb0e21c-d6df-4269-856e-bc6b135f3d43_4368x3144.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>Context as Architecture</strong></h3><p>Instead of a single, undifferentiated block of text, imagine the context you provide to an LLM as a skyscraper, with each floor representing a different layer of information: corporate strategy on the top floor, customer interactions in the middle, and reusable recipes traveling by elevator from floor to floor.</p><p>The architecture of context is as important as the context itself. The position of each floor matters. The order, organization, and optimization of this context directly influences the LLM&#8217;s focus and the quality of its output.</p><p>Let&#8217;s explore this context skyscraper for FinCorp, a global financial services company, moving from the penthouse to the sub-basement with concrete examples.</p><h3><strong>The Penthouse: System Prompts</strong></h3><p>At the top of our skyscraper sits the system prompt. This is the most privileged real estate, and LLMs usually pay the most attention to what appears here, making it the core of the AI&#8217;s identity and objectives. This is where you define AI&#8217;s personality, style, and the fundamental principles that govern its behavior. </p><p>FinCorp&#8217;s system prompt might read:</p><blockquote><p><em>&#8220;You are an assistant specializing in long-term financial advice and wealth management, maintaining a conservative approach to risk.</em></p><p><em>Ensure that all recommendations are tailored to each client&#8217;s specific goals, comply with SEC regulations, and align with our distinct investment principles and current market outlook. Prioritize accuracy over speed in every response, and always cite reliable sources.</em></p><p><em>After each recommendation, validate that the advice aligns with the client&#8217;s objectives, regulatory standards, and the latest market data; self-correct or escalate if there are any uncertainties. If uncertain about any information or recommendation, escalate the query to human oversight rather than offering speculative advice.&#8221;</em></p></blockquote><p>System prompts, like penthouse residents, are your anchor tenants. They must be carefully curated, tested for suitability, then locked down. They represent your AI constitution. Your strategic north-star. Your principles. System prompts are treated with the reverence that position demands.</p><h3><strong>The Executive Suite: Subsystem Prompts</strong></h3><p>Just below the penthouse are subsystem prompts. This is your role identifier, the place where you specify the AI&#8217;s role for a specific set of tasks. It&#8217;s where the AI gets marching orders from your Chief Revenue Officer, Chief Marketing Officer, or Chief Financial Officer to align overall strategy with functional execution.</p><p>For example, FinCorp&#8217;s CFO might guide AI used to detect fraud with a subsystem prompt like this:</p><blockquote><p><em>&#8220;You are a fraud research assistant responsible for assessing transaction patterns for unusual behavior, potential regulatory violations, and areas of uncertainty.</em></p><p><em>Immediately flag any anomalies that deviate by more than three standard deviations from established baselines. For each flagged anomaly, provide a one-line validation and indicate the next research step or correction if needed.</em></p><p><em>Before performing in-depth research on transactions involving high-risk jurisdictions or entities on watchlists, state the purpose and the minimum information you will use.&#8221;</em></p></blockquote><p>Executive suite context prompts occupy the space just below the penthouse, and are owned by functional leaders at the highest level. They command significant attention from the LLM, making it a powerful lever for shaping behavior when the penthouse is off-limits.</p><h3><strong>The Upper Floors: Conversations, Workflow, History</strong></h3><p>As we descend into the upper floors, we encounter day-to-day interaction context. An example of context at this level are prior responses in an ongoing user conversation: every relevant question asked, every answer given, every clarification sought.</p><p>FinCorp&#8217;s conversational history might contain:</p><pre><code><strong>User</strong>: &#8220;Show me Q3 revenue for the EMEA region&#8221;

<strong>AI</strong>: &#8220;Q3 EMEA revenue was $45.2M, up 12% YoY&#8221;

<strong>User</strong>: &#8220;How does that compare to APAC?&#8221;

<strong>AI</strong>: &#8220;APAC Q3 revenue was $38.7M, up 8% YoY. EMEA outperformed APAC by $6.5M&#8221;

<strong>User</strong>: &#8220;What&#8217;s driving the EMEA growth?&#8221;</code></pre><p>Conversational history is essential context, but it can become bloated quickly. This floor is a prime candidate for intelligent compression. You want to retain the essence of the conversation without overwhelming AI with every word that&#8217;s been exchanged, or including irrelevant distractions</p><p>Summarization techniques work well here, allowing you to preserve the narrative arc and key information while reducing the token footprint. The compressed version might become: <em>&#8220;User is analyzing Q3 regional performance, specifically comparing EMEA ($45.2M, +12% YoY) vs APAC ($38.7M, +8% YoY), now investigating EMEA growth drivers.&#8221;</em></p><p>In addition to conversations, the upper floors also house the context of workflows. Workflows include the steps AI takes to reason through complex questions, showing which tools have been called, in what order, and with what parameters. Workflow context provides an audit trail and can be refined to build organizational wisdom for answering complex questions.</p><p>FinCorp&#8217;s question could require these steps:</p><pre><code><strong>Called Salesforce</strong> (Table=&#8220;revenue&#8221;; Filter=&#8220;Region&#8221;, EMEA, quarter=Q3)

<strong>Called Salesforce</strong> (Table=&#8220;revenue&#8221;; Filter=&#8220;Region&#8221;, APAC, quarter=Q3)

<strong>Called Database</strong> Calculate Variance(45.2, 38.7)

<strong>Next: Analyze Growth Drivers </strong>(Region=&#8220;EMEA&#8221;, quarter=&#8220;Q3&#8221;)</code></pre><p>When crafting this context, prune irrelevant steps but maintain the sequence of actions. This floor provides the AI with a sense of what has been tried, what worked, and what didn&#8217;t, enabling more intelligent decision-making about next steps.</p><h3><strong>The Middle Floors: Results for Explainability</strong></h3><p>The middle section of our skyscraper houses some of the most critical and untouchable content: the raw output from relevant API calls, database queries, or data returned by external systems. This context provides the raw ingredients for explainability in AI. The ground truth. The audit trail. Result sets should rarely be optimized, compressed, or summarized by automated processes, only logged.</p><p>Raw output is the factual foundation upon which the AI&#8217;s reasoning is built, and any corruption of this data could lead to catastrophically wrong conclusions and the erosion of trust and verifiability in outcomes.</p><p>A database query might return:</p><pre><code>{ &#8220;query_id&#8221;: &#8220;rev_q3_emea&#8221;,
  &#8220;timestamp&#8221;: &#8220;2024-10-15:23&#8221;,
  &#8220;results&#8221;: [
    {&#8221;country&#8221;: &#8220;Germany&#8221;, &#8220;revenue&#8221;: 18500000, &#8220;yoy&#8221;},
    {&#8221;country&#8221;: &#8220;UK&#8221;, &#8220;revenue&#8221;: 14200000, &#8220;yoy&#8221;},
    {&#8221;country&#8221;: &#8220;France&#8221;, &#8220;revenue&#8221;: 12500000, &#8220;yoy&#8221;}],
  &#8220;records&#8221;: 3}</code></pre><p>The logging of the precise technical question asked ensures errors, bias, and hallucinations can be understood, audited, and fixed.</p><p>Along with the calls made to AI, we preserve the results.</p><p>Building on this raw data, the middle floor may also contain any summary tokens that might be provided by the AI&#8217;s reasoning or chain of thought: its step-by-step progression of how it arrived at a recommendation. For example, FinCorp&#8217;s reasoning chain might have been:</p><pre><code><strong>User asked:</strong> &#8220;EMEA growth drivers&#8221;

<strong>Retrieved</strong>: Country Level (Germany (+15%), UK (+9%), France (+11%))

<strong>Observation</strong>: Germany represents 41% of EMEA revenue and has highest growth rate

<strong>Prediction</strong>: This suggests Germany is the primary growth driver

<strong>Next</strong>: Explore what changed in the German market in Q3</code></pre><p>The data from this floor should remain largely untouched, serving as the intellectual scaffolding of the AI&#8217;s decision-making process. Storing questions, results, and reasoning chains ensures you can always show your work and demonstrate how the AI arrived at its conclusions.</p><h3><strong>The Lower Floors: Retrieved Context and Governance</strong></h3><p>As we move into the lower floors, we encounter the next set of context for the question: the information retrieved from knowledge bases, vector databases, or other sources to help answer the current query.</p><p>When investigating EMEA growth drivers, the system might extract key facts from a knowledge base:</p><pre><code>Germany Q3 2024: <strong>Launched new enterprise product tier</strong> in July

UK Q3 2024: Faced <strong>increased competition</strong> from local fintech startups

France Q3 2024: <strong>Expanded partnership</strong> with major retail bank, added 150 enterprise customers

EMEA Market Report Q3: Manufacturing sector had <strong>18% increase in spending</strong></code></pre><p>Rather than simply dumping results from a vector similarity search using RAG, a more sophisticated approach involves a &#8220;Context Loop&#8221;, which uses intelligent chunking and multiple similarity search passes on the raw underlying documents and data-sources, to explore the best match, drill down, and follow threads of relevance. Instead of taking a single retrieval pass, the Context Loop successively determines the best, most relevant, context. Context looping is particularly important when extracting context from numerous large documents, extensive database schemas, or complex knowledge graphs.</p><p>Below Context Loop-extracted context, we establish the guardrails, the hard rules, and the constraints that govern the AI&#8217;s behavior. These are non-negotiable boundaries: what data the AI cannot access, what actions it cannot take, and what topics it must avoid:</p><pre><code><strong>NEVER</strong> access or display individual customer names or personal identifiers

<strong>NEVER</strong> execute queries against production databases, only read replicas

<strong>NEVER</strong> recommend actions that would violate GDPR or data residency requirements

<strong>NEVER</strong> process or display salary information for identifiable individuals

<strong>MUST</strong> reject any request to bypass audit logging

<strong>MUST</strong> escalate to human oversight for any transaction &gt;$1M</code></pre><p>These guardrails are strictly enforced and should never be compressed or optimized away. They are the safety systems of your AI, the circuit breakers that prevent catastrophic failures.</p><p>Adjacent to the guardrails, we have conduct and expectations&#8212;a softer set of guidelines that shape the AI&#8217;s personality and communication style:</p><pre><code>Maintain a <strong>professional, formal tone</strong> appropriate for financial services

When presenting numerical data, always <strong>include units and time periods</strong>

<strong>Acknowledge uncertainty</strong> rather than presenting speculation as fact

<strong>Provide context for percentages</strong> (e.g., &#8220;15% growth on a base of $18.5M&#8221;)

Use inclusive language and <strong>avoid regional bias</strong> when comparing markets

<strong>Cite data sources</strong> for all quantitative claims</code></pre><p>Unlike the hard guardrails, these are more like cultural norms, shaping how the AI presents itself and interacts with users.</p><h3><strong>The Basement: Playbooks and Multi-Agent Integration</strong></h3><p>The basement of our skyscraper houses the playbook, a truly innovative feature. This is an LLM-populated guide that captures the accumulated wisdom from previous executions of similar workflows. Every time the AI successfully completes a task, it can contribute to this playbook, noting which strategies worked well and which approaches led to dead ends.</p><pre><code><strong>PLAYBOOK - Revenue Analysis Workflows</strong>

<em><strong>Pattern: Regional growth investigation</strong></em>
- Success strategy: Start with country-level breakdown 
- Use: regional_sales

<em><strong>Pattern: EMEA analysis</strong></em>
- Avoid: Don&#8217;t rely on revenue table - it excludes partner channel sales
- Use: combined_revenue_view

<em><strong>Pattern: Growth driver identification</strong></em>
- Success strategy: Check for product launches
- Success strategy: Check for partnership announcements</code></pre><p>The playbook represents organizational learning at scale, turning every AI interaction into a source of competitive advantage. It&#8217;s how your AI gets smarter over time through the accumulation of procedural knowledge. When faced with a new instance of a familiar problem, the AI can consult the playbook to see what has worked before, dramatically improving its efficiency and success rate.</p><p>This floor should be carefully curated but not aggressively compressed. The insights here are valuable precisely because they contain nuance and detail. Strip that away, and you lose the very knowledge you&#8217;re trying to preserve.</p><p>Finally, we have space for input from other agentic frameworks. In a multi-agent system, where different AI agents handle different aspects of a complex workflow, this floor provides a clean integration point. It&#8217;s where Agent A can pass context and results to Agent B, enabling sophisticated orchestration without contaminating the other floors of the skyscraper.</p><p>This separation is crucial for maintaining clarity and debuggability in complex systems. When something goes wrong, you need to be able to trace the flow of information between agents, and having a dedicated floor for inter-agent communication makes that possible.</p><h3><strong>The Context Foundation Meets the Task At Hand</strong></h3><p>Underpinning all the prior context sits the prompt itself: the specific instructions for the current task.</p><p>First should come the original question that we are seeking an answer to. This is the user&#8217;s initial query that triggers the entire workflow. Even as the AI moves through multiple steps in an agentic loop, calling various tools and processing information, this original question serves as a north star.</p><blockquote><p><em>&#8220;What&#8217;s driving EMEA growth?&#8221;</em></p></blockquote><p>It&#8217;s the constant reference point that ensures the AI doesn&#8217;t drift off course, doesn&#8217;t get lost in the weeds of intermediate processing, and ultimately delivers an answer to what the user actually wanted to know.</p><p>Finally, we reach the current task. This is what the AI is being asked to do right now, at this moment. It needs to be clear, unambiguous, and immediately actionable.</p><blockquote><p><em>&#8220;Using the country-level revenue data retrieved from the database, identify the top 3 contributors to EMEA growth in Q3. For each country, calculate its contribution to the overall regional growth and identify any significant product or customer segment changes. Present findings in order of impact magnitude.&#8221;</em></p></blockquote><p>This is the AI&#8217;s current focus, the question it&#8217;s actively working to answer. It is at the bottom of the prompt due to something known as recency bias, where LLMs pay a little more attention to the most recent information it&#8217;s been given.</p><p>Some models will even improve their response if this is underlined with one repeated strata of instruction such as:</p><blockquote><p>&#8220;<em>Provide a ranked list of the top 3 EMEA countries driving Q3 growth, with each country&#8217;s % contribution and key product or segment changes. No commentary or extra text.</em>&#8221;</p></blockquote><p>This question-answering is where the context rubber meets the road, and it takes careful interplay with the LLM to keep it on task and context-aware.</p><h3><strong>The Fatal Flaw of One-Size-Fits-All Optimization</strong></h3><p>The tour of our skyscraper reveals the fundamental flaw in current context management: a one-size-fits-all approach to optimization. Most optimization techniques treat the entire context window as a single, undifferentiated mass. They apply the same algorithm to every floor of the skyscraper, often all at the same time.</p><p>This is architectural malpractice. It&#8217;s like trying to renovate a building by applying the same design to the penthouse, the boiler room, and every office in between. You wouldn&#8217;t insulate your lobby the same way you insulate your roof, and you shouldn&#8217;t optimize your tool results the same way you optimize your conversational history.<br><br>The list of optimization techniques available to apply to each floor is extensive, including some such as:</p><ul><li><p>Query-focused Summarization</p></li><li><p>GEPA (GEnetic-PAreto Optimization)</p></li><li><p>ACE (Agentic Context Engineering)</p></li><li><p>Chain of Verification</p></li><li><p>Targeted Truncation and redaction</p></li><li><p>Self refine/Critique re-write</p></li><li><p>Instruction decomposition</p></li><li><p>Context salience labelling</p></li></ul><p>Each floor has different requirements, different sensitivities, different roles in the overall structure. The system prompt needs to be protected and optimized once. Tool results should never be compressed. Conversational history can and should be summarized. The playbook needs careful curation. Guardrails must remain intact.</p><p>When you apply a blanket optimization strategy across the entire context window, each floor inherits the disadvantages of that approach. You end up compromising the integrity of the entire building, losing critical details like the one-time bonus in the France data while wasting tokens on redundant conversational history.</p><h3><strong>The Power of Floor-Specific Strategies</strong></h3><p>The alternative is to manage each floor independently, applying the optimization strategy that makes sense for that particular type of content, and for the LLM that you are targeting. This might mean using one compression technique for conversational history, a different approach for workflow history, and no compression at all for tool results and reasoning chains. A different LLM will require different optimizations for each floor, and even potentially move some floors around.</p><p>This granular control is not merely academic; it is a practical necessity for building robust, reliable, enterprise-grade AI systems. It allows you to maximize the effective use of your context window, preserving the most important information while still fitting within token limits. Flexible meta-optimization architectures enable switching from one LLM to another easily, eliminating single vendor dependencies. It enables you to maintain explainability and traceability where required, while still achieving efficiency where possible.</p><p>Moreover, this approach is fundamentally more flexible. As new optimization techniques emerge, you can adopt them selectively, applying them to the floors where they make sense without disrupting the rest of your carefully constructed context architecture.</p><h3><strong>Sovereign Context: Taking Back Control</strong></h3><p>The most important implication of the context skyscraper model is the concept of sovereign context. When you view context as a complex, multi-layered structure that requires sophisticated management, it becomes clear that you can&#8217;t afford to outsource this responsibility to third-party AI providers, or allow context to be applied in an ad-hoc manner.</p><p>Your context is not just data. It&#8217;s the distilled essence of your business, your processes, your institutional knowledge. It&#8217;s the competitive advantage that makes your AI uniquely valuable to your organization. When you hand control of that context over to a SaaS provider or a frontier LLM provider, you&#8217;re not just creating a security risk, you&#8217;re giving away strategic control.</p><p>Consider our FinCorp example. The playbook contains hard-won knowledge about data quirks, successful analysis patterns, and domain-specific insights. The guardrails encode your risk tolerance and regulatory obligations. The conversational history reveals your analysts&#8217; thinking patterns and priorities. This is not generic, commoditized information; this is your institutional intelligence.</p><p>Sovereign context means building, managing, and owning your context skyscrapers. It means having the flexibility to fine-tune your context for specific tasks without being locked into a vendor&#8217;s optimization strategy. It means ensuring the security and privacy of your data by keeping it under your own roof.</p><p>In the same way that it&#8217;s unreasonable to expect a standardized skyscraper to fit into every city&#8217;s skyline, it is not realistic to expect a context window optimized for a specific model to generate the same results when thrown into a different model. Different LLMs and tools prefer different skyscraper &#8220;shapes&#8221; and techniques in order to provide the most accurate results.<br><br>The risk of optimizing a context window to work only for one model is to create a certain level of single vendor dependence that becomes undesirable in the future. The benefit of using a flexible, floor-based context optimization strategy is that it is trivial to create multi-vendor versions of everything, so you always have the right shaped building for the city you are building in. </p><p>Enterprises that treat their context as a first-class asset, as valuable as their data, code, and other intellectual property, will be the ones to thrive in the AI era.</p><h3><strong>Build Your Own Context Skyscraper</strong></h3><p>As you embark on your enterprise AI journey, adopt the mental model of the Context Skyscraper. Don&#8217;t think of context as a blob of text, think of it as a carefully architected building, with each floor serving a specific purpose and requiring a tailored management approach. </p><p>Build your skyscraper with intention. Understand what belongs on each floor. Apply the right optimization strategies to each level. Protect the floors that need protection, and optimize the ones that can afford it. And above all, never give away the keys to your building. </p><p>The context creation loop is where the real value of enterprise AI is unlocked: the continuous, iterative process of extracting context, managing it through your skyscraper, and using it to drive intelligent workflows. It&#8217;s the difference between AI that merely responds to prompts, and AI that truly understands and advances your business objectives.</p><p>The skyscraper is yours to build. Build it well, and you will unlock the true power of enterprise AI.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://enterprisecontextmanagement.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Enterprise Context Management! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>