<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The AI Frontier]]></title><description><![CDATA[Lessons from building an AI product at RunLLM and the latest AI research at Cal.]]></description><link>https://frontierai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!pSsQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F560e42de-7257-45ac-a22e-9b5517514828_1280x1280.png</url><title>The AI Frontier</title><link>https://frontierai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 12:30:10 GMT</lastBuildDate><atom:link href="/__u/frontierai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Joseph E. Gonzalez]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[frontierai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[frontierai@substack.com]]></itunes:email><itunes:name><![CDATA[Joseph E. Gonzalez]]></itunes:name></itunes:owner><itunes:author><![CDATA[Joseph E. Gonzalez]]></itunes:author><googleplay:owner><![CDATA[frontierai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[frontierai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Joseph E. Gonzalez]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Model inference, model products, and AI applications]]></title><description><![CDATA[Why model choice is more optimizable than ever]]></description><link>https://frontierai.substack.com/p/model-inference-model-products-and-a49</link><guid isPermaLink="false">https://frontierai.substack.com/p/model-inference-model-products-and-a49</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 13 Aug 2026 18:53:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vVf9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is the last repost of our summer break from the blog. Back next week with a new post!</em></p><p><em>Ramp published a report this week showing that there is comparatively limited usage of Fable, while Opus&#8217; usage has jumped very quickly. Simultaneously, with the flood of open-source model releases, we&#8217;re hearing more and more about teams interested in domain-specific post-training. </em></p><p><em>That reminded of us this post from last year: The direction an increasingly intelligent and abstracted model goes in might be different from the direction that an application needs models to go in. How you handle that divergence is critical question for application builders, and the argument is increasingly leaning in favor of open-source intelligence. </em></p><div><hr></div><p><span>GPT-5 is the first major model release from OpenAI that we didn&#8217;t immediately jump to put in production at </span><a href="http://www.runllm.com/">RunLLM</a><span>. To be clear, we tried, but the models that were released simply didn&#8217;t make sense for us to prioritize amongst the many other things that we could spend our time on, and the models that we would have normally prioritized didn&#8217;t show meaningful improvements on our benchmarks.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vVf9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vVf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1534748,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/172181644?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vVf9!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188ca428-4653-4d21-a2e0-e031bd8f82c5_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>Reflecting on why, we have a few key takeaways about the evolution of LLM providers over the last couple years and where the API-first products are headed. Our headline takeaway is that there are early signs of a divergence between pure LLM and inference and smarter LLM-based &#8220;products.&#8221; As application builders, we have a distinct need for control over what the underlying models are doing, so LLM inference is more appealing to us than a more tightly coupled product. GPT-5 is the first major model release that feels like a product rather than a pure LLM, which is part of the reason why we feel GPT-5 doesn&#8217;t quite meet our needs.</p><ul><li><p><strong>Consumer apps vs. developer APIs.</strong><span> We &#8212; like many others &#8212; complained about OpenAI&#8217;s product UX in the days of the model picker with 7+ models. Giving a user that many options when they wanted to run a simple query didn&#8217;t make much sense. The new UX is dramatically better, even if it had a rocky rollout (and even if GPT-5 has some room for improvement compared to previous iterations like o3). Developer APIs on the other hand are different: Application builders are going to make thoughtful, informed, often empirical choices about what options to pick, so having a variety of options is good. The options let an application builder trade off cost, latency, and quality depending on the complexity of each task they&#8217;re doing. That said, the previous model line ups were more intuitive &#8212; thinking about the reasoning effort in </span><code>gpt-5</code><span> and how that compares to a smaller model size&#8217;s output is not immediately obvious to us, and reasoning models in general seem to be higher variance (more on that below). The migration guide on the OpenAI docs helps, but we didn&#8217;t find the analogs to be drop-in replacement in our experiments. As you can tell, there&#8217;s no simple answer here, but there&#8217;s definitely room for improvement.</span></p></li><li><p><strong>Reasoning applications vs. agentic applications.</strong><span> We noticed a wave of GPT-5 based product announcements after the launch, though fewer than before. What sticks out to us is that there are applications which benefit from throwing more compute in a single LLM call for a particular task &#8212; examples like the ones Box CEO Aaron Levie posted around document processing fall into this bucket. For these use cases, switching to a more powerful reasoning model makes perfect sense. However, for agentic applications like RunLLM, improving quality often comes from composing multiple LLM calls rather than having one big call to a powerful model. In fact, we&#8217;ve found that reasoning models &#8212; dating back to </span><code>o1</code><span> &#8212; tend to increase variance in our performance in a way that our customers (and therefore our team) find unnerving. As a result, we&#8217;ve yet to put a reasoning model in our production inference pipeline; instead, we primarily rely on non-reasoning LLMs that are scoped to very narrow tasks and composed into more powerful pipelines. Some of this might be fixed with better prompt tuning, but the variance in latency and output has led us to be more conservative than before with these models.</span></p></li><li><p><strong>Open-weight vs. closed-weight.</strong><span> With all of the quality open-weight models being released over the last 6 months, it feels increasingly viable to construct high-quality applications that rely on open models. The main motivation is a sense that there isn&#8217;t any confusing logic hiding under the hood. Again, this wasn&#8217;t a concern 1-2 years ago, but as the model providers try to build more verticalized model products, it feels like a concern that might begin to emerge over the next few model releases. This will especially feel true if OpenAI moves to deprecate the GPT-4.1 family of models, which we still rely on heavily and would be quite sad to lose. Being forced to switch to a reasoning model for a simple task like filtering out irrelevant questions is not our first choice. The non-reasoning version of GPT-5 is too one-note to replace everything we do, especially since the smaller model is not available via API. For the reasons described above, we find that reasoning models aren&#8217;t (yet) the best fit for us. In that world, relying on open-source models from unbiased infrastructure feels like a safer bet &#8212; or at least moving to lean more heavily on model providers like Google, that seem to be closer to exposing &#8220;just the model.&#8221; We&#8217;re not necessarily confident that this will be an issue for us, but we&#8217;re cognizant that this is a risk area for our product today. If Google and Anthropic feel the need to follow suit with OpenAI&#8217;s latest releases, there will be a strong motivation to switch to more decentralized inference services that are transparent with model behavior rather than more abstracted services.</span></p></li><li><p><strong>Model inference vs. model products</strong><span>. Our bigger picture observation on this front is that pure inference seems to be diverging from a model as a &#8220;product.&#8221; In 2023 and 2024, an open-source inference product provided a similar output to a vertically integrated model provider API &#8212; just with different weights. It&#8217;s starting to feel like that is changing, with the model providers trying to own more of the stack and include more functionality in the system itself. Reasoning and tool use are the two most obvious examples of this. Both are great from a consumer perspective, but they don&#8217;t fit our needs as application builders. That is going to affect the direction of the development of future models. If the major model providers are in an arms race to develop the best model for consumer applications, that might start to diverge from what a model that&#8217;s useful as a composable building block in applications would look like. We don&#8217;t know enough about the development roadmaps at the frontier labs to say for certain, but we&#8217;ve certainly heard engineers from these teams say things like &#8220;a good enough model will solve enterprise problems&#8221; (which we strongly disagree with). Products becoming increasingly vertically integrated and opinionated would undoubtedly push us towards open models.</span></p></li></ul><p>As with most things in AI over the last few years, we&#8217;re speedrunning the hype cycle of past platform shifts. What&#8217;s unique about this situation is that OpenAI has managed to build both a consumer and enterprise business over just a few years, and unlike Google or Amazon as they built out their cloud businesses, both sides of OpenAI&#8217;s business rely on the same foundation.</p><p>OpenAI might very well see its consumer product as the end-all-be-all that will replace other applications. We disagree, and so does likely everyone building AI applications. The question is whether OpenAI enables its enterprise customers to continue to rely on its API products or pushes them away with too much abstraction.</p>]]></content:encoded></item><item><title><![CDATA[The SaaS Extinction Test]]></title><description><![CDATA[Why your vibe-coded weekend project won't kill Salesforce]]></description><link>https://frontierai.substack.com/p/the-saas-extinction-test-f6e</link><guid isPermaLink="false">https://frontierai.substack.com/p/the-saas-extinction-test-f6e</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 06 Aug 2026 17:56:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KPzg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Apologies for the unplanned summer pause for the last couple weeks &#8212; Joey&#8217;s been on family vacation and Vikram had his first kid. We&#8217;re slowly getting back online.</em></p><p><em>With the news of Airtable&#8217;s acquisition this past week, analyses of the previous generation of SaaS companies have started flying around the internet. Our opinions haven&#8217;t changed dramatically, so we thought we&#8217;d bring back this post from 6 months ago.</em></p><p><em>Stay tuned &#8212; we&#8217;ll be ramping back to regular content in the next couple weeks!</em></p><div><hr></div><p>Projections about the demise of the SaaS industry have reached a fever pitch over the last few weeks. The general belief seems to be increasingly that there&#8217;s no moat in software anymore. Consequently, we have seen the stock prices of massive, entrenched incumbents take a significant beating.</p><p>The bear case for software-as-a-service goes something like this: As coding agents improve, the build-versus-buy calculus shifts permanently. An engineer at any company can now pick up a coding agent, build a prototype that meets the company&#8217;s specific needs, iterate on feedback, and deploy a bespoke tool that provides more value than a generic SaaS platform at a fraction of the cost.</p><p>The source of this concern is our collective recent experience with the improving quality of coding agents. We have all had Claude Code or Cursor scaffold an entire application from scratch for a few dollars in tokens. When you see something genuinely useful built in minutes, it is easy to assume the multi-billion-dollar incumbents are doomed. And if an individual can build a prototype quickly, surely a startup can build a strong offering with a few months and a few engineers?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KPzg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KPzg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1558029,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/190740985?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>As with most hype cycles, however, the truth lands in the middle. While there are massive opportunities for disruption, the SaaS business model is not going to evaporate overnight. In our view, the distinction comes down to a quick survival checklist. If you meet one of these criteria, you are likely safe. If you miss on all of them, you should be worried:</p><ol><li><p>Are you a system of record?</p></li><li><p>Do you do more than help humans automate a single workflow?</p></li><li><p>Are you mission-critical?</p></li></ol><h3><strong>The Physics of Data Gravity</strong></h3><p>Data moats have long been the holy grail of enterprise software. Consumer products like Google or Instagram win on user behavior data, and enterprise titans like Snowflake or Datadog are powerful because they are systems of record. These companies do not just power workflows; they house years of historical data that companies have imported and structured.</p><p>Despite all the massive technological shifts in the last few years, the physics of data have not changed. Moving massive amounts of data is expensive, risky, and slow. This means that, as has been the case for fifteen years, the major cloud providers will continue to make their margins on networking and egress. The cost of moving data in and out of a third-party service &#8212; or your own cloud &#8212; remains incredibly high.</p><p>Beyond the operational cost, there is the issue of operational stickiness. If a product is collecting data for a core operational purpose, it becomes a load-bearing wall in the company&#8217;s architecture. You do not pay Datadog or Snowflake eight figures a year because storing logs is a nice-to-have feature. You pay them because when a mission-critical issue happens at 3:00 AM, you need a proven way to investigate, pinpoint, and fix the issue.</p><p>But realistically, your business will continue to exist if Datadog goes down for a little while. Things get much more difficult when we&#8217;re talking about truly mission-critical software.</p><h3><strong>The Mission-Critical Tautology</strong></h3><p>Even when companies don&#8217;t have the most interesting data, they&#8217;re likely safe if they are critical to the operations of their customers &#8211; the kinds of products that your business literally wouldn&#8217;t exist without. The obvious examples are Workday and Salesforce.</p><p>There is a certain recursive logic here: These companies are safe because they are big, and they are big because they are safe. No matter how shiny a rapidly scaffolded payroll system looks, a VP of HR is not going to risk payroll failing on a Friday morning. A CRO does not care how many new automations a startup offers if it means moving away from a Salesforce instance they have spent a decade customizing to their exact sales and revops motions.</p><p>When enterprises make these decisions, they are not just buying software &#8211; they are also offloading risk. A tech-forward company like Google or Meta certainly has the technical talent to build an internal payroll system. The question is not whether they can build it, but whether they want to own the risk of running it. Once they do the calculus, dedicating hundreds of engineers to a non-core area of expertise rarely makes sense.</p><p>By contrast, non-mission-critical software is in the danger zone. Analytics is the prime example. Traditionally, writing code for dashboards and plumbing database queries was prohibitively expensive to democratize.</p><p>Just last week at RunLLM, we built an internal dashboarding system following a conversation about product metrics. We used Python and Plotly to connect to internal systems and pull customer data via our CRM&#8217;s API. If this dashboard goes down for a day, it is annoying, but it does not halt our operations. More importantly, because the cost of building it with a coding agent was so low, and the upside of allowing our VP of Sales to modify it himself was so high, the option to buy a traditional BI tool never even entered the conversation.</p><h3><strong>The Workflow Moat Narrows</strong></h3><p>Pure workflow software is where we&#8217;re currently the most bearish. By workflow software, we mean anything that&#8217;s focused on connecting the dots between existing tools that either have data gravity or mission-criticality. The poster child on X for this is PagerDuty. At its core, PagerDuty processes data stored in another product&#8217;s telemetry store, determines if a threshold was met, and alerts an on-call engineer. This is connecting the dots between Datadog and Slack. While PagerDuty does much more in practice, that primary workflow is what most customers buy.</p><p>These are exactly the kinds of integration code workflows being replaced by agents today. The threat here is not just that agents will automate incident management or sales outbound; it is that new players or internal teams can build superior user experiences that blend human and agent capabilities from the ground up. The most interesting part of this comes from the fact that the solutions can (and should) very easily be customized to match each team&#8217;s workflow. These agent-first systems will be dramatically more valuable because they save human hours, easily displacing legacy systems that are human-first and currently scrambling to bolt on AI features as an afterthought.</p><h3><strong>The Ops Gap</strong></h3><p><span>The undiscussed concern underlying all </span><em>build it with Claude Code</em><span> software is operations. The reality of software operations and reliability is unfortunately still incredibly complex and oftentimes dwarfs the time taken to build something useful. The gap between building and operating is actually growing because as coding becomes a commodity, the relative cost of infrastructure and security grows.</span></p><p>The internal dashboard we built recently is a perfect example of this. It took about an hour to iterate on the core dashboards themselves with Claude Code. It then took about four hours to deploy it in Google Cloud so that it was properly authenticated behind our company&#8217;s SSO, set up to auto-update with new commits, and connected with the right credentials.</p><p>This gap will likely persist. The simple reason why is that Python looks the same no matter where you work, but infrastructure is dramatically different at every company. There are not yet great ways to deploy cloud software without getting into the weeds of Kubernetes clusters, networking permissions, and security + compliance. None of those things matter when you are prototyping locally, but they are the only things that matter when you are operating production software at scale.</p><h3><strong>Wrapping Up</strong></h3><p>The rumors of the demise of modern software are greatly exaggerated, but the nature of what makes software valuable is quickly changing. For the last decade, SaaS companies could survive by being a slightly better UI for a human workflow. That&#8217;s no longer defensible.</p><p>The defensibility remains exactly where it has always been: at the intersection of data and risk. If you own the system of record, the physics of networking and the high cost of data egress will protect you. If you own a mission-critical operational process, the corporate aversion to risk will protect you.</p><p>The real danger is for the middle layer of the stack &#8212; the products that have built businesses around the friction of human work. As the cost of generating code and automating workflows drops toward zero, the value of that software must move elsewhere.</p><p>This does not mean incumbents are invincible. It just means that the disruption will not come from a simple internal tool or a weekend project built with a coding agent. To win in this new environment, startups must figure out how to best manage the data, reduce the operational risk, and close the gap between a prototype and production-grade reliability.</p>]]></content:encoded></item><item><title><![CDATA[Can you know if coding agents are worth the cost?]]></title><description><![CDATA[How to think about managing the economics of software production]]></description><link>https://frontierai.substack.com/p/can-you-know-if-coding-agents-are</link><guid isPermaLink="false">https://frontierai.substack.com/p/can-you-know-if-coding-agents-are</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 16 Jul 2026 18:45:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2iPJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>A recent theme in our conversations with VPs of Engineering has been a concern about (and perhaps a mild skepticism about) the cost of tokens being sunk into coding agents. A VP of a team of ~100 engineers said to us, &#8220;We&#8217;re spending about five figures a month on coding agents. I </span><em><span>think</span></em><span> it&#8217;s making the team more productive overall. I&#8217;m not sure, but I think.&#8221;</span></p><p><span>As concerns about coding agent overuse, budget overruns, and the ultimate impact on productivity mount, we&#8217;ve found ourselves wondering about the economics of coding agents. Coding agents can generate code faster than engineers working alone, but faster production is not the same as better economics. The fact that you can generate more code doesn&#8217;t mean that you should. To be clear, this isn&#8217;t a question of writing code by hand &#8211; all code will be generated by agents &#8211; but what tasks you pick, what models you use, and how you manage your agents.</span></p><p><span>To illustrate the point, let&#8217;s take a somewhat obvious example.  A group of four can drive from SF to LA in 6 hours and pay, perhaps $100 total for gasoline. That same group of people can fly from SF to LA in 1 hour and pay $200 each for airfare. Both are valid choices and depend on the urgency of the trip, what they&#8217;ll do once they arrive, and their price sensitivity.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2iPJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2iPJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2090876,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/207328475?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2iPJ!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cc62f6-155e-4896-939e-599b3d6ac1f7_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: GPT Image.</figcaption></figure></div><p><span>Engineering leaders are grappling with the fact that the economics of producing software, managing engineering time, and ultimately building products have changed dramatically. If you get your parameters wrong, you&#8217;ll be paying a lot of money to build the wrong thing very fast.</span></p><h2><span>The economics of software production</span></h2><p><span>If we rewind to the halcyon days of 2022, the way most engineering teams budgeted was by having a set headcount and allocating those cycles to whatever the highest priority projects were. Assuming the team delivered as expected, you could roughly say that you were spending $N a quarter to produce the software you shipped that quarter.</span></p><p><span>Today, you can spend something like $2-3N to create the same software in, perhaps, a week. While it&#8217;s hard to quantify, this is definitely a more efficient process &#8211; a cost increase of 2-3x and time improvement of about 10x is some pretty easy math for any project that you can look at.</span></p><p><span>For a real world example, consider the </span><a href="https://bun.com/blog/bun-in-rust"><span>recent port of Bun from Zig to Rust</span></a><span>. The blog post detailing the post says that the port took about 11 days and cost $165k at API pricing for work that would have taken a &#8220;small team of engineers a full year.&#8221;</span></p><p><span>This introduces a totally foreign dimension to software budgeting. To be clear, this is a huge </span><em><span>net improvement</span></em><span>. A small team, say engineers, working for a year will easily cost you $600k. The fact that you can do this task for ~25% the cost doesn&#8217;t mean that you </span><em><span>should</span></em><span> do it. If the task was done with the same number of tokens on Sonnet but required 2x as much human input, that would cost ~$50k of tokens combined with an extra 10 days of salary.  The reality probably isn&#8217;t that simple, but it&#8217;s illustrative of the new dimension of cost in producing software.</span></p><p><span>Until recently, we haven&#8217;t really had to ask this question about software production, but coding agents have completely changed the game. Engineers and engineering teams now have to grapple with questions about which model they should use, what level of speed and accuracy they need, traded off against a willingness to use a cheaper model that requires more guidance.</span></p><p><span>Ultimately, the economics of software production are going to increasingly have to be tied to the value of what&#8217;s being produced, not just the time it takes to produce it. Which brings us to the humans using the models&#8230;</span></p><h2><span>The economics of engineering time</span></h2><p><span>Measuring engineering productivity has historically been extremely hard. As an industry, we long ago ruled out silly metrics like lines of code, commits made, or tasks closed. In their place, we&#8217;ve mostly just agreed that the good engineers are the ones that everyone else on the team agrees are the good ones. DORA metrics can point at the function of an engineering team but don&#8217;t tell us much about individual engineering performance.</span></p><p><span>Recently, engineering leaders have started measuring engineers on token usage. </span><a href="/__u/frontierai.substack.com/p/ai-is-not-a-line-item"><span>This is unequivocally a bad idea</span></a><span>, but it highlighted a key concern around how you think about whether engineers are using their time well and delivering on the right things. As is now becoming common wisdom, what you build matters dramatically more in a world where you could theoretically build anything.</span></p><p><span>In practice, the 2x2 between good vs. bad productivity and low vs. high token use has </span><a href="https://herald.dev/blog/winter-is-coming-for-ai-engineering"><span>4 valid quadrants</span></a><span>. In other words, token use does not correlate positively or negatively with engineering productivity. Just because an engineer burns through a lot of tokens, that doesn&#8217;t tell you very much about </span><em><span>what&#8217;s getting done</span></em><span> &#8211; they might be mismanaging context windows, regularly patching issues created by poor prompting, or undoing and redoing a lot of work because of poor prompting. Just recently, an engineer on our team accidentally burned through ~$300 tokens on Claude very quickly by mistakenly pinning Fable. That engineer is generally very good but obviously didn&#8217;t suddenly get 10x better than everyone on that day. Simultaneously, we have plenty of examples of engineers who are very judicious about token use because they plan carefully, give precise instructions, and invoke the agent when they&#8217;re confident in which direction they want to head.</span></p><p><span>There&#8217;s no easy answer here. Measuring engineering productivity has genuinely gotten dramatically harder in the past 6 months because we&#8217;ve introduced a whole new dimension that simply didn&#8217;t exist before. You shouldn&#8217;t crown an engineer because they burned through a bunch of tokens, and you shouldn&#8217;t fire them either.</span></p><h2><span>Engineering decisions &amp; product decisions</span></h2><p><span>The technology is new, and these aren&#8217;t easy decisions to make. In the stylized example above, we suggested that doing the same task with Sonnet instead of Fable would have been possible with 2x as much human input &#8211; that might be true, or Sonnet might have very quickly gone off the rails and generated totally useless results. Perhaps $165k + 10 days of time is the cheapest possible implementation with the current state of the art. It&#8217;s impossible to know otherwise until you actually try. Unfortunately, trying is expensive.</span></p><p><span>While token costs dropped rapidly in 2023-24, the rate of change has decreased over the last year &#8211; and the rate of token consumption has skyrocketed. Managing this tradeoff has gotten harder as we&#8217;re now all balancing model choice with effort level and task scope. Given that it&#8217;s new, every engineer is going to find this challenging, but this will be a skill that engineering teams develop over time with practice.</span></p><p><span>In the meantime, you need to keep a close eye on your token use.</span></p>]]></content:encoded></item><item><title><![CDATA[Data is your only moat]]></title><description><![CDATA[How different adoption models drive better applications]]></description><link>https://frontierai.substack.com/p/data-is-your-only-moat-884</link><guid isPermaLink="false">https://frontierai.substack.com/p/data-is-your-only-moat-884</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 09 Jul 2026 18:33:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MSwO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>In the post-holiday catch-up, we didn&#8217;t have a chance to get this week&#8217;s post together, so we&#8217;re bringing back our favorite post so far from this year which breaks down how problem complexity and ease of adoption affect business growth. With new model releases flying at us in the past few weeks, understanding where on this 2x2 you fall is more important than ever. </em></p><div><hr></div><p>Theoretically, we should have a stellar AI agent for every problem in our lives by now. The talent is there, the capital is certainly there, and the models are increasingly capable. And yet, the results are lopsided. Why is it that we have agents that can prospect for sales leads and answer support tickets accurately, but we don&#8217;t seem to be able to consistently generate high quality slides?</p><p>The simplest explanation might be complexity. Easier problems (e.g., answer a support question) naturally get solved first, and more open-ended problems like slide generation require more effort. That doesn&#8217;t quite hold up: Coding is obviously not a simple application area, and yet coding agents are some of the best that we have today &#8211; in fact, they are improving faster than any other single agent use case.</p><p><span>How did this happen? Ease of adoption enabled data collection at scale that in turn helped coding agents improve rapidly.Every developer could switch to Cursor in 5 minutes without any approval. That created a data flywheel (more on this below) that allowed the Cursor team to build a better application experience over time &#8211; to the point where our whole team now swears by Cursor&#8217;s </span><a href="https://cursor.com/blog/composer">Composer model</a><span> for code generation.</span></p><p>The combination of technical complexity and adoption difficulty creates an interesting 2x2:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!MSwO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!MSwO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MSwO!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3d039b-c0af-4458-838c-27a2c172ebdb_1600x900.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>You might be tempted to think that being in one of the &#8220;easy to adopt&#8221; quadrants is the holy grail &#8211; after all, who doesn&#8217;t want more data to build better models? That is certainly a valid way to build a business, but the trap is that easy to adopt </span><em>also </em><span>means easy to displace. Hard to adopt products have their own data moat: Once you&#8217;re embedded in an enterprise, you learn about how </span><em>that company</em><span> works in a way that makes your product incredibly hard to replace.</span></p><p>Whichever quadrant you fall into, data is your only moat.</p><h2><strong>Easy to adopt, easy to solve</strong></h2><p><span>Easy to adopt and easy to solve is the most obvious quadrant to work in. It didn&#8217;t take an incredible amount of foresight to see back in 2023 that consumer search on Google would be replaced by custom answers to every question that a user has &#8211; whether it was finding a nice fact or providing healthcare advice. This has been the bread-and-butter use case for the foundation model providers and plenty of new entrants (e.g., Perplexity, </span><a href="http://you.com/">You.com</a><span>) flocked to these use cases as well.</span></p><p><span>The &#8220;easy to solve, easy to adopt&#8221; quadrant is a </span><strong>value trap</strong><span>. If the barrier to entry is low for you, it&#8217;s non-existent for frontier labs (or more likely, they&#8217;ve already built it). Given that these are the &#8220;obvious&#8221; use cases, they&#8217;re the ones for which the existing chat applications will see the highest volume of usage. That means that &#8211; whatever the use case is &#8211; OpenAI, Google, and Anthropic are gathering millions of data points to improve their models in these areas. Last week&#8217;s release of </span><a href="https://openai.com/index/introducing-chatgpt-health/">ChatGPT Health</a><span> feels like an obvious step in this direction. Beyond data access, the model providers can also subsidize costs and leverage their massive user bases to learn any new application area quite quickly. In short, you very likely will get crushed by the model providers.</span></p><p><span>An interesting side note is that loyalty is quite low in this quadrant &#8211; we all </span><a href="https://www.interconnects.ai/p/use-multiple-models">use multiple chat agents</a><span> depending on the use case, and unlike with the web search market, everyone seems to be on relatively equal footing. If dominant brand leaders do emerge, we&#8217;d place our bets on the model providers.</span></p><h2><strong>Easy to adopt, hard to solve</strong></h2><p>Why did coding &#8211; ostensibly one of the hardest problems to solve! &#8211; see such rapid progress? Most importantly, it is because adoption was easy &#8211; you could see a ton of value by pasting a snippet of code into ChatGPT back in 2023, and Cursor quickly made that much easier even though quality was limited early on. Since every engineer typically has the freedom to choose their own IDE, switching from IntellIJ or VSCode to Cursor wasn&#8217;t a crazy lift. Once it was in place, it also had a very fast feedback loop &#8211; a software engineer might generate code with Cursor tens or hundreds of times a day. That created a data flywheel: Every accepted or rejected suggestion adds to training data for future model improvements. With this data in hand, it was inevitable that model quality would improve dramatically over time. Notably, other markets in this quadrant (e.g., slide generation) that don&#8217;t have the same fine-grained feedback loop have seen much slower improvements.</p><p>Anything in the &#8220;hard to solve&#8221; category is going to require significant investment &#8211; across token usage, technical talent, and likely eventually model training and RL. The ease of adoption is a powerful data acquisition flywheel that enables that deeper investment. The frontier model labs seem to view these kinds of widely used productivity agents as being in their domain. They&#8217;re already competing heavily on coding agents, and we would not be surprised to see them launch more office-suite productivity tools beyond the document editors they already have. In other words, our prediction is that these markets are going to have heavyweight fights &#8211; smaller players will struggle to compete without huge capital outlays.</p><p>Stickiness, however, continues to be low here. Many of us run multiple coding agents, and as office productivity tools improve, there&#8217;s no reason that you wouldn&#8217;t jump to whichever app makes you the prettiest slides. The argument for stickiness is company-specific customization (e.g., Cursor rules, brand templates), but it&#8217;s possible we will see interoperability or a single standard emerge to enable migration.</p><h2><strong>Hard to adopt, easy to solve</strong></h2><p>This is the area where enterprise adoption of AI has really taken off in the last two years. When we say easy to solve, we&#8217;re not implying that there&#8217;s no product depth, but it&#8217;s easy to imagine how an LLM can execute a playbook for an e-commerce return or a password reset. Given that most enterprises are looking for wins from AI, the &#8220;obvious&#8221; problems are where they&#8217;ve turned for immediate adoption. That&#8217;s enabled an incredible pace of revenue growth for the leaders in these markets.</p><p>Two key things differentiate this quadrant. First, these products are not individually adoptable &#8211; buying an agent to handle support tickets or IT helpdesk requests is an organization-level decision that likely has a buying committee. Second, the comparative simplicity of the use case is offset by the difficult and tedious reality of enterprise integrations. The teams that can navigate legacy enterprise systems have a huge leg up.</p><p><span>That integration story is where there&#8217;s a data moat. While the data you get from these agents is less broadly applicable &#8211; and enterprises will likely restrict your ability to train models with it &#8211; you&#8217;re gathering data about how </span><em>each customer</em><span> works. Over time, that will help you make your product at large better, but most importantly, your product will become stickier for each customer. The next agent that comes along will have a hard time recreating that learned expertise.</span></p><p><span>In this area, investors are treating the larger startups as </span><em>de facto</em><span> incumbents. That&#8217;s not to say that there isn&#8217;t product innovation left to be done &#8211; there very likely is! &#8211; but it&#8217;s not immediately obvious why a smaller startup would be able to compete with the likes of Sierra and Decagon, for example. What&#8217;s less clear is whether the capital these companies are raising is primarily being used to drive GTM or whether there is a clear technical moat that&#8217;s emerging, &#224; la coding-specific models. If it&#8217;s only the former, then startups might have to resort to competing on cost.</span></p><h2><strong>Hard to adopt, hard to solve.</strong></h2><p><em>Example apps: SRE, security ops</em></p><p>Hard to adopt, hard to solve problems have received (comparatively) the least attention out of all four quadrants. The potential value of solving complex engineering or operations workflows can be incredibly high, as these are tasks that typically take humans hours or days. Unfortunately, these are workflows that are also fairly custom on a company-by-company basis, which means evaluation and implementation are much more cumbersome than &#8220;easy to solve, hard to adopt&#8221; products.</p><p><span>We&#8217;ve placed </span><a href="https://www.runllm.com/product">our bet</a><span> in the hard-hard quadrant, and this is where we expect to see the next phase of growth. The hard-hard markets will grow very quickly in the next couple years for a handful of reasons. First, reasoning models are now capable of planning to handle more complex tasks, which will help grapple with multi-step solutions. Second, a lot of the complexity in solving these problems comes from the steps outside of AI &#8211; building and configuring workflows; that will get easier and faster as coding agents get better. Finally, enterprises are already actively plucking the low-hanging fruit and will look to harder problems once those run out.</span></p><p>The data moat here is the most complex and potentially the most valuable. If you build expertise in one company&#8217;s workflows, that becomes very difficult to replicate &#8211; switching products would be akin to firing an experienced engineer and replacing them with a new person. There&#8217;s potentially an opportunity to build expertise in core capabilities (e.g., an SRE agent that&#8217;s an expert in AWS). However, this improvement cycle will be significantly slower than it was with coding agents because the quantity of data is lower and verifiability is less obvious.</p><p>While every one of these markets has a company that has raised astronomical amounts of money (often well ahead of revenue growth), we have a hard time imagining that these companies are as entrenched as their equivalents in the &#8220;easy to solve, hard to adopt&#8221; category. There&#8217;s a very long game left to be played in this market.</p><h2><strong>Wrapping up</strong></h2><p><span>This map isn&#8217;t set in stone; both boundaries will change. On the complexity front, we&#8217;ve seen dramatic improvements in model capabilities every few months. However, model improvements </span><a href="/__u/frontierai.substack.com/p/model-inference-model-products-and?utm_source=publication-search">seem to be plateauing</a><span>, so there&#8217;s less interest on this axis.</span></p><p><span>The real excitement is around UX. We&#8217;ve long believed that the UX aspects of AI applications are underexplored. We would not be surprised to see new UX paradigms developed that change the way users adopt products. </span><a href="https://code.claude.com/docs/en/claude-code-on-the-web">Claude Code on the web</a><span> is probably the best recent example of this &#8211; by making a coding agent available in a web browser available to everyone, it&#8217;s allowed users who might be scared away by an IDE or terminal to access these tools.</span></p><p>Regardless of what path they take, our bet is that the next 12-24 months will see the rise of winners in the hard-hard quadrant. It won&#8217;t look as seamless as the growth of Sierra and Decagon has been &#8211; there will be longer evaluation cycles, more complex implementations, and likely an overall lower success rate. But as companies improve their process and data enables improved models, this is where incredible amounts of revenue can be generated.</p>]]></content:encoded></item><item><title><![CDATA[Open models don't need to be OpenAI]]></title><description><![CDATA[Why smart enough, fast enough, and cheap enough is good enough]]></description><link>https://frontierai.substack.com/p/open-models-dont-need-to-be-openai</link><guid isPermaLink="false">https://frontierai.substack.com/p/open-models-dont-need-to-be-openai</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 25 Jun 2026 20:49:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X_FJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Today&#8217;s post is a re-post from May of 2024. 2+ years later, the narrative is still around whether open models will catch up to the frontier, but recent model releases like GLM 5.2 &#8212; now supported in Cursor! &#8212; show that that&#8217;s the wrong question. Open models don&#8217;t need to beat the frontier. Having models that are close to frontier intelligence is incredibly useful, especially as application builders start to have concerns about frontier model access with Mythos/Fable. While you&#8217;ll need to swap in some new model names since we published this post, the argument still holds.</em></p><p><em>We&#8217;ll be off next week for a summer break and back the week after! </em></p><div><hr></div><p><span>Last fall, we wrote that </span><a href="/__u/generatingconversation.substack.com/p/openai-is-too-cheap-to-beat">OpenAI is too cheap to beat</a><span>. To date, that&#8217;s still our most popular post with over 30k views on Substack. With a title like that, it generated the amount of strong opinions you&#8217;d generally expect &#8212; both agreeing and disagreeing with us. The general gist of that post is that the cost performance tradeoff that OpenAI was offering at the time was as close as to optimal as you were going to get.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!X_FJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 424w, /__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 848w, /__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!X_FJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:266044,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 424w, /__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 848w, /__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!X_FJ!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f5394f-fa88-4a65-bf3b-bd54d33da02d_1024x1024.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: DALL-E 3.</figcaption></figure></div><p>Looking at the market today &#8212; about 6 months later &#8212; we have two distinct observations:</p><ol><li><p><span>OpenAI still sells the best top-end models. Claude 3 Opus is the only model that&#8217;s achieved comparable </span><a href="https://leaderboard.lmsys.org/">Elo</a><a href="/__u/frontierai.substack.com/p/open-llms-dont-need-to-beat-openai?utm_source=publication-search#footnote-1"><sup><span>1</span></sup></a><span>, and it&#8217;s about 3x as expensive as GPT-4 as of this writing. In that top class of model, GPT-4 is still your best cost performance tradeoff.</span></p></li><li><p><span>The </span><em>gap</em><span> between the top tier and second tier of models has shrunk dramatically. That tier is still primarily comprised of proprietary models (Sonnet, Haiku, and older versions of GPT-4), but critically, Llama 3 has squarely made its way into that second tier.</span></p></li></ol><p>Another post we wrote last fall was about how open-source LLMs shouldn&#8217;t try to win; instead, we argued, they should serve as the bases for efficient fine-tunes for task-specific experts. There, we argued that these open models should get smaller and faster at their existing quality rather than getting bigger and better to try to compete with GPT-4.</p><p><span>Looking back, we got some things very right, and we got some things </span><em>very wrong</em><span>. At a high-level, the direction open LLMs are headed is incredibly promising. Looking at a breakdown of the LMSys Elo for the top models, you can see that the gaps have started to close.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!R45E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 424w, /__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 848w, /__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 1272w, /__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!R45E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png" width="1456" height="644" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:644,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:459681,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 424w, /__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 848w, /__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 1272w, /__u/substackcdn.com/image/fetch/$s_!R45E!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda37b080-7462-4bc8-a09e-858a9a37cf97_3470x1534.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Elo scores with error bars from the LMSys Chatbot Arena.</figcaption></figure></div><p>Here&#8217;s where we think things stand today:</p><ul><li><p><strong>Open LLMs are </strong><em><strong>very, very</strong></em><strong> good.</strong><span> Even just 3 months ago, we experimented with Mixtral 8x7B and Llama 2 and discarded both because their results weren&#8217;t good enough. Today, the quality we&#8217;re able to achieve with Llama 3 and Mixtral 8x22B, given their parameter size, is quite impressive. We&#8217;ve been increasing our experimentation with them recently and are considering replacing parts of our production inference stack with Llama 3.</span></p></li><li><p><strong>Open LLMs don&#8217;t need to get that much smaller.</strong><span> What we were really arguing for last fall was that open LLMs needed to get more efficient at the inference step. What we didn&#8217;t account for was that the inference process itself would get dramatically more efficient in the last 6 months &#8212; which is embarrassing given that we&#8217;re systems people! Today, you can get at least 10 tokens/s on a MacBook for Llama 3 and production-level inference on cloud deployments.</span></p></li><li><p><strong>This makes fine-tuning open LLMs incredibly attractive.</strong><span> With the increasing quality of models and the plummeting costs (plus </span><a href="https://twitter.com/Suhail/status/1787314445682889184">the fact that GPU availability is increasing quickly</a><span>), fine-tuning open models becomes very attractive once more. There are obviously some details to work out across different model providers and the features they support. But as a team that relies on fine-tuned LLMs as a core part of our product, we can&#8217;t wait to find time to experiment with replacing fine-tuned GPT-3.5 with fine-tuned Llama 3. Our bet is that it will be higher quality, faster, and cheaper.</span></p></li><li><p><strong>But open models still won&#8217;t catch OpenAI.</strong><span> All that said, we still believe that open LLMs aren&#8217;t going to catch OpenAI. The scale advantage the proprietary model builders have &#8212; combined with the positive consumer feedback loop &#8212; is daunting. If Meta can&#8217;t get close with all its resources, it&#8217;s unlikely anyone else will &#8212; but that&#8217;s okay! Open LLMs don&#8217;t need to be the best models around to survive.</span></p></li></ul><p><span>We&#8217;ve gone back-and-forth many times over the last year about whether RAG or fine-tuning will win. The answer today squarely seem to be &#8220;both&#8221; &#8212; but the path forward for efficient, scalable fine-tuning was always murky</span><a href="/__u/frontierai.substack.com/p/open-llms-dont-need-to-beat-openai?utm_source=publication-search#footnote-2"><sup><span>2</span></sup></a><span>. That&#8217;s started to change in the last month, and that&#8217;s incredibly exciting. As we work out the details in fine-tuning Llama 3, we&#8217;ll very reasonably start to see a proliferation of these narrow, expert LLMs out in the world.</span></p>]]></content:encoded></item><item><title><![CDATA[AI changed distribution. Can you keep up?]]></title><description><![CDATA[Why every company needs to show value now &#8212; or lose]]></description><link>https://frontierai.substack.com/p/ai-changed-distribution-can-you-keep</link><guid isPermaLink="false">https://frontierai.substack.com/p/ai-changed-distribution-can-you-keep</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 18 Jun 2026 19:15:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UgZK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>There&#8217;s a temptation to try to make agents as impressive as possible up front: Give them all the data, do all the configuration, then brag about how much value you can add. That&#8217;s the wrong approach: You need to show people what your product can do immediately &#8211; even if the functionality is limited. If you don&#8217;t, you risk losing to someone who does, whether it&#8217;s an enterprising engineer building in-house or another vendor.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!UgZK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!UgZK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:966569,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/202627129?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!UgZK!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F321b328b-c12e-4fa8-adde-006b76ae7621_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>At Herald (recently rebranded from RunLLM, in case you haven&#8217;t seen), we&#8217;ve flipped our distribution model in the past couple months from a top-down, enterprise-oriented approach to a bottoms-up approach that enables developers to try our product quickly and with minimal change. To do that, we&#8217;ve rolled out a </span><a href="http://herald.dev/product"><span>new CLI</span></a><span> that&#8217;s up and running on your laptop in minutes &#8211; like Claude Code or Cursor &#8211; without requiring you to upload any credentials to the cloud.</span></p><p><span>We&#8217;d obviously love for you to </span><a href="https://herald.dev/try"><span>install our CLI</span></a><span>, try it out, and share any feedback that you have &#8211; but we also think that what we&#8217;ve learned reflects something fundamental about distribution of AI products.</span></p><h2><span>What we learned</span></h2><p><span>It&#8217;s worth starting with why we were previously in a different mode of distribution. Our product &#8211; an AI DevOps Agent &#8211; understands a team&#8217;s product, infrastructure, and business and helps engineers answer questions, debug issues, and discover incidents before customers complain. By its nature, this kind of agent requires access to a lot of different data: code, telemetry, infrastructure, CI/CD, documentation, and so on. To show off the full value of the agent, we encouraged our customers to connect as much data as possible &#8211;  but for understandable reasons, engineering teams won&#8217;t blindly connect all of those data sources to a new agent without the appropriate security and legal review. The process we set up was to engage with a company&#8217;s leadership, get buy-in, go through legal &amp; security review, get access to all the data, then show teams what we could do.</span></p><p><span>That works to an extent, but it leaves a lot of opportunity on the table. While getting the necessary approvals, we found that a number of teams &#8211; especially the most AI-enabled ones &#8211; would quickly move on to other priorities or go with an in-house solution. We could talk about why DIY solutions are worse until we&#8217;re blue in the face &#8211; most vendors are right now! &#8211; but the fact that teams can make themselves successful more quickly than ever is incontrovertible.</span></p><p><span>That&#8217;s what led us to conclude that we had to show engineers value immediately. Once they understood why we were better than what they could build in-house &#8211; even if with limited access &#8211; then we could justify the time and effort required to get comprehensive access. Of course, to do that, we had to address security concerns too. Long story short, we built a version of our agent that has a familiar form-factor (Claude Code-style CLI) but, critically, keeps all of your credentials local. An engineer can now install the agent with </span><span data-color="rgb(24, 128, 56)" style="color: rgb(24, 128, 56);">npm install -g @herald-ai/herald</span><span>, point at the tools that they already use locally (e.g., AWS CLI, Grafana MCP, etc.), and see what Herald can do immediately.</span></p><h2><span>The implications</span></h2><p><span>None of this is to say that everyone should go build a CLI. We&#8217;re working with developers, so we set out to build something that was familiar to them, but your target audience, their ideal UX, and what value you can show in a limited capacity might very well be different. In many ways, building a CLI is actually harder than traditional product-led adoption because you lose the ability to instrument many of the adoption funnel signals that you can capture on a web app.</span></p><p><span>The traditional wisdom for a lot of developer-led adoption is that you want to show people the full value of the product up front rather than gating features. While that&#8217;s true, agents are all about data access, and an individual user or small team is never going to be able to provide full data access. Your product should be designed for that limitation. It&#8217;s reasonable if your product would be </span><em><span>better</span></em><span> with more data, but it needs to be good with limited data.</span></p><p><span>There&#8217;s unfortunately no one-size-fits-all solution to operating with limited data. In our case, we&#8217;ve customized an implementation of our predictive issue detection feature to each data source that a user might configure. The full implementation works better when there&#8217;s multiple data sources to correlate across &#8211; but the explicit compromise made it easier for us to show value to the user immediately. What compromise you can make to show value immediately is the question that should drive your implementation.</span></p><p><span>Given that we&#8217;re talking about an adoption strategy that puts product front and center, you might be tempted to think that we&#8217;re saying every agent should adopt a PLG motion. Not quite. While giving people a taste of what your agent can do up front matters, you very well might still need humans in the loop to solidify and grow usage. For our agent to truly be successful in an enterprise, we still need to go through security approvals and a team-wide deployment. Predictive issue detection will still be most valuable when the IT team connects all the relevant data sources. The difference is that we can now operate in both modes &#8211; limited data to show value quickly, and comprehensive data to show the full value.</span></p><h2><span>Is this universal?</span></h2><p><span>Our initial thought was that this was unique to developers &#8211; a lot of the people we talk to are Claude Code zealots, after all. The more we think about it and look around, the more we realize that this is a challenge that every software vendor is going to face in the next few years.</span></p><p><span>To be clear, there are exceptions. Companies like Sierra and Decagon seem to be growing at a breakneck pace with large enterprise contracts, and security products have always been bought &amp; sold differently than other software infrastructure. It&#8217;s worth noting that these </span><a href="/__u/frontierai.substack.com/p/data-is-your-only-moat?utm_source=publication-search"><span>hard-to-adopt products</span></a><span> have a unique data advantage that comes from embedding, and we&#8217;re explicitly advocating for a transition from hard- to easy-to-adopt &#8211; or at least a dual approach. That&#8217;s a topic for another post.</span></p><p><span>But again, these exceptions are noteworthy because plenty of software that you might expect to be purchased top-down &#8211; inference infrastructure, voice agents, sales tooling &#8211; is seeing a pattern where a usage-first adoption is driving growth. You should build for a world where every vendor has to show value now.</span></p>]]></content:encoded></item><item><title><![CDATA[Budgeting for AI isn't hard]]></title><description><![CDATA[The cloud already taught us how &#8212; you just need visibility]]></description><link>https://frontierai.substack.com/p/budgeting-for-ai-isnt-hard</link><guid isPermaLink="false">https://frontierai.substack.com/p/budgeting-for-ai-isnt-hard</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 11 Jun 2026 18:05:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NKuQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHKDx14kX0AAtKEI.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A couple days after publishing last week&#8217;s post &#8211; <a href="/__u/frontierai.substack.com/p/ai-is-not-a-line-item">AI is not a line item</a> &#8211; we saw this fascinating tweet from Eric Glyman (CEO of Ramp) making effectively the same point.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/eglyman/status/2062921352613425446&quot;,&quot;full_text&quot;:&quot;As I wrote this, I saw X go into meltdown over tokens.\n\nYou've seen the headlines: &#8220;Uber blows yearly AI budget in just one quarter.&#8221; &#8220;Meta employee burns 281 billion tokens in April.&#8221;\n\nBut, the problem isn't spending. Spending works. Since 2023, the top quartile of our AI&quot;,&quot;username&quot;:&quot;eglyman&quot;,&quot;name&quot;:&quot;Eric Glyman&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1911178684322439168/I2bOxs5p_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-05T15:35:52.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HKDx14kX0AAtKEI.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/aUn4S7AQJZ&quot;}],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;Today, Ramp raised $750M at a $44B valuation.\n\nLast time we grew this fast, we were 1/20th the size.\n\nFor 2000 years, business was built on two pillars. Today, a third: intelligence.\n\nIt&#8217;s your least governed cost. It&#8217;s also your single greatest opportunity.&quot;,&quot;username&quot;:&quot;eglyman&quot;,&quot;name&quot;:&quot;Eric Glyman&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1911178684322439168/I2bOxs5p_normal.jpg&quot;},&quot;reply_count&quot;:69,&quot;retweet_count&quot;:73,&quot;like_count&quot;:837,&quot;impression_count&quot;:361916,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>This line in particular stuck out to us:</p><p>Finance says, &#8220;half the budget,&#8221; engineering says, &#8220;double it&#8221; and you don&#8217;t know who&#8217;s right because there is no shared language of value. There&#8217;s no attribution, and no attribution means no allocation.</p><p>This got us thinking about what it actually means to allocate budget in AI and how enterprises should be thinking about it. What&#8217;s interesting about this challenge is that there has rarely been a truly universal resource in the enterprise. Marketing never asked for budget to run cloud infrastructure, and even if engineering built an internal tool for marketing, that was typically still considered engineering&#8217;s responsibility.</p><p>Platform plays are the closest we&#8217;ve historically gotten &#8211; plenty of companies will centralize on Salesforce, but from an invoicing perspective, you could very clearly break apart seat costs for Sales Cloud, Marketing Cloud, and Service Cloud. Now, you get a single invoice at the end of the month that shows a large number of tokens used. Who used them, to what end, and was there an ROI? Who the hell knows!</p><p>We have genuine sympathy for finance teams who are facing these challenges for the first time &#8211; and for the engineering leaders who are getting their &#8220;AI budgets&#8221; blown up because, well, everything is AI now. And that&#8217;s the most interesting thing to realize: <em>Everything is AI now</em>. There can&#8217;t be a single line item for AI because AI is in every part of your budget. So what do you do instead?</p><p>The funny thing is that it&#8217;s not all that interesting. AI budgeting is a solved problem &#8211; the industry collectively learned how to account for variable costs from cloud infrastructure, and budgeting for AI spend is the same. What needs to catch up is the infrastructure to make this happen.</p><h2>Don&#8217;t confuse applications and intelligence</h2><p>This might sound trivially obvious, but it seems to get muddled more often than we might think. It&#8217;s the most obvious place to start. The X discourse is all about tokens for obvious reasons: Software is no longer zero-marginal cost, and that means that your bills genuinely are climbing faster than you&#8217;d realize with the use of AI. But your large monthly Anthropic invoice is different from an agent that&#8217;s built for a particular application area. Spend on verticalized agents should obviously go into the budget for the specific organization that is using that agent, and the team that&#8217;s buying that agent &#8211; just like with any other software! &#8211; should be precise and empirical with evaluating what the potential ROI is. You&#8217;d be shocked how often we talk to customers who seem to have given almost no thought to how they would measure the success of the agent. If there isn&#8217;t an ROI, there isn&#8217;t going to be a budget.</p><p>The interesting thing here is that application spend is getting harder to predict. Most agents are adopting token-based billing (or some abstraction on top of tokens) with some occasional seat-based access layered on top. For most of us, the token cost dominates the seat cost by at least an order of magnitude. How do you budget when your spend can vary so much? Again, history gives us clear answers. Lest we forget, this was the exact concern with (and criticism of) serverless computing ~8-10 years ago: There was no way for enterprises to know how much a workload would cost. But companies found a way around this challenge, allocating a bucket of spend to cloud providers and then burning down that spend through their own usage or through marketplace purchases. The same model will very likely emerge for AI spend.</p><p>Budgeting for &#8220;raw&#8221; intelligence is a much harder problem, and that&#8217;s where the infrastructure needs to catch up.</p><h2>Engineering discipline is enterprise discipline</h2><p>Engineering teams have historically had to develop good abstractions for resource management, and the rest of the enterprise is going to have to catch up. Developers don&#8217;t run tests on the production cluster, and the staging and production databases have to be isolated. That makes it dramatically easier for a finance team to understand COGS, development spend, and potential sources of waste.</p><p>That same level of visibility is going to need to be adopted to meter AI spend accordingly. Today, tools like Claude and Cursor give you visibility into how many tokens each user used, but that&#8217;s just the starting point. You have no idea if those tokens were spent on useful work, new experiments, or totally useless things. On one hand, you don&#8217;t want to be metering so closely that you discourage use &#8211; it&#8217;s fine if someone spends a couple bucks looking up where to eat lunch today &#8211; but visibility matters.</p><p>This is where products like <a href="https://entire.io/">Entire</a> are particularly interesting to us. Creating a clear audit trail to help teams understand what work was done, which agents were responsible for it, and how that ties back to the ultimate goals that the company sets is a critical component in understanding what AI spend is actually worth &#8211; and where the waste is happening. While it&#8217;s very early, we have a hunch that spend management and cost optimization might end up being an underrated application area for tools like Entire.</p><p>The interesting question, though, is how that same discipline extends into other areas. Again, engineering by its nature is structured. Coding agent spend can be tracked as a part of a git commit. Some areas &#8211; e.g., customer support &#8211; have natural analogs, and you can easily imagine tracking how much was spent in resolving a particular ticket. Others like marketing and sales are dramatically less amenable to precise tracking in their current structure. While we&#8217;re not experts in exactly how you should empiricize those functional areas, we think it&#8217;s somewhat inevitable that good visibility and structure on token spend will become integral to how those areas are run.</p><h2>Get with the times</h2><p>None of this is going to be tidy in the interim. Teams will experiment, budgets will be disorganized, and spend will run over what you expect. That&#8217;s fine &#8211; that&#8217;s what the early years of cloud spend looked like too! The mistake is responding to that fuzziness by painting with a broad brush. You absolutely should not be sticking all of your AI spend into one budget and saddling the engineering team with it. That might feel like an okay short-term fix, but it&#8217;ll hold your whole company back.</p><p>None of this requires inventing a new discipline. The accounting model already exists and can be copied from cloud infrastructure. What&#8217;s missing is the visibily layer, and that&#8217;s being built right now. Eric Glyman is right that no attribution means no allocation, but attribution is an engineering problem &#8211; and engineering problems get solved incredibly fast nowadays. The finance teams that handle this transition well will be the ones who recognize that AI should be treated like the next phase of infrastructure.</p>]]></content:encoded></item><item><title><![CDATA[AI is not a line item]]></title><description><![CDATA[One number for all your AI spend is a recipe for disaster]]></description><link>https://frontierai.substack.com/p/ai-is-not-a-line-item</link><guid isPermaLink="false">https://frontierai.substack.com/p/ai-is-not-a-line-item</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 04 Jun 2026 17:36:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0Jaf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As soon as we started hearing stories of companies maintaining internal leaderboards for token counts, a backlash to AI spend felt inevitable. Just a few months later, it&#8217;s arrived.</p><p>Headlines over the last few weeks have been filled with discussion of companies like Uber dramatically <a href="https://www.forbes.com/sites/janakirammsv/2026/05/17/uber-burns-its-2026-ai-budget-in-four-months-on-claude-code/">running over their AI token budgets for the year</a>, unexpected <a href="https://x.com/Polymarket/status/2060034216906068131">Claude overages</a>, and <a href="https://www.businessinsider.com/uber-coo-andrew-macdonald-ai-token-spending-harder-justify-2026-5">whether you can measure the impact from coding agents</a>. And of course, there&#8217;s <a href="https://x.com/simonw/status/2060209010486493500">the backlash to the backlash</a>. We&#8217;ve heard stories ourselves about CTOs and VPs of Engineering freezing budgets and instituting extreme scrutiny on all new AI spend.</p><p>Token leaderboards are no more intelligent than ranking engineers based on lines of code generated. Even if you assume that everyone was behaving with the best of intentions, using more tokens doesn&#8217;t mean you got more done. You&#8217;re measuring the wrong thing and allocating your budget the wrong way.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0Jaf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0Jaf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9024489,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/200647895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0Jaf!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34870e05-0e03-4926-84d6-06fa716fa8af_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>What happens next is somewhat predictable. Enterprises will pull back on token spend, institute per-employee limits, and have teams submit budget requests for token allocations. Budget freezes will be widespread practice by the end of the year. From our perspective, this is completely the wrong way to approach AI spend. You shouldn&#8217;t be turning up spend to brag about how big the number is, and you shouldn&#8217;t be limiting AI use out of a fear of overages <em>because you were previously tokenmaxxing</em>. In fact, thinking about your AI spend as one number is the wrong way to look at it.</p><p>Different agents &#8211; and different experiments with new tools &#8211; are going to have different impact on your business. It makes no more sense to have one number for all of your AI spend than it does to have one number for all of your salary spend. Instead, you should think about each AI tool as impacting the team or the functionality buying it. When put in that frame, the rational rules for AI spend start to look pretty different. We have a few rules that we&#8217;ve started to solidify both from our own experiences and from talking to hundreds of engineering organizations.</p><h2>Limits with limits</h2><p>Enterprises need budgets, and we&#8217;re not advocating for unrestrained use or $500M Claude bills being the new norm. But you need to set limits on how you set limits. Going from no constraints on usage and massive bills to extreme scrutiny and justification for every expense both creates organizational whiplash (how often do I have to change my habits) and also discourages experimentation. Different uses of AI &#8211; proven value adds like coding agents vs. experimental new tools &#8211; should explicitly be treated differently. If you don&#8217;t have a budget for experimenting with new tools, you&#8217;re going to fall behind.</p><p>Discouraging experimentation is probably the biggest and most avoidable own-goal. While the pace of change is very high, there&#8217;s still tons of undiscovered or poorly understood application areas, and aggressively setting limits on budgets means that your team isn&#8217;t going to learn about what&#8217;s possible. We&#8217;ve already heard engineering teams share that they&#8217;re not being allowed to spend on any new AI tools even while they believe our solution would solve a real problem for them &#8211; and would cost 1-2 orders of magnitude less than what they&#8217;re spending on salaries solving the problem (or on OpenAI or Anthropic for that matter).</p><p>This also requires reframing your thinking on software vs. headcount budget. Traditionally, software budgets and salary budgets have been treated differently, but we&#8217;ve seen cases where teams have open headcount for which they can&#8217;t find high-quality talent, while they&#8217;re blocked from spending more money on software. Agents won&#8217;t replace humans 1-for-1, but they will defray the tedious work that allows people to focus on what they&#8217;re best at.</p><p>When you limit experimentation, you&#8217;re embracing a static worldview. You&#8217;re limiting your team&#8217;s ability to adapt to the latest technology, and relying only on headcount is the hardest way to solve capacity problems. You certainly don&#8217;t want to be blowing out your budgets, but giving your competitors an advantage by sitting on your hands and pretending the technology around you isn&#8217;t changing doesn&#8217;t work either.</p><h2>Build vs. buy: The calculus (isn&#8217;t all that) different</h2><p>There&#8217;s plenty of ink being spilled on whether <a href="https://x.com/buccocapital/status/2059953927840305459">you should build all your software moving forward</a>. We&#8217;ve talked about the <a href="/__u/frontierai.substack.com/p/enterprise-ais-teenage-years?utm_source=publication-search">build vs. buy calculus before</a>, so we won&#8217;t repeat the arguments in full detail, but the build vs. buy decision plays into the budget discussion. Building a product in-house is more likely to lead to budget overruns than buying something off the shelf. Humans are generally terrible at sunk costs, so once you invest into building a prototype, it&#8217;s natural to double down on that effort &#8211; but that&#8217;s where your budget overrun is most likely to get worse.</p><p>It&#8217;s a tired trope at this point that coding agents make it mind-numbingly easy to build a good-enough demo. It takes minutes and costs pennies. But the iteration and time that&#8217;s required to take that demo and turn it into a useful product is not quite so simple &#8211; or cheap. You&#8217;re both going to spend valuable time on that productionization process and you&#8217;re going to spend tons of tokens fixing all the bugs and edge cases that you didn&#8217;t anticipate in the demo that you built in just a few minutes. More importantly, when a coding agent writes most of the code, no one&#8217;s going to know what happened when it breaks. That means Claude is going to burn tons of tokens figuring out the issue.</p><p>This is where the budget argument and ROI calculus get particularly difficult. Is this engineering expense or operational cost? How much of this is expected, and how much will stability increase with time? How do you account for the time humans are spending on this problem as opposed to others? Properly estimating complexity and allocating budget for a piece of software that&#8217;s not your key expertise is inevitably going to be noisy and inaccurate &#8211; and that inaccuracy will spill over to affect the rest of your &#8220;AI budget.&#8221; The token cost will look tiny in the build phase &#8211; but when you get to productionization, maintenance, and new features, the token and salary math get very hard to justify.</p><h2>Prioritization is more important than ever</h2><p>We&#8217;re noticing that engineering teams are increasingly distracted &#8211; we shared <a href="/__u/frontierai.substack.com/p/enterprise-ais-teenage-years">an anecdote</a> a couple weeks ago about an engineering team that was jumping between possible solutions. To some extent, this is understandable because engineering teams really can do anything right now. But running tons of experiments that end up resulting in very little (especially when there are off-the-shelf solutions available) is an expensive distraction. Once again, this runs up token costs &#8211; and the budget experiments that have dubious value shouldn&#8217;t be conflated with real systems that can help your team.</p><p>This is perhaps a more subtle challenge at first glance. Each potential experiment that you run that isn&#8217;t well thought out can go from a $5 starting point to a $1,000 sinkhole very quickly. We&#8217;ve had this happen ourselves. While we&#8217;re not advocating for shutting down all agent use for the purposes of keeping a lid on budgets, knowing what you&#8217;re trying to accomplish &#8211; and how that ties into your token use &#8211; is critical. Giving your team room to experiment is valuable, but that&#8217;s not the same as something that&#8217;s directly adding value to the business. Mixing the two up leads to bad decision making about &#8220;AI.&#8221;</p><h2>AI is not a line item</h2><p>The question underlying this whole conversation is what the budget for AI tools should be, which is (literally) a million-dollar question. We&#8217;ve seen a pretty wide range of approaches &#8211; no limits at all (quickly becoming a thing of the past), everything goes to engineering, or an AI budget per-organization. But in all cases, AI is being treated as a line item on the budget &#8211; the wrong approach.</p><p>This isn&#8217;t a question about whether AI is here to stay, but about how we think about the value AI is adding. The value AI is adding to shipping product features that directly affect your top line is very different from the value it&#8217;s adding when you&#8217;re experimenting with recreating a third-party tool &#8211; or from the value it adds in writing cold email copy. Treating that all as &#8220;AI spend&#8221; is misguided. Different agents are going to add different amounts of value and in different ways.</p><p>That also doesn&#8217;t mean that you should treat every AI agent as one that&#8217;s going to have immediately measurable ROI &#8211; you should absolutely have budget allocated to experiment with things that <em>might</em> add value (or might not). How large that number is and how strictly it&#8217;s metered probably varies from business to business. Either way, putting all AI spend in one bucket and tokenmaxxing that number up or hurriedly rushing it down is a recipe for disaster.</p>]]></content:encoded></item><item><title><![CDATA[If you don’t know what good is, AI won’t tell you]]></title><description><![CDATA[Why good agents require good leadership]]></description><link>https://frontierai.substack.com/p/if-you-dont-know-what-good-is-ai</link><guid isPermaLink="false">https://frontierai.substack.com/p/if-you-dont-know-what-good-is-ai</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 28 May 2026 18:38:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BXH-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last week, we were on a call with a VP of Engineering at a large e-commerce company. We&#8217;d had 4 meetings with the engineering teams that run reliability for him, and we had a pretty detailed picture about what systems they used, where their data lived, and how the org wanted to tackle some of their key challenges. As many larger organizations tend to do, this team was operating highly manually &#8211; getting notified about new tickets for incidents and linking them to manually created channels for discussion and triage. We were caught off guard, then, when the leader claimed that a lot of the functionality we were describing was something the team already had built in house. What he was really looking for was apparently an agent that would automate end-to-end resolution without engineering involvement.</p><p>We later confirmed with the team that we weren&#8217;t crazy. We knew more about the state of his team&#8217;s operations than he did.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BXH-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BXH-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1043352,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/199639390?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BXH-!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8c34b41-ef66-447c-bd09-24d8c4d16ecd_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>This is not particularly new for anyone who&#8217;s worked at or sold to big companies before. There are plenty of leaders who aren&#8217;t connected to what their team is actually doing &#8211; enterprises have long succeeded and grown despite this disconnect. Unfortunately, the perils of this disconnect have gotten orders of magnitude more pronounced with AI.</p><p>Leaders who don&#8217;t understand their organizations are going to prioritize initiatives and buy agents that increase disorganization without driving results. Here&#8217;s why: Agents drive real value when they conform to enterprises&#8217; workflows, which means you need a clear understanding of what <em>your team</em> does today and what&#8217;s needed to improve. Without that understanding, leaders will thrash between possible solutions, and in the worst case, they will end up purchasing solutions that amplify existing issues. That means agents should be implemented where there&#8217;s a clear picture of the current state of play as well as a deep understanding of what good looks like. If you don&#8217;t know what good is, an LLM isn&#8217;t going to tell you.</p><h2>Focus still matters</h2><p>Historically, engineering teams have had to focus aggressively before resources were scarce. There&#8217;s always been a million things that <em>could</em> be worked on &#8211; niche features, tech debt, performance optimization, internal tools, etc. &#8211; but cycles to only work on a few of them. To some extent, coding agents have alleviated that pressure, and teams can now spin up new tools (and certainly flashy demos!) faster than before.</p><p>What happens after that still requires strong focus. Every prototype can&#8217;t be taken to production, because the process of refining and maintaining that prototype paradoxically has a much higher opportunity cost than before. It is definitely true that engineering teams are faster than before, but that means time spent on productionizing something that could be bought would be better spent improving core functionality that&#8217;s in the business&#8217; wheelhouse.</p><p>This is where the disconnect happened in our conversation last week: The VP was anchored on a demo he&#8217;d been shown that relied on hard-coded workflows and runbooks and wouldn&#8217;t generalize to the complexity of their stack. On top of that, he also wasn&#8217;t aware about the manual workflows the team currently followed, which would make the demo even more difficult to generalize.</p><p>To be clear, the VP wasn&#8217;t gratuitously trying to tear us down or being blithely cynical. He was operating the way that senior leaders at large enterprises have typically operated for decades. But with coding agents and flashy demos, it&#8217;s easier than ever to convince yourself that the world looks dramatically different than it does in reality.</p><p>What happens when you&#8217;re misled? You make decisions about what you think needs to be done rather than what actually needs to be done. We started a conversation last summer with a mid-sized tech company that wanted to use an AI SRE to deal with thousands of PagerDuty alerts a week. They evaluated us and one other vendor, but by the end of the year hadn&#8217;t yet come to a decision because the CEO suggested they do two <em>more</em> evaluations. Those evaluations last until the end of Q1, and by that point, an engineer had built an in-house solution the CEO got excited about because it dogfooded their internal product. Eight months of evaluation time was wasted, and now there are two engineers on the hook for maintaining a side project.</p><p>In each of these cases, an unrealistic ideal became a blocker for a good solution.</p><h2>The LLM has no clothes</h2><p>A lack of clarity means that your agent implementations lack a clear target for what&#8217;s actually being improved. In many cases, you can very much accelerate speed with agents &#8211; but speed without quality mostly just creates noise.</p><p>Take coding as an example. Adopting coding agents certainly means that you&#8217;ll have cleaner syntax and better comments. Your engineers might even prompt the agent to write more tests! That&#8217;s not going to make <em>better software</em> in general. Coding agents aren&#8217;t going to magically give you high-quality design, thoughtful system architecture, or a culture that prioritizes catching bugs before they ship. If your team is in the habit of not ensuring their implementations closely match product requirements and validating that edge cases are properly handled, that problem is likely going to be amplified with AI.</p><p>One of the first POCs we did in the AI SRE space was with a traditional software vendor &#8211; the POC ended up failing. The organization was excited about AI as a strategic initiative for them, but their stack simply was not set up for integration with agents &#8211; their existing debugging workflows were full of manual processes that required humans to use credentials with root access to access individual deployments and pull the relevant data. They were understandably hesitant about giving an agent this kind of information, but that meant the agent was operating with limited context &#8211; and ultimately, its recommendations were useless. Their processes were immature, and the agent created more noise.</p><p>The lesson is simple: Know what good looks like before you adopt an agent. If not, you&#8217;ll be impressed by the shiny new clothes that turn out not to be real.</p><h2>Adapting to reality</h2><p>This changes the way that software is evaluated and purchased in a few key ways. Teams building agents must be flexible. As we talked about <a href="/__u/frontierai.substack.com/p/enterprise-ais-teenage-years">last week</a>, every AI startup needs to be oriented towards showing value quickly and orienting pricing and commercial terms to quick adoption. If you&#8217;re building a genuinely good product, that&#8217;s how you&#8217;ll avoid leading your customers into the traps that we discussed above. Trust us, we&#8217;ve made these mistakes!</p><p>There are also clear lessons for buyers. First, it requires trusting your team. As a leader, you need to enable your team to make the call about what agents are worth spending time on and what will deliver the results that align with your priorities. The responsibility to make decisions about which agents are good or bad has to rest with the people doing the work.</p><p>Second, there will be trial and error &#8211; and you&#8217;ll have to get used to it. We&#8217;ve heard anecdotes about major law firms buying all three of the top AI legal agents and letting their organizations use them in parallel with the intention of picking one after a year or two. While this might sound wasteful, it&#8217;s a wiser approach than being beholden to one implementation up front. It&#8217;s better to learn now rather than a year from now.</p><p>Finally, we find the idea of flat organizations particularly interesting. Brian Chesky talked about this in a <a href="https://colossus.com/episode/ai-founder-mode/">recent interview</a>: The closer you can get to the ground truth &#8211; the actual work being done and the problems being faced &#8211; the more likely you are to make intelligent decisions about what&#8217;s working and what&#8217;s not. The more abstraction you have between yourself and reality, the more likely you are to think that your team is light years ahead or behind where it actually is.</p><p>Ultimately, focus is what will win. AI that lives at the intersection of possible and good is what will create value.</p>]]></content:encoded></item><item><title><![CDATA[Enterprise AI's teenage years]]></title><description><![CDATA[How a phase of rapid change is throwing a wrench in AI adoption]]></description><link>https://frontierai.substack.com/p/enterprise-ais-teenage-years</link><guid isPermaLink="false">https://frontierai.substack.com/p/enterprise-ais-teenage-years</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 21 May 2026 19:23:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Csdt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A few weeks ago, a prospective customer told us they wanted to pause their POC with us. They were in the middle of migrating their core observability stack, they&#8217;d been talking to two other vendors in our space, and someone on their team had spent a weekend at an internal hackathon building a homegrown version of what we do. They wanted to see how the hackathon project held up before deciding what to do next. As we kept talking, we learned the observability data migration had been running for a few months but other parts of the stack hadn&#8217;t been figured out yet, so they weren&#8217;t sure how to evaluate AI SRE vendors &#8212; and they didn&#8217;t know we could help with some of the key pieces. The team invested months into evaluating, building, and migrating, but there were no clear priorities or timeline.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Csdt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Csdt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Csdt!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03043603-9a28-420f-a65a-d65c3ed5069a_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini. </figcaption></figure></div><p>Enterprise AI is in a critical transition and maturation period &#8211; we&#8217;re calling these the teenage years. From 2023-2025, we were the &#8220;kid in a toy store phase&#8221; &#8211; budgets were unconstrained, and most shiny demos could find an excited buyer. The teenage years are going to be awkward and uncertain and that means that traditional enterprise sales patterns don&#8217;t work. This phase won&#8217;t last &#8212; the market will eventually settle back into something resembling traditional enterprise software &#8212; but the vendors who navigate the window well will be the ones still standing when it closes. That means giving up on the old playbook and meeting customers where they actually are.</p><h3><strong>What enterprises are actually doing</strong></h3><p>Three things define how enterprises are buying AI right now, and none of them line up with how the enterprise sales motion was designed to work.</p><p>The first is genuine uncertainty about the future, and it&#8217;s the most obviously understandable of the three. The last twelve months of model and tooling progress have made any multi-year commitment feel risky. Buyers are looking at what happened when Claude Code and Codex took a step-function leap last fall, and they&#8217;re correctly inferring that there very well might be a next equivalent sea change in six or twelve months. Locking into a three-year contract with any vendor right now feels like signing a lease on a house you might want to leave before the paint dries.</p><p>The second is a sharply increased appetite for building in-house. The quality of coding agents has shifted the build-vs-buy calculus so dramatically that &#8220;let&#8217;s also throw together an internal version and see how it compares&#8221; has become a standard step in evaluations. We&#8217;re doing this ourselves: instead of keeping a web design consultancy on retainer, we recently rebuilt our entire website from scratch and wired it into a headless CMS, and the whole thing cost us less than $100 in tokens. The math that used to make in-house an obvious act of engineering hubris is no longer obvious. Of course, this doesn&#8217;t take into account the ongoing maintenance and the fact that same DIY effort could be applied to 10 other initiatives.</p><p>The third is genuinely chaotic decision-making. VPs who missed the last big trend &#8211; cloud, mobile, etc. &#8211; don&#8217;t want to miss AI. The customer we mentioned at the top isn&#8217;t unusual. We regularly see deals where buyers are simultaneously evaluating multiple vendors, building something themselves, and migrating an underlying tool that affects the whole evaluation &#8212; none of which they bring up until something forces it into the conversation. <strong>The speed of change has compressed everyone&#8217;s planning horizon to the point that sequential decision-making has stopped working.</strong></p><h3><strong>Traditional playbooks are wrong</strong></h3><p>If you&#8217;ve sold enterprise software before, your instincts are wrong.</p><p>The default response to a build-curious customer is to argue them out of it. You walk them through the maintenance burden, the opportunity-cost of engineering cycles, and the advanced features they haven&#8217;t considered. Normally, this is pretty compelling. Today, it almost guarantees you lose the deal. Building something yourself is fun and exciting and feels like an exercise of agency. The harder you push against it, the more it seems like you&#8217;re a vendor making a desperate argument.</p><p>You have two options here, and you have to pick the one that works best for you. The first option is to build something that adds value faster than the DIY solution takes shape &#8211; your product has to be configured and onboarded immediately. Anything longer than a weekend loses to DIY. The hackathon project isn&#8217;t going to have the same quality, but if your alternative is a 6-month enterprise sales cycle, a half-good solution today is better. By the time you&#8217;ve finished your POC, the customer has either built their own muscle around the homegrown version or moved on to whatever the next exciting thing is.</p><p>The second option is to become an enabler for the DIY engineers: Let them build an interface on top of your expertise, so they both get a sense of agency from building and move faster than they would have otherwise. This isn&#8217;t necessarily something every company can do, but enabling DIY-curious engineers can be very powerful.</p><p>The second instinct that&#8217;s wrong is pushing for longer commits and more structured POCs. Traditional enterprise selling rewards discipline here: show value, propose a multi-year deal, get the procurement process running, and get the buyer locked in. In the current market, the opposite is winning. The best AI companies we see are offering zero or minimal commitment upfront, getting customers onboarded fast, and growing contracts over time as value becomes obvious. This gives the vendor a quicker win while giving the buyer the flexibility in case the world changes again in a few months. An interesting paradox we&#8217;ve seen is that the flexibility to leave sometimes produces <em>more</em> investment in making the product successful &#8211; investment feels safe because optionality exists.</p><p>Finally, there&#8217;s the chaos of enterprise decision-making. Your risk is again in timing &#8211; time kills all deals, and the solution is, again, flexibility. Land quickly and efficiently, enable your product to become entrenched, and grow from there.</p><h3><strong>From teenage to adulthood</strong></h3><p>The transitional phase isn&#8217;t permanent, and the playbook that wins it isn&#8217;t the playbook you&#8217;ll need for the next phase.</p><p>On pricing, the current vogue for usage-based and outcome-based pricing is going to run into the wall of enterprise quarterly planning. Aligning what you charge for with the unit of actual work done is conceptually clean &#8212; it&#8217;s how consulting contracts have always worked &#8212; but enterprises are going to struggle to reconcile unpredictable monthly bills with quarterly forecasts. We&#8217;re already hearing about public companies having end-of-quarter panics over engineering budgets because spend ran wildly over. The likely endpoint is a return to longer-term, more fixed commitments &#8212; but with more sophistication about right-sizing commitments. The short-term will stay messy, because enterprises are correctly prioritizing flexibility <em>now</em>. The vendors who win the transition will be the ones who use the flexibility window to get embedded, then transition into more predictable commercial structures.</p><p>On the build-vs-buy front, we&#8217;ll likely see the euphoria tamped down. In the next year or two, an enterprise is going to make headlines for trying to homegrow a CRM, ERP, or payroll system, and it&#8217;s going to backfire spectacularly. The lesson won&#8217;t be &#8220;stop building&#8221;; it&#8217;ll be &#8220;build the things that are genuinely different for your business, and buy everything else.&#8221; In other words, the conventional wisdom will be true again.</p><p>Longer term, we think the market will eventually head back toward something that looks like traditional enterprise software &#8212; well-understood buying cycles, comparable feature matrices, predictable evaluation processes. The difference is that the companies running those cycles two years from now will be the ones who survived the in-between by abandoning the old playbook.</p><h3><strong>Hello fellow kids</strong></h3><p>The temptation in a market this strange is to keep trying to make it behave the way the old market behaved &#8211; or to throw caution to the wind and do something totally different. Reality is as always somewhere in between. To push for a formal sales cycle or for the long commit doesn&#8217;t work right now. At the same time, we aren&#8217;t in an alternate dimension where the laws of physics are backwards. Enterprise software vendors like Notion, Vercel, and Datadog figured out how to deliver value to companies quickly over the past decade. That trend has spread everywhere today. What you ultimately need to understand is what your customers&#8217; incentives are &#8211; those genuinely are different from 5 years ago &#8211; and how you can best enable them.</p><p>The transition phase is going to keep being chaotic for a while. The after is going to look a lot more like the traditional enterprise software market we all know. The window between the two is open right now, and the vendors who use it well will be the ones still standing when it closes.</p>]]></content:encoded></item><item><title><![CDATA[Product-market fit is a trap]]></title><description><![CDATA[Why the old playbook for scaling a software company might work against you]]></description><link>https://frontierai.substack.com/p/product-market-fit-is-a-trap</link><guid isPermaLink="false">https://frontierai.substack.com/p/product-market-fit-is-a-trap</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 14 May 2026 17:54:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iFGr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Cursor has, by almost any reasonable measure, run the textbook play of the last two years &#8212; find a wedge with developers, build a product they love, ride a wave of bottoms-up adoption into the fastest revenue curve in software history. And yet, in the last 6 months, Cursor seems to have lost a huge amount of ground &#8211; or at least mindshare &#8211; to Claude. The product hasn&#8217;t gotten worse &#8211; it&#8217;s gotten better, if anything. But the market moved.</p><p>We&#8217;ve written a few times before about how AI is breaking the playbooks that software companies have relied on for decades. Most of those posts have been about the way the technology itself shifts what&#8217;s possible &#8212; commoditization of models,<a href="/__u/frontierai.substack.com/p/data-is-your-only-moat"> data as the only real moat</a>, the gap between<a href="/__u/frontierai.substack.com/p/ai-companies-are-building-for-the"> enthusiasts and the rest of the market</a>. We&#8217;ve begun to suspect that the changes might be even more fundamental: Even if you do everything right, the act of operationalizing what&#8217;s working is now the thing that puts you most at risk.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iFGr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iFGr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6735207,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/197732341?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iFGr!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F276283b7-1633-4771-85c3-d2e8efa84005_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>In AI, product-market fit is a trap. Not because the thing you found doesn&#8217;t work &#8212; it almost certainly does &#8212; but because the discipline of turning it into a repeatable machine is what makes you miss the next thing. And the next thing arrives every few weeks.</p><h2><strong>Operationalizing kills your signal</strong></h2><p>The traditional advice once a company finds product-market fit is to operationalize relentlessly. On the sales side, you codify your ICP, hire a large sales team, write playbooks, and instrument the funnel. On the product side, you deepen the functionality that got you to this point. The whole goal of post-product-market fit life is to systematize what you suspect works. Historically, this is what separated companies that scale from companies that have a good first year and then plateau.</p><p>The goal of operationalizing a process is to create consistency: As the saying goes, investors want to know that putting $N into sales &amp; marketing is going to yield $M in net-new revenue. It just might be the case that consistency &#8211; which is the holy grail of a growing business &#8211; is also a risk in AI.</p><p>At the early stages of creating a business, you might have a particularly interesting customer call, come up with a new idea, and change your product roadmap for the next two months. When you&#8217;re running a machine that&#8217;s meant to scale, you can&#8217;t turn on a dime. The concern is that your customers&#8217; preferences are changing faster than ever before &#8211; the cool demo they saw on Twitter last night might be their new benchmark for what good in your market is. If that signal takes weeks to spread across your sales team and work back to the product team, you might already be too late.</p><p>Cursor&#8217;s situation is a useful illustration here. They were able to scale revenue ridiculously quickly in 2024 and 2025, and they raised massive amounts of capital to deepen their moat and scale. Over the last 6 months, as Claude Code and Codex have turned into extremely formidable competition, the sheen has worn off. We don&#8217;t want to overstate the point &#8211; a potentially $60B acquisition by SpaceX would still be an incredible outcome! &#8211; but anecdotally, Cursor is not the <em>it</em> tool that it was a year ago. A combination of the UX improvements in Claude and Codex and the increased value in vertical integration between coding agent and model provider are the current advantages. The exact reason Cursor was so successful 12 months ago &#8211; being a sleek, minimal IDE that was model agnostic &#8211; might be what&#8217;s dragging them down today. Of course, things might flip next week!</p><p>This is a Catch-22. If we saw the growth that Cursor did in 2024, we too would have made the same decisions. If they didn&#8217;t, perhaps someone else would have and grown faster than them, and they would be lost to history. What we&#8217;re pretty sure about is that &#8211; more than ever &#8211; you constantly have to be replacing your own product.</p><h2><strong>The early-late paradox</strong></h2><p>The second reason PMF is now a trap is more structural, and it&#8217;s the part of this we&#8217;re least sure how to solve.</p><p><em>Crossing the Chasm</em> canonically outlined that you sold to early adopters first, used what you learned to build a product the early majority would tolerate, and then over years made your way through the late majority and the laggards. The model assumed that these segments were sequential and that they roughly corresponded to different companies. You went after the tech-forward shops first, then the mainstream, then the slow-movers. The product, the sales motion, and the messaging could evolve in step with the segment you were chasing.</p><p>That model doesn&#8217;t really describe what we see anymore. The early adopters and the late majority are now sitting at the same company &#8212; sometimes on the same team. We have one call with an engineer who has rebuilt their entire workflow around Claude Code and is impatient for the next leap; the very next call is with someone, two desks over, who tried a coding agent once, didn&#8217;t like the UX, and quietly went back to writing things by hand.</p><p>We discussed this in detail in <em><a href="/__u/frontierai.substack.com/p/ai-companies-are-building-for-the">AI companies are building for the wrong users</a></em> &#8211; an excited user might look past the flaws in your UX and ask for innovative new features while the reluctant user might stop paying attention to you at the first sign of a clumsy implementation. Neither one is wrong necessarily, but building each kind of product requires a different muscle.</p><p>You can&#8217;t build two products, because you&#8217;re a startup and you have to ship something. You can&#8217;t split the difference, because a product designed to be acceptable to both ends up being remarkable to neither. The honest answer is that you have to pick &#8212; and once you pick, you have to commit, and hope that the segment you bet on is the one that pulls the rest of the organization along.</p><p>This is what makes the operationalization problem from the previous section especially acute. The natural instinct, once you find a buyer who loves you, is to build the GTM machine around that buyer. But if that buyer is the early adopter inside an organization where the late majority controls the rollout, the machine you&#8217;ve built is selling to the wrong half of the room. And the more efficient that machine gets, the harder it becomes to even notice that the other half exists.</p><h2><strong>Are we all out of luck?</strong></h2><p>Most of the writing about AI strategy ends with some version of &#8220;stay flexible,&#8221; which is true but not actionable. The more honest version of the conclusion is this: in this market, the companies that win are not going to be the ones that find product-market fit first. In fact, we would even argue that systematizing first might be a disadvantage because it locks you into a bet while the market is still developing.</p><p>That sounds like heresy if you grew up on the SaaS playbook. But the SaaS playbook was built for a market where the workflow you sold to in year one was approximately the same workflow your customer would have in year four. In AI, that is no longer true. The playbook might not last a quarter because customers&#8217; preferences are being rebuilt every week.</p><p>There&#8217;s a lot of chatter about rethinking how organizations are built in the age of AI &#8211; to be leaner, to be flatter, and to have more ownership. This is perhaps the most compelling reason why: Information about what customers are saying, where they&#8217;re failing, and what you need to do differently <em>must</em> travel faster than before. Product roadmaps can no longer be quarters or years out, and GTM motions need to change just as fast.</p><p>The good news is that we believe this will lead to more winners, not fewer. Every previous platform shift produced a handful of dominant companies who locked in their position early and held it for a decade or more. Whatever we think works today might change tomorrow and swap back the next day. Being ahead isn&#8217;t as safe as you think because you&#8217;re doing what worked last quarter. Iif you&#8217;re behind, you can still catch up.</p>]]></content:encoded></item><item><title><![CDATA[AI applications have two growth curves]]></title><description><![CDATA[Different markets, different physics]]></description><link>https://frontierai.substack.com/p/ai-applications-have-two-growth-curves</link><guid isPermaLink="false">https://frontierai.substack.com/p/ai-applications-have-two-growth-curves</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 07 May 2026 21:52:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E0i0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every board deck in AI right now uses the same benchmarks. Companies like Cursor and Sierra   have ridiculous growth curves. The message is implicit: this is what AI growth looks like in 2026, and if your company isn&#8217;t on this trajectory, something is wrong. Founders have to explain the gap. Investors are calibrating expectations against it. Buyers are wondering whether every vendor they talk to needs to have raised $100MM.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!E0i0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!E0i0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:221632,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/196835659?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!E0i0!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e6295bb-c1cf-415c-94f4-e33c54c71bd2_1024x559.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>This is a category error, and it&#8217;s distorting how everyone in AI reads the market. Cursor and Sierra are growing fast for very different reasons &#8212; and crucially, neither of those reasons applies to the hardest and most valuable problems AI is being pointed at. The infrastructure markets that look slow by comparison are not behind. They are operating on fundamentally different physics, and in many cases, the slowness is the thing that will make the eventual companies valuable.</p><h2><strong>The market map, revisited</strong></h2><p>We mapped this out earlier this year in<a href="/__u/frontierai.substack.com/p/data-is-your-only-moat"> </a><em><a href="/__u/frontierai.substack.com/p/data-is-your-only-moat">Data is your only moat</a></em>. The 2x2  we introduced in that post separates AI markets along two axes: how hard the problem is to solve technically, and how hard the product is to adopt organizationally. Three of the four quadrants are seeing fast growth right now, and each for a different reason. Easy-to-adopt, easy-to-solve (<em>easy-easy</em>) markets &#8212; consumer chat, basic search replacement &#8212; are growing because there&#8217;s almost no friction in either direction. The model providers are dominating these and that&#8217;s mostly fine. Easy-to-adopt, hard-to-solve markets, the Cursor quadrant, are growing because individual users can adopt without organizational approval and the data flywheel that follows is incredibly powerful. The core AI capabilities are racing ahead, and it&#8217;s hard to catch up if you&#8217;re not already in the game.</p><p>The hard-to-adopt-easy-to-solve quadrant is the more interesting case for this discussion, because the growth there has been remarkable too. Sierra, Decagon, and other enterprise support agents are scaling rapidly despite needing committee buy-in to land each customer &#8211; an individual support rep can&#8217;t buy Sierra. The reason is that the underlying problem &#8212; execute a defined playbook for a known workflow &#8212; is tractable enough that vendors can demonstrate value quickly once they&#8217;re in the door. The large fundraises these companies have done signal credibility to a Fortune 500 buyer who needs to know you&#8217;ll still be around in three years.</p><p>That leaves the fourth quadrant: hard-to-adopt-hard-to-solve. SRE, security operations, complex infrastructure agents. This is the only quadrant that doesn&#8217;t have a Cursor or a Sierra growing into it at breakneck pace, and it&#8217;s where the misreading is doing the most damage. Founders in this quadrant get benchmarked against companies in the other three quadrants and asked why they&#8217;re not on the same curve. The honest answer is that they were never going to be &#8212; and more importantly, that the slowness is fundamental to the problems being solved here.</p><h2><strong>What slowness buys you</strong></h2><p>When a market doesn&#8217;t let you move at breakneck speed &#8212; because the buyer is a committee, because the integration is genuinely complex, because the customer needs weeks to evaluate &#8212; you are forced to spend that time figuring out what the right product actually is. In a fast market, you don&#8217;t get that time. The competitive pressure is too high, the customers are too willing to start immediately, and the only correct move is to ship the obvious solution as quickly as possible and iterate from there.</p><p>Certainly for hard problems, the obvious solution is rarely the best one. Take AI SRE, which is where we spend our time. Almost every startup in the space (and there are many!) is building an AI-powered root cause analysis agent that is triggered by incident management alerts and requires customer-maintained runbooks. It&#8217;s the obvious solution: humans use runbooks, so the agent should too. Except alert thresholds are noisy, no engineering team actually maintains their runbooks well, and the agent inherits all the gaps. By trying to chase the Cursor growth curve, the players in this space haven&#8217;t asked the more interesting question: how do you use agents to detect early warning signs of an incident, validate them, and figure out the root cause from that structure &#8212; before any threshold alert ever fires?</p><p>To be transparent, we didn&#8217;t see this initially either. When we started, we set out to build the same banal runbook-driven RCA agent everyone else was building. The reason we ended up somewhere different is that the market gave us time. If customers had been willing to write us $250K checks on day one for a generic RCA agent, we would have shipped that and spent the next two years executing on the wrong product. The slowness of the market is what pushed us to keep asking whether the obvious solution was the right one. It turned out it wasn&#8217;t.</p><p>This is the first-order benefit of a slow market: it forces you to solve the actual problem, not the one that&#8217;s easiest to monetize.</p><p>Understanding this dynamic also helps you tailor your approach to the market you&#8217;re in. When you try to grow too fast for the type of market you&#8217;re in, the pressure forces you to take POCs you shouldn&#8217;t take, ship to customers you can&#8217;t actually serve well, and chase logos you don&#8217;t have the product depth to make successful. The cost isn&#8217;t just churn. It&#8217;s reputational damage that extends to the entire category. We&#8217;ve heard plenty of stories about AI agents in our market and adjacent ones that were rushed into enterprise environments, failed in visible ways, and left buyers with the strong sense that this whole category of product doesn&#8217;t work yet. That impression is sticky. It poisons the well not just for the vendor that failed but for everyone behind them trying to sell into the same buyers.</p><h2><strong>Education as a moat</strong></h2><p>The other key benefit is the one we underestimated coming into this experience: the education process itself becomes a moat. Most customers in hard-hard markets don&#8217;t fully know what an agent in their space is supposed to do. They have intuitions and pain points &#8212; but the category isn&#8217;t well-defined in their heads yet. That sounds like a problem, and in a fast market it would be. In a slow market, it&#8217;s an opportunity.</p><p>Look at finance agents for closing books, a market which perhaps is a few steps ahead of AI SRE on the same curve. Walk into any controller&#8217;s office and you&#8217;ll find a long list of vendors all positioned around roughly the same pitch: We automate your close. Some are reconciling against templates, some are categorizing transactions, some are flagging variance. The buyer&#8217;s job is to figure out which of these is actually solving their problem and which is selling them a glorified macro.</p><p>A vendor with a real point of view &#8211; we&#8217;re not experts, but <a href="https://digits.com/">Digits</a> seems interesting &#8211; can stand out from the noise. The idea behind digits isn&#8217;t to put a pretty UI on top of your existing, messy general ledger data: They&#8217;re rebuilding the concept of a general ledger from scratch in order to support quicker, more efficient books that are constantly kept up to date. You might not agree with the approach, but it expresses a clear point of view that&#8217;s cohesive &#8211; not just AI slapped on top of whatever you had before. A differentiated point of view earns trust that competitors with vague positioning can&#8217;t easily replicate.</p><p>You don&#8217;t get to do this in a fast market. By the time Cursor was teaching anyone what coding agents could do, every developer on Twitter had already figured it out themselves. The market did the education for them. In hard-to-adopt markets, you do the education yourself, and the customers who learn from you tend to remember who taught them.</p><h2><strong>How to read these markets</strong></h2><p>Taken together &#8212; time to find the right solution, protection from category-poisoning failures, and education-as-differentiation &#8212; and you start to see why the slow curve isn&#8217;t just a consolation prize. It&#8217;s the trajectory that builds companies competitors can&#8217;t easily copy in a quarter when the funding cycle turns. The fast curve produces companies that grow incredibly quickly and then have to fight off model providers, incumbents, and a long tail of well-funded copycats with similar products. The slow curve produces companies that, by the time they&#8217;re visible, have a product nobody else can replicate without going through the same multi-year process.</p><p>This has implications for how you interpret business in these markets. The signals worth tracking are whether a company has a non-obvious technical thesis, whether their POC win rate is improving over time, whether their existing customers are deepening usage rather than just renewing. These are the signals that actually predict which companies in hard-hard markets will still be standing in five years.</p><p>The companies winning the fast AI markets right now deserve their growth. We&#8217;re not arguing otherwise. What we are arguing is that the curve they&#8217;re on is not the only curve, and treating it as the benchmark for all of AI is going to lead to a lot of misallocated capital, misread pipelines, and eventually, a lot of surprise when the slow-market companies turn out to be the ones with the deepest moats.</p><p>The next phase of AI isn&#8217;t going to look like Cursor&#8217;s growth chart. For the most valuable problems, it isn&#8217;t supposed to.</p>]]></content:encoded></item><item><title><![CDATA[The Inference Economy: Token Use]]></title><description><![CDATA[Why demand might matter more than supply]]></description><link>https://frontierai.substack.com/p/the-inference-economy-token-use</link><guid isPermaLink="false">https://frontierai.substack.com/p/the-inference-economy-token-use</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 30 Apr 2026 17:54:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!py7U!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>For boring logistical reasons, our blog post for this week didn&#8217;t come together. This was one of our posts from last year, and it feels more relevant than ever: As model providers <a href="/__u/frontierai.substack.com/p/model-inference-model-products-and?utm_source=publication-search">layer more functionality above inference</a> and move towards closed API models like Mythos, it is critical that application builders have a plan for managing costs <strong>and</strong> model access. Back next week!</em></p><div><hr></div><p>Last week, we wrote about some of the <a href="/__u/frontierai.substack.com/p/the-inference-economy">trends in the inference economy</a>: plateauing token costs, using the right model for the right task, managing different modes of compute, and the impacts on pricing. Our perspective for the purposes of that that post was derived from observing what was going on with LLM inference in general in recent months. Of course, we&#8217;re always thinking about how that affects our decision-making as application builders, and we touched on that briefly but mostly focused on what was actually happening to token costs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!py7U!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!py7U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!py7U!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66e68069-88e6-41f1-834c-7846d4ca40e6_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: GPT-5.</figcaption></figure></div><p>Looking back on that post, it occurred to us that the underlying assumption was that the dynamics of token consumption &#8212; from the perspective of application builders &#8212; is changing as well. Those changes are likely what&#8217;s driving the data center buildouts in the first place and are driving the demand side of the equation, which is forcing supply to keep up. As a result, today&#8217;s post is about what&#8217;s driving token token demand and how you should think about managing your demand as a token consumer.</p><h2><strong>Changes in token demand</strong></h2><p>What do we mean when we say that there&#8217;s a change in the demand side of the equation? Simply put, <strong>we&#8217;re all using </strong><em><strong>more tokens to process more data</strong></em>. The increased token volume is partially driven by increased usage, but that&#8217;s far from the whole story. Yes, of course AI applications have more usage, but what&#8217;s much more interesting &#8212; and challenging &#8212; is that we&#8217;re seeing a trend in increased token consumption <em>per-request</em>, not just an increase in the aggregate number of requests processed. The driver behind that is ultimately quality.</p><p>As we&#8217;ve said many times on this blog, getting the behavior you want out of LLMs is all about providing the right information at the right time to a model. If you have the wrong context, you&#8217;re going to get poor results. That means the million dollar question is how you find the right information.</p><p>Search was the first solution that we all turned to &#8212; first, vector search, then returning back to more traditional text search mechanisms. Very quickly, however, we all turned to having an LLM read an input and evaluate how relevant it was to the problem we were solving (&#8221;reranking&#8221;). Presciently, Vik Singh, now a CVP at Microsoft, <a href="https://www.youtube.com/watch?v=DtS5FWwCW6A&amp;t=16s">said this to us over 2 years ago</a>: &#8220;If LLMs were fast enough&#8230; why not use the LLM to do a much more advanced similarity search&#8230; I think that&#8217;s what people actually want.&#8221;</p><p>The LLM-preprocesses-data paradigm is pervasive in our systems today. At RunLLM, we pre-read data at ingestion time to organize it properly, we read the results of text + vector search to analyze their relevance to a question, we analyze logs and dashboards in real-time with LLMs, and so on. Each one of those tasks is done by a model call <em>in isolation</em> to understand whether that data should be used for future decision-making. Without the involvement of LLMs in these stages, it would be almost impossible for us to provide high-quality results to our customers. We have a joke internally at RunLLM: <a href="https://en.wikipedia.org/wiki/Fundamental_theorem_of_software_engineering">The solution to every problem in computer science is another layer of abstraction</a>, and <a href="/__u/frontierai.substack.com/p/throw-more-ai-at-your-problems">the solution to every problem in AI is another layer of LLM calls</a>.</p><p>That means that median &#8212; and perhaps more importantly, p99 &#8212; token consumption (and therefore request costs) are going up very quickly. We&#8217;re all solving harder problems, which means we&#8217;re throwing more data into LLMs and ultimately consuming dramatically more tokens. In our minds, this is one of the key drivers of increased token demand.</p><p>Luckily for the data center builders, this trend is not going anywhere. We might get more efficient and cheaper inference (though we&#8217;re skeptical, as we talked about last week), but as LLMs get more integrated into every application and workflow, per-request token use is only going to go up, not down. As a single data point, we have tons of ideas for how we can throw more LLMs at some of the challenges we face in a single investigation at RunLLM, but we&#8217;re primarily limited by cost, latency, or evaluations at the moment.</p><p>If you&#8217;re going to inevitably use more tokens, it&#8217;s worth thinking about how to be as thoughtful as possible about those tokens.</p><h2><strong>How to manage your token demand</strong></h2><p>We&#8217;re pretty confident token demand is going up, and as we discussed last week, token costs are plateauing. Depending on how long it takes to build and power these new datacenters, that means we should all be thinking about how to manage our token usage, especially as models get better and more expensive. We&#8217;ve been experimenting with many of these techniques for a while now at RunLLM, so we thought we&#8217;d share some early lessons.</p><ul><li><p><strong>Model size is your best friend.</strong> Not all models are created equal, and neither are all tasks. Throwing your largest model at every task will probably maximize quality, but it will burn through your budget faster than you can imagine. (We mistakenly spent $63 on a single investigation at RunLLM last month. &#128561;) There are plenty of things that we do &#8212; gating questions, filtering documents, synthesizing logs &#8212; that aren&#8217;t hard but just require processing data efficiently. For simple tasks, there&#8217;s really no reason to use a state-of-the-art-model &#8212; GPT-4.1 Mini (one of our current favorites) or a smaller open-source equivalent will get the job done just fine. Unfortunately, we don&#8217;t have a cut-and-dry rule for when to use what model. It&#8217;s more of an art than a science right now, but evaluation frameworks for specific tasks certainly will help guide you in the right direction.</p></li><li><p><strong>Be flexible with your providers.</strong> We&#8217;ve long believed that LLM inference is a race to the bottom. If models get better, the main question becomes who can give you that model for as cheap as possible &#8212; especially with open-weight models. However, we touched on the fact that switching model providers is harder than it used to be last week because model providers are making stronger assumptions, but tools like DSPy make prompt optimization easier than ever, which should alleviate some of that tension. While you might not want to be ready to switch between every model provider on the market (there are a lot!), it&#8217;s probably worth your time to be ready to use one of a few different providers &#8212; or even using features like batch mode within individual providers &#8212; when you have the opportunity. The biggest issue with this is actually security &amp; compliance: More data subprocessors create more data exposure risks and make your vendor approvals that much harder. But if you&#8217;re in an area where this is less of a concern, keeping your options open is definitely an option to reduce costs.</p></li><li><p><strong>Do you need reasoning?</strong> Reasoning models use a lot more tokens than regular LLMs, and it&#8217;s correspondingly quite difficult to control the output costs. It&#8217;s worth asking whether and when you need a reasoning model. For daily personal use, we tend to default to ChatGPT 5 Thinking, but <em>we actually are not currently using any reasoning models in production at RunLLM</em>. We&#8217;ve had much better luck with breaking problems down into fine-grained steps, using regular Python for orchestration and tool calling, and picking the right model for the right task (see above). Interestingly, this mirrors some of the reasoning-based task planning workflows we see in our daily usage, but with much stronger guardrails. Of course, we&#8217;re not solving problems with the generality of a consumer app like ChatGPT, so we have narrower score and much stronger guardrails. But for many workflow/task automation-oriented applications, you might be able to be much more efficient than you realize.</p></li><li><p><strong>Don&#8217;t run straight to fine-tuning/post-training.</strong> RL is a hot topic right now. As we mentioned last week, <a href="https://cursor.com/blog/composer">Cursor&#8217;s new custom autocomplete model</a> has rekindled the excitement around fine-tuning models for custom tasks. It&#8217;s certainly appealing: You take a smaller model, feed it a bunch of data, and voil&#224; &#8212; cheaper, faster inference. Unfortunately, reality is not quite that simple. For one thing, <a href="https://x.com/natolambert/status/1907530344925970907?utm_source=chatgpt.com">RL and post-training are hard</a>. The promise of the recent wave of RL environment startups for post-training is that they will help remove the complexity here, but we&#8217;re not convinced that that&#8217;s a viable solution. The hard part isn&#8217;t running an algorithm to update weights &#8212; it&#8217;s framing the problem in a way that&#8217;s actually going to yield the results you want and getting enough data that fits that problem framing. (This is <a href="https://www.alexirpan.com/2018/02/14/rl-hard.html">not a new challenge</a> in RL.) What&#8217;s lost in the hype around Cursor is that they collected an immense amount of data well-suited to RL from the natural use of the product &#8212; every tab complete suggestion was accepted or rejected, which is a very friendly RL framing. By way of contrast, we have over 1MM question-answer pairs from RunLLM but a tiny fraction of those have actual feedback, and only a fraction of those have actionable feedback &#8212; e.g., we tend to get a most of our negative feedback on &#8220;I don&#8217;t know&#8221; answers, which are usually good because there&#8217;s not sufficient data to answer. If you&#8217;re in a domain where you have enough data and the expertise + resources to use post-training, it&#8217;s certainly viable from a unit margins perspective. But it&#8217;s not the panacea everyone&#8217;s claiming.</p></li></ul><p>What&#8217;s interesting about these dynamics at the moment is that we&#8217;re all focused on costs but not as focused on our own pricing power. Of course, any business is always going to want to reduce its COGS &#8212; the more efficient you can be, the better your business scales. This is probably the right place to be given the fierce competitive dynamics in many AI markets. At the same time, while we&#8217;re working on technical solutions to reducing COGS, we&#8217;re also mindful of the fact that as AI applications mature and the ROI becomes more obvious, we&#8217;re likely going to see a corresponding increase in pricing power. The best applications will possibly command a significant premium. That won&#8217;t apply in every market of course &#8212; only in the ones where quality matters most.</p><p>Guesses aside, it&#8217;s clear that the economics of AI are changing faster than we would have expected. The sudden deceleration of per-token cost changes has coincided with more mature applications that require more tokens &#8212; a double-whammy to increase costs. There will be other solutions (technical and non-technical) that will change these dynamics, but for the foreseeable future, we&#8217;re all going to be keeping a close eye on our OpenAI bills.</p>]]></content:encoded></item><item><title><![CDATA[Agents can't choose between structure and flexibility]]></title><description><![CDATA[Why maximizing in either direction is a failure mode]]></description><link>https://frontierai.substack.com/p/agents-cant-choose-between-structure</link><guid isPermaLink="false">https://frontierai.substack.com/p/agents-cant-choose-between-structure</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 23 Apr 2026 16:17:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PaIB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I think it&#8217;s safe to say that when the LLM hype cycle started a few years ago, no one expected one of the great debates of our time would be between Python and Markdown as agent specification languages. But here we are, and this has quickly turned into one of the most consequential architectural questions in AI.</p><p>Before we dive into the consequences of this debate, we&#8217;ll take a moment to define our terms.</p><p>The Python camp uses code to express strict requirements for the steps an agent should take to accomplish a task. The Markdown camp uses English to express broad goals and constraints and lets the agent plan its way to the outcome. The tradeoffs are fairly straightforward. Code creates strong guardrails and reduces the chance that the agent&#8217;s plan goes off the rails. Markdown gives powerful models the freedom to explore, adapts flexibly across tools and models, but risks the agent doing something unexpected and undesirable.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!PaIB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!PaIB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2361180,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/195255759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!PaIB!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea294116-4f94-4277-8017-c87721a42b8d_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: ChatGPT.</figcaption></figure></div><p>Most of the debate treats this as a choice between two defensible positions. It isn&#8217;t. Both maximalist positions are, in fact, failure modes, and the reason is the same: Neither one is actually agent-native. Agents, like humans, are increasingly being given complex tasks, and that requires the flexibility to choose the right tool for the right task (or subtask). Code-maximalism forces agents to follow deterministic workflows and strips out the reasoning that makes them useful. Markdown-maximalism abdicates control and produces systems you can&#8217;t debug, correct, or improve. Picking a side is how you avoid the hard work of designing an agent.</p><p>We&#8217;re publishing this as part of the Agent Native series because these two approaches increasingly define how agent interactions get built &#8212; and because both maximalist versions end up in the same place we wrote about last week: copy-pasting what a human would do, just in different syntax.</p><p><strong>What code-maximalism gets wrong</strong></p><p>The code-maximalist pitch is reliability. You tell the agent exactly what to do in specific cases, surface errors when things break, and get tightly scoped results. Given that LLMs make mistakes, misunderstand intent, and generally do all sorts of weird things, this sounds appealing in theory. Enforce correctness at the code layer. Don&#8217;t trust the model to do the right thing.</p><p>We&#8217;re intimately familiar with where this can go wrong in the AI SRE space. Almost every vendor in our space tells customers they have to write runbooks. The product then encodes those runbooks as workflows and has the agent execute them in response to specific alerts. The results are trustworthy in the narrow sense: the agent does roughly what you expected. It&#8217;s also useless the moment an alert looks different from anything that&#8217;s come before or the underlying architecture changes. We started down this misguided path ourselves in the early days and quickly learned that it would rarely work in practice.</p><p>This approach fails to be agent-native in three ways. First, it copy-pastes what a human does. A human picks one hypothesis &#8212; the most likely based on experience &#8212; and runs it down. That works when the human is confident, but when the initial hypothesis is wrong, it creates a lot of wasted work. An agent doesn&#8217;t have to fall into that trap. It can evaluate multiple hypotheses in parallel, and some will be dead ends, but the chance it lands on the right answer goes up dramatically. That&#8217;s the architecture we&#8217;ve built RunLLM around, and it&#8217;s consistently how we see real incidents get resolved.</p><p>Second, the runbook approach gives humans no meaningful visibility. SREs don&#8217;t need to confirm that the agent executed Step 3 of the runbook. They need to know what the agent tried, what it ruled out, and why. A well-worn path automates some tedious work, but it doesn&#8217;t let the human trust or learn from the agent&#8217;s reasoning.</p><p>Third, encoded workflows don&#8217;t evolve &#8211; they lose the intelligence that agents promise. When the underlying system changes or requirements shift, every encoding has to change with it. There&#8217;s no way for the agent to take feedback, understand that the expected behavior has changed, and adapt on its own without someone going back into the harness.</p><p><strong>What Markdown-maximalism gets wrong</strong></p><p>The Markdown-maximalist is optimized for flexibility. Describe the goal, hand it to a capable model, let it figure things out. This is portable, expressive, and gets you something working quickly. Where creativity or open-ended problem-solving matters, it can be dramatically more useful than a fixed workflow.</p><p>The degenerate version of this is AI slide generation. We don&#8217;t know the exact architecture behind these tools, but from the outside they read as &#8220;let the LLM do everything&#8221; applications &#8212; one prompt in, a whole slide deck out. The failure mode is familiar to anyone who&#8217;s used one. Something is off. The layout is weird on slide 7, the chart doesn&#8217;t match the claim, the flow of the argument is scrambled. You want to say: &#8220;On slide 7, make the flow vertical instead of horizontal and move the chart to the bottom.&#8221; You usually can&#8217;t get this to work the way you expect. There&#8217;s no discrete layout logic to adjust, no separable step for chart placement, no addressable unit smaller than the whole generation. You re-prompt, get a new deck that&#8217;s wrong in a different way, and start over.</p><p>It would be easy to write this off as a strawman. Serious Markdown-maximalists aren&#8217;t arguing for one-shotting every single application. The sophisticated version of the position is skills.md plus a basic agent loop &#8212; rich context, thoughtful instructions, and a capable model reasoning its way through. Guide the agent through context, the argument goes, rather than constraining it with fine-grained LLM calls.</p><p>Complex applications expose the gap. When you&#8217;re grappling with reality, there are plenty of engineering decisions that still require strict constraints: Context management and summarization, model selection, cost management, and cross-agent coordination to name a few. In each one of these cases, the challenge is not trusting the model to reason intelligently. It is building the tooling and infrastructure that allows a thoughtful model to execute these tasks efficiently and reliably.</p><p>In production, this results in a code harness that manages context, routes between models, orchestrates sub-agents, and handles the predictable places where pure prompting breaks down. That ends up being a hybrid architecture with markdown doing the guidance work and code doing the structural work &#8212; which is exactly the position the debate was supposed to be between.</p><p>If you start with a Markdown-maximalist architecture, you&#8217;re probably going to end up building plenty of narrow, harness-like capabilities &#8211; content management, model routing, etc. &#8211; to enforce constraints whether you like it or not. The question is just whether you design those hooks intentionally or let the code component grow organically. You should be intentional about the design.</p><p><strong>The hybrid isn&#8217;t a compromise</strong></p><p>The teams building serious agents have, largely independently, landed in the same place: Markdown for intent and domain guidance, code for enforcement, tool execution, and anything that must not fail silently. Claude Code works this way. We built RunLLM this way.</p><p>It&#8217;s tempting to read this as an unopinionated compromise. That&#8217;s the wrong framing. The whole point of agents is that &#8211; unlike traditional software &#8211; they have an understanding of the problem to be solved and can use the right tools to get there. Code-maximalism compromises on the planning and Markdown-maximalism compromises on execution and learning.</p><p>The reason hybrid architectures are winning is because they&#8217;re the only architectures that support what agents are actually supposed to do. An agent needs reasoning flexibility to handle situations it hasn&#8217;t seen before, and it needs deterministic guardrails so humans can trust it and intervene when needed. Neither extreme gives you both, which means neither maximalist position gives you a truly flexible agent. It gives you either a workflow with aspirations or a wish with nothing to execute it.</p><p>The architectural work is figuring out, for each part of your system, which layer it belongs to. What needs to be expressed as intent and reasoned about? What needs to be enforced and checked? Where does the agent need creativity, and where does it need constraints? This is the hard part, and it&#8217;s the part that picking a side lets you avoid.</p><p><strong>What agent-native actually requires</strong></p><p>When you stop treating Python vs. Markdown as the debate, the architectural priorities come into focus. Can your agent evaluate multiple hypotheses in parallel, or does it march down one? Can a human see what the agent tried and why, or do they just get a final answer? Can the agent adapt when the underlying system changes, or does someone need to go edit the harness? Can a user correct the output at the level of granularity they care about, or is it all-or-nothing?</p><p>The maximalist debates are a symptom of an industry still thinking about agents as workflow automators &#8212; either very rigid ones, or very loose ones. The teams building agent-native products are past that argument, because they&#8217;ve figured out that the argument was never really about Python or Markdown. It was about whether you were willing to do the work to build something that actually behaves like an agent.</p>]]></content:encoded></item><item><title><![CDATA[Build agents for humans]]></title><description><![CDATA[Why copying human workflows is an anti-pattern]]></description><link>https://frontierai.substack.com/p/build-agents-for-humans</link><guid isPermaLink="false">https://frontierai.substack.com/p/build-agents-for-humans</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 16 Apr 2026 17:30:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!40sp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We&#8217;ve been running an experiment since the beginning of the year with a new AI SDR product that we bought. The tool uses signals from activity out in the world to qualify prospects and automate email outreach &#8211; things like new hires, website visits, or in our case, recent incidents. Last week, we were reviewing some of the outreach it had sent, and we noticed titles that were wildly outside our target market.</p><p>We&#8217;re not sure exactly how the system is doing it &#8211; probably a small LLM call or maybe a glorified regex. It turns out that the system was looking at a title like &#8220;VP of Construction Engineering&#8221; and deciding that it was close enough to &#8220;VP of Engineering.&#8221; The model wasn&#8217;t given enough context about what kind of engineering we cared about in our ICP. It saw two titles that shared words and pattern-matched them together. We had been emailing VPs of Construction Engineering, which &#8212; for a developer tools company &#8212; is absolutely insane. It&#8217;s safe to say that we weren&#8217;t very happy with these results.</p><p>A human with even a shred of context handles this trivially &#8211; if someone on your SDR team was reaching out to construction companies regularly, you would seriously wonder if they were qualified for the role. But the agent had been designed the way most agents are designed today: take a task a human does, write down the steps, and hand each one to an LLM, then chain the results together. This step looked something like &#8220;compare the title to the list of accepted titles&#8221; &#8211; but without any context. The result is a workflow that works when everything fits the expected pattern and falls apart the moment anything requires judgment or is mildly under-specified.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!40sp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!40sp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9403794,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/194431367?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!40sp!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1edf49d9-fc50-4af4-8010-26b3956290d5_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>This is the dominant approach to building agents right now, and we think it&#8217;s a legitimate anti-pattern. Most agents are built to copy-paste human workflows into automated pipelines. That&#8217;s destined to fail, because the things that make humans good at a task &#8212; context, judgment, the ability to notice when something feels off &#8212; are exactly the things that get lost in translation. The best agents aren&#8217;t built to act like humans. They&#8217;re built around a clear understanding of what agents do well that humans don&#8217;t, and vice versa.</p><h2><strong>The interactive agent wins</strong></h2><p>The clearest illustration of this is the difference in approach between coding agents that work interactively and the ones that tried to present themselves as autonomous software engineers. Products like Cursor started with tab completion and narrow task agents. That put the burden on software engineers to break tasks into small chunks, check in on the agent&#8217;s progress, and verify results. The agent did the tedious, high-volume part of code generation; the human brought context, taste, and judgment. Even today, as coding agents have dramatically improved, one of Claude&#8217;s strengths is <em>not trying to do too much</em> &#8211; focus on doing what the user asked, pause, and ask for more input.</p><p>On the other end of the spectrum were products that presented themselves as AI-powered software engineers, like Cognition&#8217;s Devin. You&#8217;d assign a whole task up front, and the agent would figure out the plan, gather context, and execute. The models simply weren&#8217;t ready for that. Without sufficient specification, the agent would regularly run off the rails and produce work that didn&#8217;t accomplish what anyone wanted. The scope was too broad, the feedback loop was too slow, and the human had no good way to course-correct until the damage was already done.</p><p>There&#8217;s a reason that the products that won were optimized for interactive, incremental workflows &#8212; and notably, Cursor and Claude Code do not call themselves &#8220;AI-powered software engineers.&#8221; Anyone who&#8217;s worked in software knows there&#8217;s far more to the job than code generation. Coding agents have made engineers who adopt them dramatically more productive, but that productivity comes from the collaboration pattern, not from replacing the engineer.</p><h2><strong>The same pattern shows up everywhere</strong></h2><p>This isn&#8217;t unique to coding. Consider slide generation &#8212; one of the most frustrating agent experiences out there. We hate making slides and would love it if an agent could take it off of our hands. Unfortunately, it just doesn&#8217;t work because the dominant approach is to hand an agent a prompt and get back a finished deck. The result is almost always unusable. The layouts are generic, the content is vaguely relevant but never quite right (and usually a mangled version of what you put in), and by the time you&#8217;ve fixed everything that&#8217;s wrong with the output, you might as well have started from scratch.</p><p>Contrast that with working Claude&#8217;s integration inside the Powerpoint app. The collaboration mode is that the agent makes incremental improvements based on your feedback. Change this layout. Tighten this bullet. Make this chart clearer. That workflow actually produces something useful, because the human is driving the decisions that require taste &#8212; what story the deck tells, what level of detail matters, what the audience cares about &#8212; and the agent is handling the mechanical execution that&#8217;s tedious and slow for a human. That balance of responsibilities is critical to creating useful outcomes.</p><p>This principle is incredibly important. It&#8217;s one thing if you&#8217;re doing a trivial task like summarizing a Slack thread and creating a ticket for the engineering team to execute on &#8211; you don&#8217;t need human taste and judgment most of the time. But when it comes to cold outbound into a selective customer pool, making customer-facing slide decks, or programming, the value of human taste has not diminished in the slightest. If anything, when the cost of generating most work is zero, results that have good taste stand out more than ever.</p><p>This is what separates agents that people actually rely on from the ones that get tried once and forgotten. The difference has a name &#8212; or at least, we think it should.</p><h2><strong>What it means to be agent-native</strong></h2><p>Being agent-native means building agents that effectively automate the tedious parts of humans&#8217; jobs while letting human creativity and judgment shine. It means starting from a precise understanding of what your agent is uniquely good at and what your users are uniquely good at, and designing the interaction to take advantage of both. It means being disciplined about what your product should do and what it should absolutely not be doing.</p><p>That sounds simple, but it&#8217;s actually the opposite of how most agents get built. The default instinct is to automate as much as possible &#8212; to position the agent as a replacement for the human workflow rather than a tool that reshapes the workflow around each party&#8217;s strengths. The result is agents that try too hard to be like humans: they take on too much scope, they lack the context that makes human judgment valuable, and they give users too few opportunities to steer.</p><p>Taking a step back, this instinct is understandable. For as long as we&#8217;ve been building software, the goal has been to do what humans do, but faster. Email sequencing tools automated tedious button-pressing. Tab complete helped engineers type what they were going to type anyway. Those tools worked because the task was narrow and the right answer was obvious &#8212; there was no judgment involved, just speed. The mistake is assuming that the same philosophy scales to agents operating over tasks where context and taste actually matter. When you ask an agent to write a cold email to a prospect or generate a slide deck for a board meeting, you&#8217;re not asking it to do the same thing faster. You&#8217;re asking it to make decisions that require context it doesn&#8217;t have.</p><p>The best agent builders we&#8217;ve talked to share a different philosophy. They&#8217;ve understood that modern LLMs have an incredibly deep but somewhat narrow intelligence, and they&#8217;ve thought carefully about where the handoff between agent and human should happen. They&#8217;ve designed their products so that when things go wrong &#8212; and things always go wrong &#8212; the human can course-correct quickly rather than discovering the problem after the agent has gone down a dead-end path. They&#8217;ve built trust through narrow, reliable capabilities before expanding scope.</p><p>When you build something that&#8217;s agent-native, you have the opportunity to genuinely delight your customers &#8212; not because the agent does everything, but because it does the right things exceptionally well and makes its users better at everything else.</p><h2><strong>The Agent Native series</strong></h2><p>We&#8217;re spending time talking with founders and product leaders at agent-native companies across a range of product verticals. The Agent Native series will share what we&#8217;ve learned from those conversations &#8212; how the best builders think about the division of labor between agents and humans, how they design for trust, and what they&#8217;ve learned about where agents should and shouldn&#8217;t operate.</p><p>Over the next few weeks, we&#8217;ll be publishing posts in collaboration with each of these teams. We&#8217;ll share a little bit of background about the problem space, their underlying product philosophy, and what lessons we&#8217;ve taken away for our own day to day work. Stay tuned!</p>]]></content:encoded></item><item><title><![CDATA[Agents can't check their own work]]></title><description><![CDATA[How agents made creation cheap and validation impossible]]></description><link>https://frontierai.substack.com/p/agents-cant-check-their-own-work</link><guid isPermaLink="false">https://frontierai.substack.com/p/agents-cant-check-their-own-work</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 09 Apr 2026 18:44:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ibdf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine you ask Claude to build you a financial model in Excel. The spreadsheet you get back probably has the right structure, reasonable assumptions, and formulas that link together in a way that looks correct. But now you have to check it. Do you open every cell and inspect every formula? If you choose to do that, you might as well have built the model yourself. If you ship it as-is, you&#8217;re putting your trust in a junior employee who works at superhuman speed but might have mistakenly encoded some very strange assumptions that didn&#8217;t stand out at first glance.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ibdf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ibdf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/183cface-485b-4619-91ca-e7305189534d_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9071678,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/193719714?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ibdf!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F183cface-485b-4619-91ca-e7305189534d_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini. </figcaption></figure></div><p>The core challenge we face is how we can efficiently validate the work that an agent has done. This problem is lost amongst all the discourse around agents curing cancer and taking jobs. Independent of any future advances in model and agent quality, it is already true today that coding agents, document generators, and AI co-workers have made work cheap. What they <em>haven&#8217;t</em> done is make it any easier to know whether the thing that was created is actually right. We&#8217;re not talking about code quality or best practices &#8212; we&#8217;re talking about the basic question of whether the output does what you intended. And right now, the answer is that most people don&#8217;t have a good way to check.</p><p>The bottleneck has moved. It used to be creation; now it&#8217;s validation. And if you&#8217;re not careful, you&#8217;re just replacing one form of slow, expensive work with another.</p><h2>Why validation is so hard</h2><p>The obvious explanation is that validation takes time &#8212; someone has to review the output, and at the volume agents produce, there simply aren&#8217;t enough hours in the day to do that. The rule of thumb used to be that one person might manage 5-7 employees &#8211; now you can manage as many as you can keep track of. Except those employees start from a blank slate with every task and only get better when new model updates are released.</p><p>That&#8217;s painful enough, but it&#8217;s perhaps not the deepest problem with validation. The more insidious challenge is that most people can&#8217;t validate agent output even in principle, because they don&#8217;t have enough clarity about what they wanted in the first place.</p><p>Think about how UI work used to happen before LLMs. A designer would spend days thinking through a feature &#8212; mapping out interactions, edge cases, and state transitions. They&#8217;d produce detailed mocks and hand them to an engineer. When the engineer delivered an implementation, validation was straightforward: does this match the mocks? The mocks were a checklist. If the designer had missed something, the engineer would probably realize that during implementation. The designer had done much of the hard thinking about UX upfront, and that clarity made it possible to evaluate the result quickly and confidently.</p><p>Now imagine you&#8217;re building the same feature, but instead of going through that process, someone types a loose description with missing details into a coding agent and gets a working UI back in minutes. There are no mocks to check against. The agent made dozens of decisions about interaction patterns, edge cases, and visual hierarchy that nobody specified. The output looks complete &#8212; it runs, it renders, you can click through it &#8212; but the completeness is an illusion. The agent didn&#8217;t resolve any ambiguity thoroughly; it just papered over it with plausible defaults. And now you&#8217;re stuck trying to figure out whether those defaults were good ones, without any reference point for what &#8220;good&#8221; was supposed to look like.</p><p>Before agents, ambiguity in your thinking got resolved <em>during</em> the work. When you mocked features or wrote code, you encountered edge cases and made decisions about them as you went. When you built a financial model, you were forced to confront your own assumptions cell by cell. The work itself was a forcing function for clarity. Agents have removed that forcing function, and nothing has replaced it yet.</p><p>This is why clarity of intent has become so critical &#8212; both because it makes agents produce better output and because without it, you can&#8217;t evaluate the quality of what you get back. If you knew exactly what you wanted, validation becomes something like a checklist: did the agent match these assumptions, handle these edge cases, produce these outputs? You can move fast. But if you went in with a vague sense of direction and let the agent make a hundred small decisions on your behalf, you&#8217;re now stuck trying to reverse-engineer whether each of those decisions was a good one. If you don&#8217;t start with clear thinking, it can often be harder than doing the work yourself.</p><p>The analogy we keep coming back to is delegating to an extremely intelligent college grad with opaque judgment. Whether an agent has good or bad judgment is subjective, but what matters is ensuring that the agent understands exactly what it&#8217;s supposed to do &#8211; and executes on it. When it&#8217;s only humans working on something, the back and forth between a group of people helps discover and test assumptions in the specification. Now, the model is going that all internally &#8211; the end result might be great, but not great in the way that you need.</p><h2>The tools aren&#8217;t ready</h2><p>Even for people who do have clarity of intent, the tools and workflows we rely on aren&#8217;t designed for agent-speed validation. Think about the financial model example again. You know what assumptions the model should encode. You know what the outputs should look like for a few test cases. But spreadsheets don&#8217;t give you a way to express those expectations and check them programmatically &#8211; or frankly even quickly for that matter. You&#8217;re stuck eyeballing formulas or plugging in test numbers manually, essentially doing the same work you always did.</p><p>The same pattern shows up in software engineering, where it&#8217;s even more acute. Previously, you thought about edge cases and validated your assumptions <em>while</em> writing code. That consideration happened incrementally, spread across hours or days of work. Now an agent collapses all of that into a single output, which means all of the validation that used to be distributed across the creation process gets compressed into the review phase. The result is that code review &#8212; which used to be a manageable check on largely human-produced work &#8212; has become the load-bearing wall of software quality. There&#8217;s no chance it lasts.</p><p>There&#8217;s no incremental testing when an agent does a task, because it delivers everything at once, almost instantaneously. If you skip the validation step or do it halfheartedly, you&#8217;re consigning yourself to shoddy results that you won&#8217;t discover until something breaks downstream. And if you do the validation step thoroughly, you&#8217;ve burned most of the time you saved by using the agent in the first place.</p><h2>The interaction mode has to change</h2><p>We&#8217;ve established that validation is more necessary than ever and also that it&#8217;s harder than ever. We&#8217;ve also established that existing modes of validating work aren&#8217;t going to work as we dramatically scale how much there is to check. This creates a question we haven&#8217;t seen anyone answer well yet: How should validation actually work?</p><p>Consider the financial model example again. You know what assumptions the model should encode, and you know what the outputs should look like for a handful of test cases. But the way you interact with a spreadsheet today is by opening cells and mucking around with value &#8212; a mode designed for a world where a person carefully built the model and you&#8217;re spot-checking their work. That mode collapses when an agent built the entire thing at once and every cell is a potential surprise.</p><p>What you actually want to do is interrogate the model. Ask it questions &#8211; what happens to revenue if churn doubles? What&#8217;s driving the margin assumption in Q3? If the answers match your expectations, you build confidence. If they don&#8217;t, you&#8217;ve found the problem without having to reverse-engineer deeply linked formulas. The validation step becomes a conversation with the output, not an inspection of its internals.</p><p>Not every validation step is going to look like a conversation, but similar principles apply to code, documents, or designs. When an agent generates a complete feature, you need to see whether the UX that was implemented matches what you had in mind. When an agent writes a report, you need to understand what the thesis of the document was and how that argument was validated. The UX of agent-assisted work needs to be built around this kind of fast, targeted verification, not around the line-by-line review workflows we inherited from a world where humans did the creating.</p><p>Today, almost none of our tools support this. Spreadsheets don&#8217;t let you express expectations and test them. Code review tools are built for diffing human-authored changes, not for interrogating agent-generated systems (where PRs are quickly ballooning in size). Document editors assume you&#8217;ll read and redline, not ask and verify. The interaction mode hasn&#8217;t caught up to the speed of creation, and until it does, validation is going to remain the thing that eats all the time agents save.</p><h2>What this means</h2><p>We think this is among the most important unsolved problems in AI tooling right now &#8211; and it&#8217;s barely even being discussed. The industry has invested enormously in making generation faster and cheaper, and it&#8217;s worked. But generation without validation is just moving the bottleneck, not eliminating it. People are going to be overwhelmed by agent output, and they&#8217;re either going to be disappointed by the quality of what ships or they&#8217;re going to stop checking and be surprised when things break.</p><p>The UX of AI-assisted work has to change to match this reality. It can&#8217;t just be about faster creation; it has to be about faster verification. That might mean agents that validate their own output against explicit expectations. It might mean production tooling that watches for the consequences of bad agent work in real time. It will almost certainly mean both, and probably things we haven&#8217;t imagined yet.</p><p>But one thing is clear: Without solving validation, we haven&#8217;t actually solved the productivity problem. We&#8217;ve just moved it.</p>]]></content:encoded></item><item><title><![CDATA[AI agents shouldn't have a job title]]></title><description><![CDATA[Agents are better when they stop pretending to be people]]></description><link>https://frontierai.substack.com/p/ai-agents-shouldnt-have-a-job-title</link><guid isPermaLink="false">https://frontierai.substack.com/p/ai-agents-shouldnt-have-a-job-title</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 02 Apr 2026 18:20:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iL-S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s a mental model we keep coming back to: humans are deep and vertical, agents are shallow and wide. A human owns a role end to end &#8212; the judgment, the context, the edge cases, the cross-team coordination that never makes it into a job description. An agent, by contrast, handles the learnable tasks that exist across multiple adjacent roles simultaneously. The shared context across those roles can actually make it better at each one. However, agents often fail at the messy edge cases that take up lots of human time</p><p>This distinction sounds simple, but it has profound implications for how you build agents &#8212; and how you talk about them. The right way to design an agent is to start from what humans struggle to do at scale: process large volumes of data without fatigue, consider many possibilities in parallel, or find an interesting detail to personalize sales outreach.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iL-S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iL-S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9208357,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/192989070?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iL-S!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48734e44-4b51-4816-bd2c-7d0f6ba640d7_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>This is a critical distinction in how you explain, build, and iterate on your product. You can choose to focus on cooperating with humans &#8211; an agent designed around what humans can&#8217;t do and built to hand off gracefully when things get messier than expected. Or you can choose to copy-paste a human job description into a product spec with a rigid set of requirements, something that will inevitably disappoint customers with mismatched expectations.</p><p>The AI industry has, almost universally, chosen the second path. You can&#8217;t take a step in San Francisco without seeing a billboard for an AI SDR, an AI SOC analyst, or an AI SRE. We understand why &#8212; customers are searching for these terms, and if your website doesn&#8217;t speak their language, you&#8217;ll lose the SEO battle before you get to make your pitch. But we think the race to anthropomorphize agents is setting the entire industry up to fail.</p><h2><strong>Why job-title naming creates a trust crisis</strong></h2><p>When you name your agent after a job title, you&#8217;re making an implicit promise: that your product can do what that person does. That includes everything that never appears in the formal job description.</p><p>Take an AI SDR. The pitch is intuitive: identify the right prospects, send them the right message, and generate pipeline. In our experience, agents are genuinely good at account research, ICP matching, and identifying the right person to contact at a given company. Where they fall short is the second part. What a customer buying an AI SDR expects is to turn it on, feed it some messaging, and come back to a healthy pipeline. What they mostly get is good targeting attached to generic outreach that doesn&#8217;t convert.</p><p>The gap is the messy bits that humans often intuit well. Good outbound sales requires creativity and the ability to build rapport with a stranger in a single sentence. Humans who are good at sales have spent years developing that instinct. Agents, today, produce messaging that sounds automated, and we&#8217;re all inundated with that nonsense. It&#8217;s easy to imagine an agent that does all the research and hands off to a human to get the prospect over the line. But that&#8217;s not the expectation the name set.</p><p>We could make this argument across every anthropomorphized agent market. In our own experience, the phrase &#8220;AI SRE&#8221; is frankly a poor description of what our product actually does &#8212; but it&#8217;s what customers are searching for, which creates a genuine tension we&#8217;ll come back to.</p><h2><strong>What good naming looks like</strong></h2><p>The most instructive counterexample is coding agents. Cursor didn&#8217;t call itself an AI software engineer. It started off as an AI-powered IDE &#8212; a tool that generates code effectively across a complex codebase. They didn&#8217;t set the expectation of being an autonomous engineer that ships features end to end.</p><p>This wasn&#8217;t just semantic humility. The people building these products had a good understanding of software engineering to know that everything a software engineer did was not immediately in scope for automation. Humans would still be responsible for the architectural decisions, the cross-functional coordination, the customer conversations that shape what gets built.</p><p>By scoping the promise to a task &#8211; generating code &#8211; rather than the role, they gave themselves the opportunity to match (and beat!) users&#8217; expectations. Engineers who adopted these tools didn&#8217;t feel like they were being replaced. They felt like they were getting faster at the part of their job they liked least. Over time, trust built through thousands of small interactions where the tool did what it said it would do. That trust is what eventually opened the door to agents taking on more.</p><p>It&#8217;s not a coincidence that coding agents have the deepest adoption, the strongest data flywheels, and the most widespread acknowledgment of quality in the entire AI application landscape.</p><h2><strong>What agents should actually be built for</strong></h2><p>The deeper issue with anthropomorphizing isn&#8217;t just the naming &#8212; it&#8217;s that job-title thinking constrains product design in ways that limit what agents can actually become.</p><p>Most agent products today are built around a simple idea: observe what a human does, then automate it. For an AI SRE, most products are asking their customers to write a runbook for every alert type. Then they build an agent that follows the runbook, posts a summary in Slack, and pretends to have capabilities that auto-remediate an issue. This approach produces passable results but creates tons of process overhead that users quickly learn to ignore &#8211; and it doesn&#8217;t create the opportunities for users to actually trust the end-to-end resolution workflow. This is marginally better than the human workflow it replaces, but you&#8217;ve left most of the agent&#8217;s actual capabilities on the table.</p><p>The more productive question to ask is what can an agent do that a human fundamentally can&#8217;t? Not something that an agent is slightly faster or cheaper at &#8212; what is truly different about an agent&#8217;s capabilities? Continuing the AI SRE analogy, an agent can monitor thousands of signals simultaneously without fatigue. It can run dozens of hypotheses in parallel and discard the ones that don&#8217;t pan out. It can surface patterns across an entire customer base that no individual human would ever accumulate enough exposure to recognize.</p><p>This realization meaningfully changed how we&#8217;re building RunLLM. Rather than following fixed alert thresholds the way a human on-call engineer would, we detect incidents by analyzing the frequency and content of observability data to surface anomalies before they&#8217;ve been formally flagged. Rather than working through a runbook linearly, we spin up multiple subagents in parallel to evaluate many hypotheses simultaneously. And rather than claiming that we&#8217;re going to automate end-to-end resolution, we give users the option to pick the handoff point that they prefer. The agent does not try to map a human workflow; it focuses on automating things that humans can&#8217;t do and letting humans handle the rest. That&#8217;s why our product delivers outcomes that a copy-pasted human workflow never could.</p><h2><strong>Where this is all going</strong></h2><p>The implication of this argument extends beyond product design and naming. If agents are genuinely shallow and wide &#8212; good at the learnable, scalable slice of many adjacent roles &#8212; then the natural endpoint isn&#8217;t one agent per job function. The boundaries between traditional roles will naturally become less fixed over time. Humans already manage multiple agents across overlapping domains, and that will only increase.</p><p>A software engineer who uses an SRE agent to monitor production isn&#8217;t doing an SRE&#8217;s job. They&#8217;re doing their own job better, with more visibility. An SRE who uses a coding agent to patch a bug during an incident isn&#8217;t moving into software engineer. They&#8217;re resolving the incident faster. The agent isn&#8217;t replacing either person &#8212; it&#8217;s dissolving a boundary that only existed because humans have limited bandwidth.</p><p>The companies that stop asking &#8220;which human job does this agent replace?&#8221; and start asking &#8220;what can this agent do that the human couldn&#8217;t?&#8221; are going to build products that are genuinely differentiated, not just incrementally better. The rest will keep fighting over SEO terms while quietly underdelivering on the promises those terms imply.</p>]]></content:encoded></item><item><title><![CDATA[AI Companies Are Building for the Wrong Users]]></title><description><![CDATA[Why building for your biggest fans might not work forever]]></description><link>https://frontierai.substack.com/p/ai-companies-are-building-for-the</link><guid isPermaLink="false">https://frontierai.substack.com/p/ai-companies-are-building-for-the</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 26 Mar 2026 18:29:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YJVx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A close friend is a software engineer at one of the most established technology companies in the world &#8212; a perennial member of whatever your favorite big tech acronym is. They&#8217;ve been there for close to a decade, and recently, they&#8217;ve been working to boost their own productivity using an internal AI coding agent. As they&#8217;ve shared what&#8217;s working and where the limitations of the agent are within their team, interest quickly spread in their broader org. They ended up giving a talk to hundreds of engineers inside the company.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YJVx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YJVx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:10279758,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/192235582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YJVx!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b5480e2-cbc2-409c-bdca-9fa8b740baf5_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>We were fascinated to hear about their experience. Despite being an engineer at one of the most technically sophisticated companies in the world &#8211; one that builds its own foundation models! &#8211; they encountered significant and widespread reluctance to adopt coding agents. Some engineers simply liked writing code by hand &#8212; it was part of how they thought about their craft. Others couldn&#8217;t figure out how the tool was supposed to fit into their workflow. Still others had tried it, found the UX confusing, and quietly moved on without sharing any feedback.</p><p>This isn&#8217;t a story about backwards-thinking people who aren&#8217;t ready for change. These are excellent engineers at a company that&#8217;s been at the center of technology for decades &#8211; if you saw this name on a resume, you would immediately consider interviewing the person. And yet, for most of them, AI coding tools remain something they&#8217;ve heard about but haven&#8217;t really integrated into how they work.</p><p>If that&#8217;s true there, imagine how much of a challenge this is everywhere else.</p><h2>The bubble is smaller than you think</h2><p>If you&#8217;re reading this blog, you&#8217;re almost certainly deep in the AI bubble. You&#8217;ve probably used Claude Code or Cursor to build something from scratch on a weekend and told everyone you know about it. You&#8217;ve definitely seen the LinkedIn posts &#8212; the ones breathlessly explaining that a repo with Claude Code skills for go-to-market is going to revolutionize the way you do sales. We get it. We write this blog, after all.</p><p>The problem is that spending time inside the bubble makes it easy to mistake the bubble for the world. The users who are most excited about AI &#8212; the ones tolerating rough edges, experimenting constantly, building their own workflows &#8212; are a small fraction of the people who are eventually going to need to use your product. And crucially, their enthusiasm can paper over a lot of gaps. If someone is genuinely excited to use your product, they&#8217;ll push through a confusing onboarding flow, forgive a clunky UX, and figure out the right prompting pattern through trial and error. You can build a real business on these true believers for the first little while &#8211; and they&#8217;re genuinely wonderful people to work with.</p><p>But you will eventually have to account for the rest. And we think most AI companies aren&#8217;t building for them.</p><h2>Why the gap is a product problem, not a people problem</h2><p>The instinct when faced with reluctant users is to assume they&#8217;ll come around &#8212; that adoption is just a matter of time, and the product doesn&#8217;t need to change if its working well for the true believers. AI adoption has a different quality from previous technology changes, and the reason is rooted in something more fundamental than UX.</p><h3>User identity</h3><p>AI tools attack user identity in a way that most software doesn&#8217;t. Think about what it means to be a software engineer who has spent years taking pride in the quality of the code they write. Now an agent is generating code at 10x the speed, and most of it looks like slop to them. For many senior engineers, it&#8217;s probably true &#8211; they could write better code if they took their time. More importantly, that code is an affront to how they see themselves professionally. They&#8217;re never going to merge this code without a close inspection, no matter how many benchmark improvement graphs you put in front of them.</p><p>But here&#8217;s the flip side, and it&#8217;s where the real product opportunity lies: The best AI products have the opportunity to actively <em>reinforce</em> user identity too. Done right, your product doesn&#8217;t make the careful engineer feel replaced; it makes them the most forward-thinking engineer on their team. It doesn&#8217;t make the skeptical PM feel like they&#8217;re cutting corners; it makes them the person who shipped a working prototype before anyone else had finished writing the spec. The goal isn&#8217;t to neutralize the identity question. It&#8217;s to make your user the superstar of their organization, in a way that feels consistent with who they already are.</p><p>Getting there requires meeting users where they are and bringing them along gradually. The coding agent world offers the clearest model of how this works in practice. The trust that made widespread agent adoption possible wasn&#8217;t built by agents &#8212; it was built by tab completion. Engineers who would never have handed off a function to an autonomous agent were perfectly comfortable accepting a 1-5 line suggestion. That comfort, built up over thousands of small interactions, is what created the foundation for people to trust agents with more. It&#8217;s not a coincidence that the products with the deepest agent adoption today are the ones that also had the best tab completion two years ago.</p><h3>Editorial judgment</h3><p>The corollary to meeting users where they are is not overwhelming them with everything at once. This has always been true in product, but AI has made it a more acute problem. The ability to ship new features quickly &#8212; which AI has genuinely accelerated, and AI products are taking full advantage of &#8212; makes it easier than ever to pile on functionality before users have developed trust in the core experience.</p><p>A reluctant adopter who opens your product and finds a sprawling set of capabilities they don&#8217;t understand isn&#8217;t going to dig in and figure it out. They&#8217;re going to confirm their suspicion that this thing isn&#8217;t for them, and they&#8217;re not coming back for a long while. The same logic applies at a feature level: if a user is still getting comfortable with tab completion, immediately introducing an autonomous agent mode isn&#8217;t going to accelerate their adoption &#8212; it&#8217;s going to spook them. Doing less, more reliably, for the right person at the right time is how you earn the trust that eventually lets you do more.</p><p>This concern applies in many other domains. We see every day that most engineering leaders are only now coming around to the idea of having an AI SRE help them manage production software &#8211; and while these folks are figuring out what they want, vendors are promising that they&#8217;ll autonomously make changes to your production infrastructure while you&#8217;re still asleep. That might sound exciting to the most ardent believer, but it sounds like absolute insanity to most people. Most products in the space haven&#8217;t figured out the balance yet.</p><h2>The maturation moment</h2><p>There&#8217;s a useful historical parallel here. In the early days of cloud computing, the customers were startups and tech-forward companies who were willing to figure things out as they went. AWS in 2008 was powerful but rough, and the people using it had enough technical sophistication and tolerance for ambiguity to make it work. Then came the maturation phase: large enterprises, regulated industries, banks, healthcare systems. These customers had different expectations, different risk tolerances, and different definitions of &#8220;good enough.&#8221; The cloud providers that thrived were the ones that understood the product had to change &#8212; not just get cheaper or faster, but fundamentally adapt to a different kind of user.</p><p>AI is entering the same phase. The early adopters &#8212; the ones giving talks to thousands of engineers, the ones building apps on weekends and posting about it &#8212; have gotten the industry here. The next phase of growth is the middle 50%: the users who are curious but skeptical, willing to try but quick to bounce, not interested in figuring out the right prompting strategy on their own. That is the audience that&#8217;s going to determine whether AI actually delivers on the productivity promises that have been made on its behalf.</p><p>What makes this moment unusual is that unlike the cloud, AI seems to be compressing the early adoption and maturation phases together. The gap between <em>this is a tool for enthusiasts</em> and <em>this needs to work for everyone</em> is closing faster than anyone expected. For product builders, that&#8217;s both an opportunity and a warning. The companies that figure out how to cross the chasm early are going to have a significant head start on the ones that are still optimizing for the bubble when the rest of the world shows up.</p>]]></content:encoded></item><item><title><![CDATA[Generic enterprise AI is indefensible]]></title><description><![CDATA[Why foundation models will own enterprise intelligence]]></description><link>https://frontierai.substack.com/p/generic-enterprise-ai-is-indefensible</link><guid isPermaLink="false">https://frontierai.substack.com/p/generic-enterprise-ai-is-indefensible</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 19 Mar 2026 18:08:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!m8Yz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There are a lot of companies right now trying to establish themselves as the intelligence layer for the enterprise. The pitch is appealing: large companies generate enormous amounts of data across dozens of SaaS tools, struggle to know what lives where, and need something that can reason across all of it at the right moment. The holy grail is an agent that makes sense of everything that you have and&#8230; just figures it out. The agent should pull the right data, analyze it in the right way, and even resolve different identifiers across systems. Any given task might require data across sales, marketing, product, and engineering, and the agent should be capable of handling that complexity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!m8Yz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!m8Yz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:731600,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/191500370?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!m8Yz!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae609bf5-b0ba-4b49-aa60-ef6ea42cf2a0_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>While the promise is appealing, the reality is not. We think the companies building in this space are walking into a trap &#8212; and it&#8217;s worth understanding why.</p><h2>The promise of horizontal intelligence</h2><p>Imagine a finance team trying to assess the risk profile of outstanding invoices. Correctly determining what the aggregate risk of these unpaid invoices means pulling together bank data, invoicing tools, customer contracts, historical payment records, and the latest usage data. Doing that by hand is tedious, error-prone, and slow &#8212; exactly the kind of workflow that a well-built agent should be able to automate. And any given company has tens or hundreds of these tasks a day, each one of which is bespoke and repeats only rarely. It&#8217;s obvious why you&#8217;d want to build a startup around this problem.</p><p>The horizontal intelligence pitch is essentially this: We&#8217;ll connect to all your data, understand who you are and what you have access to, and surface the right information at the right time. That&#8217;s a genuinely useful tool, and we&#8217;re not arguing otherwise.</p><p>What we&#8217;re arguing is that it&#8217;s an extremely difficult place to build a defensible startup.</p><h2>Why horizontal intelligence is a trap</h2><p>The risk that&#8217;s becoming increasingly hard to ignore is that the foundation model providers are coming for this space themselves &#8212; and they have a built in head start.</p><p>Each of the major providers has a slightly different wedge. Google owns Workspace, and while the Gemini integration isn&#8217;t yet great, odds are that it will improve significantly. OpenAI, through their apps framework, is increasingly able to integrate across a diverse set of tools and has built a developer framework that encourages application providers to do the heavy lifting of integration themselves. But it&#8217;s arguably Anthropic that&#8217;s furthest along: the Claude desktop app is genuinely impressive, and the integration of Claude Code and Cowork into the local file system &#8212; editing documents, analyzing data, generating artifacts &#8212; clearly indicates where this is going.</p><p>The uncomfortable question for horizontal intelligence startups is <strong>why will your product handle enterprise data better than the major model providers?</strong> We have yet to hear a compelling answer to that.</p><p>There are two interesting things about this kind of task that give the model providers a clear advantage. First, there&#8217;s relatively little specialization to this kind of intelligence. Every task is unique, and it&#8217;s often the case that once you have the right data, it&#8217;s relatively easy to solve the problem &#8211; the data is just all spread all over the place. Second, pre-built templates like slide formats and document structures are starting to become standard agent inputs. Once these are agreed on, all that&#8217;s left is data connectors, search, and general intelligence &#8211; all things that the model providers have in spades.</p><p>The products that hold the actual underlying data are going to feel increasing pressure to make it available to the model providers&#8217; agents &#8212; or risk being displaced by something that is more LLM-friendly. That&#8217;s not a dynamic that favors a startup sitting in the middle.</p><h2>Why specialized intelligence wins</h2><p>None of this means there&#8217;s no room for startups in the agent space. Quite the opposite &#8212; we think there&#8217;s a significant opportunity. It&#8217;s just not in the horizontal middle ground between model providers and data stores. The opportunity exists in workflows that require specific expertise and customization that are genuinely difficult to recreate without encoding that intelligence into the architecture of a purpose-built agent.</p><p>The reason specialized vertical agents can win where horizontal agents can&#8217;t boils down to a few related ideas.</p><p>First, when an agent makes a strong set of assumptions about the workflow it&#8217;s solving, it becomes much easier to encode intelligence into how the agent actually operates. There&#8217;s a meaningful difference between &#8220;summarize this document&#8221; and &#8220;execute this specific workflow for this engineering team in this company.&#8221; The former is something a general-purpose model will get better at over time. The latter requires a level of domain-specificity that can&#8217;t be learned from generic usage data &#8211; a combination of company-specific context, general engineering principles, and the ability to learn and improve over time.</p><p>Second, that narrower focus allows you to solve increasingly hard problems. We&#8217;ve all had the experience of asking a general-purpose agent to do something complex, only to find that it didn&#8217;t quite understand what we wanted and produced something generic and likely useless. A specialized agent can enforce much stronger guardrails on what gets generated &#8212; it knows what good looks like, and it can be built to enforce that rather than relying on general-purpose reasoning and hoping for the best.</p><p>Third, and most importantly, this specialization creates a data flywheel that the model providers can&#8217;t easily replicate. The model providers have enormous amounts of generic usage data. But in order to bring that data to bear on complex, domain-specific workflows, they also need to know what a good outcome looks like &#8212; and for enterprise-specific workflows, that ground truth simply isn&#8217;t available to them at scale. The more complex the task, the harder it is for a foundation model to build useful generic intelligence around it. That&#8217;s where the moat lives.</p><p>We&#8217;ve written before that <a href="/__u/frontierai.substack.com/p/data-is-your-only-moat?utm_source=publication-search">data is the only moat available to most AI startups</a>. Vertical agents are the clearest expression of that thesis. The data flywheel a vertical agent builds takes two forms: a large volume of broadly applicable data about one particular kind of workflow (which generalizes across customers), and deep data about how one specific company does that workflow (which makes the agent increasingly irreplaceable within that account). Both are difficult for a generalist to compete with.</p><h2>Picking a lane</h2><p>Our conviction, built from a few years of building in this space, is that the companies that are going to win are the ones that pick a lane and go deep. The less an LLM can solve your problem off the shelf, the better your defensibility over time.</p><p>That&#8217;s exactly why we&#8217;ve focused on what we&#8217;ve called hard-hard problems &#8212; the enterprise workflows that are technically complex and organizationally difficult to adopt. These markets are daunting. In many of them, startups have raised significant capital without anyone running away with the space yet. That&#8217;s not a sign that the problems can&#8217;t be solved. It&#8217;s a sign that the model providers can&#8217;t solve them generically, and that the engineering required to do it well is genuinely hard. That difficulty is exactly the moat you want.</p><p>If you&#8217;re building an agent today and you find yourself describing your product as an intelligence layer, or as a way to connect enterprise data to a model, we&#8217;d encourage you to stress test that framing. The foundation model providers are building the same thing &#8212; and they have more data, more compute, and more distribution than you do. The only place that argument doesn&#8217;t apply is in the workflows that are too complex, too specialized, and too company-specific for them to tackle. That&#8217;s where we are building.</p>]]></content:encoded></item><item><title><![CDATA[The SaaS Extinction Test]]></title><description><![CDATA[Why your vibe-coded weekend project won't kill Salesforce]]></description><link>https://frontierai.substack.com/p/the-saas-extinction-test</link><guid isPermaLink="false">https://frontierai.substack.com/p/the-saas-extinction-test</guid><dc:creator><![CDATA[Vikram Sreekanti]]></dc:creator><pubDate>Thu, 12 Mar 2026 17:09:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KPzg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Projections about the demise of the SaaS industry have reached a fever pitch over the last few weeks. The general belief seems to be increasingly that there&#8217;s no moat in software anymore. Consequently, we have seen the stock prices of massive, entrenched incumbents take a significant beating.</p><p>The bear case for software-as-a-service goes something like this: As coding agents improve, the build-versus-buy calculus shifts permanently. An engineer at any company can now pick up a coding agent, build a prototype that meets the company&#8217;s specific needs, iterate on feedback, and deploy a bespoke tool that provides more value than a generic SaaS platform at a fraction of the cost.</p><p>The source of this concern is our collective recent experience with the improving quality of coding agents. We have all had Claude Code or Cursor scaffold an entire application from scratch for a few dollars in tokens. When you see something genuinely useful built in minutes, it is easy to assume the multi-billion-dollar incumbents are doomed. And if an individual can build a prototype quickly, surely a startup can build a strong offering with a few months and a few engineers?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KPzg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_webp, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KPzg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1558029,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://frontierai.substack.com/i/190740985?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_424, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_848, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1272, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KPzg!, /__u/frontierai.substack.com/w_1456, /__u/frontierai.substack.com/c_limit, /__u/frontierai.substack.com/f_auto, /__u/frontierai.substack.com/q_auto:good, /__u/frontierai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78724e5e-1c1d-4c29-9646-15a3ff5a63d5_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini.</figcaption></figure></div><p>As with most hype cycles, however, the truth lands in the middle. While there are massive opportunities for disruption, the SaaS business model is not going to evaporate overnight. In our view, the distinction comes down to a quick survival checklist. If you meet one of these criteria, you are likely safe. If you miss on all of them, you should be worried:</p><ol><li><p>Are you a system of record?</p></li><li><p>Do you do more than help humans automate a single workflow?</p></li><li><p>Are you mission-critical?</p></li></ol><h3><strong>The Physics of Data Gravity</strong></h3><p>Data moats have long been the holy grail of enterprise software. Consumer products like Google or Instagram win on user behavior data, and enterprise titans like Snowflake or Datadog are powerful because they are systems of record. These companies do not just power workflows; they house years of historical data that companies have imported and structured.</p><p>Despite all the massive technological shifts in the last few years, the physics of data have not changed. Moving massive amounts of data is expensive, risky, and slow. This means that, as has been the case for fifteen years, the major cloud providers will continue to make their margins on networking and egress. The cost of moving data in and out of a third-party service &#8212; or your own cloud &#8212; remains incredibly high.</p><p>Beyond the operational cost, there is the issue of operational stickiness. If a product is collecting data for a core operational purpose, it becomes a load-bearing wall in the company&#8217;s architecture. You do not pay Datadog or Snowflake eight figures a year because storing logs is a nice-to-have feature. You pay them because when a mission-critical issue happens at 3:00 AM, you need a proven way to investigate, pinpoint, and fix the issue.</p><p>But realistically, your business will continue to exist if Datadog goes down for a little while. Things get much more difficult when we&#8217;re talking about truly mission-critical software.</p><h3><strong>The Mission-Critical Tautology</strong></h3><p>Even when companies don&#8217;t have the most interesting data, they&#8217;re likely safe if they are critical to the operations of their customers &#8211; the kinds of products that your business literally wouldn&#8217;t exist without. The obvious examples are Workday and Salesforce.</p><p>There is a certain recursive logic here: These companies are safe because they are big, and they are big because they are safe. No matter how shiny a rapidly scaffolded payroll system looks, a VP of HR is not going to risk payroll failing on a Friday morning. A CRO does not care how many new automations a startup offers if it means moving away from a Salesforce instance they have spent a decade customizing to their exact sales and revops motions.</p><p>When enterprises make these decisions, they are not just buying software &#8211; they are also offloading risk. A tech-forward company like Google or Meta certainly has the technical talent to build an internal payroll system. The question is not whether they can build it, but whether they want to own the risk of running it. Once they do the calculus, dedicating hundreds of engineers to a non-core area of expertise rarely makes sense.</p><p>By contrast, non-mission-critical software is in the danger zone. Analytics is the prime example. Traditionally, writing code for dashboards and plumbing database queries was prohibitively expensive to democratize.</p><p>Just last week at RunLLM, we built an internal dashboarding system following a conversation about product metrics. We used Python and Plotly to connect to internal systems and pull customer data via our CRM&#8217;s API. If this dashboard goes down for a day, it is annoying, but it does not halt our operations. More importantly, because the cost of building it with a coding agent was so low, and the upside of allowing our VP of Sales to modify it himself was so high, the option to buy a traditional BI tool never even entered the conversation.</p><h3><strong>The Workflow Moat Narrows</strong></h3><p>Pure workflow software is where we&#8217;re currently the most bearish. By workflow software, we mean anything that&#8217;s focused on connecting the dots between existing tools that either have data gravity or mission-criticality. The poster child on X for this is PagerDuty. At its core, PagerDuty processes data stored in another product&#8217;s telemetry store, determines if a threshold was met, and alerts an on-call engineer. This is connecting the dots between Datadog and Slack. While PagerDuty does much more in practice, that primary workflow is what most customers buy.</p><p>These are exactly the kinds of integration code workflows being replaced by agents today. The threat here is not just that agents will automate incident management or sales outbound; it is that new players or internal teams can build superior user experiences that blend human and agent capabilities from the ground up. The most interesting part of this comes from the fact that the solutions can (and should) very easily be customized to match each team&#8217;s workflow. These agent-first systems will be dramatically more valuable because they save human hours, easily displacing legacy systems that are human-first and currently scrambling to bolt on AI features as an afterthought.</p><h3><strong>The Ops Gap</strong></h3><p>The undiscussed concern underlying all <em>build it with Claude Code</em> software is operations. The reality of software operations and reliability is unfortunately still incredibly complex and oftentimes dwarfs the time taken to build something useful. The gap between building and operating is actually growing because as coding becomes a commodity, the relative cost of infrastructure and security grows.</p><p>The internal dashboard we built recently is a perfect example of this. It took about an hour to iterate on the core dashboards themselves with Claude Code. It then took about four hours to deploy it in Google Cloud so that it was properly authenticated behind our company&#8217;s SSO, set up to auto-update with new commits, and connected with the right credentials.</p><p>This gap will likely persist. The simple reason why is that Python looks the same no matter where you work, but infrastructure is dramatically different at every company. There are not yet great ways to deploy cloud software without getting into the weeds of Kubernetes clusters, networking permissions, and security + compliance. None of those things matter when you are prototyping locally, but they are the only things that matter when you are operating production software at scale.</p><h3><strong>Wrapping Up</strong></h3><p>The rumors of the demise of modern software are greatly exaggerated, but the nature of what makes software valuable is quickly changing. For the last decade, SaaS companies could survive by being a slightly better UI for a human workflow. That&#8217;s no longer defensible.</p><p>The defensibility remains exactly where it has always been: at the intersection of data and risk. If you own the system of record, the physics of networking and the high cost of data egress will protect you. If you own a mission-critical operational process, the corporate aversion to risk will protect you.</p><p>The real danger is for the middle layer of the stack &#8212; the products that have built businesses around the friction of human work. As the cost of generating code and automating workflows drops toward zero, the value of that software must move elsewhere.</p><p>This does not mean incumbents are invincible. It just means that the disruption will not come from a simple internal tool or a weekend project built with a coding agent. To win in this new environment, startups must figure out how to best manage the data, reduce the operational risk, and close the gap between a prototype and production-grade reliability.</p>]]></content:encoded></item></channel></rss>