<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Brains & Bots]]></title><description><![CDATA[Written by Cindi Howson, Chief Data & AI Strategy Officer, alongside Sonny Rivera & Jane Smith. We sit at the intersection of enterprise data, practical analytics, and modern AI execution. Unvarnished takes, no fluff & practical execution at scale.]]></description><link>https://brainsandbots.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!LK4T!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbeff929-396b-4d3e-b1d7-4d5384febb8a_500x500.png</url><title>Brains &amp; Bots</title><link>https://brainsandbots.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 12:32:36 GMT</lastBuildDate><atom:link href="/__u/brainsandbots.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Brains & Bots]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[brainsandbots@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[brainsandbots@substack.com]]></itunes:email><itunes:name><![CDATA[Brains & Bots]]></itunes:name></itunes:owner><itunes:author><![CDATA[Brains & Bots]]></itunes:author><googleplay:owner><![CDATA[brainsandbots@substack.com]]></googleplay:owner><googleplay:email><![CDATA[brainsandbots@substack.com]]></googleplay:email><googleplay:author><![CDATA[Brains & Bots]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Desperately Seeking a Single Semantic Layer ]]></title><description><![CDATA[You can&#8217;t trust AI without a good context layer]]></description><link>https://brainsandbots.substack.com/p/desperately-seeking-a-single-semantic</link><guid isPermaLink="false">https://brainsandbots.substack.com/p/desperately-seeking-a-single-semantic</guid><dc:creator><![CDATA[Brains & Bots]]></dc:creator><pubDate>Sat, 22 Aug 2026 12:31:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LK4T!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbeff929-396b-4d3e-b1d7-4d5384febb8a_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Fish and Chips. The iconic dish of any British pub. Also, my daughter&#8217;s favorite birthday dinner! In fact, my English husband personally deep fries the fish and chips for her in an industrial- strength deep fryer that blows our fragile electric out every time.</span></p><p><span>I am in charge of resetting the breakers. And slicing the potatoes.</span></p><p><span>The thing is, in my early years of living in Europe, this British dish got me very confused. Why do they call them chips? Why not french fries?</span></p><p><span>And why are potato chips called crisps in England? When dining in a restaurant, a burger and chips is potato chips, but fish and chips, is, well, french fries.</span></p><p><span>Context  - and culture - is the difference between a correct meal and a mistake, whether a disappointed birthday girl or disgruntled patron.</span></p><p><span>The same is true with context for AI agents.</span></p><div class="callout-block" data-callout="true"><p><span>According to a recent </span><a href="https://venturebeat.com/data/57-of-enterprises-traced-a-wrong-ai-answer-to-missing-business-context-credible-bets-portable-open-source-semantic-code-beats-proprietary-metadata"><span>Venture Beat survey,</span></a><span> 57% of AI-generated insights are wrong due to bad context. Woa!</span></p></div><p><span>Humans are adept at answering ambiguous questions, AI agents are not.</span></p><p><span>Semantic layers are one way AI agents go from probabilistic results to deterministic, trusted insights. But this also depends on how the semantic layer is used and how much context is leveraged. Given the criticality of this in an agentic world, data teams are pursuing semantic layers as a foundation for AI.</span></p><p><span>They are right in treating semantic layers with the utmost importance. But they are wrong in thinking there should be one semantic layer to rule them all.</span></p><h2><span>The Fallacy of One Semantic Layer</span></h2><p><span>I understand why data teams want one semantic layer - it goes back to our industry&#8217;s wish for a single version of the truth. The answer to our quest for one truth, initially, was a data warehouse. Put all the data in one place and users could trust the data! I will spare you the history lesson in spread marts, data lakes, data swamps, and data mesh.</span></p><p><span>The bottom line is that each function has their own version of truth. Also, sometimes business definitions change.</span></p><p><span>When a CFO refers to revenue, the CFO is referring to GAAP revenue, based on invoices, net of returns. In a SaaS world, this is annual recurring revenue.</span></p><p><span>When a sales person refers to revenue, they most often are referring to order value, whether invoiced or not, and excluding professional services as there is no commission here..</span></p><p><span>Both are right. The difference is in context, here based on the person&#8217;s job function. With my food example, the definition of chips changes based on the main course (Fish or Burger) or country (British versus American).</span></p><p><span>In this regard, context must be part of your semantic layer strategy and is why organizations should be pursuing a mesh of layers, federated across the organization.</span></p><h2><span>Where Does Knowledge Exist?</span></h2><p><span>In a workshop </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Francois Lopitaux&quot;,&quot;id&quot;:5579762,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6544b0f0-52b5-413d-8557-0d53195e077d_144x144.png&quot;,&quot;uuid&quot;:&quot;eda5263e-70f1-4b8a-a512-85a28edecfcc&quot;}" data-component-name="MentionToDOM"></span> <span>and I led at the Gartner data and analytics summit, we asked people to write down where business definitions most often reside. Many wrote, &#8220;in someone&#8217;s head.&#8221;  A few wrote, &#8220;in our data catalog.&#8221;  Others, &#8220;in a spreadsheet.&#8221;</span></p><p><span>As </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Jessica Talisman, MLS&quot;,&quot;id&quot;:24176542,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zEsI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18f1fe4e-779e-4a27-be92-71fac460ee01_935x935.jpeg&quot;,&quot;uuid&quot;:&quot;47f49895-68e0-4bec-87c2-c5494f10e3aa&quot;}" data-component-name="MentionToDOM"></span> and Tony Seale <span>shared on </span><a href="https://www.thoughtspot.com/data-chief/ep135/how-semantic-layers-and-ontologies-create-trusted-ai"><span>The Data and AI Chief podcast,</span></a><span> how your business operates is your unique IP. How your business operates may be partially documented in employee manuals or corporate intranets, but more often it is part of the tribal knowledge of a company.</span></p><p><span>The idea that we can get everyone first to document, then to agree, then to centralize, then to maintain business definitions seems laughable to me. I think of one customer I have advised who told me they have 33 different definitions of customer! I could easily understand three (prospective, active, former), but not 33.</span></p><p><span>And yet, if we do not have one single semantic layer, how else do we stop people from talking past one another or debating whose number is right?</span></p><h2><span>An Operating Model for Semantic and Context Layers</span></h2><p><span>I can envision a central information steward documenting these 33 different definitions, but unlikely that the same data steward would maintain such meta data as definitions are rationalized or refined.  Data catalogs and information glossaries started in this vein but have met with mixed success. Business users and data experts recognize the need. The organizational muscle to share cross functionally and commit to maintaining is weak.</span></p><p><span>Hence, my recommendation for the agentic AI era is to pursue a layered, federated approach:</span></p><ol><li><p><span>Centralize commonly used elements that require company-wide alignment</span></p></li><li><p><span>Decentralize and federate domain and business-specific terms</span></p></li><li><p><span>Leverage AI to capture existing context</span></p></li><li><p><span>Allow business users to dynamically update context</span></p></li><li><p><span>Periodically synchronize and merge centralized context with federated context</span></p></li></ol><p><span>You will note data mesh principles here such as domain ownership. Vanguard shared on The Data and AI Chief podcast why they are federating their semantic layer approach (tune in </span><a href="https://youtu.be/9fZ-5QZNiUs?si=SPJU1xUYT_SlJenp&amp;t=867"><span>here</span></a><span>).</span></p><p><span>I suspect the biggest discomfort proponents of a single semantic layer will have are with points 4 and 5. Here, I am being a pragmatist rather than a purist. The control freak, single-version-of-the-truth seeker in me wants to tell a business user that they MUST tell the data steward when a business definition changes. The pragmatist in me knows this will not happen.</span></p><p><span>Just last week, as I was looking at podcast listener stats, I told Spotter (ThoughtSpot&#8217;s agentic analytics agent), whenever I ask for podcast listeners by time, use the download field. If I ask about particular episodes, use the listener field (a strange way the source system tracks data).   Should I go back to our Snowflake DBA and ask him to annotate the data data model with more precise definitions? What about the owner of that ThoughtSpot Spotter semantic model?</span></p><p><span>I don&#8217;t do either. I am too busy (so is the data engineer and ThoughtSpot admin). And I am grateful that the context I gave Spotter while I asked a question is automatically shared with my podcast producers and preserved in </span><a href="https://www.thoughtspot.com/blog/spotter-memory"><span>Spotter Memory</span></a><span>. I consider this a reasonable, federated approach in which I am largely the domain owner of podcast data.</span></p><p><span>Context can be provided during the analytics workflow but it also exists in a variety of systems, both structured and unstructured. Databricks recently previewed Genie Ontology that will crawl tables, queries, and dashboards for context. Glean differentiates itself on its context layer. Atlan has repositioned itself from data catalogue to open context layer. Such context is critical in ensuring trusted insights and efficient usage of</span><a href="https://www.thoughtspot.com/blog/token-maxxing-and-inference-ops-the-new-finops-frontier"><span> tokens.</span></a></p><p><span>I want these definitions to be domain controlled and dynamically enhanced by AI.</span></p><p><span>On the other hand, for something as widely used and shared as &#8220;customer&#8221;, I want a different approach. Likewise for publicly reported KPIs and shared data in regulated industries, I want greater centralization but one that is flexible.</span></p><h1><span>A Medallion Model for Semantics</span></h1><p><span>With data modeling, we have gotten comfortable with multiple levels of data quality, with the gold layer being the highest quality in a lakehouse.</span></p><p><span>The same should happen with certain company-wide KPIs and dimensions. For a publicly traded company, GAAP revenue should be in that gold semantic layer.  The most widely used dimensions such as customer and product, should also be in a gold semantic layer. Federated semantic models may be considered silver and user-specific definitions, bronze.</span></p><div class="callout-block" data-callout="true"><p><span>This gold layer could be in a database semantic view, a metrics layer, a data catalogue, a third-party semantic layer such as dbt, or in an AI semantic layer such as ThoughtSpot.</span></p></div><p><span>Where it should not be is in any layer that is closed and proprietary. That includes many legacy BI tools.</span></p><p><span>We also need a mechanism by which definitions enriched by domain users - while asking for insights - can be promoted back to this gold layer.  Efforts such as </span><a href="https://ossie.apache.org/"><span>Apache OSSIE</span></a><span> (an evolution from OSI) are promising in allowing such interoperability. But not all database vendors are supporting OSSIE, nor are a range of AI and BI tools. So the onus is on each customer to define their semantic layer strategy - the ownership, the update process, the sources of truth.</span></p><p><span>But what about those three (or 33) different customer definitions? The definitions should be explicit to contain the context. Active Customer. Free Trial Customer. Inactive Customer. It&#8217;s okay if the generally accepted default for &#8220;customer&#8221; is &#8220;active customer.&#8221;</span></p><p><span>An intelligent AI agent will also use information such role, function, region, or frequently used data models as context to reduce ambiguity.</span></p><h1><span>What&#8217;s The Best Place for a Gold Semantic Layer?</span></h1><p><span>As I work with customers on developing their semantic and context layer strategy, I have noticed differences in reasoning here:</span></p><ul><li><p><span>Some want it closer to the database for re-usability.</span></p></li><li><p><span>Others are deciding based on the ease of updating and whoever is closest to the business domain.</span></p></li><li><p><span>DBT customers like the code-based approach for auditability and version control. The same is said of ThoughtSpot customers who export their models to TML.</span></p></li></ul><p><span>What I have yet to hear is how organizations envision maintaining these models as context is enriched. Nobody wants to dumb down a context layer that is more mature than a particular database or further ahead than standards within OSSIE. </span></p><p><span>Back to my fish and chips analogy. Humans can deal with such subtleties and ambiguity. We automatically factor in context. AI agents are more rigid. Quality semantics with robust context is the answer to ensure trusted insights.  If we hope to scale agents for trusted insights, the tech has to evolve and so does an organization&#8217;s semantic layer strategy and operating model.</span></p><p><span> I look forward to hearing how you are approaching this critical part of the AI stack.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to be a continuous AI learner. </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h3><span>A Word About My Biases (aka Experiences)</span></h3><p><span>I am a business person at heart who is technically savvy. In deploying Business Objects and authoring three Complete Reference books, my universes were well designed, complete with descriptions for every field - imported from spreadsheets, the only source anyone remotely maintainted. In using and deploying Tableau, I did many one off analyses and never bothered to document any metric. At ThoughtSpot, I advise our top customers on data and AI strategy with a business-first mindset and right-sized governance approach. I am delighted to see that </span><a href="https://www.linkedin.com/posts/biscorecard_datamodeling-semanticlayer-metricslayer-ugcPost-7376252413059678208-5b-1/?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAAFX10BZ-96_gDDByblilJgpSW3xET8gbk"><span>Semantic Layers are Suddenly Sexy Again</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Stop Counting Prompts. Start Counting Outcomes.]]></title><description><![CDATA[Observe, Optimize, Operate: A playbook for the AI token economy]]></description><link>https://brainsandbots.substack.com/p/stop-counting-prompts-start-counting</link><guid isPermaLink="false">https://brainsandbots.substack.com/p/stop-counting-prompts-start-counting</guid><dc:creator><![CDATA[Sonny Rivera]]></dc:creator><pubDate>Fri, 21 Aug 2026 19:59:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nAQF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nAQF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 424w, /__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 848w, /__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nAQF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png" width="1456" height="722" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:722,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1299949,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://brainsandbots.substack.com/i/212158566?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 424w, /__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 848w, /__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nAQF!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbac38f7c-e1f7-4453-9155-bf70bca65300_1915x949.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In May, an AI consultant told Axios that one of their enterprise clients ran up a $500 million Claude bill in a single month.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> The company had put no limit on how many licenses employees could use. People ran long agentic workflows that compounded token spend quietly in the background, and the consultant noted expensive tooling getting pointed at things a person could do in seconds, like checking the weather. Uber burned its entire 2026 AI budget by April, four months in, after Claude Code adoption jumped from 32 percent of engineers in February to 84 percent by March. One Uber executive ran up $1,200 in a single two-hour coding session. Microsoft pulled Claude Code licenses stating engineers were costing $500 to $2,000 a month.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>That is the base rate. The individual stories underneath it are far worse. So I can understand why I&#8217;ve been getting questions about what specific actions should be taken to reduce AI costs. And what can make the costs more predictable? These are hard questions that can often end with &#8220;it depends.&#8221;  So, I&#8217;m going to do my best to leave you a framework for controlling costs and some specific actions you can take right now.</p><h3></h3><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">&#129504;&#129302; Enjoying <strong>Brains &amp; Bots? </strong>Subscribe for weekly takes on data, AI, and where they collide!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>Observe, Optimize, and Operate (O<sup>3</sup>)</h3><p>In my mind, I think of these as visibility, optimization, and continuous improvement. I&#8217;ve cleverly named it O<sup>3</sup> or &#8220;O Cubed&#8221;.  In each of these areas, I&#8217;ll lay out some specifics that will help you right now.  It&#8217;s worth noting that one-time reviews and adjustments to your agentic stack aren&#8217;t going to control your AI costs and that&#8217;s why you need a system or a framework. So let&#8217;s dive in.</p><h3>O1: Observe</h3><p>You can&#8217;t optimize what you can&#8217;t see, and most organizations can&#8217;t see very much. So, observation comes first because every recommendation later in the framework depends on having something or a number to compare against.</p><p>Before we begin observing, let&#8217;s make sure we understand what we are looking at.  For our use case, let&#8217;s first understand tokens. </p><h4>A token is four line items, not one</h4><p><strong>Input tokens</strong> are what you send into the model. It&#8217;s your prompt, the system instructions, all the documents and data you stuffed into the context. Note: I find that is the line item that gets inflated by design choices that nobody ever revisits. Recently, Anthropic&#8217;s own engineering team found they could delete repeated tool-use examples from a system prompt entirely.  They just included the same instructions in the tool descriptions but only once, with no loss of capability.  This removed redundant tokens from every single call.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p><strong>Cache read tokens</strong> are the input you have sent before.  They are served back at a steep discount, as much as 90 percent. That discount compounds fast at volume for relevant questions but can be costly for unrelated questions.  Cache is the single easiest win on this list but it&#8217;s not full-proof.</p><p><strong>Thinking tokens</strong> represent the reasoning that a model does before it answers. We&#8217;ve all seen Claude &#8216;thinking&#8217;, &#8216;ruminating&#8217;, and &#8216;crunching&#8217;. Instead, Claude should say &#8216;generating more tokens and costs&#8217;. It&#8217;s important because those reasoning models are billed at output token rates which is the most expensive tier. It can also wreck your forecast because the amount is not fixed or bounded. Meaning, the same question you sent to the same model twice, can result in very different amounts of reasoning.</p><p><strong>Output tokens</strong> cost the most per token. So your output costs are driven by verbosity, and by result sets that get echoed back through a context window instead of rendered somewhere cheaper.</p><p>To put real numbers on it: Claude Sonnet 5 runs $2 per million input tokens and $10 per million output, Opus 5 runs $5 in and $25 out.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> A model doing heavy reasoning can burn several times its visible output in thinking tokens you never see, all billed at the output rate. If your team runs agentic workflows with multi-step tool calls, every tool definition, every intermediate result, and every retry adds tokens you did not budget for.</p><p>There is one more level we need to observe; the tokenizer. While the rates for tokens may have held steady between model releases, the tokenizer may have changed. Claude 4.7 token generation jumped by 30% according to Anthropic.  Call it &#8216;<em>Shrinikflation</em>&#8217; if you like.  Meaning, you can&#8217;t just watch the dollars billed, you have to look at the tokens count too. </p><h4>The unit economics nobody models</h4><p>In my experience, unit economics of AI and agentic workloads is one of the hardest capabilities to get right. Yes, it&#8217;s even more difficult than it was for cloud data platforms.  Why, because gap in understanding between AI token usage and a measurable business outcome is huge.  It was much easier to connect a compute-hour to a completed job or pipeline.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7A1-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 424w, /__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 848w, /__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7A1-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png" width="1456" height="888" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:888,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:120116,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://brainsandbots.substack.com/i/212158566?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 424w, /__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 848w, /__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7A1-!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4b7b4ee-a28c-4b50-8cfc-e7cfaf68e56b_1800x1098.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Let&#8217;s break this down. I see three layers between what is billed and the business value delivered. API pricing is easy; it&#8217;s dollars per million tokens. Application cost, in the middle, is harder. It&#8217;s how many tokens one real user interaction actually consumes including multi-turn interactions, retries, context growth, and tool calls.  For the middle layer, consider the following: </p><ol><li><p>Cost as a distribution. Because token usage for a single business question is unpredictable, the cost arrives as a distribution rather than a single point.</p></li><li><p>There is a language tax. Verbose prompts, bloated context windows, and multilingual requests ( yes, non-english tokenization just cost more), are all driving up cost without anyone deciding. </p></li><li><p>Shrinkflation is real. Your token cost may surge upward in new releases while value delivery remains flat.</p></li></ol><p>Business outcomes are at the top. You can imagine a deferred support call, a closed ticket, or an hour saved by an analyst. Note: before you finalize the forecast, find every token meter in your AI stack.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><h4>Report outcomes instead of activities</h4><p>Let&#8217;s be clear, the frontier labs bill for activity not business value. They have no incentive to help you measure value, because consumption is the product. So now you know that it&#8217;s on you to show value. Here&#8217;s what to you shouldn&#8217;t report:</p><p>Activity - &#8220;We executed 10,000 prompts this quarter.&#8221;</p><p>Instead report the outcomes, &#8220;We deferred 500 support calls, recovered $50,000 in subscription revenue at risk, and saved 200 hours of analyst work.&#8221; Which would you want to hear?</p><p>Cloud data teams have taught us a lot here. I was on those teams when we had this same approach 5 or 6 years ago. A executive leader would ask what the data team had accomplished and we would proudly claim, &#8220;we ran 40K dbt models&#8221;. Not good. Today, we are seeing the same maturity gap with agents; only now the tools are different.</p><h4>&#128161; Best practice: tag every agent</h4><p>Categorize every agent by team, use case, and environment. Note whether it&#8217;s a background agent or an interactive agent before it ever gets deployed. We all know that teams are facing pressure to deliver quickly and that often leads skipping perceived unimportant details like tagging. But without we&#8217;ve lost visibility and puts limits on cost visibility, allocation, forecasting, and ultimately the unit economics.</p><p><strong>Why it saves money</strong>: you can&#8217;t cut what you can&#8217;t attribute to a team, a product, or a workflow.  This is a first principle in cost controls.  Without proper tagging, your team is turning every cost metric into a guess.</p><h4>&#128161; Best practice: audit business value on a cadence</h4><p>Survey the business teams that are using running agents weekly, monthly, or quarterly. Ask whether the work still produces value, separately from whether the agent is still running. Do it need to run every hour, every day, on weekends? The there&#8217;s lots of quick wins here. Be ruthless! Retire every agent fails to show value.</p><p><strong>Why it saves money</strong>: usage always outlives value. For example, a workflow built for a product launch keeps consuming tokens long after the launch is over, and nobody notices, because nothing is broken.</p><h3>O2: Optimize</h3><p>Visibility and surveys implemented, check. The next move is architecture. One instinct is to route everything to your most capable model, which is often the expensive model and usually not necessary. The other instinct is to tell people to prompt less, which may be even worse. It&#8217;s worse because it trades a cost problem for an adoption problem. In my opinion, the most expensive prompt is the one your domain experts are afraid ask due to ai rationing.</p><h4>Task routing equals warehouse re-sizing</h4><p>We&#8217;ve done this before with Snowflake compute. Right-sizing the model with the complexity of the task is what&#8217;s important here. For example, classification and extraction go to small, cheap models just like we did with simple data ingestions using Snowflake x-small compute. If you have genuinely hard multi-step reasoning tasks, those deserve the expensive models and should be governed appropriately.  After all, not every user needs to solve that 100 year old unsolved math problem. </p><p>Cache aggressively wherever context repeats. Design orchestration so failures happen cheap and fast instead of retrying an expensive call three times. It&#8217;s easy to say and I recognize that it&#8217;s not always easy to implement.  So, think big, start small, move fast, and iterate!</p><blockquote><p>Note: Right-sizing alone won&#8217;t solve the unpredictable cost problems.  Again, that&#8217;s why we need a framework.</p></blockquote><h4>&#9888;&#65039; Caution: right-size to the task, then verify against realized cost</h4><p>The same query doesn&#8217;t cost the same thing twice. Identical question, identical model, repeated runs: the most expensive run cost close to ten times the cheapest. Databricks reached the same conclusion independently, that per-token pricing is an unreliable indicator of total task cost. Cheaper-per-token models routinely cost more per completed task. <a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p>Why it saves money: capability, rather than price, determines how many turns a task takes to successfully. The rule here is to route based on the measured cost per completed task. Use the list price as just one input among many.</p><h4>&#128161; Best practice: decompose into multi-step workflows and treat session length as a cost control</h4><p>Break a task into steps, each handled by the cheapest model that is <strong>capable</strong> of completing it. Rather than sending full context to a single expensive call. Start new sessions for new topics instead of letting one session accumulate results across unrelated questions.</p><p>Why it saves money: a well-orchestrated pipeline can lower total spend even as it adds steps. This is because each step&#8217;s input context remains small. Conversely, session length compounds costs in the wrong direction, since results re-sent on every turn cost way more than results generated just once.</p><h4>Vendors are shipping the tooling, and that covers half the job</h4><p>FinOps for AI is becoming a software category of its own. And it&#8217;s filling in fast, which is a good sign that the problem is real.</p><ul><li><p>Ramp launched AI Token Spend Management in July. It provides one dashboard across AI providers, weekly briefings on usage trends and efficiency opportunities.  It also has real-time controls and alerts to stop an overrun while it is happening.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a></p></li><li><p>Nvidia and Snowflake moved into the model router market in August. Routers productize the task-to-model matching this section describes, sending each request to the cheapest model capable of handling it.</p></li><li><p>Linux Foundation has launched a Tokenomics Foundation.</p></li></ul><p>My recommendation is to evaluate what&#8217;s best for your organization and buy them. They are genuinely useful, and building your own meter from scratch is a poor use of an engineering budget. If coding agents have made us all builders, the build vs buy question now is &#8220;Could We&#8221; vs &#8220;Should We&#8221;.</p><div class="pullquote"><p>Observe and Optimize only pay off when you apply them to your own business processes. </p></div><p>You have now reached the area where most cost control efforts stall. Observe and Optimize only pay off when you apply them to your own business processes. No vendor can answer the hard questions for you, like which department&#8217;s work justifies more spend, what counts as an accepted output in your context, which agent deserves to still be running next month, and what a deflected support call is actually worth on your P&amp;L. You and your domain experts rely on experience and judgement to answer those questions and then determine where to apply your resources. Don&#8217;t delegate that part.</p><h3>O3: Operate</h3><p>You can think of Observe and Optimize stages as constrained projects with end dates and deliverables. But, Operate is different.  It&#8217;s the part that runs forever.  It&#8217;s also the part that needs continual care and feeding before it takes root. Once you set the cadence and maturity take hold, you can sit back and harvest the value. Your process will have become a FinOps practice.</p><h4>&#9888;&#65039; Callout: set a time-to-live on every background agent</h4><p>Every background or scheduled agent gets an expiration date or a review date which should be set on deployment of the agent. When it expires, someone reviews the ongoing value and renews it deliberately. Otherwise, it stops running and accumulating costs.</p><p>Why it saves money: background agents are easy to forget especially when the creator doesn&#8217;t pay the bill. Nobody sees them, nobody complains about them, yet they bill every night. Setting TTL keeps the business user engaged, increases accountability, and makes value decisions explicit.</p><h4>Cost per accepted task</h4><p>If you aren&#8217;t familiar with <em>cost per accepted task</em>, it&#8217;s the metric that matters most. The main problems are that few organizations track it and the frontier labs can&#8217;t track it. It&#8217;s not cost per query or cost per prompt. It&#8217;s the cost per output your team actually kept and used.</p><p>I break acceptance rate into three categories and track them. The categories are <em>as-is</em>, <em>edited and used</em>, and <em>discarded</em>. Now track those.  You&#8217;ll soon see examples of workflows that cost three times more per call and get accepted nine times more often.</p><div id="datawrapper-iframe" class="datawrapper-wrap outer" data-attrs="{&quot;url&quot;:&quot;https://datawrapper.dwcdn.net/iT5ht/3/&quot;,&quot;thumbnail_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ce437cd-5716-4a3a-b00a-7b733d7e73b9_1220x1016.png&quot;,&quot;thumbnail_url_full&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19a2579d-dd71-4d43-999e-671e077b9228_1220x1136.png&quot;,&quot;height&quot;:560,&quot;title&quot;:&quot;Cost per accepted task&quot;,&quot;description&quot;:&quot;Illustrative example: the same 100 requests, run through two workflows&quot;,&quot;belowTheFold&quot;:true}" data-component-name="DatawrapperToDOM"><iframe id="iframe-datawrapper" class="datawrapper-iframe" src="https://datawrapper.dwcdn.net/iT5ht/3/" width="730" height="560" frameborder="0" scrolling="no" loading="lazy"></iframe><script type="text/javascript">!function(){"use strict";window.addEventListener("message",(function(e){if(void 0!==e.data["datawrapper-height"]){var t=document.querySelectorAll("iframe");for(var a in e.data["datawrapper-height"])for(var r=0;r<t.length;r++){if(t[r].contentWindow===e.source)t[r].style.height=e.data["datawrapper-height"][a]+"px"}}}))}();</script></div><p>Let&#8217;s compare Workflow A, that premium workflow, with Workflow B, the cheaper workflow.  Workflow B has a 50 percent discarded rate. So, the more expensive workflow wins on cost every time.  Why? Because you pay twice for the rejected half, once for the bad output and again for the human hour spent redoing it.</p><p>Lastly, don&#8217;t forget to measure the failure rate. Track all three numbers together. How often you get a correct answer, what an average correct answer costs, and what the worst runs cost. Think of these as best case, average case, and worst case scenarios. The failure rate or worst case is usually ignored but it is the one that wrecks your forecast.</p><h4>Demand determinism for anything governed</h4><p>One last operating principle, and it is the one I feel most strongly about. Trusted answers are non-negotiable! Using probabilistic SQL against any governed metric is a problem for your budget and your compliance audit. Wherever a number has to be defensible, push the resolution into a deterministic layer and accept that the layer will sometimes fail closed and tell you it cannot answer. Some people will find that infuriating but I think it&#8217;s the entire point.</p><h3>Wrap up</h3><p>Frontier labs will keep billing based on tokens, because tokens are what they can meter. Your job is to build the layer above the meter. That&#8217;s the layer that translates consumption into value, so own it. You&#8217;ll be using the same concepts that we used to build FinOps on top of cloud consumption. And your team probably has the skills already, because you have been doing it for the cloud data warehouse for years.</p><p>Observe what a token actually costs and who is spending it. Optimize the routing and the orchestration, with realized cost as your evidence rather than a pricing page. Operate with TTLs, acceptance rates, and one outcome metric per workflow; ensure somebody owns that metric.</p><p>So, the bill arrives in tokens, full stop. But, you control whether leadership sees the number of tokens or a story about the business value driven by those tokens. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/p/stop-counting-prompts-start-counting/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/brainsandbots.substack.com/p/stop-counting-prompts-start-counting/comments"><span>Leave a comment</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">&#129504;&#129302; Enjoying <strong>Brains &amp; Bots? </strong>Subscribe for weekly takes on data, AI, and where they collide!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p> Axios, <a href="https://www.axios.com/2026/05/28/ai-spending-roi-enterprise-costs">AI sticker shock hits corporate America</a>, May 28, 2026. Worth noting the account is secondhand: an unnamed consultant describing an unnamed client, with no confirmation from Anthropic or the company. The Uber and Microsoft figures in the same paragraph are separately and directly reported.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I suspect this has more to do with Microsoft strategy and usage of their large install based, but I digress.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p> <a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">The New Rules of Context Engineering for Claude 5 Generation Models</a>, claude.com.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><a href="https://platform.claude.com/docs/en/about-claude/pricing">Claude Pricing</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p> I don&#8217;t want to complicate the story that everything is billed on tokens. But, Snowflake Cortex Analyst prices per message, a fixed number of credits regardless of how many tokens the message contains. That looks like a cleaner unit until you notice the message fee covers only the text-to-SQL generation. It drives token cost by data stuffing and by running compute services.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Databricks (2026), <a href="https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase">Why Estimating Prompt Cost is Hard</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p><a href="https://www.prnewswire.com/news-releases/ramp-launches-ai-token-spend-controls-302827389.html">Ramp Launches AI Token Spend Controls</a>, July 2026.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Saving $$$ with MCP]]></title><description><![CDATA[Less chisel, more mortise! 3 concrete ways MCP cuts token costs and engineering debt]]></description><link>https://brainsandbots.substack.com/p/saving-with-mcp</link><guid isPermaLink="false">https://brainsandbots.substack.com/p/saving-with-mcp</guid><dc:creator><![CDATA[Jane Smith]]></dc:creator><pubDate>Tue, 11 Aug 2026 14:01:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f70310ce-467a-48ae-b41a-aa8900e713c8_1358x744.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>When I grew up in the 1980s, my father was a builder. This experience has influenced how I think about things (what do you mean that&#8217;s a pile of timber? It&#8217;s a pile of cash!) and also given me a shed-load of useful analogies for my work as a Field Chief Data &amp; AI Officer. (The only danger is that analogies about houses and foundations can very quickly resemble a retelling of the Three Little Pigs.)</span></p><p><span>As anxiety around token costs becomes real for any data and AI leader with budgets scrutinised and inference fin-ops becoming real, I think about the power tool my dad would never allow me to use&#8230; the mortising machine - the ultimate way to offload cost.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>If offloading AI cost via MCP,  or, cutting mortises into timber are of interest to you - here are three concrete (see what I did there?) ways to offload AI cost with MCP.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uayL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uayL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg" width="780" height="532" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:532,&quot;width&quot;:780,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Master the mortise-and-tenon joint&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Master the mortise-and-tenon joint" title="Master the mortise-and-tenon joint" srcset="/__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!uayL!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc390d24-00b4-4581-9f35-85364cf08bd4_780x532.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>When Refactoring becomes a Factory</span></h4><p><span>A headache I had as a Chief Data Officer was my convoluted mess of dbt models, SQL scripts, and legacy stored procedures, all so dependent on each other that a single column definition change could break multiple dashboards downstream.</span></p><p><span>Even a simple metric change was a drama... if leadership wanted to change for example, the loss ratio formula from claims paid / gross written premium to claims paid / earned premium, (I worked at an insurance company) you had to find everywhere that loss ratio was calculated; dbt staging, intermediate SQL, the semantic layer etc, make the change then retest downstream dashboards, models and reports.</span></p><p><span>LLMs didn&#8217;t really help... sure you could give an LLM an instruction such as: &#8216;refactor loss ratio to use earned premium not gross written, find every definition across the repo and list all reports affected.&#8217; But to understand that instruction the LLM would need the full repo inputted into the context, easily ~500k tokens. Also take into account that agents, as they search, reason, fail, retry, tend to &#8216;loop&#8217; as they work. So a 20 person team changing metrics regularly can run to five figures a month just reading code. So much context stuffing also raises hallucination risks, not great in any organisation but especially not in something as regulated as insurance!</span></p><p><span>MCP handles this more efficiently. The agent calls a dbt/Git MCP server for the metric and gets back only the 3 models where loss ratio is defined. A lineage call returns the 4 affected downstream models as JSON. The LLM inspects the 3 files, updates the logic and opens a PR for human review costing ~1.5k tokens, not ~500k! Deterministic tools handle the indexing and the LLM handles reasoning. </span></p><h4><span>Structured Data - Hard and Expensive for LLMs</span></h4><p><span>Here&#8217;s where we get to the mortises! If you&#8217;re wondering what mortises are, they&#8217;re slots in timber that joins fit into. You can certainly pay a joiner to hand chisel the 80 or so mortises you&#8217;ll need for a wooden staircase, or, you canhave the mortise machine do it for you.</span></p><p><span>When you ask Claude (or any LLM running on your structured data) a question like; &#8216;what is net revenue by region?&#8217; it must read your entire semantic layer metadata to understand what &#8216;net revenue&#8217; and &#8216;region&#8217; mean. It then runs the query on the warehouse, formats the raw JSON output into text and charts. A single query can burn ~13k tokens. Multiplied across teams, a prototype built in an afternoon can become a $75,000 a month beast Yes, you can build in prompt caching and that will cut what you&#8217;re billed for repeat context if subsequent questions are asked. It will require engineers to build and oversee it though so whilst token cost will go down, engineering cost will not.</span></p><p><span>A tool that is NOT text to SQL but tokenised search will do this very differently. I&#8217;ll use ThoughtSpot&#8217;s Spotter here (yes, thank you for asking, I do work for ThoughtSpot!) Rather than context stuffing raw schemas into the LLM,  a lightweight text to token knowledge graph maps queries directly to database schemas, using a small fraction of the tokens a text to LLM tool would use. The SQL compilation and chart rendering then happen deterministically on the warehouse at no token cost at all.</span></p><p><span>The tools can easily be used together so that when the question is asked to the LLM, it makes an MCP to Spotter which will then do the heavy lifting of the compute (via it&#8217;s deterministic mode of action already explained). Spotter sends a response back to Claude to format. I haven&#8217;t formally benchmarked the saving, but as a principle you&#8217;re replacing a large, unpredictable token spend with a small, mostly fixed one.</span></p><h4><span>Tokens are the Tip of the Iceberg: Custom Code for Integrations</span></h4><p><span>When you have to write code for each new integration to oversee authentication, error handling etc. it quickly becomes a large engineering burden. I know because I&#8217;ve been there and it&#8217;s very&#8230;. multiplicative&#8230; e.g. 3 x AI assistance for 4 x enterprise systems becomes 12 custom integrations.</span></p><p><span>And this is really only the start. Every time lets say, SAP updates their API, every time you upgrade your agents, the connections break requiring engineering support. Last time I looked at a P&amp;L I did not find engineers to be cheap. Worse that that, this approach keeps engineers in this zone of &#8216;invisible work&#8217; where they&#8217;re busy on work items neither seen nor understood by execs. &#8216;But what do your team do all day?&#8217; is a question I was asked as a CDO and having to justify your team&#8217;s existence is not a good position to be in. However if you build a single MCP server that any MCP compatible agent can plug into and this problem largely goes away.</span></p><p><span>If we take a scenario where a company wants its internal AI agent to perform weekly sales health checks by calling data from ThoughtSpot (revenue metrics), Salesforce (pipeline deals), and Jira (to see delivery blockers). Without MCP engineers need to write 3 python wrappers, security layers and converters taking about 6 weeks to build. MCP would require a few days only to configure permissions.</span></p><p><span>This is essentially Total Cost of Ownership. Execs currently are edgy about token costs but will often not see nor understand ~&#163;10,000 of engineering costs. Be mindful of what&#8217;s visible to them and what isn&#8217;t&#8230; and educate them.</span></p><p>Just as TCO is important, so is security with MCP don&#8217;t miss the opportunity to build trust by engaging your security team early</p><p><span>MCP should be a no-brainer but I still see a lot of point connections. Organisational silos, prototypes that become full production solutions etc. it&#8217;s essentially a new type of tech debt.</span></p><h4><span>Conclusion: Offload the Load</span></h4><p><span>If the timber was cash in my dad&#8217;s eyes, the tokens certainly are cash in any data &amp; AI leader's eyes&#8230; and things get expensive when you&#8217;re not using the right power tools. MCP is your mortise machine, so use it to offload cost and token anxiety.</span></p><p><span>As my dad would say&#8230; the right tools for the right job.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What are the must-have AI skills for High School Students? ]]></title><description><![CDATA[And what should you not outsource to AI?]]></description><link>https://brainsandbots.substack.com/p/what-are-the-must-have-ai-skills</link><guid isPermaLink="false">https://brainsandbots.substack.com/p/what-are-the-must-have-ai-skills</guid><dc:creator><![CDATA[Cindi Howson]]></dc:creator><pubDate>Thu, 06 Aug 2026 14:26:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ncF8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ncF8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ncF8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:10318268,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://brainsandbots.substack.com/i/209944026?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ncF8!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F966299cb-1fe5-41c4-854c-86a99e9b6611_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I had an angry mom recently blast me on why she hates AI, fearing for her son&#8217;s job prospects. Is college even worth it if students are just regurgitating what LLMs tell them? </p><p><span>Teens are using ChatGPT and Claude as study guides, essay writers, vacation planners, poster-makers and more.  AI is now a life skill. Just as high schoolers learn essential skills such as driving, algebra, biology, and world history, AI has joined the ranks of foundational skills. There are AI skills to acquire, but there also are skills that we should consider NOT outsourcing to AI.</span></p><p><span>This applies to teens as well as adults.</span></p><p><span>We do not want to outsource the skills that make us uniquely human. When we outsource too much to AI, those skills atrophy and risk disappearing altogether. Let&#8217;s take map reading as an example. Without a map on a student&#8217;s smart phone, could they find their way while visiting a new country?  Paper maps and the ability to read them seems to be a lost skill for GenZ. Most teens and young adults are fine with this dying skill as long as they have cell service, wifi, a well charged phone, and there is no mass power outage. Drop them on a safari &#8230; or hiking in remote parts &#8230; and it&#8217;s a challenge!</span></p><p><span>Here are the three skills I would not concede to AI:</span></p><ol><li><p><strong><span>Critical thinking.</span></strong><span> Newer AI models have advanced reasoning modes. But critical thinking is a uniquely human skill that goes back to how we are coded for survival. Critical thinking involves solving complex problems, reflecting on what are the most important problems to solve, and interpreting assumptions and alternatives. This includes asking better questions of AI.  When AI starts spewing answers that are confidently inaccurate, critical thinking skills allow you to challenge the results.  That vacation you planned using AI in which the proposed flight connection does not exist? &#8211; critical thinking skills come into play here. AI has been trained on the knowns, not on the new, novel, and possible.</span></p></li></ol><ol start="2"><li><p><strong><span>Writing.</span></strong><span> Phew, Substack lit up last week with the introduction of AI checker Pangram. If you have used Claude or ChatGPT all year long to draft your papers, without first thinking about the points you want to make, your writing skills have already atrophied. Anecdotally, people have told me the decline is steep in a matter of weeks. Further, the ability to organize your thoughts into a persuasive argument or problem solving strategy is weakened. </span><a href="https://www.media.mit.edu/projects/your-brain-on-chatgpt/overview/"><span>Brain research</span></a><span> confirms this decline. So, then, imagine you are now in your first job or hoping to land the most desirable internship. You are in a meeting and your boss poses a problem, asks you to assess the situation and recommend what to do. Your ability to articulate and think on the spot will be weaker than the co-workers who have only been using AI as a thought partner. In this regard, you should use AI primarily as a thought partner and a super spell checker, but not as the writer of your first drafts. Draft it yourself first, think critically about the points that are most important, then use AI help you make it better such as for fact checking or critiquing your arguments.  Further, research shows that </span><a href="https://journalistsresource.org/education/longhand-versus-laptop-note-taking/"><span>handwriting </span></a><span>allows for further memory retention and creative sparks.</span></p></li></ol><p></p><p><span>Now, that weekly summary of a status report that requires little thought &#8211; that is a good task to outsource to AI. Students in which English is a second language have also shared how letting AI write for them is helpful. For sure, my rusty German writing I absolutely butcher the language. But I would apply the same principles. How much do you want to hone your English or foreign language skills? Draft first,then use AI to correct.</span></p><ol start="3"><li><p><strong><span>Emotional Intelligence and Empathy.</span></strong><span> Being smart is often associated with high IQ. EQ is a high emotional intelligence, defined as the ability to recognize and regulate your emotions, while also understanding and empathizing with others. Empathy is a uniquely human skill. Empathy is the ability to put yourself in someone else&#8217;s shoes &#8211; to truly feel what another person is feeling emotionally. The idea that AI has feelings or conscience has repeatedly been refuted. Sycophancy &#8211; AI telling you what you want to hear &#8211; is a known problem with AI. Will I Am framed this so eloquently at the Davos World Economic Forum, </span><a href="https://www.youtube.com/watch?v=HPF4I-R4sUo"><span>Imagination in Action break out</span></a><span>, earlier this year, saying &#8220;AI will f*** up relationships.&#8221; AI girlfriends and boyfriends will tell you exactly what you want to hear, setting up unrealistic expectations for healthy relationships. As humans, we are wired for connection. Deeper connections are enhanced by empathy. As the amount of AI generated content explodes, humans are looking for more authentic content and relationships.  The most successful leaders have high EQ and use EQ to motivate others.</span></p></li></ol><p><span>Now that you know what skills not to outsource to AI, here are the AI skills to acquire and hone.</span></p><ol><li><p><strong><span>AI privacy and safety</span></strong><span>. Every AI company has their own privacy settings that includes how they use your prompts, the degree they retain information about you, and what can be done with that information. Deepfakes are increasingly used by scammers to trick someone into sending money.  The emotional damage of a photo of a real person manipulated into a nude photo is profound. </span><a href="https://www.latimes.com/california/story/2024-04-02/laguna-beach-high-school-investigating-creation-of-ai-generated-images-of-students"><span>Schools also may suspend students for such behavior.</span></a><span> Certain AI platforms take AI safety seriously and ban such manipulation, while others allow it.</span></p></li><li><p><strong><span>Prompt engineering aka asking good questions</span></strong><span>. Prompt engineering is really  about asking good questions and giving AI the most context that will yield optimal output at the least cost. Also use the right LLM for the right task, leveraging cheaper models for simple tasks (see t</span><a href="https://www.thoughtspot.com/blog/token-maxxing-and-inference-ops-the-new-finops-frontier"><span>his blog</span></a><span> on tokenomics).</span></p></li><li><p><strong><span>Training data</span></strong><span>. Large language models are trained on vast sources of public data. This may be news articles, social media posts, and the like. It becomes a game of numbers and skews heavily to western cultures. As you apply your critical thinking, consider gaps in data that may skew answers. For example, one </span><a href="https://markcubanai.org/"><span>Mark Cuban Foundation AI bootcamp</span></a><span> student asked AI to create an image of a successful business woman in Singapore. It defaulted to a male business person, even though Singapore has the highest percentage of female C level executives. Models also may not be trained on the most recent data, so rapidly changing news and sports results may cause AI to give you wrong results.</span></p></li><li><p><strong><span>Hallucinations.</span></strong><span>  AI </span><a href="https://www.business-reporter.com/industrial-innovations/preventing-hallucinations-in-generative-ai"><span>hallucinates,</span></a><span> which basically means it will make up answers and give false answers.  Think of AI as one big Mad Libs game where it is largely predicting the next word in a sentence. Sometimes it gets that predicted word wrong, and even more so, early models were terrible at basic math. In other words, you cannot blindly trust AI.  There are some things you must fact check.</span></p></li><li><p><strong><span>Semantic layers and context</span></strong><span>. </span><a href="https://www.techtarget.com/data-technologies/news/366646162/Potential-consequences-are-severe-when-AI-agents-lack-context"><span>Semantic layers</span></a><span> are a layer that sits above the data sources to tell AI where to find data, what particular terms mean, and how that definition varies by role or region. It ensures accurate AI. Humans are good at ambiguity, AI is not. So when you are at a restaurant and you order &#8220;Fisn and Chips,&#8221; the waiter knows you want french fries with that fish (Chips is the British phrase for what Americans call french fries).  Get the context wrong and AI</span></p></li></ol><p><span>The teens that will come out ahead in the next decade are not the ones who let AI think for them, or those that refuse to use it at all.  They are the students who know what is still important: the questions to wrestle with before you begin to format a prompt, the first draft that you struggled through by hand, and the human relationships that develop your experience. AI can be a thought partner, but only for a teen who still thinks for themselves. That distinction, more than any particular skill, is what separates agency from dependence and success from mediocrity.</span></p><p><em><span>Thank you to </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Charlotte Dungan&quot;,&quot;id&quot;:478205145,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/03cc4f91-fd5b-4cce-94a8-c9edaadabdeb_1024x1024.jpeg&quot;,&quot;uuid&quot;:&quot;18cdc573-c9a8-4ad9-b30f-abea6754b279&quot;}" data-component-name="MentionToDOM"></span>, Director of Education at MCF, <span>for posing this question to me and thank you to both Charlotte and </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Jane Smith&quot;,&quot;id&quot;:327687457,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4a9a1fe-6f43-4cb2-94fa-d23bc236149c_896x896.png&quot;,&quot;uuid&quot;:&quot;774613f0-f3dc-4a61-99cc-0965fed49efe&quot;}" data-component-name="MentionToDOM"></span> <span>for feedback as I wrote this, sometimes on post-it-notes, but mostly by computer! </span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Nobody Asks One Question]]></title><description><![CDATA[Why drill-down breaks, and why the most expensive thing in your architecture is the one you can't remove.]]></description><link>https://brainsandbots.substack.com/p/nobody-asks-one-question</link><guid isPermaLink="false">https://brainsandbots.substack.com/p/nobody-asks-one-question</guid><dc:creator><![CDATA[Sonny Rivera]]></dc:creator><pubDate>Wed, 05 Aug 2026 15:46:40 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/142a9c21-8aec-4b61-a6c8-77d06e0d253a_1043x578.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Part 2 of 2. [Part 1: "<a href="/__u/brainsandbots.substack.com/p/you-cant-budget-a-probability">You Can't Budget a Probability</a>" covers who writes the SQL, and why probabilistic generation gives you a distribution instead of a number.]</em></p><p>In Part 1, I left off wondering why our optimized Claude to Cortex Analyst text-to-sql solution only yielded us 15% improvement in token usage.  Recall what I called it Pillar one, Claude calling MCP with Cortex Analyst using semantic views to get deterministic SQL, governed joins, and metric definitions that don&#8217;t drift between sessions. </p><p>I said that number was a ceiling, not an average. You still don&#8217;t get is predictable costs and a budget you can forecast. This is the post where I show you why. Let&#8217;s start with asking, what happens when your analyst asks the 2nd and 3rd questions and how much that really costs.</p><h3>Why state is the expensive part</h3><p>In an architecture where the model does the orchestration, there are two conversation histories running in parallel, and they are vastly different sizes.</p><div id="datawrapper-iframe" class="datawrapper-wrap outer" data-attrs="{&quot;url&quot;:&quot;https://datawrapper.dwcdn.net/IfJ13/2/&quot;,&quot;thumbnail_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58ed4fc8-2d3d-41f6-8e90-507bee15a252_1220x744.png&quot;,&quot;thumbnail_url_full&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73c3b765-7d40-434f-a57f-33a7d5031bba_1220x868.png&quot;,&quot;height&quot;:424,&quot;title&quot;:&quot;A. What each system carries (tokens)&quot;,&quot;description&quot;:&quot;Both system grow linearly. Only one carries result sets.&quot;,&quot;belowTheFold&quot;:false}" data-component-name="DatawrapperToDOM"><iframe id="iframe-datawrapper" class="datawrapper-iframe" src="https://datawrapper.dwcdn.net/IfJ13/2/" width="730" height="424" frameborder="0" scrolling="no"></iframe><script type="text/javascript">!function(){"use strict";window.addEventListener("message",(function(e){if(void 0!==e.data["datawrapper-height"]){var t=document.querySelectorAll("iframe");for(var a in e.data["datawrapper-height"])for(var r=0;r<t.length;r++){if(t[r].contentWindow===e.source)t[r].style.height=e.data["datawrapper-height"][a]+"px"}}}))}();</script></div><p><strong>Snowflake&#8217;s history</strong> carries your questions and the generated SQL (see the Cortex Analyst line on the chart above). Cortex Analyst does support multi-turn Q&amp;A but you pass the prior prompts and response back in a <span data-color="#6aa84f" style="color: rgb(106, 168, 79);">messages</span> array. So, Snowflake built a dedicated agent whose only job is to take that history plus your new question and reframe it into one fully-specified standalone question.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Elegant design. But look at what is in that array: questions and SQL. Ten turns is maybe two thousand tokens.</p><p><strong>The model&#8217;s history</strong> carries your questions, the generated SQL, and every result set you have looked at.  That&#8217;s the Claude line from above. Ten turns at a few thousand tokens per result set is roughly fifty thousand tokens or more, <strong>re-sent on every turn</strong>.</p><p>Same conversation. Two histories. One of them is an order of magnitude larger than the other. And, it&#8217;s the one being metered per token.</p><p>This is the single most useful reframe I can offer: <strong>the expensive context is not the reasoning. It is the data</strong>. And in architectures 1 and 2, only one of the two systems is holding the data.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>That explains the limitation we opened with. Cortex Analyst cannot answer &#8220;what&#8217;s the revenue of the second product&#8221; because result rows were never in its history to begin with. It can refine a question so that &#8220;now just the last 7 days&#8221; works fine, because that only needs the prior question. It can never reference the result or the answer.</p><p>So who does resolve it, when the composite system gets it right? The model does. The rows are sitting in the model&#8217;s context window. It reads them, works out that &#8220;the second product&#8221; means <span data-color="#45818e" style="color: rgb(69, 129, 142);">Widget Pro</span>. Next the model rewrites the follow-up question as a self-contained question, and sends that to the semantic service.</p><p>Read that paragraph again, because there is something important in it. <strong>The frontier model has become your state layer.</strong> Not by design. By default, because it was the only component holding the results. You didn&#8217;t explicitly make this decision, its a consequence of the architecture.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">&#129504;&#129302; Enjoying <strong>Brains &amp; Bots? </strong>Subscribe for weekly takes on data, AI, and where they collide!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>The optimization that isn&#8217;t</h3><p>Once a team sees the token bill, someone usually proposes an obvious fix: stop sending all those rows to the model. Let&#8217;s aggregate the answer and just send the top 10.</p><p>Wait what? That breaks everything. The moment the aggregated rows leave the context, every follow-up question breaks, because the model no longer holds the facts it needs to resolve a reference. Yes you have optimized but at what cost. You have taken away the very thing that makes drill-down work at all.</p><p>Unfortunately that means that the most expensive component in your text-to-sql architecture is also load-bearing. You cannot remove it, and you cannot shrink it without losing the conversational behavior you built the system for. WTH!</p><div class="pullquote"><p>The expensive context is not the intent resolution. It is the data<span>.</span></p></div><h3>Tokens are not outcomes</h3><p>Let&#8217;s step back from the mechanics of our text-to-sql for a second and look at what you are actually getting for you money.</p><p>You are not paying for insights. You are paying for tokens processed. Insights and tokens align reasonably well if you can get them on turn one. But&#8230; they come apart badly after that.</p><p>By turn ten, most of what you are paying for is re-reading data you already have. The model is not learning anything new about your business when it re-ingests the same six category totals for the ninth time.  It&#8217;s being given the same rows again and again because it has no memory. At this point, it&#8217;s just overhead with a metered rate.</p><p>Worse, is what I call Value Inversion, the relationship between the business value and the tokens used actually inverts. Let me explain, the questions that generate the most tokens are often the questions that generated the least value. Meaning, that failed query that got retried, the ambiguous join that took three attempts, the long exploratory session that ended in &#8220;never mind&#8221;,  those cost the most and provided the least. You end up paying the most per unit of insight for your worst outcomes.</p><p>Compare that to how the rest of your data stack is priced. A warehouse query costs what it costs based on the work it does (bytes scanned, compute consumed, etc). Run the same query twice and you pay roughly twice, which is annoying but honest.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> Nothing about that model charges you more for re-reading a result you already have.</p><div class="pullquote"><p>That is the rub. Token metering serves the provider&#8217;s business model, not yours.</p></div><h3>What happens to the cost curve</h3><p>Because each turn re-sends everything before it, cost grows faster than the number of turns. Turn one is cheap. Turn ten is expensive. Turn thirty is punishing. </p><p>Recent research auditing frontier model costs across multi-turn agentic benchmarks found exactly that. The replayed message history is the largest single cost component in multi-turn agent workloads, not the model&#8217;s reasoning.</p><div id="datawrapper-iframe" class="datawrapper-wrap outer" data-attrs="{&quot;url&quot;:&quot;https://datawrapper.dwcdn.net/OEhsC/2/&quot;,&quot;thumbnail_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/07078b5f-f3cb-4a27-af57-36f5be8f0b44_1220x830.png&quot;,&quot;thumbnail_url_full&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/24ce15b7-f03d-4f13-9364-0bd4cc2fcb21_1220x954.png&quot;,&quot;height&quot;:467,&quot;title&quot;:&quot;B. What you actually pay for&quot;,&quot;description&quot;:&quot;Stateless replay makes cost quadratic for multiple turns.&quot;,&quot;belowTheFold&quot;:true}" data-component-name="DatawrapperToDOM"><iframe id="iframe-datawrapper" class="datawrapper-iframe" src="https://datawrapper.dwcdn.net/OEhsC/2/" width="730" height="467" frameborder="0" scrolling="no" loading="lazy"></iframe><script type="text/javascript">!function(){"use strict";window.addEventListener("message",(function(e){if(void 0!==e.data["datawrapper-height"]){var t=document.querySelectorAll("iframe");for(var a in e.data["datawrapper-height"])for(var r=0;r<t.length;r++){if(t[r].contentWindow===e.source)t[r].style.height=e.data["datawrapper-height"][a]+"px"}}}))}();</script></div><p>Prompt caching helps, and you should absolutely use it. It cuts the cost of your system prompts and tool schemas to roughly a tenth. But caching protects the scaffolding of your system, not the data itself. Every new result set is new content. Cache reads soften the curve. They do not flatten it.</p><p>This is why &#8220;start a new session&#8221; turns out to be real cost advice and not a joke. Yes, it good for cost management and it&#8217;s real accuracy advice too. Both model providers and the broader research note that long or topic-shifting conversations degrade model performance. Cool! The same hygiene helps both problems, which should tell you something about how tightly coupled cost and correctness are here.</p><h3>The thing underneath all of this</h3><p>Every architecture pays for semantic resolution. There is no configuration where that work disappears. The only question is when you pay.</p><p>You pay once, at design time, by modeling your business logic into a governed layer that generates deterministic SQL. Or you pay every turn, at inference time, by having a probabilistic system re-derive that logic from scratch. That means pay the expensive output-token cost; coupled with a variance you cannot forecast and an answer that might differ from the last one.</p><p>That&#8217;s the actual build-versus-buy decision. It was never really about a chat interface or natural language input those are the easy parts now. The chat interface has been commoditized. What has not been commoditized is self-service analytics, knowing what your data means, holding onto that meaning between questions, and deterministically answering a question.</p><p>This is why I have come around to thinking of the semantic layer as a FinOps concern and not only a governance one. I have spent most of my career arguing for semantic layers on correctness and trust grounds, and those arguments are good, but they have not always won the budget fight. This one just might, because it shows up on an invoice.</p><p>For disclosure: I work at ThoughtSpot, and our architecture sits in that third category from Part 1. ThoughtSpot utilizes reasoning models for semantic resolution and determining intent but doesn&#8217;t meter AI tokens. It&#8217;s deterministic SQL generation from a governed semantic layer, orchestration that selects best fit models rather than exposing you to any single model&#8217;s cost profile, and result sets that never enter a context window. You should evaluate that claim as skeptically as any other vendor claim, and you should hold anyone else in this category to exactly the same standard.</p><p>But test the architecture, not the demo. Ask the second question. Then ask the third. Then the forth. Lastly measure cost per insight not cost per million tokens.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/p/nobody-asks-one-question?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! If you enjoy it, share it with friend or leave a comment</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/p/nobody-asks-one-question?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/brainsandbots.substack.com/p/nobody-asks-one-question?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/p/nobody-asks-one-question/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/brainsandbots.substack.com/p/nobody-asks-one-question/comments"><span>Leave a comment</span></a></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="https://www.snowflake.com/en/engineering-blog/cortex-analyst-multi-turn-conversations-support/"><span>Multi-Turn Conversations in Cortex Analyst</span></a><span> (on the dedicated agent that reframes follow-ups into fully-specified standalone questions). </span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Feel free to look back at <a href="/__u/brainsandbots.substack.com/p/you-cant-budget-a-probability">Part 1</a>, for a reminder.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>You may actually pay less! Caching!</p></div></div>]]></content:encoded></item><item><title><![CDATA[You Can't Budget a Probability ]]></title><description><![CDATA[Part 1 of 2: What DIY text-to-SQL actually costs and why you can't forecast it.]]></description><link>https://brainsandbots.substack.com/p/you-cant-budget-a-probability</link><guid isPermaLink="false">https://brainsandbots.substack.com/p/you-cant-budget-a-probability</guid><dc:creator><![CDATA[Sonny Rivera]]></dc:creator><pubDate>Wed, 29 Jul 2026 12:09:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/389822b0-3f3a-467e-b3e6-24c8dcab75b9_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I have been trying to answer a simple question for a few a while now: what does one business question cost?</p><p>I know what the cost per million tokens is for a frontier model like Claude Sonnet. That number is on Anthropic&#8217;s pricing page and it tells you almost nothing. I&#8217;m talking about a business question like &#8220;What are sales by category for the last 30 days.&#8221; What did that cost the organization, and what will thousands of them cost next quarter?</p><p>It&#8217;s a trick question, I can&#8217;t give you a number because there isn&#8217;t just one. The answer is it&#8217;s a distribution, and it&#8217;s rather wide. And the shape is determined by two architectural decisions that most teams never think about.</p><p>Here is the conversation that sent me looking.</p><blockquote><p><strong><span>&#8220;What are my products?&#8221;</span></strong></p><p><strong><span>&#8220;What&#8217;s the revenue of the second one?&#8221;</span></strong></p></blockquote><p><span>Every analyst I have ever worked with does some version of this a hundred times a week. They ask, evaluate, narrow, pivot, and drill. It&#8217;s not the nature of the question, heck it&#8217;s not even a clever question or an edge case. But it is the job.</span></p><p><span>As a data leader and Snowflake Data Superhero I know a little bit about this. It&#8217;s a two step process and you need access to the result set from the original question. Yet, Snowflake states  in their own product docs that Cortex Analyst can&#8217;t access results from previous SQL queries.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span> The example above is the  identical scenario from their documentation:  ask for a list of products, then ask about the second one to demonstrate that the service has no way to know which product you mean.</span></p><div class="callout-block" data-callout="true"><p><strong><span>Access to the results of previous SQL queries</span></strong></p><p>Cortex Analyst doesn&#8217;t have access to results from previous SQL queries. For example, if you first ask, &#8220;What are my products?&#8221; and then ask, &#8220;What is the revenue of the second product?&#8221;, Cortex Analyst cannot refer to the list of products from the first query to get the second product.&#8221;</p></div><p><span>I want to be careful here, because this is not a knock on Snowflake. It is the opposite. Snowflake is doing the right thing that most of this industry is not doing: publishing the architectural limits of their own service in plain language. I wish more vendors did it. So, I&#8217;m more concerned about </span><em><strong><span>why</span></strong></em><span> the limitation exists. The answer falls out from an architectural decision, and it&#8217;s this same decision that determines what your bill looks like at the end of the month.</span></p><p><span>Most teams evaluating natural-language analytics are asking about accuracy, I do it all the time. Can it get the SQL right? And the answer right? Those are important questions, but also incomplete. The way I see it, there are two structural questions underneath it, and you probably answer both before you write a line of code, usually without realizing you have answered them at all.</span></p><blockquote><p><strong><span>Who resolves the semantics?</span></strong><span> That is: who understands the user intent and turns &#8220;revenue by category&#8221; into a join path, a grain, a filter, and an aggregate.</span></p><p><strong><span>Who holds the state?</span></strong><span> That is: who remembers the last question, the last query, and critically, the last result set.</span></p></blockquote><p><span>The first question determines how </span><em>predictable</em><span> your cost is. The second determines how </span><em><span>fast</span></em><span> your cost grows. Your cost is those two factors multiplied together. Nearly everything expensive and surprising about DIY text-to-SQL is downstream of these two answers.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">&#129504;&#129302; Enjoying <strong>Brains &amp; Bots? </strong>Subscribe for weekly takes on data, AI, and where they collide!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong><span>Three architectures</span></strong></h3><p>Let me put some structure around this. I have been mapping the natural-language analytics landscape against those two axes, and there are essentially three configurations. I built this out concretely against a Snowflake MCP setup using a typical star schema, fact_sales with date, product, category, customer, order, city, and state dimensions, and one deliberately easy question: <em>What are sales by category, by day, for the last 30 days?</em></p><h4><strong><span>Architecture 1: The LLM writes the SQL</span></strong></h4><p><span>You connect a frontier model to your warehouse over MCP, expose a SQL execution tool, and let it rip. No semantic model. The model, let&#8217;s say Claude Sonnet 4.x, discovers your schema, infers the join paths, guesses the grain, writes the SQL, and interprets the results.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gm1x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 424w, /__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 848w, /__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gm1x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png" width="1299" height="821" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:821,&quot;width&quot;:1299,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:183442,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://brainsandbots.substack.com/i/208080050?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 424w, /__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 848w, /__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gm1x!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33b98517-2831-4f29-9505-d2c0e57eaa74_1299x821.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The frontier model, let&#8217;s say Claude Sonnet, determines every piece of business meaning. Which table is &#8220;category&#8221;? Which key ties the fact to the dimension? What &#8220;revenue&#8221; actually means in your organization? All this is re-derived by a probabilistic system, from scratch, on every single turn.<em> You pay for that derivation in output tokens, and you get a slightly different derivation every time</em>.</p><h4>Architecture 2: A service writes the SQL from a semantic model</h4><p>Turn on something like Cortex Analyst. Now Claude uses skills and tools to send your prompt to Snowflake Cortex Analyst.  The text-to-SQL happens inside Snowflake, against a declared semantic view with real join paths and metric definitions. The SQL that comes back is deterministic and governed. This is a genuine improvement in correctness, and I do not want to undersell it.</p><p>But watch what happens to the orchestration. Cortex Analyst returns SQL text, not result rows.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> So the model has to take that SQL and hand it back to a separate execution tool. It pays output tokens to relay a query it did not author and cannot improve. I have started calling this the Echo Tax, and it is significant. It basically cancels out all of the savings you get from skipping schema discovery.</p><p><strong>And the results still land in the model's context window. All the results, every row. Driving up context window size and token costs.</strong></p><h4>Architecture 3: The semantic layer resolves it, and holds the state</h4><p>The third configuration is one where semantic resolution happens deterministically in a governed layer, and result sets never enter a context window at all. The model resolves intent. The layer generates the SQL. Results go to a rendering layer instead of a context window or thinking tokens.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!M9-7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 424w, /__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 848w, /__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 1272w, /__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_webp, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!M9-7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png" width="1440" height="921" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:921,&quot;width&quot;:1440,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:196031,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://brainsandbots.substack.com/i/208080050?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_424, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 424w, /__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_848, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 848w, /__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_1272, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 1272w, /__u/substackcdn.com/image/fetch/$s_!M9-7!, /__u/brainsandbots.substack.com/w_1456, /__u/brainsandbots.substack.com/c_limit, /__u/brainsandbots.substack.com/f_auto, /__u/brainsandbots.substack.com/q_auto:good, /__u/brainsandbots.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa140d9e-799b-4a4d-930a-8157a71dc081_1440x921.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>This is where ThoughtSpot&#8217;s deterministic semantic layer and API resides. Yes, ThoughtSpot is where I work. I will come back to that at the end rather than pretending I have no stake in this.</p><div id="datawrapper-iframe" class="datawrapper-wrap outer" data-attrs="{&quot;url&quot;:&quot;https://datawrapper.dwcdn.net/EFCQm/3/&quot;,&quot;thumbnail_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/623a3661-a735-4333-9021-1fb4d43ca506_1220x696.png&quot;,&quot;thumbnail_url_full&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/013b4ba5-262e-46ff-b5cd-112e59825259_1220x820.png&quot;,&quot;height&quot;:400,&quot;title&quot;:&quot;Side by Side Comparison&quot;,&quot;description&quot;:&quot;Semantic Resolution and State Management&quot;,&quot;belowTheFold&quot;:true}" data-component-name="DatawrapperToDOM"><iframe id="iframe-datawrapper" class="datawrapper-iframe" src="https://datawrapper.dwcdn.net/EFCQm/3/" width="730" height="400" frameborder="0" scrolling="no" loading="lazy"></iframe><script type="text/javascript">!function(){"use strict";window.addEventListener("message",(function(e){if(void 0!==e.data["datawrapper-height"]){var t=document.querySelectorAll("iframe");for(var a in e.data["datawrapper-height"])for(var r=0;r<t.length;r++){if(t[r].contentWindow===e.source)t[r].style.height=e.data["datawrapper-height"][a]+"px"}}}))}();</script></div><p><strong><sub>Table 1 - Side by Side Comparison</sub></strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>The thing I want you to notice is that architectures 1 and 2 differ on pillar one (SQL Generation) and are identical on pillar two (Who holds results). Both put your result rows in a context window. That is why the token savings from adding a semantic service are far smaller than anyone expects. In my modeling of that single easy question, roughly 15 percent on the happy path, and effectively nothing on output tokens.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>15 percent, what? For adding a semantic layer. That number surprised me enough that I went back and re-derived it twice.</p><p>The reason is simple once you see it: the biggest line item in both architectures is the result payload, and it is byte-identical in both. <strong>A semantic layer that only fixes who writes the SQL has fixed the correctness problem and left the cost problem completely untouched</strong>.</p><p>And that 15 percent is a best case. It&#8217;s just one question, on a fresh session, with nothing going wrong. It&#8217;s not a real business conversation, and the reason why is the second pillar: who holds the state. </p><div class="pullquote"><p>A semantic layer that <strong>only</strong> fixes who writes the SQL has fixed the correctness problem and left the cost problem completely untouched.</p></div><h4>Why probabilistic answers are unbudgetable</h4><p>Everything else up to this point is structural. You can model it, draw it, and forecast it. And pillar one is much more problematic, because it introduces variance you cannot forecast at all.</p><p>I have been going through a recent study, &#8220;The Price Reversal Phenomenon&#8221;, auditing what frontier models actually cost in practice, versus their listed per-token price.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> And here three findings should change how anyone in a FinOps role thinks about model selection and costs..</p><p><strong>The same query doesn&#8217;t cost the same thing twice</strong>. Send an identical question to an identical model over and over, and the cost swings wildly.  In some cases, the most expensive run cost almost ten times as much as the cheapest run. Internal research from Databricks says the same thing, that per-token pricing is an unreliable indicator of overall task cost.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p>Think of it this way; each new instance samples different reasoning paths, produces different levels of &#8220;thinking&#8221;, and uses fewer or more tokens. So even when the answer is correct, the reasoning paths may be different, implying that changing the prompt does not remove that randomness.</p><p>Key Takeaway : It demonstrates how these models work, and you can&#8217;t just re-prompt your way into efficient token usage.</p><p>So there is no such thing as a cost per query. It&#8217;s a range. Remember, when you hand finance a number and you&#8217;ve handed them just one roll of the dice.</p><p><strong>Cheaper-per-token models routinely cost more per task</strong>. Lower-priced models tend to need more turns. They hit more dead ends, more backtracking, more looping. And each extra turn costs more than the last. When a run fails, a &#8220;cheaper&#8221; model can end up costing 3&#8211;10&#215; more than what a stronger model would have spent to succeed. Meaning, you pay the biggest premium for the answers you don&#8217;t actually get. I can&#8217;t imagine a worse cost profile for a production system.</p><p><strong>Averages obscure the actual risk to your budget</strong>. When comparing models on average cost, the cheaper model was secretly more expensive about 32% of the time.</p><h4>The Prototype Trap</h4><p>A common scenario that I see is what I call the Prototype Trap; I spoke about this at the Data + AI Summit last month.   Teams are doing a prototype or a POC but they&#8217;re looking at only the easy questions. <br><br>On easy questions, almost every model performs well, uses very few turns, and costs about the same. Price reversals largely vanish. It&#8217;s on hard questions that have many turns, real ambiguity, and genuine reasoning that the differences.  Those are driving the cost explosion.<br><br>The typical prototype, someone asks for revenue by region, got a right answer in four seconds. The room was sold. But real production use cases run the hard ones. The ambiguous ones, the ones with three plausible join paths, the ones where &#8220;customer&#8221; means something different in billing than in CRM.</p><p>The prototype didn&#8217;t actually test the use cases that determines whether it really works or not. </p><h4>So, Here&#8217;s what I would actually do</h4><p>I&#8217;d question whether or not I should you built it myself or buy it? But, that&#8217;s probably a topic for another post. So, Let&#8217;s get back on track. <br><br>I&#8217;ll be frank about the third architecture&#8217;s failure modes, because they are real. Semantic layers are intentional by design and they demand an upfront modeling investment in order to see the greatest value. All semantic layers constrain the question space to the data product that has been modeled. When you ask something outside the model, a good semantic layer fails closed. It tells you it cannot answer rather than guessing. Some people find that infuriating. I think it is the entire point, but I understand the frustration. So, if your organization cannot commit to the modeling discipline then you will not see the benefit.</p><p>To achieve a predictable analytics budget alongside reliable metrics and genuine drill-down capabilities, focus your architectural requirements on these key demands derived from the two pillars.</p><p><strong>Aggregate in the warehouse, not in the model</strong>. Six category totals instead of a hundred and eighty daily rows is not a one-turn saving. Those rows are re-ingested on every remaining turn of the session. This is the single highest-leverage change available in a DIY architecture.</p><p><strong>Keep results out of the context window</strong>. If your architecture requires result sets to live in a model&#8217;s context for follow-up questions to work, you have accepted superlinear cost growth as a permanent condition.</p><p><strong>Treat session length as a cost control</strong>. New topic, new session. It bounds context growth and improves accuracy simultaneously.</p><p><strong>Demand determinism for anything governed.</strong> Probabilistic SQL is always wrong but you cannot budget a probability, and you cannot audit one either.</p><p><strong>Measure what a successful answer costs</strong>. Per-token pricing tells you nothing here, and neither does a clean run on an easy question. Track three numbers instead: how often you get a right answer, what an average right answer costs, and what the worst runs cost. That last one is often skipped, and it&#8217;s the one that wrecks a forecast. An evaluation that throws out the failures is measuring a system you don&#8217;t have.</p><h4>What this leaves you with</h4><p>Pillar one, where the semantics live,  is where everyone does the optimization. The most logical solution is to move the semantic resolution out of the frontier model and  into a semantic model or to a lesser extent semantic views. Great start, you get deterministic SQL, governed joins, and metric definitions that don&#8217;t drift between sessions. Take the win.</p><p>What you don&#8217;t get is predictable costs and a budget you can forecast. You narrowed the distribution with fewer retries and unambiguous joins but it&#8217;s still a distribution So, you&#8217;ll get about 15% improvement on the happy path which is only the easiest questions in a single turn.</p><p>That&#8217;s because we have not discussed the second pillar yet. Recall, in architectures 1 and 2, every row you look at lands in a context window and it gets re-sent on every turn. You are still paying the Echo tax and making the data the most expensive part of the cost.</p><p>In Part 2, &#8220;<a href="/__u/brainsandbots.substack.com/p/nobody-asks-one-question">Nobody Asks One Questions&#8221;</a>, we get to what happens when your analyst asks the 2nd and 3rd questions and how much that really costs.</p><p style="text-align: center;">Thanks for reading Brains &amp; Bots!<br>Share this essay with a friend or coworker. Your support is appreciated &#128153;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share Brains &amp; Bots&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/brainsandbots.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share Brains &amp; Bots</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://brainsandbots.substack.com/p/you-cant-budget-a-probability/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/brainsandbots.substack.com/p/you-cant-budget-a-probability/comments"><span>Leave a comment</span></a></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span>Snowflake, </span><em><a href="https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst"><span>Cortex Analyst</span></a></em><a href="https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst"><span> documentation</span></a><span>. Limitations on access to results from previous SQL queries. </span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span>Snowflake, </span><em><a href="https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-agents-mcp"><span>Cortex Agents MCP server</span></a></em><a href="https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-agents-mcp"><span> documentation</span></a><span>. The Analyst tool returns the SQL statement in its output; SQL execution tool responses are truncated at 250 KB.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>&#8220;Super&#8209;linear&#8221; is a standard math / computer science term, not just a casual phrase I made up. It means &#8220;grows faster than linearly&#8221; as something gets larger.</span><br><span>Example: if cost per task goes like n&#178; while tasks are n, that cost is super&#8209;linear in n, because it increases more quickly than a straight line.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>My own token modeling across both architectures, single question, fresh session. These are modeled estimates, not measured values so verify with /cost in a Claude Code session or the usage block on a Messages API response before treating them as benchmarks.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><a href="https://arxiv.org/html/2603.23971v2#A5">The Price Reversal Phenomenon</a>: When Cheaper Reasoning Models End Up Costing More.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Databricks (2026). <a href="https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase">Why Estimating Prompt Cost is Hard</a></p></div></div>]]></content:encoded></item></channel></rss>