<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Hamel’s Substack]]></title><description><![CDATA[My newsletter lives at https://ai.hamel.dev]]></description><link>https://hamelhusain.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!pph_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F75dffc2c-e51f-410c-97ef-d61e5c24af76_256x256.png</url><title>Hamel’s Substack</title><link>https://hamelhusain.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 12:01:37 GMT</lastBuildDate><atom:link href="/__u/hamelhusain.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Hamel Husain]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[hamelhusain@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[hamelhusain@substack.com]]></itunes:email><itunes:name><![CDATA[Hamel Husain]]></itunes:name></itunes:owner><itunes:author><![CDATA[Hamel Husain]]></itunes:author><googleplay:owner><![CDATA[hamelhusain@substack.com]]></googleplay:owner><googleplay:email><![CDATA[hamelhusain@substack.com]]></googleplay:email><googleplay:author><![CDATA[Hamel Husain]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[AI Product Engineering Notes]]></title><description><![CDATA[Notes from 13 sessions on evals, context, and systems.]]></description><link>https://hamelhusain.substack.com/p/ai-product-engineering-notes</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/ai-product-engineering-notes</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Sun, 30 Aug 2026 23:56:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1FEF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI product engineering is the work of turning a model into a useful product. Evals set the foundation for improving an AI product, but there are many knobs to turn, each with different trade-offs.</p><p>The lowest hanging fruit is to optimize retrieval and context. Once that&#8217;s done, improve your systems and harness. Consider post-training after the other approaches are exhausted.</p><p>I hosted an <a href="https://maven.com/lls/21f487">AI Product Engineering series</a> to explore these topics. I&#8217;ve summarized all 13 sessions, organized by theme, with links to the recordings and source materials. The full series is 9.5 hours. The notes take about 20 minutes to read.</p><p>All 13 sessions are outlined below. Click the table to read the <a href="https://hamel.dev/notes/llm/ai-product-engineering/">full post</a> on my website.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://hamel.dev/notes/llm/ai-product-engineering/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 424w, /__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 848w, /__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!1FEF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png" width="821" height="1050" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1050,&quot;width&quot;:821,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:154910,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://hamel.dev/notes/llm/ai-product-engineering/&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hamelhusain.substack.com/i/213470107?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 424w, /__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 848w, /__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1FEF!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76b69af1-e187-4457-833e-c0f1a4e0f940_821x1050.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you want to stay on Substack, here are links to the notes from all 13 sessions.</p><p><strong>Evals</strong></p><ul><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/evals-data-agents.html">Evals for Data Agents</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/evals-model-cascades.html">Cut Classification Costs With a Model Cascade</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/evals-error-analysis.html">Automating Error Analysis</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/evals-production.html">Case Study: Putting Evals Into Production</a></p></li></ul><p><strong>Context</strong></p><ul><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/context-multivector-retrieval.html">An Intro to Multivector Retrieval</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/context-search-agents.html">How to Improve Search Agents</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/context-embedding-models.html">How to Optimize Retrieval Embeddings</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/context-ocr.html">How to Choose an OCR Model</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/context-subtext.html">Steer AI Writing With Footnotes for Agents</a></p></li></ul><p><strong>Systems</strong></p><ul><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/systems-inference-latency.html">Debugging Inference Latency</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/systems-open-model-memory.html">Open Weight Model Economics</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/systems-agent-sandboxes.html">The Case for Agent Sandboxes</a></p></li><li><p><a href="https://hamel.dev/notes/llm/ai-product-engineering/systems-post-training.html">Turn Eval Results Into a Better Model</a></p></li></ul><p><em>Nearly every improvement in these notes starts with good evals. If you want to work through evals with a live group, the <a href="https://maven.com/parlance-labs/evals?promoCode=ai-product-eng-c6">AI Evals course</a> has hands-on exercises and office hours.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://hamelhusain.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Hamel&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Evals For Engineers & Product Managers]]></title><description><![CDATA[Resources to help you learn more about AI Evals]]></description><link>https://hamelhusain.substack.com/p/ai-evals-for-engineers-and-product-c8a</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/ai-evals-for-engineers-and-product-c8a</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Fri, 21 Aug 2026 17:11:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dluV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Because you&#8217;re subscribed to this newsletter, chances are you already know the basics of evals. Over the last two months, Shreya and I ran a free AI Product Engineering series, fifteen live sessions with guests who build AI products in production. Live talks only take you so far. The best way to get good at evals is guided practice with feedback.</p><p>That&#8217;s why I&#8217;m writing. Our next live cohort of AI Evals for Engineers &amp; PMs starts September 5. You have two paths from here. You can join us for the live course, or keep learning on your own with the free sessions below. Here&#8217;s a guide for both.</p><h2>Path 1: Join the live cohort</h2><p>After four weeks, you&#8217;ll know how to find where your AI product fails and measure each failure with evaluators you trust. You&#8217;ll also know how to put those evals in CI and improve the product on accuracy and cost. The course runs September 5 to October 3.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://maven.com/parlance-labs/evals?promoCode=substack-25-c4&quot;,&quot;text&quot;:&quot;Learn more about the course&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://maven.com/parlance-labs/evals?promoCode=substack-25-c4"><span>Learn more about the course</span></a></p><p><strong>What our students are saying:</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dluV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 424w, /__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 848w, /__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dluV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png" width="1456" height="1069" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1069,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 424w, /__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 848w, /__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dluV!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd455fcaf-d97b-4af3-ba76-9c44994b15dc_2354x1728.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>See more </span><a href="https://testimonial.to/ai-evals-course/all">testimonials here</a><span>.</span></p><p><strong>What you&#8217;ll learn</strong></p><p>The course is built around one running example. You build a real support agent, then take it all the way to production evals. It runs in five modules.</p><ol><li><p><strong>Building agents.</strong> You build a support agent and instrument it so you can measure how it behaves, then generate the synthetic data and scenarios the later modules use.</p></li><li><p><strong>Error analysis.</strong> You read the agent&#8217;s traces to find where it fails, then build a small set of evaluators you trust to measure each failure.</p></li><li><p><strong>CI/CD for agents.</strong> You turn those failures into a regression suite that runs before you deploy, and add lightweight monitoring for after you ship.</p></li><li><p><strong>Safety and adversarial evaluation.</strong> You red-team the agent for attacks like prompt injection and tool misuse, then turn each successful attack into a test you keep.</p></li><li><p><strong>Improving agents.</strong> You compare versions of the agent to raise accuracy, then cut cost without losing quality.</p></li></ol><p>There&#8217;s also a bonus session on evals interview prep.</p><p><strong>What&#8217;s included</strong></p><ul><li><p>About 10 hours of live office hours</p></li><li><p>Lifetime access to our Discord community of past students</p></li><li><p>A 150+ page course reader</p></li><li><p>Access to future cohorts, so you can keep coming to office hours and get the latest material</p></li><li><p>6 months of access to our AI Evals Assistant, trained on everything we&#8217;ve made on evals</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://maven.com/parlance-labs/evals?promoCode=substack-25-c4&quot;,&quot;text&quot;:&quot;Learn more about the course&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://maven.com/parlance-labs/evals?promoCode=substack-25-c4"><span>Learn more about the course</span></a></p><h2>Path 2: Keep learning with the free sessions</h2><p>If the live course isn&#8217;t right for you now, that&#8217;s fine. Start with the reading below, then go deeper with the sessions.</p><p><strong>Essential reading</strong></p><ul><li><p><a href="https://hamel.dev/blog/posts/evals-faq/">Evals FAQ</a>. The most common eval questions answered, from the basics to advanced.</p></li><li><p><a href="https://hamel.dev/blog/posts/eval-smell/">&#8220;It&#8217;s Hard to Eval&#8221; Is a Product Smell</a>. Your product should make it easy to verify AI outputs.</p></li><li><p><a href="https://hamel.dev/blog/posts/field-guide/">A Field Guide to Rapidly Improving AI Products</a>. The evaluation and improvement process from more than 30 production builds.</p></li><li><p><a href="https://hamel.dev/blog/posts/llm-judge/">Using LLM-as-a-Judge For Evaluation: A Complete Guide</a>. A step-by-step guide to judges you can trust.</p></li><li><p><a href="https://parlance-labs.com/blog/posts/auto-evals.html">Do Automated Evals Work?</a>. We checked 100 human-labeled traces against automated eval systems.</p></li><li><p>Shreya&#8217;s research: <a href="https://arxiv.org/abs/2404.12272">Who Validates the Validators?</a>. How to align LLM judges with human labels.</p></li></ul><p><strong>The free sessions</strong></p><p>The whole <a href="https://maven.com/lls/21f487">AI Product Engineering series</a> is free to watch. Here are the fifteen, grouped by topic.</p><p><strong>Evals</strong></p><ul><li><p>Stop Letting Your LLM &#8220;Find the Issues&#8221; In Your Data, with Shreya Shankar. How to use AI to speed up error analysis while you stay in control of what counts as a failure. <a href="https://maven.com/p/49ea8c">video</a></p></li><li><p>Evals Your Whole Team Can Run, with Lucas Machado Rocha. How to make evals simple enough for your whole team to run, using how Nova Escola did it. <a href="https://maven.com/p/447b3e">video</a></p></li><li><p>Can AI Actually Answer Your Data Questions? with Shreya Shankar. How to evaluate data agents on messy real questions, where even the best models get only about 57% right. <a href="https://maven.com/p/4cd26a">video</a></p></li><li><p>The Cold-Start Eval Problem, with Shreya Shankar. How to build evals from day one, before you have a single user. <a href="https://maven.com/p/5fb214">video</a></p></li><li><p>Turn Eval Results Into a Better Model, with Will Brown and Florian Brand. How to turn your eval results into training that improves the model. <a href="https://maven.com/p/7d699b">video</a></p></li></ul><p><strong>Retrieval and search</strong></p><ul><li><p>Scaling Late Interaction to Billions of Documents, with Marek Galovic. How to keep the accuracy of late-interaction retrieval as you scale to billions of documents, without the cost blowing up. <a href="https://maven.com/p/6544a2">video</a></p></li><li><p>What Makes a Good Search Agent? with Nandan Thakur. How to benchmark and debug a search agent, and why the retriever affects accuracy as much as the model. <a href="https://maven.com/p/cbf1b6">video</a></p></li><li><p>Why We Embed Documents, and Why You Should Go Multi-Vector, with Ben Clavi&#233;. When single-vector embeddings fail on long or unusual documents, and when multi-vector is worth the cost. <a href="https://maven.com/p/817899">video</a></p></li><li><p>Stop Picking Embedding Models off the MTEB Leaderboard, with Radu Gheorghe. How to pick the embedding model that fits your own data, when a high leaderboard rank won&#8217;t tell you. <a href="https://maven.com/p/fef4e1">video</a></p></li></ul><p><strong>Cost and open models</strong></p><ul><li><p>Why Your AI Product Feels Slow Even When the Model Is Fast, with Abi Aryan. A hands-on walk through where your latency and cost come from when the model is fast but the product feels slow. <a href="https://maven.com/p/b3f9bf">video</a></p></li><li><p>How to Choose an OCR Model, with Joe Barrow. How to decide when a managed OCR API is enough and when to host your own open model. <a href="https://maven.com/p/95c291">video</a></p></li><li><p>How to Use Open Models Effectively, with Zach Mueller. When an open model can replace a frontier model without losing quality, so you stop overpaying. <a href="https://maven.com/p/9583fd">video</a></p></li><li><p>Stop Paying Full Price for LLM Classification, with Shreya Shankar. How to cut the cost of LLM classification at scale without hurting quality. <a href="https://maven.com/p/35471f">video</a></p></li></ul><p><strong>Agents and context</strong></p><ul><li><p>No Country for Old Text: Stop Wasting Your Context, with Bryan Bischof and Adam Conway. How to keep track of who wrote each part of AI-generated text, and why, so that context stays with it across tools. <a href="https://maven.com/p/ee1f95">video</a></p></li><li><p>Don&#8217;t Build Agents, Build Environments Instead, with Adam Azzam. How to build the fast, dependable environment your agent runs in, which does more for results than the agent code. <a href="https://maven.com/p/0684ab">video</a></p></li></ul><p>Whichever path you pick, I hope this helps you build something that works.</p><p>Best, Hamel</p>]]></content:encoded></item><item><title><![CDATA["It's Hard to Eval" Is a Product Smell]]></title><description><![CDATA[Your product should make it easy to verify AI outputs.]]></description><link>https://hamelhusain.substack.com/p/its-hard-to-eval-is-a-product-smell</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/its-hard-to-eval-is-a-product-smell</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Mon, 29 Jun 2026 20:26:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tI69!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tI69!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 424w, /__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 848w, /__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tI69!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Social card for \&quot;It's Hard to Eval\&quot; Is a Product Smell&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Social card for &quot;It's Hard to Eval&quot; Is a Product Smell" title="Social card for &quot;It's Hard to Eval&quot; Is a Product Smell" srcset="/__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 424w, /__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 848w, /__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tI69!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eb234e9-f32a-4718-9f86-d80b112fa1e7_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>&#128073;<strong> For the best reading experience, read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> with interactive figures.</strong> &#128072;</em></p><p>For the past 3 years, AI evals have been my professional focus.<a href="#notes"><sup>1</sup></a> The most common objection I hear to evals is &#8220;our product is hard to eval&#8221;.</p><p>This objection is a product smell. Artifacts that are hard for you to verify are often hard for users too. In the worst case, users have to redo the work from scratch to verify the output. More importantly, designing your product for ease of verification should come before building evals.</p><p>In this post, I&#8217;ll walk through three products I advised on that faced this issue. I&#8217;ll also show before and after sketches to demonstrate design principles. After these examples, I&#8217;ll discuss how to apply this general pattern to your product.</p><h2>Example 1: the AI data agent</h2><p>Almost every company I&#8217;ve worked with builds an internal AI data agent. You ask it a business question, like what was net revenue for Product A last quarter, and it finds relevant data sources, runs the queries, and provides an answer. The goal of this agent is to reduce dependency on data analysts.</p><p>A common mistake when building AI data agents is to make the answer the only output, as illustrated below.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!aGJg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 424w, /__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 848w, /__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!aGJg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png" width="632" height="277" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cadb91da-4b09-490c-8664-025057451712_632x277.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:277,&quot;width&quot;:632,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data agent before design. The only output is the answer.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data agent before design. The only output is the answer." title="Data agent before design. The only output is the answer." srcset="/__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 424w, /__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 848w, /__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aGJg!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcadb91da-4b09-490c-8664-025057451712_632x277.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Since the only output is the answer, there is nothing here to check. This is a static screenshot for Substack. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to interact with this figure.</figcaption></figure></div><p>In the sketch above, the user has no way to verify the answer beyond redoing work.<a href="#notes"><sup>2</sup></a> A better design is to provide the user with checkable artifacts, informed by how a domain expert might validate the output. Here are techniques I use to validate metrics as a data scientist:</p><ol><li><p>Compare the quantity and any intermediate calculations against a trusted source, like a vetted dashboard or report, or a similar analysis a colleague has already vetted.<a href="#notes"><sup>3</sup></a></p></li><li><p>Confirm the metric definition precisely. A number like net revenue can include or exclude things like returns and discounts.</p></li><li><p>Sanity-check a related quantity. If I can&#8217;t verify the number directly, I pull a related number that should move with it, like units sold or unique customers, and check if the combination is plausible.</p></li><li><p>Look at what is beneath the aggregate. A total can hide problems, so I break it down by dimensions like region or time period and sanity-check the distribution.</p></li><li><p>Read the query. For an important number I look at the SQL to confirm it does what I think, and I tweak it and rerun to test my assumptions.</p></li><li><p>Note anything I could not verify. If a step has no trusted reference to check it against, I flag it instead of presenting it as settled.</p></li></ol><p>Here&#8217;s what a better interface might look like. The two tabs below show the same answer at two levels of detail. The chat reply surfaces the details worth seeing up front, and the notebook holds the full analysis behind the answer. Use the tabs to switch between them.</p><p><em>Substack shows the two tabs as static screenshots below.</em></p><h3>Chat</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-Bk8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 424w, /__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 848w, /__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-Bk8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png" width="632" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:632,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data agent after design. The answer includes sources, assumptions, and verification details.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data agent after design. The answer includes sources, assumptions, and verification details." title="Data agent after design. The answer includes sources, assumptions, and verification details." srcset="/__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 424w, /__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 848w, /__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-Bk8!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c8251f-17f9-4587-b945-6c76ec9aa240_632x630.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The agent surfaces the important details behind the answer. Select the Notebook tab above to see the full analysis generated by the agent. This is a static screenshot for Substack. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to interact with this figure.</figcaption></figure></div><h3>Notebook</h3><p>This is the notebook the agent worked in while producing its answer. Scroll within the figure to see all the cells. The sidebar has a Contents tab for jumping between sections and an Assistant tab for asking follow-ups.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!J7_I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 424w, /__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 848w, /__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!J7_I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png" width="795" height="885" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:885,&quot;width&quot;:795,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Notebook view generated by the data agent, with assumptions, SQL, outputs, and unresolved issues.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Notebook view generated by the data agent, with assumptions, SQL, outputs, and unresolved issues." title="Notebook view generated by the data agent, with assumptions, SQL, outputs, and unresolved issues." srcset="/__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 424w, /__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 848w, /__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J7_I!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc70c5da7-1a53-4f1c-bd08-0f84d9a21406_795x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Static screenshot of the notebook tab. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to scroll within this figure.</figcaption></figure></div><p><em><strong>&#128073; The <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> lets you scroll the notebook, and inspect the figures interactively. &#128072;</strong></em></p><p>There is a lot to unpack here. Here are notable changes:</p><ul><li><p>The agent optionally performs retrieval from vetted analyses, and the interface shows which one was used along with who authored it.</p></li><li><p>There is progressive disclosure of details. The chat reply shows high value items like sources, assumptions, and issues. The user can optionally open an interactive notebook to see the full context.</p></li><li><p>The AI-generated notebook (see notebook tab above) is organized to promote verification: it opens with the assumptions the agent made, like the metric definition and where it came from, then shows the queries it ran and the numbers they returned. It breaks the total down so you can sanity-check the distribution, and it closes with a list of what it could not verify, each item left as a cell you can run.</p></li><li><p>The AI agent is also available in the notebook to help with follow-ups. Finally, the user can publish the notebook back to a knowledge base, where it can be retrieved by future analyses, creating a virtuous cycle.</p></li></ul><p>This design sketch is far from perfect. The point is that the product should help the user verify the answer as a domain expert would. Compare it to the earlier approach, where the only output was the number.</p><p>Data agents like these are not science fiction. Hex<a href="#notes"><sup>4</sup></a> is my favorite product in this genre; it integrates notebooks and chat better than anything I&#8217;ve seen. Here are screenshots from their landing page:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bq8x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 424w, /__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 848w, /__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bq8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png" width="1220" height="1324" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8597218-5324-472c-bbb8-880bc968922b_1220x1324.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1324,&quot;width&quot;:1220,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chat interface.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chat interface." title="Chat interface." srcset="/__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 424w, /__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 848w, /__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bq8x!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8597218-5324-472c-bbb8-880bc968922b_1220x1324.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Chat interface.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!IHgL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 424w, /__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 848w, /__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!IHgL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png" width="1456" height="1103" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1103,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Notebook view which allows the user to see the intermediate steps and data.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Notebook view which allows the user to see the intermediate steps and data." title="Notebook view which allows the user to see the intermediate steps and data." srcset="/__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 424w, /__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 848w, /__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IHgL!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ec7b792-3456-42bf-9972-b74a62105c65_1602x1214.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Notebook view which allows the user to see the intermediate steps and data.</figcaption></figure></div><p>But what does this have to do with evals? If you design your product for verification, annotation becomes less expensive and evals will have better signals to draw from. More importantly, you&#8217;ll provide your users with a better product.</p><h2>Example 2: the PE curriculum builder</h2><p>A founder I advised was building an AI tool that writes physical education lesson plans for K-12 teachers. A teacher enters their constraints, like the grade they teach, how long the class is, whether it meets indoors or outdoors, and what equipment they have. The tool then writes a lesson plan for those constraints. The goal is to save teachers the time they spend planning and give them a plan that fits their class. Here is a sketch of what the product looks like:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HGzo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 424w, /__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 848w, /__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HGzo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png" width="831" height="494" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/39839397-80e6-4d0e-b051-5a40af807baa_831x494.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:494,&quot;width&quot;:831,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;PE lesson plan before design.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="PE lesson plan before design." title="PE lesson plan before design." srcset="/__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 424w, /__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 848w, /__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HGzo!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39839397-80e6-4d0e-b051-5a40af807baa_831x494.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The teacher enters constraints and the tool writes a plan from scratch. The only output is the plan, so its difficult to verify. This is a static screenshot for Substack. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to interact with this figure.</figcaption></figure></div><p>The founder asked me how to eval the lesson plans. I turned the question around: <strong>what does a teacher care about?</strong></p><p>The fastest way to trust a plan is to see that a teacher like them already uses it. Additionally, teachers value visibility into what others are doing so they can learn new approaches. Therefore, a better design might start from vetted lesson plans that are actively used in schools. When the tool generates a plan, it shows which vetted plan it started from, who uses that plan, and a diff of what it changed for this teacher&#8217;s constraints.</p><p>Next, the teacher can check a small set of changes against a plan they already trust, instead of judging a whole plan from scratch. Here&#8217;s a sketch of what a better interface might look like:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OWdN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 424w, /__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 848w, /__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OWdN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png" width="831" height="691" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:691,&quot;width&quot;:831,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;PE lesson plan after design.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="PE lesson plan after design." title="PE lesson plan after design." srcset="/__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 424w, /__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 848w, /__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OWdN!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380fd77c-f592-451e-9bc5-bd257ed4d265_831x691.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The plan is anchored to a vetted plan another teacher uses. The changes are marked so the teacher can check a diff against a plan they trust. This is a static screenshot for Substack. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to interact with this figure.</figcaption></figure></div><p>In this version, most of the plan is inherited from a vetted plan. The teacher&#8217;s review is scoped to a few edits, each with a reason explaining why the change was made. This is a more efficient way to review a plan because it reduces the cognitive load of judging a whole plan from scratch.</p><p>Designing for this makes the product simpler to build. Instead of stuffing hundreds of examples into a prompt, the tool captures important dimensions, retrieves a close match, and adapts it. Automated evals now become tractable because there is less surface area to test. For example, you can verify that the plan retrieval picked a sensible anchor, and each edit honors the constraints.</p><h2>Example 3: the workers-comp medical report</h2><p>The last example comes from a workers&#8217; compensation tool a founder asked me to help with. It reads a patient&#8217;s chart (intake forms, imaging reports, therapy notes, prior exams) and generates a long expert opinion report, often fifty pages or more. Here is a sketch of the product:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!QD34!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 424w, /__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 848w, /__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!QD34!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png" width="831" height="526" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:526,&quot;width&quot;:831,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Workers-comp medical report before design.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Workers-comp medical report before design." title="Workers-comp medical report before design." srcset="/__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 424w, /__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 848w, /__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QD34!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96552684-9868-4c0c-afbd-c916e4a62b0c_831x526.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The only output is a fifty-page narrative. To trust it, the doctor has to re-read the whole chart and check every claim, which can take as long as writing the report from scratch. This is a static screenshot for Substack. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to interact with this figure.</figcaption></figure></div><p>The problem is the same as the other examples, but the stakes are higher. The only output is the report, and the doctor is the one accountable for it. To trust it they have to go back through the chart and confirm the facts and inferences themselves. That can take as long as writing the report from scratch, which defeats the point of the tool.</p><p>You might object that a fifty-page opinion is hard to verify. That is true, and the product should not pretend otherwise. Helping a doctor understand the evidence is arguably more valuable than the finished document. Therefore, I advised the founder to make the product work like a research assistant instead of a report generator.</p><p>For example, the product could read every record and pull out relevant facts, with a link back to the page so the doctor can check each one. Where two exams disagree, or the chart leaves a question open, the product should surface that. The doctor can then resolve any contradictions and fill in the gaps. Finally, the product can assemble the final report from what they have already checked. Here is a sketch of what this might look like:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ParX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 424w, /__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 848w, /__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ParX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png" width="831" height="621" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d06a23f8-1960-4535-be3e-c78b98965c29_831x621.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:621,&quot;width&quot;:831,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Workers-comp medical report after design.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Workers-comp medical report after design." title="Workers-comp medical report after design." srcset="/__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 424w, /__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 848w, /__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ParX!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06a23f8-1960-4535-be3e-c78b98965c29_831x621.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The product walks the doctor through the evidence first, surfacing each finding with a link back to the chart. Afterwards, the final report is built from the facts they reviewed. This is a static screenshot for Substack. Read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a> to interact with this figure.</figcaption></figure></div><p>The research assistant version of this product allows the doctor to build trust by verifying facts as they go. Similar to the other examples, this design is easier to build and evaluate. Now there are scoped units to grade, such as whether a contradiction is real or whether a citation supports a claim.</p><h2>Generalizing the pattern</h2><p>It is important to understand how users verify your product&#8217;s AI artifacts. Sometimes, this may require assembling supporting evidence. In other cases, it could mean refactoring the entire workflow so that the user is in the loop (like the workers&#8217; comp example).</p><p>Below are questions that can guide your product&#8217;s design for verification:</p><ol><li><p>What does the user actually need to check?</p></li><li><p>What trusted thing can they compare it against?</p></li><li><p>Are there signals or heuristics that experts use to aid in verification?</p></li><li><p>What smaller units can they accept, edit, or reject?</p></li></ol><p>A common thread across these examples is provenance. The fastest way to make an output checkable is to show where each part came from, with links to see more detail. Additionally, you can use progressive disclosure so these sources don&#8217;t overwhelm the user.</p><p>What needs verifying also changes as the user&#8217;s trust grows. Early on, the data agent should make provenance obvious, like where a metric definition came from. Once the user trusts the agent gets it right, that detail can collapse by default in the card. Good design meets users where they are instead of showing everything.</p><p>These principles hold even when a product seems easy to eval. Coding is a good example: it is one of the most verifiable kinds of work, with tests, types, and diffs. Even so, some coding agents go the extra mile to make their work checkable. Cursor and Devin both record a short video of the UI changes they make, so you can confirm the work is right without reproducing it yourself.<a href="#notes"><sup>5</sup></a></p><h2>None of this is new</h2><p>Evals thinking is aligned with good product design. Gathering supporting data and breaking down workflows into smaller units makes automated grading easier. However, I don&#8217;t want to pretend like any wisdom here is new.</p><p>All of these ideas stem from well-established design principles. For example, watching an expert work to learn what they check before you build is called needfinding.<a href="#notes"><sup>6</sup></a> In research-heavy work like the medical case, there is a design goal called sensemaking, which is the work of building a structured understanding of a body of evidence you can reason over.<a href="#notes"><sup>7</sup></a> There are many other concepts, but I think you get the idea.</p><p>Even though these ideas are well established, a reminder is due in the age of AI. Before AI, verification often happened incidentally during the process of creating work product. With AI, verification is the bottleneck. It is time to think about it more explicitly.</p><p><em><strong>&#128073; For the interactive figures and best reading experience, read the <a href="https://hamel.dev/blog/posts/eval-smell/">original post</a>. &#128072;</strong></em></p><div><hr></div><p><em>Thanks to <a href="https://www.sh-reya.com/">Shreya Shankar</a> and <a href="https://isaacflath.com/">Isaac Flath</a> for feedback on this post.</em></p><p>To stay updated on my writing on AI evals, subscribe to my newsletter below.</p><h2>Notes</h2><ol><li><p>More of my writing and teaching on evals: <a href="https://hamel.dev/blog/posts/evals/">Your AI Product Needs Evals</a>, <a href="https://hamel.dev/blog/posts/field-guide/">A Field Guide to Rapidly Improving AI Products</a>, <a href="https://hamel.dev/blog/posts/llm-judge/">Using LLM-as-a-Judge for Evaluation</a>, <a href="https://hamel.dev/blog/posts/evals-faq/">LLM Evals: Everything You Need to Know</a>, <a href="https://hamel.dev/blog/posts/eval-tools/">Selecting the Right AI Evals Tool</a>, <a href="https://hamel.dev/blog/posts/evals-skills/">Evals Skills for Coding Agents</a> and <a href="https://hamel.dev/blog/posts/revenge/">The Revenge of the Data Scientist</a>. I also co-teach the <a href="https://maven.com/parlance-labs/evals?promoCode=evals-info-url">AI Evals for Engineers &amp; PMs</a> course and co-authored the O&#8217;Reilly book <a href="https://www.oreilly.com/library/view/evals-for-ai/9798341660717/">Evals for AI Engineers</a>.</p></li><li><p><a href="https://x.com/lennysan/status/2054631157191598294">Lenny Rachitsky tweeted about this recently</a>: most of one data science team&#8217;s work is now reviewing half-baked AI analysis from PMs and engineers, and half of it is wrong.</p></li><li><p>When I worked at Airbnb we had an internal tool called the <a href="https://github.com/airbnb/knowledge-repo">Knowledge Repo</a>, a place where data scientists published notebooks of their deep dives on analytics, modeling, and so on (<a href="https://medium.com/airbnb-engineering/scaling-knowledge-at-airbnb-875d73eff091">they wrote about it here</a>). It was one of the fastest ways to get context on a new project, since you could read what someone had already worked out. I don&#8217;t know if it is still in use, but the paradigm is a good one.</p></li><li><p><a href="https://www.linkedin.com/in/bryan-bischof/">Bryan Bischof</a> led the creation of the AI at <a href="https://www.hex.tech/">Hex</a>. Bryan is a data scientist himself, the kind of domain expert the product serves, and I think that is part of why it is designed well.</p></li><li><p>See a <a href="https://www.youtube.com/watch?v=XbZvC4KTH68">demo of Cursor</a> recording its UI changes, and <a href="https://docs.devin.ai/work-with-devin/testing-and-recordings">Devin&#8217;s documentation</a> on testing and recordings.</p></li><li><p>Dev Patnaik and Robert Becker, <a href="https://onlinelibrary.wiley.com/doi/10.1111/j.1948-7169.1999.tb00250.x">&#8220;Needfinding: The Why and How of Uncovering People&#8217;s Needs&#8221;</a>, Design Management Journal.</p></li><li><p>Daniel M. Russell, Mark J. Stefik, Peter Pirolli, and Stuart K. Card, <a href="https://www.markstefik.com/wp-content/uploads/2014/04/1993-Cost-Structure-of-Sensemaking1.pdf">&#8220;The Cost Structure of Sensemaking&#8221;</a>.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Evals Skills for Coding Agents]]></title><description><![CDATA[Teach your coding agent evals.]]></description><link>https://hamelhusain.substack.com/p/evals-skills-for-coding-agents</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/evals-skills-for-coding-agents</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Tue, 03 Mar 2026 18:24:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c_TO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today, I&#8217;m publishing <a href="https://github.com/hamelsmu/evals-skills">evals-skills</a>, a set of skills that for AI product evals<sup> [1]</sup>. They distill what I&#8217;ve learned helping 50+ companies and teaching 4,000+ students build evaluation systems.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c_TO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c_TO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8413067-f287-4eb6-9469-6125cad6d148_1280x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c_TO!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8413067-f287-4eb6-9469-6125cad6d148_1280x720.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Why Skills for Evals</strong></h2><p>Coding agents now instrument applications, run experiments, analyze data, and build interfaces. I&#8217;ve been pointing them at evals.</p><p>OpenAI&#8217;s Harness Engineering <a href="https://openai.com/index/harness-engineering/">article</a> makes the case well: they built a product entirely with Codex agents (~1 million lines of code, 1,500 PRs, three engineers, five months) and found that <strong>improving the infrastructure around the agent</strong> yielded better returns than improving the model. Their agents queried distributed traces to verify their own work against runtime evidence. Documentation tells the agent what to do. Telemetry tells it whether it worked. Evals apply the same principle to AI output quality.</p><p>All major eval vendors now ship an MCP server<sup> [2]</sup>. The tedious parts: instrumenting your app, orchestrating experiments, building annotation tools, now belong to coding agents.</p><p>But an agent with an eval platform still needs to know what to do with it. Say a support bot tells a customer &#8220;your plan includes free returns&#8221; when it doesn&#8217;t. Another says &#8220;I&#8217;ve canceled your order&#8221; when nobody asked. Both are hallucinations, but one gets a fact wrong and the other makes up a user action. If you lump them together in a generic &#8220;hallucination score&#8221; real problems will likely go undetected.</p><p>These skills fill in the gaps. They complement the vendor MCP servers: those give your agent access to traces and experiments, these teach it what to do with them.</p><h2><strong>The Skills</strong></h2><p>If you&#8217;re new to evals or inheriting an existing eval pipeline, start with <strong>eval-audit</strong>. It inspects your current setup (or lack of one), runs diagnostic checks across six areas (error analysis, evaluator design, judge validation, human review, labeled data, pipeline hygiene), and produces a prioritized list of problems with next steps. Install the skills and give your agent this prompt:</p><blockquote><p>Install the eval skills plugin from https://github.com/hamelsmu/evals-skills, then run /evals-skills:eval-audit on my eval pipeline. Investigate each diagnostic area using a separate subagent in parallel, then synthesize the findings into a single report. Use other skills in the plugin as recommended by the audit.</p></blockquote><p>If you&#8217;re experienced with evals, you can skip the audit and pick the skill you need:</p><ul><li><p><strong>error-analysis</strong>: Read traces, categorize failures, build a vocabulary of what&#8217;s broken</p></li><li><p><strong>generate-synthetic-data</strong>: Create diverse test inputs when real data is sparse</p></li><li><p><strong>write-judge-prompt</strong>: Design binary Pass/Fail LLM-as-Judge evaluators</p></li><li><p><strong>validate-evaluator</strong>: Calibrate judges against human labels using TPR/TNR and bias correction</p></li><li><p><strong>evaluate-rag</strong>: Evaluate retrieval and generation quality separately</p></li><li><p><strong>build-review-interface</strong>: Generate annotation interfaces for human trace review</p></li></ul><p></p><p>These skills are a starting point. They only cover parts of evals that generalize across projects. Skills grounded in your stack, your domain, and your data will outperform them. Start here, then write your own. Our <a href="https://maven.com/parlance-labs/evals">AI Evals course</a> teaches the end to end workflow.</p><p>&#128073; The repo is here: <a href="https://github.com/hamelsmu/evals-skills">github.com/hamelsmu/evals-skills</a> &#128072;</p><div><hr></div><p><em>P.S. Our next </em><a href="https://maven.com/parlance-labs/evals?promoCode=newsletter-25">&#8203;AI Evals course&#8203;</a><em> cohort starts March 16th. The one after that won&#8217;t run until fall. You&#8217;ll get lifetime access to all materials and learn with students from OpenAI, Meta, Google, Walmart, Airbnb and more. Use code </em><a href="https://maven.com/parlance-labs/evals?promoCode=newsletter-25">&#8203;</a><em><a href="https://maven.com/parlance-labs/evals?promoCode=newsletter-25">newsletter-25</a></em><a href="https://maven.com/parlance-labs/evals?promoCode=newsletter-25">&#8203;</a><em> for a discount.</em></p><h5><strong>Footnotes</strong></h5><ol><li><p>Not foundation model benchmarks like MMLU or HELM that measure general LLM capabilities. Product evals measure whether <em>your</em> pipeline works on <em>your</em> task with <em>your</em> data. If you aren&#8217;t familiar with product-specific AI evals, check out my <a href="https://hamel.dev/blog/posts/evals-faq/">AI Evals FAQ</a>.</p></li><li><p><a href="https://www.braintrust.dev/docs/reference/mcp">Braintrust</a>, <a href="https://github.com/langchain-ai/langsmith-mcp-server">LangSmith</a>, <a href="https://github.com/Arize-ai/phoenix/tree/main/js/packages/phoenix-mcp">Phoenix</a>, <a href="https://truesight.goodeyelabs.com/docs/mcp-integration">Truesight</a>, and others.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[AI Evals For Engineers & Product Managers]]></title><description><![CDATA[Resources to help you learn more about AI Evals]]></description><link>https://hamelhusain.substack.com/p/ai-evals-for-engineers-and-product</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/ai-evals-for-engineers-and-product</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Fri, 23 Jan 2026 03:06:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FQgp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Because you&#8217;re subscribed to this newsletter, chances are you&#8217;re already familiar with the basics of AI evals. However, we have only scratched the surface of this subject. The best way to master evals is through guided practice and feedback.</p><p>That&#8217;s why I&#8217;m writing: our next live cohort of <strong><a href="https://maven.com/parlance-labs/evals?promoCode=substack-25-c4">AI Evals For Engineers &amp; PMs</a></strong> starts this Monday, January 26th. You can either join us for a live, interactive course, or continue learning on your own with our free materials. Here&#8217;s a guide for both paths.</p><div><hr></div><h2><strong>Path 1: Join Our Live Cohort</strong></h2><p>After completing this course, you&#8217;ll walk away with the skills to find, diagnose, and prioritize AI errors, create data flywheels, and automate your evals with approaches you can trust (the syllabus is below). The course runs Jan 26 - Feb 21, 2026 (4 weeks).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://maven.com/parlance-labs/evals?promoCode=substack-25-c4&quot;,&quot;text&quot;:&quot;Learn more about the course&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://maven.com/parlance-labs/evals?promoCode=substack-25-c4"><span>Learn more about the course</span></a></p><p>&#8203;<strong>What our students are saying:</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FQgp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 424w, /__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 848w, /__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FQgp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png" width="1456" height="1069" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1069,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 424w, /__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 848w, /__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FQgp!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c1b7128-5ec3-42fb-9edc-a117dbea8dba_2354x1728.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>See more <a href="https://testimonial.to/ai-evals-course/all">testimonials here</a>.</p><p><strong>What you&#8217;ll learn:</strong></p><ul><li><p><strong>Collect data for evals:</strong> Understand instrumentation and observability for tracking system behavior. Learn approaches for generating synthetic data to maximize error discovery. Choose the right tools and vendors for your use case.</p></li><li><p><strong>Get clarity with error analysis:</strong> Apply data analysis techniques to rapidly find systematic issues in your product. Master the processes and tools to annotate and analyze data quickly. Learn how to analyze agentic systems (tool calls, RAG, etc.) to identify patterns and errors.</p></li><li><p><strong>Implement effective evaluations:</strong> Create evals customized to your product that provide immediate value, NOT generic off-the-shelf evals. Align evals with stakeholders and domain experts so you can scientifically trust the results. Create high-quality LLM-as-a-judge and code-based evals with a systematic, iterative process.</p></li><li><p><strong>Master architecture-specific strategies:</strong> Measure and debug RAG systems for retrieval relevance and factual accuracy. Tame multi-step pipelines to identify error propagation and root causes. Apply techniques to multi-modal settings including text, image, and audio.</p></li><li><p><strong>Run evals in production:</strong> Set up automated evaluation gates in CI/CD pipelines. Understand methods for consistent comparison across experiments. Implement safety and quality control guardrails.</p></li><li><p><strong>Ensure evals lead to high ROI:</strong> Develop strong intuition for when to write an eval and when NOT to. Design interfaces to remove friction from reviewing data. Avoid common pitfalls surrounding team organization, collaboration, tools, and metrics.</p></li></ul><p><strong>What&#8217;s included:</strong></p><ul><li><p>10+ hours of office hours.</p></li><li><p>Lifetime access to a Discord community with 3,000+ students, so you can get unstuck anytime.</p></li><li><p>A 150+ page course reader.</p></li><li><p>6 months of unlimited usage of our AI Eval Assistant that has access to everything we have ever created on the topic of evals.</p></li></ul><p>More importantly, we are offering <strong>unlimited free retakes</strong>. This means you can enroll now to lock in your spot at the current price, and join any future cohort (including the live office hours) if your schedule gets tight. You also get lifetime access to the private Discord community, so the sooner you enroll, the more value you get.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://maven.com/parlance-labs/evals?promoCode=substack-25-c4&quot;,&quot;text&quot;:&quot;Learn more about the course&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://maven.com/parlance-labs/evals?promoCode=substack-25-c4"><span>Learn more about the course</span></a></p><h2>&#8203;<strong>Path 2: Continue Learning with Free Resources</strong></h2><p>If a live course isn&#8217;t the right fit for you right now, no problem. My goal is to help everyone build better AI, so I&#8217;ve compiled some of our best free resources to help you on your journey.</p><p><strong>Essential Reading</strong></p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/index.html">LLM Evals FAQ</a> &#8212; The most common questions on implementing evals (and our answers)</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals/">Your AI Product Needs Evals</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/llm-judge/">LLM-as-a-Judge: A Complete Guide&#8203;</a></p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/field-guide/">A Field Guide To Improving AI Products</a>&#8203;</p></li><li><p>Shreya&#8217;s research: <a href="https://arxiv.org/abs/2404.12272">Who Validates the Validators?</a></p></li><li><p>Building Effective AI-Powered Data Pipelines (Shreya Shankar): <a href="https://hamel.dev/notes/llm/data-processing/shreya-data-processing.html">annotated notes</a>&#8203;</p></li></ul><p><strong>Lightning Lessons: AI Evals</strong></p><p>Free lessons on key topics:&#8203;</p><ul><li><p>Error Analysis: The AI Engineer&#8217;s Best ROI: <a href="https://docs.google.com/presentation/d/1ZNThtixnbZGPpRcjMs8aGeQbzs-Ixe0mbTfvoaPqn3c">slides</a> &#9702; <a href="https://www.youtube.com/watch?v=qH1dZ8JLLdU">video</a>&#8203;</p></li><li><p>The Evals That Made GitHub Copilot Happen: <a href="https://docs.google.com/presentation/d/13lnoKKO647FNcVpG-eKy7bP6pvzzjM_fPJocW8rx9NQ/edit?usp=sharing">slides</a>&#9702; <a href="https://maven.com/p/da8264/how-evals-made-git-hub-copilot-happen">video</a>&#8203;</p></li><li><p>Stop Managing AI Projects Like Traditional Software: <a href="https://www.figma.com/slides/m2QuDvlE8E6VQEcGnotLGH/Burn-your-Roadmap">slides</a> &#9702; <a href="https://www.youtube.com/watch?v=R_HnI9oTv3c">video</a>&#8203;</p></li><li><p>WTH is an AI PM?: <a href="https://maven.com/p/544677/wth-is-an-ai-product-manager">video</a>&#8203;</p></li></ul><p><strong>Lightning Lessons: RAG Series</strong></p><p>We recently ran a 7-part series on modern retrieval with leading researchers as part of our course. I wrote <a href="https://hamel.dev/notes/llm/rag/not_dead.html">annotated notes</a> covering all seven talks:</p><ol><li><p>I Don&#8217;t Use RAG, I Just Retrieve Documents (Ben Clavi&#233;)</p></li><li><p>Modern IR Evaluation In The RAG Era (Nandan Thakur)</p></li><li><p>Reasoning Opens Up New Retrieval Frontiers (Orion Weller)</p></li><li><p>Going Further: Late Interaction Beats Single Vector Limits (Antoine Chaffin)</p></li><li><p>The Map Is Not The Territory &#8211; Highly Multimodal Search</p></li><li><p>Context Rot: When Long Context Fails (Kelly Hong)</p></li><li><p>You Don&#8217;t Need a Graph DB (Jo Kristian Bergum)</p><p>&#8203;</p></li></ol><div><hr></div><p>Whichever path you choose, I hope these resources help you build something great.</p><p>Best,</p><p>Hamel</p>]]></content:encoded></item><item><title><![CDATA[Why AI Evals Are An Increasingly Important Skill]]></title><description><![CDATA[Free resources to help you get started.]]></description><link>https://hamelhusain.substack.com/p/why-ai-evals-are-an-increasingly</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/why-ai-evals-are-an-increasingly</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Mon, 24 Nov 2025 18:33:10 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/179841950/2e576fcab4d3f0a721936a931e233e0f.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>When I started helping companies build AI products several years ago, I noticed a consistent pattern. Everyone was stuck in vibe-check hell.  The same questions kept surfacing:<br><br>- How do I switch to a newer model and ensure everything still works? <br>- What&#8217;s the right granularity and division of labor for my agents? <br>- How do I debug a system with lots of tool calls and handoffs? What do I even focus on first? <br>- When I change my prompt, how do I know I&#8217;m not inadvertently breaking something else?<br>- How do I tune my retrieval or RAG database properly?<br><br>The answer lies in systematically measuring your AI product. Thankfully, there are existing practices from machine learning that we can bring to AI engineering that are more accessible than you think.  Here are some free resources you can use to get started:</p><h2>1. AI Evals FAQ </h2><p>These cover questions from our <a href="https://maven.com/parlance-labs/evals?promoCode=evals-info-url">course</a> where we&#8217;ve taught over 3,000 students. Some of the questions answered:<br><br>Q: What are LLM Evals?<br>Q: What&#8217;s a minimum viable evaluation setup?<br>Q: Why is &#8220;error analysis&#8221; so important in LLM evals, and how is it performed?<br>Q: What is the best approach for generating synthetic data?<br>Q: Are there scenarios where synthetic data may not be reliable?<br>Q: Why do you recommend binary (pass/fail) evaluations?<br>Q: Should I use &#8220;ready-to-use&#8221; evaluation metrics?<br>Q: How many people should annotate my LLM outputs?<br>Q: Should PMs and engineers collaborate on error analysis? How? <br>Q: What parts of evals can be automated with LLMs?<br>Q: Should I stop writing prompts manually in favor of automated tools?<br>Q: What makes a good custom interface for reviewing LLM outputs?<br>Q: What gaps in eval tooling should I be prepared to fill myself?<br>Q: How should I version and manage prompts?<br>Q: How are evaluations used differently in CI/CD vs. monitoring production?<br>Q: What&#8217;s the difference between guardrails &amp; evaluators?<br>Q: Is RAG dead?<br>Q: How should I approach evaluating my RAG system?<br>Q: How do I evaluate sessions with human handoffs? <br>Q: How do I evaluate complex multi-step workflows? <br>Q: How do I evaluate agentic workflows?</p><p><strong>You can <a href="https://maven.com/p/412512/frequently-asked-questions-and-answers-about-ai-evals?utm_medium=lead_magnet_share_link&amp;utm_source=instructor">download the FAQ here.</a></strong></p><h2>2. AI Evals Flashcards</h2><p>These are flashcards we distribute in our <a href="https://maven.com/parlance-labs/evals?promoCode=evals-info-url">course</a> with bite sized guidance on important topics. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fwBc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fwBc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:624014,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hamelhusain.substack.com/i/179841950?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fwBc!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27af03a6-6c10-4975-8b35-5ecdf63a96c8_1280x720.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://maven.com/p/15420a/ai-eval-flashcards?utm_medium=lead_magnet_share_link&amp;utm_source=instructor">Download the flashcards here</a></strong></p><div><hr></div><p><em>P.S. if you want to master AI Evals,  you may be interested in our course: AI Evals for Engineers and PMs, which has taught over 3,000 students from over 500 companies like Amazon, WalMart, Meta, OpenAI, Google and more.  The course is 25% off through tomorrow.  <a href="https://evals.info">You can learn more here</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Final call for our AI Evals course (and a question I'm getting a lot) ]]></title><description><![CDATA[Extended enrollment in our AI Evals course.]]></description><link>https://hamelhusain.substack.com/p/final-call-for-our-last-live-ai-evals</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/final-call-for-our-last-live-ai-evals</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Sun, 05 Oct 2025 18:45:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!J_Ne!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi everyone,</p><p>Over the past few weeks, we've dived into some of the toughest questions in AI Evals. </p><p>We've extended enrollment through this <strong>Tuesday, October 7th</strong>. Every lesson is recorded and you get <strong>lifetime access</strong> to all materials.</p><p>This brings you to a fork in the road. You can either join us for this cohort, or you can continue learning on your own with our free materials. Here&#8217;s a guide for both paths.</p><h2><strong>&#8203;<br>&#8203;Path 1: Join Our Next Cohort</strong></h2><p>If you're ready to build AI applications with a proven, data-driven framework, this is your last chance to do it with us live. You'll walk away with the skills to find, diagnose, and prioritize AI errors, create data flywheels, and automate your evals with approaches you can trust.</p><p><strong>&#10145;&#65039; <a href="https://maven.com/parlance-labs/evals?promoCode=substack-35">Enroll Now &amp; Get 35% Off </a></strong>&#8203;</p><p>Enrollment closes this Tuesday, October 7th.</p><p><strong>What our students are saying:</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!J_Ne!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 424w, /__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 848w, /__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!J_Ne!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png" width="1456" height="1069" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1069,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 424w, /__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 848w, /__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J_Ne!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72b74075-f34b-44d7-aa6e-a76a635b1864_2354x1728.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">See more <a href="https://testimonial.to/ai-evals-course/all">testimonials here</a>.</figcaption></figure></div><p>See more on our <a href="https://testimonial.to/ai-evals-course/all">public testimonial page</a>.</p><p><strong>Course Syllabus (2 lessons per week):</strong></p><ul><li><p><strong>Week 1:</strong> Fundamentals &amp; Lifecycle LLM Application Evaluation, Systematic Error Analysis</p></li><li><p><strong>Week 2:</strong> Implementing Effective Evaluations, Collaborative Evaluation Practices</p></li><li><p><strong>Week 3:</strong> Architecture-Specific Evaluation Strategies (RAG, Agents), Production Monitoring</p></li><li><p><strong>Week 4:</strong> Efficient Human Review Systems, Cost Optimization</p></li></ul><p><strong>Plus:</strong></p><ul><li><p>12 hours of office hours</p></li><li><p>Lifetime access to a discord channel where you can learn and connect with 1,000+ other students working on evals.</p></li><li><p>Over <strong>160 pages of detailed notes</strong> made available exclusively to students to compliment your learning.</p></li><li><p><strong>10 months of unlimited access to our new AI Eval Assistant.</strong> This is an alpha version of a private AI powered by all our course materials (lectures, the course reader, office hours, etc.) that you can have unlimited chats with 24/7.</p></li></ul><p><strong>&#10145;&#65039; <a href="https://maven.com/parlance-labs/evals?promoCode=substack-35">Enroll Now &amp; Get 35% Off </a> (</strong>Enrollment closes this Tuesday, October 7th.)</p><div><hr></div><h2><strong>Path 2: Continue Learning with Free Resources</strong></h2><p>If a live course isn't the right fit for you right now, no problem. My goal is to help everyone build better AI, so I've compiled some of our best free resources to help you on your journey:</p><ul><li><p>&#8203;<strong><a href="https://hamel.dev/blog/posts/evals-faq/">LLM Evals FAQs</a>:</strong> The most common questions on implementing evals (and our answers) from our course.</p></li><li><p><strong>Essential Reading:</strong></p><ul><li><p>Shreya's research: <a href="https://arxiv.org/abs/2404.12272">Who Validates the Validators?</a>&#8203;</p></li><li><p>Hamel's blog posts: "<a href="https://hamel.dev/blog/posts/evals/index.html">Your AI Product Needs Evals</a>", "<a href="https://hamel.dev/blog/posts/llm-judge/">Creating a LLM Judge That Drives Business Results</a>", and "<a href="https://hamel.dev/blog/posts/field-guide/">A Field Guide To Improving AI Products</a>".</p></li></ul></li><li><p>&#8203;<strong><a href="https://hamel.dev/notes/llm/rag/not_dead.html">Free Series on RAG</a>:</strong> Our special series on evaluating and optimizing RAG.</p></li><li><p><strong>Lightning Lessons:</strong> Quick, digestible lessons on key topics related to evals:</p><ul><li><p>Build your own eval tools with notebooks: <a href="https://maven.com/p/466f8a/build-your-own-eval-tools-with-notebooks?utm_medium=ll_share_link&amp;utm_source=instructor">video</a>&#8203;</p></li><li><p>Inspect, an OSS Python library for LLM evals: <a href="https://hamel.dev/notes/llm/evals/inspect.html">slides</a> | <a href="https://www.youtube.com/watch?v=_UY49Q_qFhs">video</a>&#8203;</p></li><li><p>The Evals That Made GitHub Copilot Happen: <a href="https://docs.google.com/presentation/d/13lnoKKO647FNcVpG-eKy7bP6pvzzjM_fPJocW8rx9NQ/edit?usp=sharing">slides</a> | <a href="https://youtu.be/LwLxlEwrtRA?si=TwIZUHTq8gYetFjP">recording</a>&#8203;</p></li><li><p>Stop Managing AI Projects Like Traditional Software: <a href="https://www.figma.com/slides/m2QuDvlE8E6VQEcGnotLGH/Burn-your-Roadmap?node-id=1-394">slides</a> | <a href="https://youtu.be/R_HnI9oTv3c">recording</a></p></li><li><p>How To Setup Evals For Agents: <a href="https://github.com/langchain-ai/openevals">github</a> | <a href="https://maven.com/p/a58f3f">video</a>&#8203;</p></li><li><p>Optimize Your Dev Setup For Evals w/ Cursor Rules &amp; MCP: <a href="https://kentro-tech.github.io/ai-evals-context-talk/presentation.html">slides</a> | <a href="https://maven.com/p/f24547/optimize-your-dev-setup-for-evals-w-cursor-rules-mcp">recording</a>&#8203;</p></li><li><p>Error Analysis: The AI Engineer&#8217;s Best ROI: <a href="https://docs.google.com/presentation/d/1ZNThtixnbZGPpRcjMs8aGeQbzs-Ixe0mbTfvoaPqn3c/edit?slide=id.g338d04086c1_0_0#slide=id.g338d04086c1_0_0">slides</a> | <a href="https://youtu.be/qH1dZ8JLLdU">recording</a>&#8203;</p></li><li><p>Hybrid Search Is Just The Beginning: Optimizing the R in RAG: <a href="https://docs.google.com/presentation/d/1K_2lByZ7K1jfJvm6sfZOQwGLrXK_yv9OidN6MJMNHfE/edit?slide=id.g34a7b37c3ba_0_938#slide=id.g34a7b37c3ba_0_938">slides</a> | <a href="https://youtu.be/QwFHNk6Mvu8">recording</a>&#8203;</p></li><li><p>WTH is an AI PM?: <a href="https://maven.com/p/544677/wth-is-an-ai-product-manager?utm_medium=ll_share_link&amp;utm_source=instructor">recording</a>&#8203;</p></li><li><p>Optimizing Structured Data Retrieval With Evals: <a href="https://docs.google.com/presentation/d/1WjgKBkJgk1y08K1W_xgTXBN8i-eQIu4xcc6kHaJbPUw/edit?usp=sharing">slides</a> | <a href="https://maven.com/p/79dbc3/optimize-structured-data-retrieval-with-evals?utm_medium=ll_share_link&amp;utm_source=instructor">recording</a>&#8203;</p></li><li><p>LLM Evals: Common Mistakes: <a href="https://youtu.be/GL0XhAj5LPE?si=UUO-WURw6vjKBygs">video</a>&#8203;</p></li></ul></li></ul><div><hr></div><p>Whichever path you choose, thank you for being part of this community.</p><p>Hope to see some of you in class,</p><p>Hamel</p>]]></content:encoded></item><item><title><![CDATA[Is It the Right Time to Learn AI Evals? Answering 3 Common Questions]]></title><description><![CDATA[Why I'm excited to teach you about AI Evals]]></description><link>https://hamelhusain.substack.com/p/is-it-the-right-time-to-learn-ai</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/is-it-the-right-time-to-learn-ai</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Fri, 03 Oct 2025 20:13:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/BsWxPI9UM4c" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Many teams building with AI are stuck in a &#8220;whack-a-mole&#8221; cycle. When a product fails, they blindly hack at prompts, hoping for a fix without knowing what&#8217;s truly broken or why. This isn&#8217;t a sustainable way to build reliable AI.</p><p>AI Evaluation is the engineering discipline that solves this. It provides a systematic process to move from guesswork to a structured method for building high-quality AI products. It&#8217;s the skill that separates teams that ship reliable AI from those that don&#8217;t.</p><p>With enrollment for our <a href="https://maven.com/parlance-labs/evals?promoCode=substack-c3">final course</a> of the year closing this <strong>Monday, October 6th</strong>, I know some are still deciding if now is the time to learn this skill. I&#8217;ve been getting a few common questions, and I want to address them head-on.</p><div><hr></div><p>If you&#8217;re hesitating, you might be thinking one of these things:</p><h4><strong>1. &#8220;I don&#8217;t have time in my busy schedule.&#8221;</strong></h4><p>I understand completely. The course is designed for busy professionals, which is why the lessons are professionally recorded and edited for you to watch on your own time.</p><p>More importantly, for this October cohort ONLY, we are offering <strong>unlimited free retakes</strong>. This means you can enroll now to lock in your spot and the current price, and join any future cohort if your schedule gets tight. You also get lifetime access to the private Discord community, so the sooner you enroll, the more value you get.</p><h4><strong>2. &#8220;I&#8217;m worried I don&#8217;t have the AI skills to keep up.&#8221;</strong></h4><p>We&#8217;ve got you covered. As a bonus for all students, <a href="https://elite-ai-assisted-coding.dev/">Isaac Flath</a> is teaching a 5+ hour mini-course on how to tackle the homework assignments using AI assistance.</p><p>Drawn from his popular AI coding <a href="https://maven.com/kentro/context-engineering-for-coding">course</a>, this series is designed for both developers and PMs. You will learn how to use AI agents as a partner to apply the course concepts, explore new topics, and complete the homework effectively. It&#8217;s a course on AI Evals, and we want you to use AI to learn it.</p><h4><strong>3. &#8220;It&#8217;s too expensive.&#8221;</strong></h4><p>You&#8217;re right, it&#8217;s a significant investment. We priced it this way so we can provide an unparalleled level of support. I can promise you that we answer every single question. We have 3+ TAs dedicated to making sure you&#8217;re never stuck.</p><div><hr></div><p>A final heads-up: our seats for this cohort are nearly full.</p><p>If you want to join the final cohort of the year and get all the bonuses (including unlimited retakes), now is the time.</p><p><a href="https://maven.com/parlance-labs/evals?promoCode=substack-c3">&#10145;&#65039; Enroll in the AI Evals Course</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://maven.com/parlance-labs/evals?promoCode=substack-c3" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 424w, /__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 848w, /__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!SKyB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png" width="532" height="299.25" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:532,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://maven.com/parlance-labs/evals?promoCode=substack-c3&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 424w, /__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 848w, /__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SKyB!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7097b9a9-0f28-4557-a060-a1a298672dcf_1536x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hope to see you there.</p><p>Hamel</p><p><em>P.S. If this is the first time you&#8217;ve heard about AI Evals, you may want to check out this <a href="https://youtu.be/BsWxPI9UM4c?si=nboJSY5ZwWofAtPx">end-to-end primer</a> I recently did with </em><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Lenny Rachitsky&quot;,&quot;id&quot;:1849774,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/afba5161-65bb-4d99-8d6b-cce660917fa1_1540x1540.png&quot;,&quot;uuid&quot;:&quot;43b8a3bb-4698-4df7-a096-cdc129e4c93d&quot;}" data-component-name="MentionToDOM"></span> </p><div id="youtube2-BsWxPI9UM4c" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;BsWxPI9UM4c&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/BsWxPI9UM4c?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div>]]></content:encoded></item><item><title><![CDATA[The Ultimate AI Evals FAQ (Now New & Improved)]]></title><description><![CDATA[All your AI Evals questions in one place!]]></description><link>https://hamelhusain.substack.com/p/the-ultimate-ai-evals-faq-now-new</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/the-ultimate-ai-evals-faq-now-new</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Wed, 01 Oct 2025 18:44:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ebda08fc-eca9-412f-83b3-37d74f2ce902_1566x1628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Over the past weeks, I have been publishing frequently asked questions about AI Evals. The response was incredible, and so were the follow-up questions.</p><p>I&#8217;m excited to announce a major update to that post.</p><p>The <a href="https://hamel.dev/blog/posts/evals-faq/">Evals FAQ</a> has been completely reorganized and expanded with tons of new content, making it an even more comprehensive resource for anyone building with LLMs.</p><p>I've grouped the questions into clear sections, so you can quickly find answers whether you're just getting started, designing complex evaluations, or trying to debug a tricky RAG pipeline.</p><p>Below is the full table of contents to give you a sense of the new structure and depth.</p><h2><strong>Table of Contents</strong></h2><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#getting-started-fundamentals">Getting Started &amp; Fundamentals</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-what-are-llm-evals">Q: What are LLM Evals?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-what-is-a-trace">Q: What is a trace?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-whats-a-minimum-viable-evaluation-setup">Q: What&#8217;s a minimum viable evaluation setup?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-much-of-my-development-budget-should-i-allocate-to-evals">Q: How much of my development budget should I allocate to evals?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-will-todays-evaluation-methods-still-be-relevant-in-5-10-years-given-how-fast-ai-is-changing">Q: Will today&#8217;s evaluation methods still be relevant in 5-10 years given how fast AI is changing?</a>&#8203;</p></li></ul></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#error-analysis-data-collection">Error Analysis &amp; Data Collection</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed">Q: Why is "error analysis" so important in LLM evals, and how is it performed?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-surface-problematic-traces-for-review-beyond-user-feedback">Q: How do I surface problematic traces for review beyond user feedback?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-often-should-i-re-run-error-analysis-on-my-production-system">Q: How often should I re-run error analysis on my production system?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-what-is-the-best-approach-for-generating-synthetic-data">Q: What is the best approach for generating synthetic data?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-are-there-scenarios-where-synthetic-data-may-not-be-reliable">Q: Are there scenarios where synthetic data may not be reliable?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-approach-evaluation-when-my-system-handles-diverse-user-queries">Q: How do I approach evaluation when my system handles diverse user queries?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-can-i-efficiently-sample-production-traces-for-review">Q: How can I efficiently sample production traces for review?</a>&#8203;</p></li></ul></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#evaluation-design-methodology">Evaluation Design &amp; Methodology</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-why-do-you-recommend-binary-passfail-evaluations-instead-of-1-5-ratings-likert-scales">Q: Why do you recommend binary (pass/fail) evaluations instead of 1-5 ratings (Likert scales)?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-i-practice-eval-driven-development">Q: Should I practice eval-driven development?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-i-build-automated-evaluators-for-every-failure-mode-i-find">Q: Should I build automated evaluators for every failure mode I find?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-i-use-ready-to-use-evaluation-metrics">Q: Should I use "ready-to-use" evaluation metrics?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-are-similarity-metrics-bertscore-rouge-etc.-useful-for-evaluating-llm-outputs">Q: Are similarity metrics (BERTScore, ROUGE, etc.) useful for evaluating LLM outputs?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-can-i-use-the-same-model-for-both-the-main-task-and-evaluation">Q: Can I use the same model for both the main task and evaluation?</a>&#8203;</p></li></ul></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#human-annotation-process">Human Annotation &amp; Process</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-many-people-should-annotate-my-llm-outputs">Q: How many people should annotate my LLM outputs?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-product-managers-and-engineers-collaborate-on-error-analysis-how">Q: Should product managers and engineers collaborate on error analysis? How?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-i-outsource-annotation-labeling-to-a-third-party">Q: Should I outsource annotation &amp; labeling to a third party?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-what-parts-of-evals-can-be-automated-with-llms">Q: What parts of evals can be automated with LLMs?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-i-stop-writing-prompts-manually-in-favor-of-automated-tools">Q: Should I stop writing prompts manually in favor of automated tools?</a>&#8203;</p></li></ul></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#tools-infrastructure">Tools &amp; Infrastructure</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-should-i-build-a-custom-annotation-tool-or-use-something-off-the-shelf">Q: Should I build a custom annotation tool or use something off-the-shelf?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-what-makes-a-good-custom-interface-for-reviewing-llm-outputs">Q: What makes a good custom interface for reviewing LLM outputs?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-what-gaps-in-eval-tooling-should-i-be-prepared-to-fill-myself">Q: What gaps in eval tooling should I be prepared to fill myself?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-seriously-hamel.-stop-the-bullshit.-whats-your-favorite-eval-vendor">Q: Seriously Hamel. Stop the bullshit. What&#8217;s your favorite eval vendor?</a>&#8203;</p></li></ul></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#production-deployment">Production &amp; Deployment</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-are-evaluations-used-differently-in-cicd-vs.-monitoring-production">Q: How are evaluations used differently in CI/CD vs. monitoring production?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-whats-the-difference-between-guardrails-evaluators">Q: What&#8217;s the difference between guardrails &amp; evaluators?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-can-my-evaluators-also-be-used-to-automatically-fix-or-correct-outputs-in-production">Q: Can my evaluators also be used to automatically </a><em><a href="https://hamel.dev/blog/posts/evals-faq/#q-can-my-evaluators-also-be-used-to-automatically-fix-or-correct-outputs-in-production">fix</a></em><a href="https://hamel.dev/blog/posts/evals-faq/#q-can-my-evaluators-also-be-used-to-automatically-fix-or-correct-outputs-in-production"> or </a><em><a href="https://hamel.dev/blog/posts/evals-faq/#q-can-my-evaluators-also-be-used-to-automatically-fix-or-correct-outputs-in-production">correct</a></em><a href="https://hamel.dev/blog/posts/evals-faq/#q-can-my-evaluators-also-be-used-to-automatically-fix-or-correct-outputs-in-production"> outputs in production?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-much-time-should-i-spend-on-model-selection">Q: How much time should I spend on model selection?</a>&#8203;</p></li></ul></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#domain-specific-applications">Domain-Specific Applications</a>&#8203;</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-is-rag-dead">Q: Is RAG dead?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-should-i-approach-evaluating-my-rag-system">Q: How should I approach evaluating my RAG system?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-choose-the-right-chunk-size-for-my-document-processing-tasks">Q: How do I choose the right chunk size for my document processing tasks?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-debug-multi-turn-conversation-traces">Q: How do I debug multi-turn conversation traces?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-evaluate-sessions-with-human-handoffs">Q: How do I evaluate sessions with human handoffs?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-evaluate-complex-multi-step-workflows">Q: How do I evaluate complex multi-step workflows?</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/#q-how-do-i-evaluate-agentic-workflows">Q: How do I evaluate agentic workflows?</a>&#8203;</p></li></ul></li></ul><h2><strong>Other Formats</strong></h2><p>You can read the full post on the blog, or download it for offline reading.</p><ul><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/">Read the full post online</a>&#8203;</p></li><li><p>&#8203;<a href="https://hamel.dev/blog/posts/evals-faq/evals-faq.pdf">Download as a PDF&#8203;</a></p></li><li><p>Listen to an <a href="https://soundcloud.com/hamel-husain/llm-evals-faq?utm_source=clipboard&amp;utm_campaign=wtshare&amp;utm_medium=widget&amp;utm_content=https%253A%252F%252Fsoundcloud.com%252Fhamel-husain%252Fllm-evals-faq">audio version</a>.</p></li></ul><p>I hope this becomes an invaluable resource for you. If you think someone else could benefit from this information, please send them this email!</p><p>All the best,</p><p>Hamel</p><div><hr></div><p><em>P.S. I want to let you know about this code to get <strong><a href="https://maven.com/parlance-labs/evals?promoCode=substack-35">35% off our next cohort of our evals course</a></strong>. You'll get lifetime access to the material and recordings, and learn with students from OpenAI, Meta, Google, WalMart, Airbnb and more. </em></p>]]></content:encoded></item><item><title><![CDATA[How do I evaluate agentic workflows?]]></title><description><![CDATA[Part 7 of a series of LLM eval FAQs]]></description><link>https://hamelhusain.substack.com/p/how-do-i-evaluate-agentic-workflows</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/how-do-i-evaluate-agentic-workflows</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Tue, 23 Sep 2025 18:07:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ie-8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is part 7 of a series of <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9wcmV2aWV3LmNvbnZlcnRraXQtbWFpbDIuY29tL2NsaWNrL2RwaGVoMGh6aG0vYUhSMGNITTZMeTlvWVcxbGJDNWtaWFl2WW14dlp5OXdiM04wY3k5bGRtRnNjeTFtWVhFdg==">FAQs</a> from my recent <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9iaXQubHkvZXZhbHMtYWk=">course on AI Evals</a> I've been teaching with Shreya Shankar. Let's jump right in:</p><h2><strong>Q: How do I evaluate agentic workflows?</strong></h2><p>We recommend evaluating agentic workflows in two phases:</p><p><strong>1. End-to-end task success.</strong> Treat the agent as a black box and ask &#8220;did we meet the user&#8217;s goal?&#8221;. Define a precise success rule per task (exact answer, correct side-effect, etc.) and measure with human or <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9oYW1lbC5kZXYvYmxvZy9wb3N0cy9sbG0tanVkZ2Uv">aligned LLM judges</a>. Take note of the first upstream failure when conducting <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9oYW1lbC5kZXYvYmxvZy9wb3N0cy9ldmFscy1mYXEvI3Etd2h5LWlzLWVycm9yLWFuYWx5c2lzLXNvLWltcG9ydGFudC1pbi1sbG0tZXZhbHMtYW5kLWhvdy1pcy1pdC1wZXJmb3JtZWQ=">error analysis</a>.</p><p>Once error analysis reveals which workflows fail most often, move to step-level diagnostics to understand why they&#8217;re failing.</p><p><strong>2. Step-level diagnostics.</strong> Assuming that you have sufficiently <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9oYW1lbC5kZXYvYmxvZy9wb3N0cy9ldmFscy8jbG9nZ2luZy10cmFjZXM=">instrumented your system</a> with details of tool calls and responses, you can score individual components such as: - <em>Tool choice</em>: was the selected tool appropriate? - <em>Parameter extraction</em>: were inputs complete and well-formed? - <em>Error handling</em>: did the agent recover from empty results or API failures? - <em>Context retention</em>: did it preserve earlier constraints? - <em>Efficiency</em>: how many steps, seconds, and tokens were spent? - <em>Goal checkpoints</em>: for long workflows verify key milestones.</p><p>Example: &#8220;Find Berkeley homes under $1M and schedule viewings&#8221; breaks into: parameters extracted correctly, relevant listings retrieved, availability checked, and calendar invites sent. Each checkpoint can pass or fail independently, making debugging tractable.</p><p><strong>Use transition failure matrices to understand error patterns.</strong> Create a matrix where rows represent the last successful state and columns represent where the first failure occurred. This is a great way to understand where the most failures occur.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ie-8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ie-8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png" width="1140" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1140,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ie-8!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F786a442e-657f-4dd6-840c-ad0753bc2d7a_1140x742.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Transition failure matrix showing hotspots in text-to-SQL agent workflow</figcaption></figure></div><p>Transition matrices transform overwhelming agent complexity into actionable insights. Instead of drowning in individual trace reviews, you can immediately see that GenSQL &#8594; ExecSQL transitions cause 12 failures while DecideTool &#8594; PlanCal causes only 2. This data-driven approach guides where to invest debugging effort. Here is another <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly93d3cuZmlnbWEuY29tL2RlY2svbndSbGg1cmVudTRzNG9sYUNzZjlsRy9GYWlsdXJlLWlzLWEtRnVubmVsP25vZGUtaWQ9MjAwOS05MjcmdD1HSmxUdHhROGJMSmFROTJBLTE=">example</a> from Bryan Bischof, that is also a text-to-SQL agent:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-iaf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 424w, /__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 848w, /__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-iaf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png" width="1456" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 424w, /__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 848w, /__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-iaf!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08470061-ce53-442a-9d7b-33d459e5d48d_2154x1102.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Bischof, Bryan &#8220;Failure is A Funnel - Data Council, 2025&#8221;</figcaption></figure></div><p>In this example, Bryan shows variation in transition matrices across experiments. How you organize your transition matrix depends on the specifics of your application. For example, Bryan&#8217;s text-to-SQL agent has an inherent sequential workflow which he exploits for further analytical insight. You can watch his <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly95b3V0dS5iZS9SX0huSTlvVHYzYz9zaT1oUlJoRGl5ZEhVNWs2aWtj">full talk</a> for more details.</p><p><strong>Creating Test Cases for Agent Failures</strong></p><p>Creating test cases for agent failures follows the same principles as our previous FAQ on <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9oYW1lbC5kZXYvYmxvZy9wb3N0cy9ldmFscy1mYXEvI3EtaG93LWRvLWktZGVidWctbXVsdGktdHVybi1jb252ZXJzYXRpb24tdHJhY2Vz">debugging multi-turn conversation traces</a> (i.e. try to reproduce the error in the simplest way possible, only use multi-turn tests when the failure actually requires conversation context, etc.).</p><p></p><div><hr></div><p><em>P.S. I want to let you know about this code to get <strong><a href="https://maven.com/parlance-labs/evals?promoCode=substack-35">35% off our next cohort of our evals course</a></strong>. You'll get lifetime access to the material and recordings, and learn with students from OpenAI, Meta, Google, WalMart, Airbnb and more. </em></p>]]></content:encoded></item><item><title><![CDATA[Are similarity metrics (BERTScore, ROUGE, etc.) useful for evaluating LLM outputs? ]]></title><description><![CDATA[Part 8 of a series of LLM Eval FAQs]]></description><link>https://hamelhusain.substack.com/p/are-similarity-metrics-bertscore</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/are-similarity-metrics-bertscore</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Sat, 20 Sep 2025 18:33:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pph_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F75dffc2c-e51f-410c-97ef-d61e5c24af76_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is part 8 of a series of <a href="https://hamel.dev/blog/posts/evals-faq/">FAQs</a> from my recent <a href="https://maven.com/parlance-labs/evals?promoCode=substack-35">course on AI Evals</a> I've been teaching with Shreya Shankar.</p><p>Let's jump in:</p><h2><strong>Q: Are similarity metrics (BERTScore, ROUGE, etc.) useful for evaluating LLM outputs?</strong></h2><p>Generic metrics like BERTScore, ROUGE, cosine similarity, etc. are not useful for evaluating LLM outputs in most AI applications. Instead, we recommend using <a href="https://hamel.dev/blog/posts/evals-faq/#q-why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed">error analysis</a> to identify metrics specific to your application&#8217;s behavior. We recommend designing <a href="https://hamel.dev/blog/posts/evals-faq/#q-why-do-you-recommend-binary-(passfail-evaluations-instead-of-1-5-ratings-(likert-scales)).">binary pass/fail</a> evals (using LLM-as-judge) or code-based assertions.</p><p>As an example, consider a real estate CRM assistant. Suggesting showings that aren&#8217;t available (can be tested with an assertion) or confusing client personas (can be tested with a LLM-as-judge) is problematic . Generic metrics like similarity or verbosity won&#8217;t catch this. A relevant quote from the course:</p><p>&#8220;The abuse of generic metrics is endemic. Many eval vendors promote off the shelf metrics, which ensnare engineers into superfluous tasks.&#8221;</p><p>Similarity metrics aren&#8217;t always useless. They have utility in domains like search and recommendation (and therefore can be useful for <a href="https://hamel.dev/blog/posts/evals-faq/#q-how-should-i-approach-evaluating-my-rag-system">optimizing and debugging retrieval</a> for RAG). For example, cosine similarity between embeddings can measure semantic closeness in retrieval systems, and average pairwise similarity can assess output diversity (where lower similarity indicates higher diversity).</p><p><strong>Have questions about evals that you haven't seen answered yet in <a href="https://hamel.dev/blog/posts/evals-faq/">this FAQ</a>? </strong>Hit reply and ask! I'll include the answer in a future email!</p><p>Thanks,</p><p>Hamel Husain</p>]]></content:encoded></item><item><title><![CDATA[Should I stop writing prompts manually in favor of automated tools?]]></title><description><![CDATA[Part 10 of a series of LLM Eval FAQs]]></description><link>https://hamelhusain.substack.com/p/should-i-stop-writing-prompts-manually</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/should-i-stop-writing-prompts-manually</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Fri, 19 Sep 2025 18:22:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7sqx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Feee58cd7-9a81-4ef6-b0f4-faeed62d5166_400x400.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is part 10 of a series of <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9oYW1lbC5kZXYvYmxvZy9wb3N0cy9ldmFscy1mYXEv">FAQs</a> from my recent <a href="https://preview.convertkit-mail2.com/click/dpheh0hzhm/aHR0cHM6Ly9iaXQubHkvZXZhbHMtYWk=">course on AI Evals</a> I've been teaching with Shreya Shankar.</p><p>Nowadays, it's nearly impossible to discuss LLMs without the topic of prompt optimization/automation coming up. Let's get right into it.</p><h2><strong>Q: Should I stop writing prompts manually in favor of automated tools?</strong></h2><p>Automating prompt engineering can be tempting, but you should be skeptical of tools that promise to optimize prompts for you, especially in early stages of development. When you write a prompt, you are forced to clarify your assumptions and externalize your requirements. Good writing is good thinking [1]. If you delegate this task to an automated tool too early, you risk never fully understanding your own requirements or the model&#8217;s failure modes.</p><p>This is because automated prompt optimization typically hill-climb a predefined evaluation metric. It can refine a prompt to perform better on known failures, but it cannot discover <em>new</em> ones. Discovering new errors requires <a href="https://hamel.dev/blog/posts/evals-faq/#q-why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed">error analysis</a>. Furthermore, research shows that evaluation criteria tends to shift after reviewing a model&#8217;s outputs, a phenomenon known as &#8220;criteria drift&#8221; [2]. This means that evaluation is an iterative, human-driven sensemaking process, not a static target that can be set once and handed off to an optimizer.</p><p>A pragmatic approach is to use LLMs to improve your prompt based on <a href="https://hamel.dev/blog/posts/evals-faq/#q-why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed">open coding</a> (open-ended notes about traces). This way, you maintain a human in the loop who is looking at the data and externalizing their requirements. Once you have a high-quality set of evals, prompt optimization can be effective for that last mile of performance.</p><ol><li><p>Paul Graham, <a href="https://paulgraham.com/writes.html">&#8220;Writes and Write-Nots&#8221;</a>&#8203;</p></li><li><p>Shreya Shankar, et al., <a href="https://arxiv.org/abs/2404.12272">&#8220;Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences&#8221;</a></p><p></p></li></ol><div><hr></div><p><em>P.S. I want to let you know about this code to get <strong><a href="https://maven.com/parlance-labs/evals?promoCode=substack-35">35% off our next cohort of our evals course</a></strong>. You'll get lifetime access to the material and recordings, and learn with students from OpenAI, Meta, Google, WalMart, Airbnb and more. </em></p>]]></content:encoded></item><item><title><![CDATA[Beyond Naive RAG: A new, free mini-book on advanced retrieval]]></title><description><![CDATA[A free mini-book on advanced retrieval.]]></description><link>https://hamelhusain.substack.com/p/beyond-naive-rag-a-new-free-mini</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/beyond-naive-rag-a-new-free-mini</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Mon, 25 Aug 2025 18:51:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/baec17d6-83b7-43e5-a598-5bcb33abf9bd_2150x1360.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi everyone,</p><p>The world of Retrieval-Augmented Generation (RAG) has moved far beyond the simple "embed and search" paradigm of 2023. The pace of innovation is staggering, but it has also created a confusing landscape of new terms and techniques: agentic RAG, late-interaction models, multi-modal retrieval, and more.</p><p>How do you know which of these advanced methods are just hype and which are essential for building production-ready systems?</p><p>To help answer that, we&#8217;ve compiled a new, free resource for the community.</p><p><strong>I am excited to share "</strong><a href="https://www.dropbox.com/scl/fi/t8d59wupeeb3bdldhbq6o/Beyond-Naive-RAG-Practical-Advanced-Methods.pdf?rlkey=y0aawyxocpadmq461h752xrg0&amp;st=79ss1un4&amp;dl=0">&#8203;Beyond Naive RAG: Practical Advanced Methods&#8203;</a><strong>": </strong>an open mini-book that condenses over 5 hours of expert instruction from leading researchers into an easy-to-read, annotated format.</p><p>Each chapter is based on a presentation from our AI Evals <a href="https://maven.com/parlance-labs/evals?promoCode=applied-llms-rag">&#8203;course&#8203;</a>, featuring world-class researchers who are pushing the boundaries of what&#8217;s possible with RAG systems.</p><p>We explore five key areas where the field is advancing:</p><ul><li><p><strong>The "RAG is Dead" Controversy:</strong> Ben Clavi&#233; sets the record straight on why retrieval is more important than ever.</p></li><li><p><strong>Modern IR Evaluation:</strong> Nandan Thakur explains why traditional search metrics are failing and introduces new benchmarks like FreshStack that measure what really matters for RAG: diversity, grounding, and coverage.</p></li><li><p><strong>Reasoning-Enhanced Retrieval:</strong> Orion Weller from Johns Hopkins University shows how to embed instruction-following and reasoning directly into the retrieval process with models like Promptriever and Rank1.</p></li><li><p><strong>Late Interaction Models:</strong> Antoine Chaffin dives into the limitations of single-vector search and explains how models like ColBERT overcome the information loss problem for superior performance.</p></li><li><p><strong>Multiple Representations:</strong> Bryan Bischof and Ayush Chaurasia provide a powerful framework for building flexible systems that use intelligent routing across multiple, diverse "maps" of your data to better serve user intent.</p></li></ul><p>This mini-book is designed for practitioners. It combines theory with practical recipes and timestamped references to the original videos, allowing you to explore the topics that matter most to you.</p><p>Below is the table of contents of the book.</p><h2><strong>Table of Contents</strong></h2><ul><li><p><strong>1: I Don&#8217;t Use RAG, I Just Retrieve Documents</strong></p><ul><li><p>1.1 The Title</p></li><li><p>1.2 The &#8220;RAG is Dead&#8221; Controversy</p></li><li><p>1.3 What is RAG Really?</p></li><li><p>1.4 The Standard RAG Flow</p></li><li><p>1.5 The Problem with &#8220;2023 RAG&#8221;</p></li><li><p>1.6 Single-Vector Search Limitations</p></li><li><p>1.7 Why Long Context Doesn&#8217;t Replace RAG</p></li><li><p>1.8 Retrieval is Essential</p></li><li><p>1.9 Key Takeaways</p></li><li><p>1.10 Better RAG is the Solution</p></li><li><p>1.11 The Retrieval Landscape</p></li></ul></li><li><p><strong>2: Modern IR Evaluation for RAG</strong></p><ul><li><p>2.1 Introduction and Speaker Background</p></li><li><p>2.2 The History of Information Retrieval</p></li><li><p>2.3 The Cranfield Paradigm</p></li><li><p>2.4 The BEIR Benchmark</p></li><li><p>2.5 Problems with Current Benchmarks</p></li><li><p>2.6 The RAG Era Changes Everything</p></li><li><p>2.7 Different Users, Different Goals</p></li><li><p>2.8 The Evaluation Mismatch</p></li><li><p>2.9 Introducing FreshStack</p></li><li><p>2.10 FreshStack Data Sources</p></li><li><p>2.11 The FreshStack Pipeline</p></li><li><p>2.12 FreshStack Evaluation Metrics</p></li></ul></li><li><p><strong>3: Optimizing Retrieval with Reasoning Models</strong></p><ul><li><p>3.1 LLM Capabilities: Instruction Following and Reasoning</p></li><li><p>3.2 The Search Paradigm Hasn&#8217;t Changed</p></li><li><p>3.3 Evolution of Search Paradigms</p></li><li><p>3.4 Understanding Instructions in IR</p></li><li><p>3.5 Introducing Promptriever and Rank1</p></li><li><p>3.6 Promptriever: Instruction-Trained Retrieval</p></li><li><p>3.7 Promptriever Evaluation Results</p></li><li><p>3.8 Rank1: Reasoning-Based Reranking</p></li><li><p>3.9 Rank1 Performance Results</p></li><li><p>3.10 Finding Novel Relevant Documents</p></li></ul></li><li><p><strong>4: Late Interaction Models For RAG</strong></p><ul><li><p>4.1 Dense Vector Search Architecture</p></li><li><p>4.2 Why Dense Models Became Popular</p></li><li><p>4.3 The Benchmark Problem</p></li><li><p>4.4 Hidden Limitations</p></li><li><p>4.5 Long Context Performance</p></li><li><p>4.6 Complex Retrieval Tasks</p></li><li><p>4.7 BM25&#8217;s Surprising Strength</p></li><li><p>4.8 The Pooling Problem</p></li><li><p>4.9 Late Interaction Solution</p></li><li><p>4.10 Performance Advantages</p></li><li><p>4.11 Interpretability Benefits</p></li><li><p>4.12 Barriers to Adoption</p></li><li><p>4.13 PyLate: Making Late Interaction Accessible</p></li><li><p>4.14 Future Research Directions</p></li></ul></li><li><p><strong>5: RAG with Multiple Representations</strong></p><ul><li><p>5.1 The Map is Not the Territory</p></li><li><p>5.2 Deconstructing RAG Buzzwords</p></li><li><p>5.3 A First-Principles View of RAG</p></li><li><p>5.4 The Three Responsibilities of an IR Engineer</p></li><li><p>5.5 Practical Application: Curving Space</p></li><li><p>5.6 Agents as Routers</p></li><li><p>5.7 Dynamic Representations</p></li><li><p>5.8 Demo: Semantic Dot Art</p></li><li><p>5.9 System Architecture</p></li><li><p>5.10 Integration with Other Techniques</p></li></ul></li><li><p><strong>Conclusion</strong></p><ul><li><p>Key Takeaways</p></li><li><p>Looking Forward</p></li><li><p>Resources</p></li></ul></li></ul><p>You can get your free copy here:</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.dropbox.com/scl/fi/t8d59wupeeb3bdldhbq6o/Beyond-Naive-RAG-Practical-Advanced-Methods.pdf?rlkey=y0aawyxocpadmq461h752xrg0&amp;st=79ss1un4&amp;dl=0&quot;,&quot;text&quot;:&quot;Download: Beyond Naive RAG (PDF)&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.dropbox.com/scl/fi/t8d59wupeeb3bdldhbq6o/Beyond-Naive-RAG-Practical-Advanced-Methods.pdf?rlkey=y0aawyxocpadmq461h752xrg0&amp;st=79ss1un4&amp;dl=0"><span>Download: Beyond Naive RAG (PDF)</span></a></p><p>We hope this helps you navigate the evolving landscape of RAG and build better AI systems. Know someone who might benefit from learning about advanced RAG methods? Send them this email!</p><p>Thanks,</p><p>Hamel</p><p><em>P.S. I want to let you know about this code to get </em><a href="https://maven.com/parlance-labs/evals?promoCode=evals-faq-new">&#8203;</a><em><strong><a href="https://maven.com/parlance-labs/evals?promoCode=evals-faq-new">35% off our next cohort of our evals course</a></strong></em><a href="https://maven.com/parlance-labs/evals?promoCode=evals-faq-new">&#8203;</a><em>. You'll get lifetime access to the material and recordings, and learn with students from OpenAI, Meta, Google, WalMart, Airbnb and more.</em></p>]]></content:encoded></item><item><title><![CDATA[The best public example of AI Evals I've ever seen]]></title><description><![CDATA[Think evals are just for engineers? See how a product manager transformed her prototype to a production-ready app with evals.]]></description><link>https://hamelhusain.substack.com/p/the-best-public-example-of-ai-evals</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/the-best-public-example-of-ai-evals</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Sun, 17 Aug 2025 14:58:19 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/62c0dd79-1850-4a3f-a832-76aaea3bcd22_1280x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi everyone,</p><p>I'm often asked for the best public example of an AI evals workflow for a real, production application. I finally have an answer.</p><p>If you've ever thought that evals are too "fancy," too complex, or something only deeply technical engineers can do, I want you to watch this talk.</p><p>Teresa Torres, a legendary product discovery coach and author, was a student in our first AI Evals cohort. She's not a daily coder, but she took the core principles from the course and applied them to build an AI interview coach from scratch.</p><p>What she created is, in my opinion, the single best public demonstration of how to do evals right.</p><p><a href="https://youtu.be/N-qAOv_PNPc">&#8203;Watch the full presentation: From Noob to 5 Automated Evals in 4 Weeks (as a PM)&#8203;</a></p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;39289312-cbe1-42bf-8797-12063ecd1810&quot;,&quot;duration&quot;:null}"></div><p>In this incredibly hands-on lesson, Teresa shows exactly how she:</p><ul><li><p><strong>Started with error analysis FIRST</strong> to find real, user-impacting issues instead of getting lost in generic metrics.</p></li><li><p><strong>Used Jupyter notebooks</strong> to systematically investigate failures and analyze results (it's a fantastic commercial for notebooks).</p></li><li><p><strong>Built her own custom annotation tools</strong> and widgets directly inside her notebooks to speed up her workflow.</p></li><li><p><strong>Wrote both LLM-as-a-judge and simple code-based assertions</strong> to test for the specific errors she found.</p></li><li><p><strong>Iterated relentlessly</strong> through this feedback loop, measurably squashing bugs and improving her product with each cycle.</p></li><li><p><strong>Kept things simple the whole time</strong>, proving you don't need a massive, over-engineered stack to get started.</p></li></ul><p>This isn't a theoretical talk. It's a real-world case study of someone building a production app, learning Python and new tools along the way, and using a practical evals process to drive massive improvements. It&#8217;s an empowering story that shows this is within reach for anyone willing to dive in.</p><p>Hope you enjoy it,</p><p>Hamel</p><p><strong>P.S.</strong> Teresa's journey started in our <a href="https://bit.ly/evals-ai">&#8203;</a><strong><a href="https://bit.ly/evals-ai">AI Evals course</a></strong><a href="https://bit.ly/evals-ai">&#8203;</a>. If her story inspires you to build your own systematic feedback loops, I wanted to let you know about our upcoming cohort.</p><p>You'll get lifetime access to all materials and learn the same frameworks that helped Teresa build her application. the <strong>early bird discount for our next cohort ends this Friday</strong>. If you're interested, you can learn more and register here:</p><p><strong>&#10145;&#65039; </strong><a href="https://bit.ly/evals-ai">&#8203;</a><strong><a href="https://bit.ly/evals-ai">AI Evals Course (&gt; $1,000 Discount)</a></strong><a href="https://bit.ly/evals-ai">&#8203;</a></p><p>Hope to see some of you there!</p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://hamelhusain.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Hamel&#8217;s Substack! Subscribe for free to receive new posts</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Frequently Asked Questions (And Answers) About AI Evals]]></title><description><![CDATA[FAQ from our course on AI Evals.]]></description><link>https://hamelhusain.substack.com/p/evals-faq</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/evals-faq</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Thu, 31 Jul 2025 14:38:03 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ab7a180a-eb0b-4195-816b-26d8c461d561_2081x1109.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This document curates the most common questions Shreya and I received while <a href="https://bit.ly/evals-ai">teaching</a> 2,000+ engineers &amp; PMs AI Evals. <em>Warning: These are sharp opinions about what works in most cases. They are not universal truths. Use your judgment.</em></p><div><hr></div><p><strong>&#128073; </strong><em><strong>If you want to learn more about AI Evals, check out our <a href="https://bit.ly/evals-ai">AI Evals course</a></strong></em>. Here is a <a href="https://bit.ly/evals-ai">35% discount code</a> for readers. &#128072;</p><div><hr></div><h1>Listen to the audio version of this FAQ</h1><p>If you prefer to listen to the audio version (narrated by AI), you can play it <a href="https://soundcloud.com/hamel-husain/llm-evals-faq">here</a>.</p><div class="soundcloud-wrap" data-attrs="{&quot;url&quot;:&quot;https://api.soundcloud.com/tracks/2138083206&quot;,&quot;title&quot;:&quot;&quot;,&quot;description&quot;:&quot;&quot;,&quot;thumbnail_url&quot;:&quot;&quot;,&quot;author_name&quot;:&quot;&quot;,&quot;author_url&quot;:&quot;&quot;,&quot;targetUrl&quot;:&quot;&quot;}" data-component-name="SoundcloudToDOM"><iframe src="https://w.soundcloud.com/player/?auto_play=false&amp;buying=false&amp;liking=false&amp;download=false&amp;sharing=false&amp;show_artwork=true&amp;show_comments=false&amp;show_playcount=false&amp;show_user=true&amp;hide_related=true&amp;visual=false&amp;start_track=0&amp;url=https%3A%2F%2Fapi.soundcloud.com%2Ftracks%2F2138083206" frameborder="0" gesture="media" scrolling="no" allowfullscreen="true"></iframe></div><p><a href="https://soundcloud.com/hamel-husain">Hamel Husain</a> &#183; <a href="https://soundcloud.com/hamel-husain/llm-evals-faq">LLM Evals FAQ</a></p><h1>Getting Started &amp; Fundamentals</h1><h2>Q: What are LLM Evals?</h2><p>If you are completely new to product-specific LLM evals (not foundation model benchmarks), see these posts: <a href="/__u/hamelhusain.substack.com/blog/posts/evals/index.html">part 1</a>, <a href="/__u/hamelhusain.substack.com/blog/posts/llm-judge/index.html">part 2</a> and <a href="/__u/hamelhusain.substack.com/blog/posts/field-guide/index.html">part 3</a>. Otherwise, keep reading.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://hamel.dev/evals" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Mmea!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://hamel.dev/evals&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mmea!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb595a825-47de-4df7-a56a-db83d868d34f_2081x1109.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://hamel.dev/evals">Your AI Product Needs Eval (Evaluation Systems)</a></strong></p><p><strong>Contents:</strong></p><ol><li><p>Motivation</p></li><li><p>Iterating Quickly == Success<br></p></li><li><p>Case Study: Lucy, A Real Estate AI Assistant</p></li><li><p>The Types Of Evaluation</p><ol><li><p>Level 1: Unit Tests</p></li><li><p>Level 2: Human &amp; Model Eval</p></li><li><p>Level 3: A/B Testing</p></li><li><p>Evaluating RAG</p></li></ol></li><li><p>Eval Systems Unlock Superpowers For Free</p><ol><li><p>Fine-Tuning</p></li><li><p>Data Synthesis &amp; Curation</p></li><li><p>Debugging</p></li></ol></li></ol><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://hamel.dev/llm-judge/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!QD0E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://hamel.dev/llm-judge/&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QD0E!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F556c070d-6594-4d28-8a7e-c67329ea37d5_1600x900.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://hamel.dev/llm-judge/">Creating a LLM-as-a-Judge That Drives Business Results</a></strong></p><p><strong>Contents:</strong></p><ol><li><p>The Problem: AI Teams Are Drowning in Data</p></li><li><p>Step 1: Find The Principal Domain Expert</p></li><li><p>Step 2: Create a Dataset</p></li><li><p>Step 3: Direct The Domain Expert to Make Pass/Fail Judgments with Critiques</p></li><li><p>Step 4: Fix Errors</p></li><li><p>Step 5: Build Your LLM as A Judge, Iteratively</p></li><li><p>Step 6: Perform Error Analysis</p></li><li><p>Step 7: Create More Specialized LLM Judges (if needed)</p></li><li><p>Recap of Critique Shadowing</p></li><li><p>Resources</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://hamel.dev/field-guide" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!l9UJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://hamel.dev/field-guide&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!l9UJ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f30a878-5090-4369-81bb-54b6cdf09cdf_1600x900.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://hamel.dev/field-guide">A Field Guide to Rapidly Improving AI Products</a></strong></p><p><strong>Contents:</strong></p><ol><li><p>How error analysis consistently reveals the highest-ROI improvements</p></li><li><p>Why a simple data viewer is your most important AI investment</p></li><li><p>How to empower domain experts (not just engineers) to improve your AI</p></li><li><p>Why synthetic data is more effective than you think</p></li><li><p>How to maintain trust in your evaluation system</p></li><li><p>Why your AI roadmap should count experiments, not features</p></li></ol><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/what-are-llm-evals.html">&#8599; Focus view</a></p><h2>Q: What is a trace?</h2><p>A trace is the complete record of all actions, messages, tool calls, and data retrievals from a single initial user query through to the final response. It includes every step across all agents, tools, and system components in a session: multiple user messages, assistant responses, retrieved documents, and intermediate tool interactions.</p><p><strong>Note on terminology:</strong> Different observability vendors use varying definitions of traces and spans. <a href="https://mlops.systems/posts/2025-06-04-instrumenting-an-agentic-app-with-arize-phoenix-and-litellm.html#llm-tracing-tools-naming-conventions-june-2025">Alex Strick van Linschoten&#8217;s analysis</a> highlights these differences (screenshot below):</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!3fRF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!3fRF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!3fRF!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58852dc0-0367-4e4b-b5c8-9bf81b0527c9_900x586.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Vendor differences in trace definitions as of 2025-07-02</figcaption></figure></div><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/what-is-a-trace.html">&#8599; Focus view</a></p><h2>Q: What&#8217;s a minimum viable evaluation setup?</h2><p>Start with error analysis, not infrastructure. Spend 30 minutes manually reviewing 20-50 LLM outputs whenever you make significant changes. Use one domain expert who understands your users as your quality decision maker (a &#8220;benevolent dictator&#8221;).</p><p>If possible, <strong>use notebooks</strong> to help you review traces and analyze data. In our opinion, this is the single most effective tool for evals because you can write arbitrary code, visualize data, and iterate quickly. You can even build your own custom annotation interface right inside notebooks, as shown in this <a href="https://youtu.be/aqKUwPKBkB0?si=5KDmMQnRzO_Ce9xH">video</a>.</p><div id="youtube2-aqKUwPKBkB0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;aqKUwPKBkB0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/aqKUwPKBkB0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/whats-a-minimum-viable-evaluation-setup.html">&#8599; Focus view</a></p><h2>Q: How much of my development budget should I allocate to evals?</h2><p>It&#8217;s important to recognize that evaluation is part of the development process rather than a distinct line item, similar to how debugging is part of software development.</p><p>You should always be doing <a href="https://www.youtube.com/watch?v=qH1dZ8JLLdU">error analysis</a>. When you discover issues through error analysis, many will be straightforward bugs you&#8217;ll fix immediately. These fixes don&#8217;t require separate evaluation infrastructure as they&#8217;re just part of development.</p><p>The decision to build automated evaluators comes down to cost-benefit analysis. If you can catch an error with a simple assertion or regex check, the cost is minimal and probably worth it. But if you need to align an LLM-as-judge evaluator, consider whether the failure mode warrants that investment.</p><p>In the projects we&#8217;ve worked on, <strong>we&#8217;ve spent 60-80% of our development time on error analysis and evaluation</strong>. Expect most of your effort to go toward understanding failures (i.e.&nbsp;looking at data) rather than building automated checks.</p><p>Be <a href="https://ai-execs.com/2_intro.html#a-case-study-in-misleading-ai-advice">wary of optimizing for high eval pass rates</a>. If you&#8217;re passing 100% of your evals, you&#8217;re likely not challenging your system enough. A 70% pass rate might indicate a more meaningful evaluation that&#8217;s actually stress-testing your application. Focus on evals that help you catch real issues, not ones that make your metrics look good.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-much-of-my-development-budget-should-i-allocate-to-evals.html">&#8599; Focus view</a></p><h2>Q: Will today&#8217;s evaluation methods still be relevant in 5-10 years given how fast AI is changing?</h2><p>Yes. Even with perfect models, you still need to verify they&#8217;re solving the right problem. The need for systematic error analysis, domain-specific testing, and monitoring will still be important.</p><p>Today&#8217;s prompt engineering tricks might become obsolete, but you&#8217;ll still need to understand failure modes. Additionally, a LLM cannot read your mind, and <a href="https://arxiv.org/abs/2404.12272">research shows</a> that people need to observe the LLM&#8217;s behavior in order to properly externalize their requirements.</p><p>For deeper perspective on this debate, see these two viewpoints: <a href="https://m.youtube.com/watch?si=qknrtQeITqJ7VsJH&amp;v=4dUFIRj-BWo&amp;feature=youtu.be">&#8220;The model is the product&#8221;</a> versus <a href="https://www.youtube.com/watch?v=EEw2PpL-_NM">&#8220;The model is NOT the product&#8221;</a>.</p><p><strong>&#8220;The model is the product&#8221;:</strong></p><div id="youtube2-4dUFIRj-BWo" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;4dUFIRj-BWo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/4dUFIRj-BWo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong>&#8220;The model is NOT the product&#8221;:</strong></p><div id="youtube2-EEw2PpL-_NM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;EEw2PpL-_NM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/EEw2PpL-_NM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/will-these-evaluation-methods-still-be-relevant-in-5-10-years-given-how-fast-ai-is-changing.html">&#8599; Focus view</a></p><h1>Error Analysis &amp; Data Collection</h1><h2>Q: Why is "error analysis" so important in LLM evals, and how is it performed?</h2><p>Error analysis is <strong>the most important activity in evals</strong>. Error analysis helps you decide what evals to write in the first place. It allows you to identify failure modes unique to your application and data. The process involves:</p><h3>1. Creating a Dataset</h3><p>Gathering representative traces of user interactions with the LLM. If you do not have any data, you can generate synthetic data to get started.</p><h3>2. Open Coding</h3><p>Human annotator(s) (ideally a benevolent dictator) review and write open-ended notes about traces, noting any issues. This process is akin to &#8220;journaling&#8221; and is adapted from qualitative research methodologies. When beginning, it is recommended to focus on noting the first failure observed in a trace, as upstream errors can cause downstream issues, though you can also tag all independent failures if feasible. A <a href="https://hamel.dev/blog/posts/llm-judge/#step-1-find-the-principal-domain-expert">domain expert</a> should be performing this step.</p><h3>3. Axial Coding</h3><p>Categorize the open-ended notes into a &#8220;failure taxonomy.&#8221;. In other words, group similar failures into distinct categories. This is the most important step. At the end, count the number of failures in each category. You can use a LLM to help with this step.</p><h3>4. Iterative Refinement</h3><p>Keep iterating on more traces until you reach <a href="https://delvetool.com/blog/theoreticalsaturation">theoretical saturation</a>, meaning new traces do not seem to reveal new failure modes or information to you. As a rule of thumb, you should aim to review at least 100 traces.</p><p>You should frequently revisit this process. There are advanced ways to <a href="/__u/hamelhusain.substack.com/how-can-i-efficiently-sample-production-traces-for-review.html">sample data more efficiently</a>, like clustering, sorting by user feedback, and sorting by high probability failure patterns. Over time, you&#8217;ll develop a &#8220;nose&#8221; for where to look for failures in your data.</p><p>Do not skip error analysis. It ensures that the evaluation metrics you develop are supported by real application behaviors instead of counter-productive generic metrics (which most platforms nudge you to use). For examples of how error analysis can be helpful, see <a href="https://www.youtube.com/watch?v=e2i6JbU2R-s">this video</a>, or this <a href="https://hamel.dev/blog/posts/field-guide/">blog post</a>.</p><p>Here is a visualization of the error analysis process by one of our students, <a href="https://www.linkedin.com/in/pawel-huryn/">Pawel Huryn</a> - including how it fits into the overall evaluation process:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!eF3v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 424w, /__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 848w, /__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 1272w, /__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!eF3v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 424w, /__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 848w, /__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 1272w, /__u/substackcdn.com/image/fetch/$s_!eF3v!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eb5da16-2286-44f8-b164-b26084bbdf74_1200x1500.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed.html">&#8599; Focus view</a></p><h2>Q: How do I surface problematic traces for review beyond user feedback?</h2><p>While user feedback is a good way to narrow in on problematic traces, other methods are also useful. Here are three complementary approaches:</p><h3>Start with random sampling</h3><p>The simplest approach is reviewing a random sample of traces. If you find few issues, escalate to stress testing: create queries that deliberately test your prompt constraints to see if the AI follows your rules.</p><h3>Use evals for initial screening</h3><p>Use existing evals to find problematic traces and potential issues. Once you&#8217;ve identified these, you can proceed with the typical evaluation process starting with error analysis.</p><h3>Leverage efficient sampling strategies</h3><p>For more sophisticated trace discovery, use outlier detection, metric-based sorting, and stratified sampling to find interesting traces. Generic metrics can serve as exploration signals to identify traces worth reviewing, even if they don&#8217;t directly measure quality.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-surface-problematic-traces-for-review-beyond-user-feedback.html">&#8599; Focus view</a></p><h2>Q: How often should I re-run error analysis on my production system?</h2><p>Re-run error analysis when making significant changes: new features, prompt updates, model switches, or major bug fixes. A useful heuristic is to set a goal for reviewing <em>at least</em> 100+ fresh traces each review cycle. Typical review cycles we&#8217;ve seen range from 2-4 weeks. See this FAQ on how to sample traces effectively.</p><p>Between major analyses, review 10-20 traces weekly, focusing on outliers: unusually long conversations, sessions with multiple retries, or traces flagged by automated monitoring. Adjust frequency based on system stability and usage growth. New systems need weekly analysis until failure patterns stabilize. Mature systems might need only monthly analysis unless usage patterns change. Always analyze after incidents, user complaint spikes, or metric drift. Scaling usage introduces new edge cases.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-often-should-i-re-run-error-analysis-on-my-production-system.html">&#8599; Focus view</a></p><h2>Q: What is the best approach for generating synthetic data?</h2><p>A common mistake is prompting an LLM to <code>"give me test queries"</code> without structure, resulting in generic, repetitive outputs. A structured approach using dimensions produces far better synthetic data for testing LLM applications.</p><p><strong>Start by defining dimensions</strong>: categories that describe different aspects of user queries. Each dimension captures one type of variation in user behavior. For example:</p><ul><li><p>For a recipe app, dimensions might include Dietary Restriction (<em>vegan</em>, <em>gluten-free</em>, <em>none</em>), Cuisine Type (<em>Italian</em>, <em>Asian</em>, <em>comfort food</em>), and Query Complexity (<em>simple request</em>, <em>multi-step</em>, <em>edge case</em>).</p></li><li><p>For a customer support bot, dimensions could be Issue Type (<em>billing</em>, <em>technical</em>, <em>general</em>), Customer Mood (<em>frustrated</em>, <em>neutral</em>, <em>happy</em>), and Prior Context (<em>new issue</em>, <em>follow-up</em>, <em>resolved</em>).</p></li></ul><p><strong>Start with failure hypotheses</strong>. If you lack intuition about failure modes, use your application extensively or recruit friends to use it. Then choose dimensions targeting those likely failures.</p><p><strong>Create tuples manually first</strong>: Write 20 tuples by hand&#8212;specific combinations selecting one value from each dimension. Example: (<em>Vegan</em>, <em>Italian</em>, <em>Multi-step</em>). This manual work helps you understand your problem space.</p><p><strong>Scale with two-step generation</strong>:</p><ol><li><p><strong>Generate structured tuples</strong>: Have the LLM create more combinations like (<em>Gluten-free</em>, <em>Asian</em>, <em>Simple</em>)</p></li><li><p><strong>Convert tuples to queries</strong>: In a separate prompt, transform each tuple into natural language</p></li></ol><p>This separation avoids repetitive phrasing. The (<em>Vegan</em>, <em>Italian</em>, <em>Multi-step</em>) tuple becomes: <code>"I need a dairy-free lasagna recipe that I can prep the day before."</code></p><h3>Generation approaches</h3><p>You can generate tuples two ways:</p><p><strong>Cross product then filter</strong>: Generate all dimension combinations, then filter with an LLM. Guarantees coverage including edge cases. Use when most combinations are valid.</p><p><strong>Direct LLM generation</strong>: Ask the LLM to generate tuples directly. More realistic but tends toward generic outputs and misses rare scenarios. Use when many dimension combinations are invalid.</p><p><strong>Fix obvious problems first</strong>: Don&#8217;t generate synthetic data for issues you can fix immediately. If your prompt doesn&#8217;t mention dietary restrictions, fix the prompt rather than generating specialized test queries.</p><p>After iterating on your tuples and prompts, <strong>run these synthetic queries through your actual system to capture full traces</strong>. Sample 100 traces for error analysis. This number provides enough traces to manually review and identify failure patterns without being overwhelming.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/what-is-the-best-approach-for-generating-synthetic-data.html">&#8599; Focus view</a></p><h2>Q: Are there scenarios where synthetic data may not be reliable?</h2><p>Yes: synthetic data can mislead or mask issues. For guidance on generating synthetic data when appropriate, see What is the best approach for generating synthetic data?</p><p>Common scenarios where synthetic data fails:</p><ol><li><p><strong>Complex domain-specific content</strong>: LLMs often miss the structure, nuance, or quirks of specialized documents (e.g., legal filings, medical records, technical forms). Without real examples, critical edge cases are missed.</p></li><li><p><strong>Low-resource languages or dialects</strong>: For low-resource languages or dialects, LLM-generated samples are often unrealistic. Evaluations based on them won&#8217;t reflect actual performance.</p></li><li><p><strong>When validation is impossible</strong>: If you can&#8217;t verify synthetic sample realism (due to domain complexity or lack of ground truth), real data is important for accurate evaluation.</p></li><li><p><strong>High-stakes domains</strong>: In high-stakes domains (medicine, law, emergency response), synthetic data often lacks subtlety and edge cases. Errors here have serious consequences, and manual validation is difficult.</p></li><li><p><strong>Underrepresented user groups</strong>: For underrepresented user groups, LLMs may misrepresent context, values, or challenges. Synthetic data can reinforce biases in the training data of the LLM.</p></li></ol><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/are-there-scenarios-where-synthetic-data-may-not-be-reliable.html">&#8599; Focus view</a></p><h2>Q: How do I approach evaluation when my system handles diverse user queries?</h2><blockquote><p>Complex applications often support vastly different query patterns&#8212;from &#8220;What&#8217;s the return policy?&#8221; to &#8220;Compare pricing trends across regions for products matching these criteria.&#8221; Each query type exercises different system capabilities, leading to confusion on how to design eval criteria.</p></blockquote><p><em><strong><a href="https://youtu.be/e2i6JbU2R-s?si=8p5XVxbBiioz69Xc">Error Analysis</a> is all you need.</strong></em> Your evaluation strategy should emerge from observed failure patterns (e.g.&nbsp;error analysis), not predetermined query classifications. Rather than creating a massive evaluation matrix covering every query type you can imagine, let your system&#8217;s actual behavior guide where you invest evaluation effort.</p><p>During error analysis, you&#8217;ll likely discover that certain query categories share failure patterns. For instance, all queries requiring temporal reasoning might struggle regardless of whether they&#8217;re simple lookups or complex aggregations. Similarly, queries that need to combine information from multiple sources might fail in consistent ways. These patterns discovered through error analysis should drive your evaluation priorities. It could be that query category is a fine way to group failures, but you don&#8217;t know that until you&#8217;ve analyzed your data.</p><p>To see an example of basic error analysis in action, <a href="https://youtu.be/e2i6JbU2R-s?si=8p5XVxbBiioz69Xc">see this video</a>.</p><div id="youtube2-e2i6JbU2R-s" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;e2i6JbU2R-s&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/e2i6JbU2R-s?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-approach-evaluation-when-my-system-handles-diverse-user-queries.html">&#8599; Focus view</a></p><h2>Q: How can I efficiently sample production traces for review?</h2><p>It can be cumbersome to review traces randomly, especially when most traces don&#8217;t have an error. These sampling strategies help you find traces more likely to reveal problems:</p><ul><li><p><strong>Outlier detection:</strong> Sort by any metric (response length, latency, tool calls) and review extremes.</p></li><li><p><strong>User feedback signals:</strong> Prioritize traces with negative feedback, support tickets, or escalations.</p></li><li><p><strong>Metric-based sorting:</strong> Generic metrics can serve as exploration signals to find interesting traces. Review both high and low scores and treat them as exploration clues. Based on what you learn, you can build custom evaluators for the failure modes you find.</p></li><li><p><strong>Stratified sampling:</strong> Group traces by key dimensions (user type, feature, query category) and sample from each group.</p></li><li><p><strong>Embedding clustering:</strong> Generate embeddings of queries and cluster them to reveal natural groupings. Sample proportionally from each cluster, but oversample small clusters for edge cases. There&#8217;s no right answer for clustering&#8212;it&#8217;s an exploration technique to surface patterns you might miss manually.</p></li></ul><p>As you get more sophisticated with how you sample, you can incorporate these tactics into the design of your annotation tools.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-can-i-efficiently-sample-production-traces-for-review.html">&#8599; Focus view</a></p><h1>Evaluation Design &amp; Methodology</h1><h2>Q: Why do you recommend binary (pass/fail) evaluations instead of 1-5 ratings (Likert scales)?</h2><blockquote><p>Engineers often believe that Likert scales (1-5 ratings) provide more information than binary evaluations, allowing them to track gradual improvements. However, this added complexity often creates more problems than it solves in practice.</p></blockquote><p>Binary evaluations force clearer thinking and more consistent labeling. Likert scales introduce significant challenges: the difference between adjacent points (like 3 vs 4) is subjective and inconsistent across annotators, detecting statistical differences requires larger sample sizes, and annotators often default to middle values to avoid making hard decisions.</p><p>Having binary options forces people to make a decision rather than hiding uncertainty in middle values. Binary decisions are also faster to make during error analysis - you don&#8217;t waste time debating whether something is a 3 or 4.</p><p>For tracking gradual improvements, consider measuring specific sub-components with their own binary checks rather than using a scale. For example, instead of rating factual accuracy 1-5, you could track &#8220;4 out of 5 expected facts included&#8221; as separate binary checks. This preserves the ability to measure progress while maintaining clear, objective criteria.</p><p>Start with binary labels to understand what &#8216;bad&#8217; looks like. Numeric labels are advanced and usually not necessary.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/why-do-you-recommend-binary-passfail-evaluations-instead-of-1-5-ratings-likert-scales.html">&#8599; Focus view</a></p><h2>Q: Should I practice eval-driven development?</h2><p><strong>Generally no.</strong> Eval-driven development (writing evaluators before implementing features) sounds appealing but creates more problems than it solves. Unlike traditional software where failure modes are predictable, LLMs have infinite surface area for potential failures. You can&#8217;t anticipate what will break.</p><p>A better approach is to start with error analysis. Write evaluators for errors you discover, not errors you imagine. This avoids getting blocked on what to evaluate and prevents wasted effort on metrics that have no impact on actual system quality.</p><p><strong>Exception:</strong> Eval-driven development may work for specific constraints where you know exactly what success looks like. If adding &#8220;never mention competitors,&#8221; writing that evaluator early may be acceptable.</p><p>Most importantly, always do a cost-benefit analysis before implementing an eval. Ask whether the failure mode justifies the investment. Error analysis reveals which failures actually matter for your users.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-i-practice-eval-driven-development.html">&#8599; Focus view</a></p><h2>Q: Should I build automated evaluators for every failure mode I find?</h2><p>Focus automated evaluators on failures that persist after fixing your prompts. Many teams discover their LLM doesn&#8217;t meet preferences they never actually specified - like wanting short responses, specific formatting, or step-by-step reasoning. Fix these obvious gaps first before building complex evaluation infrastructure.</p><p>Consider the cost hierarchy of different evaluator types. Simple assertions and reference-based checks (comparing against known correct answers) are cheap to build and maintain. LLM-as-Judge evaluators require 100+ labeled examples, ongoing weekly maintenance, and coordination between developers, PMs, and domain experts. This cost difference should shape your evaluation strategy.</p><p>Only build expensive evaluators for problems you&#8217;ll iterate on repeatedly. Since LLM-as-Judge comes with significant overhead, save it for persistent generalization failures - not issues you can fix trivially. Start with cheap code-based checks where possible: regex patterns, structural validation, or execution tests. Reserve complex evaluation for subjective qualities that can&#8217;t be captured by simple rules.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-i-build-automated-evaluators-for-every-failure-mode-i-find.html">&#8599; Focus view</a></p><h2>Q: Should I use "ready-to-use" evaluation metrics?</h2><p><strong>No.&nbsp;Generic evaluations waste time and create false confidence.</strong> (Unless you&#8217;re using them for exploration).</p><p>One instructor noted:</p><blockquote><p>&#8220;All you get from using these prefab evals is you don&#8217;t know what they actually do and in the best case they waste your time and in the worst case they create an illusion of confidence that is unjustified.&#8221;<sup>1</sup></p></blockquote><p>Generic evaluation metrics are everywhere. Eval libraries contain scores like helpfulness, coherence, quality, etc. promising easy evaluation. These metrics measure abstract qualities that may not matter for your use case. Good scores on them don&#8217;t mean your system works.</p><p>Instead, conduct error analysis to understand failures. Define binary failure modes based on real problems. Create custom evaluators for those failures and validate them against human judgment. Essentially, the entire evals process.</p><p>Experienced practitioners may still use these metrics, just not how you&#8217;d expect. As Picasso said: &#8220;Learn the rules like a pro, so you can break them like an artist.&#8221; Once you understand why generic metrics fail as evaluations, you can repurpose them as exploration tools to find interesting traces (explained in the next FAQ).</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-i-use-ready-to-use-evaluation-metrics.html">&#8599; Focus view</a></p><h2>Q: Are similarity metrics (BERTScore, ROUGE, etc.) useful for evaluating LLM outputs?</h2><p>Generic metrics like BERTScore, ROUGE, cosine similarity, etc. are not useful for evaluating LLM outputs in most AI applications. Instead, we recommend using error analysis to identify metrics specific to your application&#8217;s behavior. We recommend designing binary pass/fail.) evals (using LLM-as-judge) or code-based assertions.</p><p>As an example, consider a real estate CRM assistant. Suggesting showings that aren&#8217;t available (can be tested with an assertion) or confusing client personas (can be tested with a LLM-as-judge) is problematic . Generic metrics like similarity or verbosity won&#8217;t catch this. A relevant quote from the course:</p><blockquote><p>&#8220;The abuse of generic metrics is endemic. Many eval vendors promote off the shelf metrics, which ensnare engineers into superfluous tasks.&#8221;</p></blockquote><p>Similarity metrics aren&#8217;t always useless. They have utility in domains like search and recommendation (and therefore can be useful for optimizing and debugging retrieval for RAG). For example, cosine similarity between embeddings can measure semantic closeness in retrieval systems, and average pairwise similarity can assess output diversity (where lower similarity indicates higher diversity).</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/are-similarity-metrics-bertscore-rouge-etc-useful-for-evaluating-llm-outputs.html">&#8599; Focus view</a></p><h2>Q: Can I use the same model for both the main task and evaluation?</h2><p>For LLM-as-Judge selection, using the same model is usually fine because the judge is doing a different task than your main LLM pipeline. The judges we recommend building do scoped binary classification tasks. Focus on achieving high True Positive Rate (TPR) and True Negative Rate (TNR) with your judge on a held out labeled test set rather than avoiding the same model family. You can use these metrics on the test set to understand how well your judge is doing.</p><p>When selecting judge models, start with the most capable models available to establish strong alignment with human judgments. You can optimize for cost later once you&#8217;ve established reliable evaluation criteria. We do not recommend using the same model for open ended preferences or response quality (but we don&#8217;t recommend building judges this way in the first place!).</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/can-i-use-the-same-model-for-both-the-main-task-and-evaluation.html">&#8599; Focus view</a></p><div><hr></div><p><strong>&#128073; </strong><em><strong>If you want to learn more about AI Evals, check out our <a href="https://bit.ly/evals-ai">AI Evals course</a></strong></em>. Here is a <a href="https://bit.ly/evals-ai">35% discount code</a> for readers. &#128072;</p><div><hr></div><h1>Human Annotation &amp; Process</h1><h2>Q: How many people should annotate my LLM outputs?</h2><p>For most small to medium-sized companies, appointing a single domain expert as a &#8220;benevolent dictator&#8221; is the most effective approach. This person&#8212;whether it&#8217;s a psychologist for a mental health chatbot, a lawyer for legal document analysis, or a customer service director for support automation&#8212;becomes the definitive voice on quality standards.</p><p>A single expert eliminates annotation conflicts and prevents the paralysis that comes from &#8220;too many cooks in the kitchen&#8221;. The benevolent dictator can incorporate input and feedback from others, but they drive the process. If you feel like you need five subject matter experts to judge a single interaction, it&#8217;s a sign your product scope might be too broad.</p><p>However, larger organizations or those operating across multiple domains (like a multinational company with different cultural contexts) may need multiple annotators. When you do use multiple people, you&#8217;ll need to measure their agreement using metrics like Cohen&#8217;s Kappa, which accounts for agreement beyond chance. However, use your judgment. Even in larger companies, a single expert is often enough.</p><p>Start with a benevolent dictator whenever feasible. Only add complexity when your domain demands it.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-many-people-should-annotate-my-llm-outputs.html">&#8599; Focus view</a></p><h2>Q: Should product managers and engineers collaborate on error analysis? How?</h2><p>At the outset, collaborate to establish shared context. Engineers catch technical issues like retrieval issues and tool errors. PMs identify product failures like unmet user expectations, confusing responses, or missing features users expect.</p><p>As time goes on you should lean towards a benevolent dictator for error analysis: a domain expert or PM who understands user needs. Empower domain experts to evaluate actual outcomes rather than technical implementation. Ask &#8220;Has an appointment been made?&#8221; not &#8220;Did the tool call succeed?&#8221; The best way to empower the domain expert is to give them custom annotation tools that display system outcomes alongside traces. Show the confirmation, generated email, or database update that validates goal completion. Keep all context on one screen so non-technical reviewers focus on results.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-product-managers-and-engineers-collaborate-on-error-analysis-how.html">&#8599; Focus view</a></p><h2>Q: Should I outsource annotation &amp; labeling to a third party?</h2><p>Outsourcing error analysis is usually a big mistake (with some exceptions). The core of evaluation is building the product intuition that only comes from systematically analyzing your system&#8217;s failures. You should be extremely skeptical of this process being delegated.</p><h3><strong>The Dangers of Outsourcing</strong></h3><p>When you outsource annotation, you often break the feedback loop between observing a failure and understanding how to improve the product. Problems with outsourcing include:</p><ul><li><p>Superficial Labeling: Even well-defined metrics require nuanced judgment that external teams lack. A critical misstep in error analysis is excluding domain experts from the labeling process. Outsourcing this task to those without domain expertise, like general developers or IT staff, often leads to superficial or incorrect labeling.<br></p></li><li><p>Loss of Unspoken Knowledge: A principal domain expert possesses tacit knowledge and user understanding that cannot be fully captured in a rubric. Involving these experts helps uncover their preferences and expectations, which they might not be able to fully articulate upfront.<br></p></li><li><p>Annotation Conflicts and Misalignment: Without a shared context, external annotators can create more disagreement than they resolve. Achieving alignment is a challenge even for internal teams, which means you will spend even more time on this process.</p></li></ul><h3><strong>The Recommended Approach: Build Internal Capability</strong></h3><p>Instead of outsourcing, focus on building an efficient internal evaluation process.</p><p>1. Appoint a &#8220;Benevolent Dictator&#8221;. For most teams, the most effective strategy is to appoint a single, internal domain expert as the final decision-maker on quality. This individual sets the standard, ensures consistency, and develops a sense of ownership.</p><p>2. Use a collaborative workflow for multiple annotators. If multiple annotators are necessary, follow a structured process to ensure alignment: * Draft an initial rubric with clear Pass/Fail definitions and examples. * Have each annotator label a shared set of traces independently to surface differences in interpretation. * Measure Inter-Annotator Agreement (IAA) using a chance-corrected metric like Cohen&#8217;s Kappa. * Facilitate alignment sessions to discuss disagreements and refine the rubric. * Iterate on this process until agreement is consistently high.</p><h3><strong>How to Handle Capacity Constraints</strong></h3><p>Building internal capacity does not mean you have to label every trace. Use these strategies to manage the workload:</p><ul><li><p>Smart Sampling: Review a small, representative sample of traces thoroughly. It is more effective to analyze 100 diverse traces to find patterns than to superficially label thousands.<br></p></li><li><p>The &#8220;Think-Aloud&#8221; Protocol: To make the most of limited expert time, use this technique from usability testing. Ask an expert to verbalize their thought process while reviewing a handful of traces. This method can uncover deep insights in a single one-hour session.<br></p></li><li><p>Build Lightweight Custom Tools: Build custom annotation tools to streamline the review process, increasing throughput.</p></li></ul><h3><strong>Exceptions for External Help</strong></h3><p>While outsourcing the core error analysis process is not recommended, there are some scenarios where external help is appropriate:</p><ul><li><p>Purely Mechanical Tasks: For highly objective, unambiguous tasks like identifying a phone number or validating an email address, external annotators can be used after a rigorous internal process has defined the rubric.<br></p></li><li><p>Tasks Without Product Context: Well-defined tasks that don&#8217;t require understanding your product&#8217;s specific requirements can be outsourced. Translation is a good example: it requires linguistic expertise but not deep product knowledge.<br></p></li><li><p>Engaging Subject Matter Experts: Hiring external SMEs to act as your internal domain experts is not outsourcing; it is bringing the necessary expertise into your evaluation process. For example, <a href="https://www.ankihub.net/">AnkiHub</a> hired 4th-year medical students to evaluate their RAG systems for medical content rather than outsourcing to generic annotators.</p></li></ul><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-i-outsource-annotation-and-labeling-to-a-third-party.html">&#8599; Focus view</a></p><h2>Q: What parts of evals can be automated with LLMs?</h2><p>LLMs can speed up parts of your eval workflow, but they can&#8217;t replace human judgment where your expertise is essential. For example, if you let an LLM handle all of <a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed.html">error analysis</a> (i.e., reviewing and annotating traces), you might overlook failure cases that matter for your product. Suppose users keep mentioning &#8220;lag&#8221; in feedback, but the LLM lumps these under generic &#8220;performance issues&#8221; instead of creating a &#8220;latency&#8221; category. You&#8217;d miss a recurring complaint about slow response times and fail to prioritize a fix.</p><p>That said, LLMs are valuable tools for accelerating certain parts of the evaluation workflow <em>when used with oversight</em>.</p><h3>Here are some areas where LLMs can help:</h3><ul><li><p><strong>First-pass axial coding:</strong> After you&#8217;ve open coded 30&#8211;50 traces yourself, use an LLM to organize your raw failure notes into proposed groupings. This helps you quickly spot patterns, but always review and refine the clusters yourself. <em>Note: If you aren&#8217;t familiar with axial and open coding, see <a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/why-is-error-analysis-so-important-in-llm-evals-and-how-is-it-performed.html">this faq</a>.</em></p></li><li><p><strong>Mapping annotations to failure modes:</strong> Once you&#8217;ve defined failure categories, you can ask an LLM to suggest which categories apply to each new trace (e.g., &#8220;Given this annotation: [open_annotation] and these failure modes: [list_of_failure_modes], which apply?&#8221;).<br></p></li><li><p><strong>Suggesting prompt improvements:</strong> When you notice recurring problems, have the LLM propose concrete changes to your prompts. Review these suggestions before adopting any changes.<br></p></li><li><p><strong>Analyzing annotation data:</strong> Use LLMs or AI-powered notebooks to find patterns in your labels, such as &#8220;reports of lag increase 3x during peak usage hours&#8221; or &#8220;slow response times are mostly reported from users on mobile devices.&#8221;</p></li></ul><h3>However, you shouldn&#8217;t outsource these activities to an LLM:</h3><ul><li><p><strong>Initial open coding:</strong> Always read through the raw traces yourself at the start. This is how you discover new types of failures, understand user pain points, and build intuition about your data. Never skip this or delegate it.<br></p></li><li><p><strong>Validating failure taxonomies:</strong> LLM-generated groupings need your review. For example, an LLM might group both &#8220;app crashes after login&#8221; and &#8220;login takes too long&#8221; under a single &#8220;login issues&#8221; category, even though one is a stability problem and the other is a performance problem. Without your intervention, you&#8217;d miss that these issues require different fixes.<br></p></li><li><p><strong>Ground truth labeling:</strong> For any data used for testing/validating LLM-as-Judge evaluators, hand-validate each label. LLMs can make mistakes that lead to unreliable benchmarks.<br></p></li><li><p><strong>Root cause analysis:</strong> LLMs may point out obvious issues, but only human review will catch patterns like errors that occur in specific workflows or edge cases&#8212;such as bugs that happen only when users paste data from Excel.</p></li></ul><p>In conclusion, start by examining data manually to understand what&#8217;s actually going wrong. Use LLMs to scale what you&#8217;ve learned, not to avoid looking at data.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/what-parts-of-evals-can-be-automated-with-llms.html">&#8599; Focus view</a></p><h2>Q: Should I stop writing prompts manually in favor of automated tools?</h2><p>Automating prompt engineering can be tempting, but you should be skeptical of tools that promise to optimize prompts for you, especially in early stages of development. When you write a prompt, you are forced to clarify your assumptions and externalize your requirements. Good writing is good thinking <sup>2</sup>. If you delegate this task to an automated tool too early, you risk never fully understanding your own requirements or the model&#8217;s failure modes.</p><p>This is because automated prompt optimization typically hill-climb a predefined evaluation metric. It can refine a prompt to perform better on known failures, but it cannot discover <em>new</em> ones. Discovering new errors requires error analysis. Furthermore, research shows that evaluation criteria tends to shift after reviewing a model&#8217;s outputs, a phenomenon known as &#8220;criteria drift&#8221; <sup>3</sup>. This means that evaluation is an iterative, human-driven sensemaking process, not a static target that can be set once and handed off to an optimizer.</p><p>A pragmatic approach is to use LLMs to improve your prompt based on open coding (open-ended notes about traces). This way, you maintain a human in the loop who is looking at the data and externalizing their requirements. Once you have a high-quality set of evals, prompt optimization can be effective for that last mile of performance.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-i-stop-writing-prompts-manually-in-favor-of-automated-tools.html">&#8599; Focus view</a></p><h1>Tools &amp; Infrastructure</h1><h2>Q: Should I build a custom annotation tool or use something off-the-shelf?</h2><p><strong>Build a custom annotation tool.</strong> This is the single most impactful investment you can make for your AI evaluation workflow. With AI-assisted development tools like Cursor or Lovable, you can build a tailored interface in hours. I often find that teams with custom annotation tools iterate ~10x faster.</p><p>Custom tools excel because:</p><ul><li><p>They show all your context from multiple systems in one place</p></li><li><p>They can render your data in a product specific way (images, widgets, markdown, buttons, etc.)</p></li><li><p>They&#8217;re designed for your specific workflow (custom filters, sorting, progress bars, etc.)</p></li></ul><p>Off-the-shelf tools may be justified when you need to coordinate dozens of distributed annotators with enterprise access controls. Even then, many teams find the configuration overhead and limitations aren&#8217;t worth it.</p><p><a href="https://youtu.be/fA4pe9bE0LY">Isaac&#8217;s Anki flashcard annotation app</a> shows the power of custom tools&#8212;handling 400+ results per query with keyboard navigation and domain-specific evaluation criteria that would be nearly impossible to configure in a generic tool.</p><div id="youtube2-fA4pe9bE0LY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;fA4pe9bE0LY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/fA4pe9bE0LY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/should-i-build-a-custom-annotation-tool-or-use-something-off-the-shelf.html">&#8599; Focus view</a></p><h2>Q: What makes a good custom interface for reviewing LLM outputs?</h2><p>Great interfaces make human review fast, clear, and motivating. We recommend building your own annotation tool customized to your domain. The following features are possible enhancements we&#8217;ve seen work well, but you don&#8217;t need all of them. The screenshots shown are illustrative examples to clarify concepts. In practice, I rarely implement all these features in a single app. It&#8217;s ultimately a judgment call based on your specific needs and constraints.</p><h3><strong>1. Render Traces Intelligently, Not Generically</strong>:</h3><p>Present the trace in a way that&#8217;s intuitive for the domain. If you&#8217;re evaluating generated emails, render them to look like emails. If the output is code, use syntax highlighting. Allow the reviewer to see the full trace (user input, tool calls, and LLM reasoning), but keep less important details in collapsed sections that can be expanded. Here is an example of a custom annotation tool for reviewing real estate assistant emails:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!jnYp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 424w, /__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 848w, /__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!jnYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 424w, /__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 848w, /__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jnYp!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0e4cf13-9a3f-4830-9ef4-6aed16818833_1984x1736.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">A custom interface for reviewing emails for a real estate assistant.</figcaption></figure></div><h3><strong>2. Show Progress and Support Keyboard Navigation</strong>:</h3><p>Keep reviewers in a state of flow by minimizing friction and motivating completion. Include progress indicators (e.g., &#8220;Trace 45 of 100&#8221;) to keep the review session bounded and encourage completion. Enable hotkeys for navigating between traces (e.g., N for next), applying labels, and saving notes quickly. Below is an illustration of these features:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Wxi-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 424w, /__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 848w, /__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Wxi-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 424w, /__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 848w, /__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Wxi-!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57bb4f01-e303-42e1-bfe3-22f0839fe147_1362x1098.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">An annotation interface with a progress bar and hotkey guide</figcaption></figure></div><h3><strong>3. Trace navigation through clustering, filtering, and search</strong>:</h3><p>Allow reviewers to filter traces by metadata or search by keywords. Semantic search helps find conceptually similar problems. Clustering similar traces (like grouping by user persona) lets reviewers spot recurring issues and explore hypotheses. Below is an illustration of these features:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Fy05!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Fy05!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fy05!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24b8f8be-f7b2-4a8e-92ff-50df096d6152_1564x1730.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Cluster view showing groups of emails, such as property-focused or client-focused examples. Reviewers can drill into a group to see individual traces.</figcaption></figure></div><h3><strong>4. Prioritize labeling traces you think might be problematic</strong>:</h3><p>Surface traces flagged by guardrails, CI failures, or automated evaluators for review. Provide buttons to take actions like adding to datasets, filing bugs, or re-running pipeline tests. Display relevant context (pipeline version, eval scores, reviewer info) directly in the interface to minimize context switching. Below is an illustration of these ideas:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LXpj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 424w, /__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 848w, /__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LXpj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 424w, /__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 848w, /__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LXpj!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ceb30f5-e779-40c3-bd4e-b5d239903775_2070x1152.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">A trace view that allows you to quickly see auto-evaluator verdict, add traces to dataset or open issues. Also shows metadata like pipeline version, reviewer info, and more.</figcaption></figure></div><h3>General Principle: Keep it minimal</h3><p>Keep your annotation interface minimal. Only incorporate these ideas if they provide a benefit that outweighs the additional complexity and maintenance overhead.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/what-makes-a-good-custom-interface-for-reviewing-llm-outputs.html">&#8599; Focus view</a></p><h2>Q: What gaps in eval tooling should I be prepared to fill myself?</h2><p>Most eval tools handle the basics well: logging complete traces, tracking metrics, prompt playgrounds, and annotation queues. These are table stakes. Here are four areas where you&#8217;ll likely need to supplement existing tools.</p><p>Watch for vendors addressing these gaps: it&#8217;s a strong signal they understand practitioner needs.</p><h3>1. Error Analysis and Pattern Discovery</h3><p>After reviewing traces where your AI fails, can your tooling automatically cluster similar issues? For instance, if multiple traces show the assistant using casual language for luxury clients, you need something that recognizes this broader &#8220;persona-tone mismatch&#8221; pattern. We recommend building capabilities that use AI to suggest groupings, rewrite your observations into clearer failure taxonomies, help find similar cases through semantic search, etc.</p><h3>2. AI-Powered Assistance Throughout the Workflow</h3><p>The most effective workflows use AI to accelerate every stage of evaluation. During error analysis, you want an LLM helping categorize your open-ended observations into coherent failure modes. For example, you might annotate several traces with notes like &#8220;wrong tone for investor,&#8221; &#8220;too casual for luxury buyer,&#8221; etc. Your tooling should recognize these as the same underlying pattern and suggest a unified &#8220;persona-tone mismatch&#8221; category.</p><p>You&#8217;ll also want AI assistance in proposing fixes. After identifying 20 cases where your assistant omits pet policies from property summaries, can your workflow analyze these failures and suggest specific prompt modifications? Can it draft refinements to your SQL generation instructions when it notices patterns of missing WHERE clauses?</p><p>Additionally, good workflows help you conduct data analysis of your annotations and traces. I like using notebooks with AI in-the-loop like <a href="https://julius.ai/">Julius</a>,<a href="https://hex.tech">Hex</a> or <a href="https://solveit.fast.ai/">SolveIt</a>. These help me discover insights like &#8220;location ambiguity errors spike 3x when users mention neighborhood names&#8221; or &#8220;tone mismatches occur 80% more often in email generation than other modalities.&#8221;</p><h3>3. Custom Evaluators Over Generic Metrics</h3><p>Be prepared to build most of your evaluators from scratch. Generic metrics like &#8220;hallucination score&#8221; or &#8220;helpfulness rating&#8221; rarely capture what actually matters for your application&#8212;like proposing unavailable showing times or omitting budget constraints from emails. In our experience, successful teams spend most of their effort on application-specific metrics.</p><h3>4. APIs That Support Custom Annotation Apps</h3><p>Custom annotation interfaces work best for most teams. This requires observability platforms with thoughtful APIs. I often have to build my own libraries and abstractions just to make bulk data export manageable. You shouldn&#8217;t have to paginate through thousands of requests or handle timeout-prone endpoints just to get your data. Look for platforms that provide true bulk export capabilities and, crucially, APIs that let you write annotations back efficiently.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/what-gaps-in-eval-tooling-should-i-be-prepared-to-fill-myself.html">&#8599; Focus view</a></p><h2>Q: Seriously Hamel. Stop the bullshit. What&#8217;s your favorite eval vendor?</h2><p>Eval tools are in an intensely competitive space. It would be futile to compare their features. If I tried to do such an analysis, it would be invalidated in a week! Vendors I encounter the most organically in my work are: <a href="https://www.langchain.com/langsmith">Langsmith</a>, <a href="https://arize.com/">Arize</a> and <a href="https://www.braintrust.dev/">Braintrust</a>.</p><p>When I help clients with vendor selection, the decision weighs heavily towards who can offer the best support, as opposed to purely features. This changes depending on size of client, use case, etc. Yes - it&#8217;s mainly the human factor that matters, and dare I say, vibes.</p><p>I have no favorite vendor. At the core, their features are very similar - and I often build <a href="https://hamel.dev/blog/posts/evals/#q-should-i-build-a-custom-annotation-tool-or-use-something-off-the-shelf">custom tools</a> on top of them to fit my needs.</p><p>My suggestion is to explore the vendors and see which one you like the most.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/seriously-hamel-stop-the-bullshit-whats-your-favorite-eval-vendor.html">&#8599; Focus view</a></p><h1>Production &amp; Deployment</h1><h2>Q: How are evaluations used differently in CI/CD vs.&nbsp;monitoring production?</h2><p>The most important difference between CI vs.&nbsp;production evaluation is the data used for testing.</p><p>Test datasets for CI are small (in many cases 100+ examples) and purpose-built. Examples cover core features, regression tests for past bugs, and known edge cases. Since CI tests are run frequently, the cost of each test has to be carefully considered (that&#8217;s why you carefully curate the dataset). Favor assertions or other deterministic checks over LLM-as-judge evaluators.</p><p>For evaluating production traffic, you can sample live traces and run evaluators against them asynchronously. Since you usually lack reference outputs on production data, you might rely more on on more expensive reference-free evaluators like LLM-as-judge. Additionally, track confidence intervals for production metrics. If the lower bound crosses your threshold, investigate further.</p><p>These two systems are complementary: when production monitoring reveals new failure patterns through error analysis and evals, add representative examples to your CI dataset. This mitigates regressions on new issues.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-are-evaluations-used-differently-in-cicd-vs-monitoring-production.html">&#8599; Focus view</a></p><h2>Q: What&#8217;s the difference between guardrails &amp; evaluators?</h2><p>Guardrails are <strong>inline safety checks</strong> that sit directly in the request/response path. They validate inputs or outputs <em>before</em> anything reaches a user, so they typically are:</p><ul><li><p><strong>Fast and deterministic</strong> &#8211; typically a few milliseconds of latency budget.</p></li><li><p><strong>Simple and explainable</strong> &#8211; regexes, keyword block-lists, schema or type validators, lightweight classifiers.</p></li><li><p><strong>Targeted at clear-cut, high-impact failures</strong> &#8211; PII leaks, profanity, disallowed instructions, SQL injection, malformed JSON, invalid code syntax, etc.</p></li></ul><p>If a guardrail triggers, the system can redact, refuse, or regenerate the response. Because these checks are user-visible when they fire, false positives are treated as production bugs; teams version guardrail rules, log every trigger, and monitor rates to keep them conservative.</p><p>On the other hand, evaluators typically run <strong>after</strong> a response is produced. Evaluators measure qualities that simple rules cannot, such as factual correctness, completeness, etc. Their verdicts feed dashboards, regression tests, and model-improvement loops, but they do not block the original answer.</p><p>Evaluators are usually run asynchronously or in batch to afford heavier computation such as a <a href="https://hamel.dev/blog/posts/llm-judge/">LLM-as-a-Judge</a>. Inline use of an LLM-as-Judge is possible <em>only</em> when the latency budget and reliability targets allow it. Slow LLM judges might be feasible in a cascade that runs on the minority of borderline cases.</p><p>Apply guardrails for immediate protection against objective failures requiring intervention. Use evaluators for monitoring and improving subjective or nuanced criteria. Together, they create layered protection.</p><p>Word of caution: Do not use llm guardrails off the shelf blindly. Always <a href="https://hamel.dev/blog/posts/prompt/">look at the prompt</a>.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/whats-the-difference-between-guardrails-evaluators.html">&#8599; Focus view</a></p><h2>Q: Can my evaluators also be used to automatically <em>fix</em> or <em>correct</em> outputs in production?</h2><p>Yes, but only a specific subset of them. This is the distinction between an <strong>evaluator</strong> and a <strong>guardrail</strong> that we previously discussed. As a reminder:</p><ul><li><p><strong>Evaluators</strong> typically run <em>asynchronously</em> after a response has been generated. They measure quality but don&#8217;t interfere with the user&#8217;s immediate experience.<br></p></li><li><p><strong>Guardrails</strong> run <em>synchronously</em> in the critical path of the request, before the output is shown to the user. Their job is to prevent high-impact failures in real-time.</p></li></ul><p>There are two important decision criteria for deciding whether to use an evaluator as a guardrail:</p><ol><li><p><strong>Latency &amp; Cost</strong>: Can the evaluator run fast enough and cheaply enough in the critical request path without degrading user experience?</p></li><li><p><strong>Error Rate Trade-offs</strong>: What&#8217;s the cost-benefit balance between false positives (blocking good outputs and frustrating users) versus false negatives (letting bad outputs reach users and causing harm)? In high-stakes domains like medical advice, false negatives may be more costly than false positives. In creative applications, false positives that block legitimate creativity may be more harmful than occasional quality issues.</p></li></ol><p>Most guardrails are designed to be <strong>fast</strong> (to avoid harming user experience) and have a <strong>very low false positive rate</strong> (to avoid blocking valid responses). For this reason, you would almost never use a slow or non-deterministic LLM-as-Judge as a synchronous guardrail. However, these tradeoffs might be different for your use case.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/can-my-evaluators-also-be-used-to-automatically-fix-or-correct-outputs-in-production.html">&#8599; Focus view</a></p><h2>Q: How much time should I spend on model selection?</h2><p>Many developers fixate on model selection as the primary way to improve their LLM applications. Start with error analysis to understand your failure modes before considering model switching. As Hamel noted in office hours, &#8220;I suggest not thinking of switching model as the main axes of how to improve your system off the bat without evidence. Does error analysis suggest that your model is the problem?&#8221;</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-much-time-should-i-spend-on-model-selection.html">&#8599; Focus view</a></p><h1>Domain-Specific Applications</h1><h2>Q: Is RAG dead?</h2><p>Question: Should I avoid using RAG for my AI application after reading that <a href="/__u/pashpashpash.substack.com/p/why-i-no-longer-recommend-rag-for">&#8220;RAG is dead&#8221;</a> for coding agents?</p><blockquote><p>Many developers are confused about when and how to use RAG after reading articles claiming &#8220;RAG is dead.&#8221; Understanding what RAG actually means versus the narrow marketing definitions will help you make better architectural decisions for your AI applications.</p></blockquote><p>The viral article claiming RAG is dead specifically argues against using <em>naive vector database retrieval</em> for autonomous coding agents, not RAG as a whole. This is a crucial distinction that many developers miss due to misleading marketing.</p><p>RAG simply means Retrieval-Augmented Generation - using retrieval to provide relevant context that improves your model&#8217;s output. The core principle remains essential: your LLM needs the right context to generate accurate answers. The question isn&#8217;t whether to use retrieval, but how to retrieve effectively.</p><p>For coding applications, naive vector similarity search often fails because code relationships are complex and contextual. Instead of abandoning retrieval entirely, modern coding assistants like Claude Code <a href="https://x.com/pashmerepat/status/1926717705660375463?s=46">still uses retrieval</a> &#8212;they just employ agentic search instead of relying solely on vector databases, similar to how human developers work.</p><p>You have multiple retrieval strategies available, ranging from simple keyword matching to embedding similarity to LLM-powered relevance filtering. The optimal approach depends on your specific use case, data characteristics, and performance requirements. Many production systems combine multiple strategies or use multi-hop retrieval guided by LLM agents.</p><p>Unfortunately, &#8220;RAG&#8221; has become a buzzword with no shared definition. Some people use it to mean any retrieval system, others restrict it to vector databases. Focus on the ultimate goal: getting your LLM the context it needs to succeed. Whether that&#8217;s through vector search, agentic exploration, or hybrid approaches is a product and engineering decision.</p><p>Rather than following categorical advice to avoid or embrace RAG, experiment with different retrieval approaches and measure what works best for your application. For more info on RAG evaluation and optimization, see <a href="/__u/hamelhusain.substack.com/notes/llm/rag/not_dead.html">this series of posts</a>.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/is-rag-dead.html">&#8599; Focus view</a></p><h2>Q: How should I approach evaluating my RAG system?</h2><p>RAG systems have two distinct components that require different evaluation approaches: retrieval and generation.</p><p>The retrieval component is a search problem. Evaluate it using traditional information retrieval (IR) metrics. Common examples include Recall@k (of all relevant documents, how many did you retrieve in the top k?), Precision@k (of the k documents retrieved, how many were relevant?), or MRR (how high up was the first relevant document?). The specific metrics you choose depend on your use case. These metrics are pure search metrics that measure whether you&#8217;re finding the right documents (more on this below).</p><p>To evaluate retrieval, create a dataset of queries paired with their relevant documents. Generate this synthetically by taking documents from your corpus, extracting key facts, then generating questions those facts would answer. This reverse process gives you query-document pairs for measuring retrieval performance without manual annotation.</p><p>For the generation component&#8212;how well the LLM uses retrieved context, whether it hallucinates, whether it answers the question&#8212;use the same evaluation procedures covered throughout this course: error analysis to identify failure modes, collecting human labels, building LLM-as-judge evaluators, and validating those judges against human annotations.</p><p>Jason Liu&#8217;s <a href="https://jxnl.co/writing/2025/05/19/there-are-only-6-rag-evals/">&#8220;There Are Only 6 RAG Evals&#8221;</a> provides a framework that maps well to this separation. His Tier 1 covers traditional IR metrics for retrieval. Tiers 2 and 3 evaluate relationships between Question, Context, and Answer&#8212;like whether the context is relevant (C|Q), whether the answer is faithful to context (A|C), and whether the answer addresses the question (A|Q).</p><p>In addition to Jason&#8217;s six evals, error analysis on your specific data may reveal domain-specific failure modes that warrant their own metrics. For example, a medical RAG system might consistently fail to distinguish between drug dosages for adults versus children, or a legal RAG might confuse jurisdictional boundaries. These patterns emerge only through systematic review of actual failures. Once identified, you can create targeted evaluators for these specific issues beyond the general framework.</p><p>Finally, when implementing Jason&#8217;s Tier 2 and 3 metrics, don&#8217;t just use prompts off the shelf. The standard LLM-as-judge process requires several steps: error analysis, prompt iteration, creating labeled examples, and measuring your judge&#8217;s accuracy against human labels. Once you know your judge&#8217;s True Positive and True Negative rates, you can correct its estimates to determine the actual failure rate in your system. Skip this validation and your judges may not reflect your actual quality criteria.</p><p>In summary, debug retrieval first using IR metrics, then tackle generation quality using properly validated LLM judges.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-should-i-approach-evaluating-my-rag-system.html">&#8599; Focus view</a></p><h2>Q: How do I choose the right chunk size for my document processing tasks?</h2><p>Unlike RAG, where chunks are optimized for retrieval, document processing assumes the model will see every chunk. The goal is to split text so the model can reason effectively without being overwhelmed. Even if a document fits within the context window, it might be better to break it up. Long inputs can degrade performance due to attention bottlenecks, especially in the middle of the context. Two task types require different strategies:</p><h3>1. Fixed-Output Tasks &#8594; Large Chunks</h3><p>These are tasks where the output length doesn&#8217;t grow with input: extracting a number, answering a specific question, classifying a section. For example:</p><ul><li><p>&#8220;What&#8217;s the penalty clause in this contract?&#8221;</p></li><li><p>&#8220;What was the CEO&#8217;s salary in 2023?&#8221;</p></li></ul><p>Use the largest chunk (with caveats) that likely contains the answer. This reduces the number of queries and avoids context fragmentation. However, avoid adding irrelevant text. Models are sensitive to distraction, especially with large inputs. The middle parts of a long input might be under-attended. Furthermore, if cost and latency are a bottleneck, you should consider preprocessing or filtering the document (via keyword search or a lightweight retriever) to isolate relevant sections before feeding a huge chunk.</p><h3>2. Expansive-Output Tasks &#8594; Smaller Chunks</h3><p>These include summarization, exhaustive extraction, or any task where output grows with input. For example:</p><ul><li><p>&#8220;Summarize each section&#8221;</p></li><li><p>&#8220;List all customer complaints&#8221;</p></li></ul><p>In these cases, smaller chunks help preserve reasoning quality and output completeness. The standard approach is to process each chunk independently, then aggregate results (e.g., map-reduce). When sizing your chunks, try to respect content boundaries like paragraphs, sections, or chapters. Chunking also helps mitigate output limits. By breaking the task into pieces, each piece&#8217;s output can stay within limits.</p><h3>General Guidance</h3><p>It&#8217;s important to recognize <strong>why chunk size affects results</strong>. A larger chunk means the model has to reason over more information in one go &#8211; essentially, a heavier cognitive load. LLMs have limited capacity to <strong>retain and correlate details across a long text</strong>. If too much is packed in, the model might prioritize certain parts (commonly the beginning or end) and overlook or &#8220;forget&#8221; details in the middle. This can lead to overly coarse summaries or missed facts. In contrast, a smaller chunk bounds the problem: the model can pay full attention to that section. You are trading off <strong>global context for local focus</strong>.</p><p>No rule of thumb can perfectly determine the best chunk size for your use case &#8211; <strong>you should validate with experiments</strong>. The optimal chunk size can vary by domain and model. I treat chunk size as a hyperparameter to tune.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-choose-the-right-chunk-size-for-my-document-processing-tasks.html">&#8599; Focus view</a></p><h2>Q: How do I debug multi-turn conversation traces?</h2><p>Start simple. Check if the whole conversation met the user&#8217;s goal with a pass/fail judgment. Look at the entire trace and focus on the first upstream failure. Read the user-visible parts first to understand if something went wrong. Only then dig into the technical details like tool calls and intermediate steps.</p><h3>Multi-agent trace logging</h3><p>For multi-agent flows, assign a session or trace ID to each user request and log every message with its source (which agent or tool), trace ID, and position in the sequence. This lets you reconstruct the full path from initial query to final result across all agents.</p><h3>Annotation strategy</h3><p>Annotate only the first failure in the trace initially&#8212;don&#8217;t worry about downstream failures since these often cascade from the first issue. Fixing upstream failures often resolves dependent downstream failures automatically. As you gain experience, you can annotate independent failure modes within the same trace to speed up overall error analysis.</p><h3>Simplify when possible</h3><p>When you find a failure, reproduce it with the simplest possible test case. Here&#8217;s an example: suppose a shopping bot gives the wrong return policy on turn 4 of a conversation. Before diving into the full multi-turn complexity, simplify it to a single turn: &#8220;What is the return window for product X1000?&#8221; If it still fails, you&#8217;ve proven the error isn&#8217;t about conversation context - it&#8217;s likely a basic retrieval or knowledge issue you can debug more easily.</p><h3>Test case generation</h3><p>You have two main approaches. First, simulate users with another LLM to create realistic multi-turn conversations. Second, use &#8220;N-1 testing&#8221; where you provide the first N-1 turns of a real conversation and test what happens next. The N-1 approach often works better since it uses actual conversation prefixes rather than fully synthetic interactions, but is less flexible.</p><p>The key is balancing thoroughness with efficiency. Not every multi-turn failure requires multi-turn analysis.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-debug-multi-turn-conversation-traces.html">&#8599; Focus view</a></p><h2>Q: How do I evaluate sessions with human handoffs?</h2><p>Capture the complete user journey in your traces, including human handoffs. The trace continues until the user&#8217;s need is resolved or the session ends, not when AI hands off to a human. Log the handoff decision, why it occurred, context transferred, wait time, human actions, final resolution, and whether the human had sufficient context. Many failures occur at handoff boundaries where AI hands off too early, too late, or without proper context.</p><p>Evaluate handoffs as potential failure modes during error analysis. Ask: Was the handoff necessary? Did the AI provide adequate context? Track both handoff quality and handoff rate. Sometimes the best improvement reduces handoffs entirely rather than improving handoff execution.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-evaluate-sessions-with-human-handoffs.html">&#8599; Focus view</a></p><h2>Q: How do I evaluate complex multi-step workflows?</h2><p>Log the entire workflow from initial trigger to final business outcome. Include LLM calls, tool usage, human approvals, and database writes in your traces. You will need this visibility to properly diagnose failures.</p><p>Use both outcome and process metrics. Outcome metrics verify the final result meets requirements: Was the business case complete? Accurate? Properly formatted? Process metrics evaluate efficiency: step count, time taken, resource usage. Process failures are often easier to debug since they&#8217;re more deterministic, so tackle them first.</p><p>Segment your error analysis by workflow stages. Early stage failures (understanding user input) differ from middle stage failures (data processing) and late stage failures (formatting output). Early stage improvements have more impact since errors cascade in LLM chains.</p><p>Use transition failure matrices to analyze where workflows break. Create a matrix showing the last successful state versus where the first failure occurred. This reveals failure hotspots and guides where to invest debugging effort.</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-evaluate-complex-multi-step-workflows.html">&#8599; Focus view</a></p><h2>Q: How do I evaluate agentic workflows?</h2><p>We recommend evaluating agentic workflows in two phases:</p><p><strong>1. End-to-end task success.</strong> Treat the agent as a black box and ask &#8220;did we meet the user&#8217;s goal?&#8221;. Define a precise success rule per task (exact answer, correct side-effect, etc.) and measure with human or <a href="https://hamel.dev/blog/posts/llm-judge/">aligned LLM judges</a>. Take note of the first upstream failure when conducting error analysis.</p><p>Once error analysis reveals which workflows fail most often, move to step-level diagnostics to understand why they&#8217;re failing.</p><p><strong>2. Step-level diagnostics.</strong> Assuming that you have sufficiently <a href="https://hamel.dev/blog/posts/evals/#logging-traces">instrumented your system</a> with details of tool calls and responses, you can score individual components such as: - <em>Tool choice</em>: was the selected tool appropriate? - <em>Parameter extraction</em>: were inputs complete and well-formed? - <em>Error handling</em>: did the agent recover from empty results or API failures? - <em>Context retention</em>: did it preserve earlier constraints? - <em>Efficiency</em>: how many steps, seconds, and tokens were spent? - <em>Goal checkpoints</em>: for long workflows verify key milestones.</p><p>Example: &#8220;Find Berkeley homes under $1M and schedule viewings&#8221; breaks into: parameters extracted correctly, relevant listings retrieved, availability checked, and calendar invites sent. Each checkpoint can pass or fail independently, making debugging tractable.</p><p><strong>Use transition failure matrices to understand error patterns.</strong> Create a matrix where rows represent the last successful state and columns represent where the first failure occurred. This is a great way to understand where the most failures occur.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!PM1N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 424w, /__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 848w, /__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!PM1N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 424w, /__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 848w, /__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!PM1N!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F102ec47f-0661-4bef-8755-36b161bd9470_1140x742.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Transition failure matrix showing hotspots in text-to-SQL agent workflow</figcaption></figure></div><p>Transition matrices transform overwhelming agent complexity into actionable insights. Instead of drowning in individual trace reviews, you can immediately see that GenSQL &#8594; ExecSQL transitions cause 12 failures while DecideTool &#8594; PlanCal causes only 2. This data-driven approach guides where to invest debugging effort. Here is another <a href="https://www.figma.com/deck/nwRlh5renu4s4olaCsf9lG/Failure-is-a-Funnel?node-id=2009-927&amp;t=GJlTtxQ8bLJaQ92A-1">example</a> from Bryan Bischof, that is also a text-to-SQL agent:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iuoS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 424w, /__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 848w, /__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iuoS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 424w, /__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 848w, /__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iuoS!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc35f0da4-ada7-4e8d-a79c-e96dd421ae7b_2154x1102.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Bischof, Bryan &#8220;Failure is A Funnel - Data Council, 2025&#8221;</figcaption></figure></div><p>In this example, Bryan shows variation in transition matrices across experiments. How you organize your transition matrix depends on the specifics of your application. For example, Bryan&#8217;s text-to-SQL agent has an inherent sequential workflow which he exploits for further analytical insight. You can watch his <a href="https://youtu.be/R_HnI9oTv3c?si=hRRhDiydHU5k6ikc">full talk</a> for more details.</p><div id="youtube2-R_HnI9oTv3c" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;R_HnI9oTv3c&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/R_HnI9oTv3c?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong>Creating Test Cases for Agent Failures</strong></p><p>Creating test cases for agent failures follows the same principles as our previous FAQ on debugging multi-turn conversation traces (i.e.&nbsp;try to reproduce the error in the simplest way possible, only use multi-turn tests when the failure actually requires conversation context, etc.).</p><p><a href="/__u/hamelhusain.substack.com/blog/posts/evals-faq/how-do-i-evaluate-agentic-workflows.html">&#8599; Focus view</a></p><div><hr></div><p><strong>&#128073; </strong><em><strong>If you want to learn more about AI Evals, check out our <a href="https://bit.ly/evals-ai">AI Evals course</a></strong></em>. Here is a <a href="https://bit.ly/evals-ai">35% discount code</a> for readers. &#128072;</p><div><hr></div><h2>Footnotes</h2><ol><li><p><a href="https://www.linkedin.com/in/intellectronica/">Eleanor Berger</a>, our wonderful TA.&#8617;&#65038;</p></li><li><p>Paul Graham, <a href="https://paulgraham.com/writes.html">&#8220;Writes and Write-Nots&#8221;</a>&#8617;&#65038;</p></li><li><p>Shreya Shankar, et al., <a href="https://arxiv.org/abs/2404.12272">&#8220;Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences&#8221;</a>&#8617;&#65038;</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Stop Saying RAG Is Dead]]></title><description><![CDATA[Why the future of RAG lies in better retrieval, not bigger context windows.]]></description><link>https://hamelhusain.substack.com/p/hameldev</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/hameldev</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Fri, 11 Jul 2025 07:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OHTi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m tired of hearing &#8220;RAG is dead.&#8221; That&#8217;s why <a href="https://ben.clavie.eu/">Ben Clavi&#233;</a> and I put together this <a href="https://hamel.dev/notes/llm/rag/">open 5-part series</a> to discuss why RAG is not dead, and what the future of RAG looks like.</p><h2><strong>What&#8217;s Actually Dead</strong></h2><p><a href="https://hamel.dev/notes/llm/rag/p1-intro.html">Ben Clavi&#233;&#8217;s opener</a> nailed it: what&#8217;s dead is the 2023 marketing version of RAG. Chuck documents into a vector database, do cosine similarity, call it a day. This approach fails because compressing entire documents into single vectors loses critical information.</p><p>But retrieval is more important than ever. LLMs are frozen at training time. <strong>Million-token context windows don&#8217;t change the economics or efficiency of stuffing everything into every query.</strong></p><h2><strong>Takeaways From the Series</strong></h2><p><strong>We&#8217;ve been measuring wrong.</strong> <a href="https://hamel.dev/notes/llm/rag/p2-evals.html">Nandan Thakur showed</a> that traditional IR metrics optimize for finding the #1 result. RAG needs different goals: coverage (getting all the facts), diversity (corroborating facts), and relevance. Models that ace BEIR benchmarks often fail at real RAG tasks.</p><p><strong>Retrieval can reason.</strong> <a href="https://hamel.dev/notes/llm/rag/p3_reasoning.html">Orion Weller&#8217;s models</a> understand instructions like &#8220;find documents about data privacy using metaphors.&#8221; His Rank1 system generates explicit reasoning traces about relevance. These models find documents that traditional systems never surface.</p><p><strong>Single vectors lose information.</strong> <a href="https://hamel.dev/notes/llm/rag/p4_late_interaction.html">Antoine Chaffin demonstrated</a> how late-interaction models like ColBERT preserve token-level information. No more forcing everything into one conflicted representation. Result: 150M parameter models outperforming 7B parameter alternatives on reasoning tasks.</p><p><strong>One map isn&#8217;t enough.</strong> <a href="https://hamel.dev/notes/llm/rag/p5_map.html">Bryan and Ayush&#8217;s finale</a> showed why we need multiple representations. Their art search demo finds the same painting through literal descriptions, poetic interpretations, or similar images&#8212;each using different indices. Stop searching for the perfect embedding. Build specialized representations and route intelligently.</p><h2><strong>What&#8217;s Next</strong></h2><p>I think a path forward is to combine these ideas:</p><ul><li><p>Evaluation systems that measure what matters for your use case</p></li><li><p>Retrievers that understand instructions and reason about relevance</p></li><li><p>Representations that preserve information instead of compressing it away</p></li><li><p>Multiple specialized indices with intelligent routing</p></li></ul><div><hr></div><h2><strong>Annotated Notes From the Series</strong></h2><p>Each post includes timestamped annotations of slides, saving you hours of video watching. We&#8217;ve highlighted the most important bits and provided context for quickly grokking the material.</p><p>TitleDescription</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OHTi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 424w, /__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 848w, /__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OHTi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 424w, /__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 848w, /__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OHTi!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb164011-194d-42d6-9b85-3aee671bcbf0_8000x4500.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="https://hamel.dev/notes/llm/rag/p1-intro.html">Part 1</a>: <strong>I don&#8217;t use RAG, I just retrieve documents</strong>Ben Clavi&#233; explains why naive single-vector search is dead, not RAG itself</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LcAC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LcAC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 424w, /__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 848w, /__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LcAC!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73ed5e5b-9b72-4d90-bb78-4e759f7a8546_1500x1125.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="https://hamel.dev/notes/llm/rag/p2-evals.html">Part 2</a>: <strong>Modern IR Evals For RAG</strong>Nandan Thakur shows why traditional IR metrics fail for RAG and introduces FreshStack</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vMFb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 424w, /__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 848w, /__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vMFb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 424w, /__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 848w, /__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vMFb!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa946f56a-49ab-4512-b442-1cf80d233787_4000x2250.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="https://hamel.dev/notes/llm/rag/p3_reasoning.html">Part 3</a>: <strong>Optimizing Retrieval with Reasoning Models</strong>Orion Weller demonstrates retrievers that understand instructions and reason about relevance</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Mu6L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Mu6L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mu6L!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da4ff5b-3239-47aa-a64e-b63f5fac4fb2_1500x844.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="https://hamel.dev/notes/llm/rag/p4_late_interaction.html">Part 4</a>: <strong>Late Interaction Models For RAG</strong>Antoine Chaffin reveals how ColBERT-style models preserve information that single vectors lose</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fOjf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 424w, /__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 848w, /__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fOjf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 424w, /__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 848w, /__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fOjf!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d9a380-53d7-47fb-9d27-4246d696552e_4000x2250.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="https://hamel.dev/notes/llm/rag/p5_map.html">Part 5</a>: <strong>RAG with Multiple Representations</strong>Bryan Bischof and Ayush Chaurasia show why multiple specialized indices beat one perfect embedding</p>]]></content:encoded></item><item><title><![CDATA[A Field Guide to Rapidly Improving AI Products]]></title><description><![CDATA[Most AI teams focus on the wrong things.]]></description><link>https://hamelhusain.substack.com/p/field-guide</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/field-guide</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Mon, 24 Mar 2025 07:00:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2c022271-2d97-4d20-82fa-46db291948a5_1600x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most AI teams focus on the wrong things. Here&#8217;s a common scene from my consulting work:</p><p><strong> AI TEAM</strong></p><blockquote><p>Here&#8217;s our agent architecture &#8211; we&#8217;ve got RAG here, a router there, and we&#8217;re using this new framework for&#8230;</p></blockquote><p><strong> ME</strong></p><blockquote><p><em>[Holding up my hand to pause the enthusiastic tech lead.]</em></p><p>&#8220;Can you show me how you&#8217;re measuring if any of this actually works?&#8221;</p></blockquote><p><em> &#8230; Room goes quiet</em></p><p>This scene has played out dozens of times over the last two years. Teams invest weeks building complex AI systems, but can&#8217;t tell me if their changes are helping or hurting.</p><p>This isn&#8217;t surprising. With new tools and frameworks emerging weekly, it&#8217;s natural to focus on tangible things we can control &#8211; which vector database to use, which LLM provider to choose, which agent framework to adopt. But after helping 30+ companies build AI products, I&#8217;ve discovered the teams who succeed barely talk about tools at all. Instead, they obsess over measurement and iteration.</p><p>In this post, I&#8217;ll show you exactly how these successful teams operate. You&#8217;ll learn:</p><ol><li><p>How error analysis consistently reveals the highest-ROI improvements</p></li><li><p>Why a simple data viewer is your most important AI investment</p></li><li><p>How to empower domain experts (not just engineers) to improve your AI</p></li><li><p>Why synthetic data is more effective than you think</p></li><li><p>How to maintain trust in your evaluation system</p></li><li><p>Why your AI roadmap should count experiments, not features</p></li></ol><p>I&#8217;ll explain each of these topics with real examples. While every situation is unique, you&#8217;ll see patterns that apply regardless of your domain or team size.</p><p>Let&#8217;s start by examining the most common mistake I see teams make &#8211; one that derails AI projects before they even begin.</p><h2>1. The Most Common Mistake: Skipping Error Analysis</h2><p>The &#8220;tools first&#8221; mindset is the most common mistake in AI development. Teams get caught up in architecture diagrams, frameworks, and dashboards while neglecting the process of actually understanding what&#8217;s working and what isn&#8217;t.</p><p>One client proudly showed me this evaluation dashboard:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!RHOJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 424w, /__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 848w, /__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!RHOJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 424w, /__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 848w, /__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RHOJ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F173c5a2d-8b0f-4a00-a809-47775c572865_1600x800.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">The kind of dashboard that foreshadows failure.</figcaption></figure></div><p>This is the &#8220;tools trap&#8221; &#8211; the belief that adopting the right tools or frameworks (in this case, generic metrics) will solve your AI problems. Generic metrics are worse than useless &#8211; they actively impede progress in two ways:</p><p>First, they create a <strong>false sense of measurement and progress</strong>. Teams think they&#8217;re data-driven because they have dashboards, but they&#8217;re tracking vanity metrics that don&#8217;t correlate with real user problems. I&#8217;ve seen teams celebrate improving their &#8220;helpfulness score&#8221; by 10% while their actual users were still struggling with basic tasks. It&#8217;s like optimizing your website&#8217;s load time while your checkout process is broken &#8211; you&#8217;re getting better at the wrong thing.</p><p>Second, too many metrics fragment your attention. Instead of focusing on the few metrics that matter for your specific use case, you&#8217;re trying to optimize multiple dimensions simultaneously. When everything is important, nothing is.</p><p>The alternative? Error analysis - the single most valuable activity in AI development and consistently the highest-ROI activity. Let me show you what effective error analysis looks like in practice.</p><h3>The Error Analysis Process</h3><p>When Jacob, the founder of <a href="https://nurtureboss.io/">Nurture Boss</a>, needed to improve their apartment-industry AI assistant, his team built a simple viewer to examine conversations between their AI and users. Next to each conversation was a space for open-ended notes about failure modes.</p><p>After annotating dozens of conversations, clear patterns emerged. Their AI was struggling with date handling &#8211; failing 66% of the time when users said things like &#8220;let&#8217;s schedule a tour two weeks from now.&#8221;</p><p>Instead of reaching for new tools, they: 1. Looked at actual conversation logs 2. Categorized the types of date-handling failures 3. Built specific tests to catch these issues 4. Measured improvement on these metrics</p><p>The result? Their date handling success rate improved from 33% to 95%.</p><p>Here&#8217;s Jacob explaining this process himself:</p><div id="youtube2-e2i6JbU2R-s" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;e2i6JbU2R-s&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/e2i6JbU2R-s?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h3>Bottom-Up vs.&nbsp;Top-Down Analysis</h3><p>When identifying error types, you can take either a &#8220;top-down&#8221; or &#8220;bottom-up&#8221; approach.</p><p>The <strong>top-down</strong> approach starts with common metrics like &#8220;hallucination&#8221; or &#8220;toxicity&#8221; plus metrics unique to your task. While convenient, it often misses domain-specific issues.</p><p>The more effective <strong>bottom-up</strong> approach forces you to look at actual data and let metrics naturally emerge. At NurtureBoss, we started with a spreadsheet where each row represented a conversation. We wrote open-ended notes on any undesired behavior. Then we used an LLM to build a taxonomy of common failure modes. Finally, we mapped each row to specific failure mode labels and counted the frequency of each issue.</p><p>The results were striking - just three issues accounted for over 60% of all problems:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tT9U!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 424w, /__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 848w, /__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tT9U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a80600b-ace0-44f0-8449-de36f5158132_780x478.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 424w, /__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 848w, /__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tT9U!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a80600b-ace0-44f0-8449-de36f5158132_780x478.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Excel Pivot Tables are a simple tool, but they work!</figcaption></figure></div><ul><li><p>Conversation flow issues (missing context, awkward responses)</p></li><li><p>Handoff failures (not recognizing when to transfer to humans)</p></li><li><p>Rescheduling problems (struggling with date handling)</p></li></ul><p>The impact was immediate. Jacob&#8217;s team had uncovered so many actionable insights that they needed several weeks just to implement fixes for the problems we&#8217;d already found.</p><p>If you&#8217;d like to see error analysis in action, we recorded a <a href="https://youtu.be/qH1dZ8JLLdU">live walkthrough here</a>.</p><p>This brings us to a crucial question: How do you make it easy for teams to look at their data? The answer leads us to what I consider the most important investment any AI team can make&#8230;</p><h2>2. The Most Important AI Investment: A Simple Data Viewer</h2><p>The single most impactful investment I&#8217;ve seen AI teams make isn&#8217;t a fancy evaluation dashboard &#8211; it&#8217;s building a customized interface that lets anyone examine what their AI is actually doing. I emphasize <em>customized</em> because every domain has unique needs that off-the-shelf tools rarely address. When reviewing apartment leasing conversations, you need to see the full chat history and scheduling context. For real estate queries, you need the property details and source documents right there. Even small UX decisions &#8211; like where to place metadata or which filters to expose &#8211; can make the difference between a tool people actually use and one they avoid.</p><p>I&#8217;ve watched teams struggle with generic labeling interfaces, hunting through multiple systems just to understand a single interaction. The friction adds up: clicking through to different systems to see context, copying error descriptions into separate tracking sheets, switching between tools to verify information. This friction doesn&#8217;t just slow teams down &#8211; it actively discourages the kind of systematic analysis that catches subtle issues.</p><p>Teams with thoughtfully designed data viewers iterate 10x faster than those without them. And here&#8217;s the thing: <strong>these tools can be built in hours using AI-assisted development</strong> (like Cursor or Loveable). The investment is minimal compared to the returns.</p><p>Let me show you what I mean. Here&#8217;s the data viewer built for NurtureBoss (which we discussed earlier):</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Aawa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 424w, /__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 848w, /__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Aawa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25004440-b173-413d-a5e3-e9b503015659_2868x1568.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 424w, /__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 848w, /__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Aawa!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25004440-b173-413d-a5e3-e9b503015659_2868x1568.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Search and filter sessions</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!B1tZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 424w, /__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 848w, /__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 1272w, /__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!B1tZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 424w, /__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 848w, /__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 1272w, /__u/substackcdn.com/image/fetch/$s_!B1tZ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7005617-63e1-4388-a085-6088e4bcab4d_2876x1576.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Annotate and add notes</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!1NtV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 424w, /__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 848w, /__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!1NtV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 424w, /__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 848w, /__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 1272w, /__u/substackcdn.com/image/fetch/$s_!1NtV!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f7e4e7-f27b-4091-867b-aeb7b17609b2_2878x1576.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Aggregate and count errors</figcaption></figure></div><p>Here&#8217;s what makes a good data annotation tool:</p><ol><li><p>Show all context in one place. Don&#8217;t make users hunt through different systems to understand what happened.<br></p></li><li><p>Make feedback trivial to capture. One-click correct/incorrect buttons beat lengthy forms.</p></li><li><p>Capture open-ended feedback. This lets you capture nuanced issues that don&#8217;t fit into a pre-defined taxonomy.</p></li><li><p>Enable quick filtering and sorting. Teams need to easily dive into specific error types. In the example above, NurtureBoss can quickly filter by the channel (voice, text, chat) or the specific property they want to look at quickly.</p></li><li><p>Have hotkeys that allow users to navigate between data examples and annotate without clicking.</p></li></ol><p>It doesn&#8217;t matter what web frameworks you use - use whatever you are familiar with. Because I&#8217;m a python developer, my current favorite web framework is <a href="https://fastht.ml/docs/">FastHTML</a> coupled with <a href="https://www.answer.ai/posts/2025-01-15-monsterui.html">MonsterUI</a>, because it allows me to define the back-end and front-end code in one small python file.</p><p>The key is starting somewhere, even if it&#8217;s simple. I&#8217;ve found custom web apps provide the best experience, but if you&#8217;re just beginning, a spreadsheet is better than nothing. As your needs grow, you can evolve your tools accordingly.</p><p>This brings us to another counter-intuitive lesson: the people best positioned to improve your AI system are often the ones who know the least about AI.</p><h2>3. Empower Domain Experts To Write Prompts</h2><p>I recently worked with an education startup building an interactive learning platform with LLMs. Their product manager, a learning design expert, would create detailed PowerPoint decks explaining pedagogical principles and example dialogues. She&#8217;d present these to the engineering team, who would then translate her expertise into prompts.</p><p>But here&#8217;s the thing: prompts are just English. Having a learning expert communicate teaching principles through PowerPoint, only for engineers to translate that back into English prompts, created unnecessary friction. The most successful teams flip this model by giving domain experts tools to write and iterate on prompts directly.</p><h3>Build Bridges, Not Gatekeepers</h3><p>Prompt playgrounds are a great starting point for this. Tools like Arize, Langsmith and Braintrust let teams quickly test different prompts, feed in example datasets, and compare results. Here are some screenshots of these tools:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FohN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 424w, /__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 848w, /__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FohN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 424w, /__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 848w, /__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FohN!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70f508a1-a1f5-4746-ba49-a17e7cde2db1_3140x1724.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Arize Phoenix</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ig02!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ig02!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ig02!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cd7f496-3628-4c37-b650-85cf55e714f8_2816x1676.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">LangSmith</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XfhK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 424w, /__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 848w, /__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XfhK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 424w, /__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 848w, /__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XfhK!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c24b01c-e439-4cee-8351-2165f116f97f_2566x1310.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Braintrust</figcaption></figure></div><p>But there&#8217;s a crucial next step that many teams miss: integrating prompt development into their application context. Most AI applications aren&#8217;t just prompts &#8211; They commonly involve RAG systems pulling from your knowledge base, agent orchestration coordinating multiple steps, and application-specific business logic. The most effective teams I&#8217;ve worked with go beyond standalone playgrounds. They build what I call <em><strong>integrated prompt environments</strong></em> &#8211; essentially admin versions of their actual user interface that expose prompt editing.</p><p>Here&#8217;s an illustration of what an integrated prompt environment might look like for a real estate AI assistant:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fqKO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 424w, /__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 848w, /__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fqKO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 424w, /__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 848w, /__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fqKO!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad297a32-1d57-4a48-94db-1dbaf12538fa_2356x2284.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">The UI that users (real estate agents) see.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!b1WA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 424w, /__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 848w, /__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 1272w, /__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!b1WA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 424w, /__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 848w, /__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 1272w, /__u/substackcdn.com/image/fetch/$s_!b1WA!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b1149a7-9aff-4f90-91ad-c2e2d1b294d4_2358x2290.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">The same UI, but with an &#8220;admin mode&#8221;used by the engineering &amp; product team to iterate on the prompt and debug issues.</figcaption></figure></div><h3>Tips For Communicating With Domain Experts</h3><p>There&#8217;s another barrier that often prevents domain experts from contributing effectively: unnecessary jargon. I was working with an education startup where engineers, product managers, and learning specialists were talking past each other in meetings. The engineers kept saying, &#8220;We&#8217;re going to build an agent that does XYZ,&#8221; when really the job to be done was writing a prompt. This created an artificial barrier &#8211; the learning specialists, who were the actual domain experts, felt like they couldn&#8217;t contribute because they didn&#8217;t understand &#8220;agents.&#8221;</p><p>This happens everywhere. I&#8217;ve seen it with lawyers at legal tech companies, psychologists at mental health startups, and doctors at healthcare firms. The magic of LLMs is that they make AI accessible through natural language, but we often destroy that advantage by wrapping everything in technical terminology.</p><p>Here&#8217;s a simple example of how to translate common AI jargon:</p><p>Instead of saying&#8230; Say&#8230; &#8220;We&#8217;re implementing a RAG approach&#8221; &#8220;We&#8217;re making sure the model has the right context to answer questions&#8221; &#8220;We need to prevent prompt injection&#8221; &#8220;We need to make sure users can&#8217;t trick the AI into ignoring our rules&#8221; &#8220;Our model suffers from hallucination issues&#8221; &#8220;Sometimes the AI makes things up, so we need to check its answers&#8221;</p><p>This doesn&#8217;t mean dumbing things down &#8211; it means being precise about what you&#8217;re actually doing. When you say &#8220;we&#8217;re building an agent,&#8221; what specific capability are you adding? Is it function calling? Tool use? Or just a better prompt? Being specific helps everyone understand what&#8217;s actually happening.</p><p>There&#8217;s nuance here. Technical terminology exists for a reason &#8211; it provides precision when talking with other technical stakeholders. The key is adapting your language to your audience.</p><p>The challenge many teams raise at this point is: &#8220;This all sounds great, but what if we don&#8217;t have any data yet? How can we look at examples or iterate on prompts when we&#8217;re just starting out?&#8221; That&#8217;s what we&#8217;ll talk about next.</p><h2>4. Bootstrapping Your AI With Synthetic Data Is Effective (Even With Zero Users)</h2><p>One of the most common roadblocks I hear from teams is: &#8220;We can&#8217;t do proper evaluation because we don&#8217;t have enough real user data yet.&#8221; This creates a chicken-and-egg problem &#8211; you need data to improve your AI, but you need a decent AI to get users who generate that data.</p><p>Fortunately, there&#8217;s a solution that works surprisingly well: synthetic data. LLMs can generate realistic test cases that cover the range of scenarios your AI will encounter.</p><p>As I wrote in my <a href="https://hamel.dev/blog/posts/llm-judge/#generating-data">LLM-as-a-Judge blog post</a>, synthetic data can be remarkably effective for evaluation. <a href="https://www.linkedin.com/in/bryan-bischof/">Bryan Bischof</a>, the former Head of AI at Hex, put it perfectly:</p><blockquote><p>&#8220;LLMs are surprisingly good at generating excellent - and diverse - examples of user prompts. This can be relevant for powering application features, and sneakily, for building Evals. If this sounds a bit like the Large Language Snake is eating its tail, I was just as surprised as you! All I can say is: it works, ship it.&#8221;</p></blockquote><h3>A Framework for Generating Realistic Test Data</h3><p>The key to effective synthetic data is choosing the right dimensions to test. While these dimensions will vary based on your specific needs, I find it helpful to think about three broad categories:</p><ol><li><p><strong>Features</strong>: What capabilities does your AI need to support?</p></li><li><p><strong>Scenarios</strong>: What situations will it encounter?</p></li><li><p><strong>User Personas</strong>: Who will be using it and how?</p></li></ol><p>These aren&#8217;t the only dimensions you might care about &#8211; you might also want to test different tones of voice, levels of technical sophistication, or even different locales and languages. The important thing is identifying dimensions that matter for your specific use case.</p><p>For a real estate CRM AI assistant I worked on with <a href="https://www.rechat.com/">Rechat</a>, we defined these dimensions like this:</p><pre><code>features = [
    "property search",      # Finding listings matching criteria
    "market analysis",      # Analyzing trends and pricing
    "scheduling",          # Setting up property viewings
    "follow-up"           # Post-viewing communication
]

scenarios = [
    "exact match",         # One perfect listing match
    "multiple matches",    # Need to help user narrow down
    "no matches",         # Need to suggest alternatives
    "invalid criteria"     # Help user correct search terms
]

personas = [
    "first_time_buyer",    # Needs more guidance and explanation
    "investor",           # Focused on numbers and ROI
    "luxury_client",      # Expects white-glove service
    "relocating_family"   # Has specific neighborhood/school needs
]</code></pre><p>But having these dimensions defined is only half the battle. The real challenge is ensuring your synthetic data actually triggers the scenarios you want to test. This requires two things:</p><ol><li><p>A test database with enough variety to support your scenarios</p></li><li><p>A way to verify that generated queries actually trigger intended scenarios</p></li></ol><p>For Rechat, we maintained a test database of listings that we knew would trigger different edge cases. Some teams prefer to use an anonymized copy of production data, but either way, you need to ensure your test data has enough variety to exercise the scenarios you care about.</p><p>Here&#8217;s an example of how we might use these dimensions with real data to generate test cases for the property search feature (this is just pseudo-code, and very illustrative):</p><pre><code>def generate_search_query(scenario, persona, listing_db):
    """Generate a realistic user query about listings"""
    # Pull real listing data to ground the generation
    sample_listings = listing_db.get_sample_listings(
        price_range=persona.price_range,
        location=persona.preferred_areas
    )
    
    # Verify we have listings that will trigger our scenario
    if scenario == "multiple_matches" and len(sample_listings) &lt; 2:
        raise ValueError("Need multiple listings for this scenario")
    if scenario == "no_matches" and len(sample_listings) &gt; 0:
        raise ValueError("Found matches when testing no-match scenario")
    
    prompt = f"""
    You are an expert real estate agent who is searching for listings. You are given a customer type and a scenario.
    
    Your job is to generate a natural language query you would use to search these listings.
    
    Context:
    - Customer type: {persona.description}
    - Scenario: {scenario}
    
    Use these actual listings as reference:
    {format_listings(sample_listings)}
    
    The query should reflect the customer type and the scenario.

    Example query: Find homes in the 75019 zip code, 3 bedrooms, 2 bathrooms, price range $750k - $1M for an investor.
    """
    return generate_with_llm(prompt)</code></pre><p>This produced realistic queries like:</p><p>Feature Scenario Persona Generated Query property search multiple matches first_time_buyer &#8220;Looking for 3-bedroom homes under $500k in the Riverside area. Would love something close to parks since we have young kids.&#8221; market analysis no matches investor &#8220;Need comps for 123 Oak St.&nbsp;Specifically interested in rental yield comparison with similar properties in a 2-mile radius.&#8221;</p><p>The key to useful synthetic data is grounding it in real system constraints. For the real-estate AI assistant, this means:</p><ol><li><p>Using real listing IDs and addresses from their database</p></li><li><p>Incorporating actual agent schedules and availability windows</p></li><li><p>Respecting business rules like showing restrictions and notice periods</p></li><li><p>Including market-specific details like HOA requirements or local regulations</p></li></ol><p>We then feed these test cases through Lucy and log the interactions. This gives us a rich dataset to analyze, showing exactly how the AI handles different situations with real system constraints. This approach helped us fix issues before they affected real users.</p><p>Sometimes you don&#8217;t have access to a production database, especially for new products. In these cases, use LLMs to generate both test queries and the underlying test data. For a real estate AI assistant, this might mean creating synthetic property listings with realistic attributes &#8211; prices that match market ranges, valid addresses with real street names, and amenities appropriate for each property type. The key is grounding synthetic data in real-world constraints to make it useful for testing. The specifics of generating robust synthetic databases are beyond the scope of this post.</p><h3>Guidelines for Using Synthetic Data</h3><p>When generating synthetic data, follow these key principles to ensure it&#8217;s effective:</p><ol><li><p><strong>Diversify your dataset</strong>: Create examples that cover a wide range of features, scenarios, and personas. As I wrote in my <a href="https://hamel.dev/blog/posts/llm-judge/">LLM-as-a-Judge post</a>, this diversity helps you identify edge cases and failure modes you might not anticipate otherwise.</p></li><li><p><strong>Generate user inputs, not outputs</strong>: Use LLMs to generate realistic user queries or inputs, not the expected AI responses. This prevents your synthetic data from inheriting the biases or limitations of the generating model.</p></li><li><p><strong>Incorporate real system constraints</strong>: Ground your synthetic data in actual system limitations and data. For example, when testing a scheduling feature, use real availability windows and booking rules.</p></li><li><p><strong>Verify scenario coverage</strong>: Ensure your generated data actually triggers the scenarios you want to test. A query intended to test &#8220;no matches found&#8221; should actually return zero results when run against your system.</p></li><li><p><strong>Start simple, then add complexity</strong>: Begin with straightforward test cases before adding nuance. This helps isolate issues and establish a baseline before tackling edge cases.</p></li></ol><p>This approach isn&#8217;t just theoretical &#8211; it&#8217;s been proven in production across dozens of companies. What often starts as a stopgap measure becomes a permanent part of the evaluation infrastructure, even after real user data becomes available.</p><p>Let&#8217;s look at how to maintain trust in your evaluation system as you scale&#8230;</p><h2>5. Maintaining Trust In Evals Is Critical</h2><p>This is a pattern I&#8217;ve seen repeatedly: teams build evaluation systems, then gradually lose faith in them. Sometimes it&#8217;s because the metrics don&#8217;t align with what they observe in production. Other times, it&#8217;s because the evaluations become too complex to interpret. Either way, the result is the same &#8211; the team reverts to making decisions based on gut feeling and anecdotal feedback, undermining the entire purpose of having evaluations.</p><p>Maintaining trust in your evaluation system is just as important as building it in the first place. Here&#8217;s how the most successful teams approach this challenge:</p><h3>Understanding Criteria Drift</h3><p>One of the most insidious problems in AI evaluation is &#8220;criteria drift&#8221; &#8211; a phenomenon where evaluation criteria evolve as you observe more model outputs. In their paper <a href="https://arxiv.org/abs/2404.12272">&#8220;Who Validates the Validators?&#8221;</a>, Shankar et al.&nbsp;describe this phenomenon:</p><blockquote><p>&#8220;To grade outputs, people need to externalize and define their evaluation criteria; however, the process of grading outputs helps them to define that very criteria.&#8221;</p></blockquote><p>This creates a paradox: you can&#8217;t fully define your evaluation criteria until you&#8217;ve seen a wide range of outputs, but you need criteria to evaluate those outputs in the first place. In other words, <strong>it is impossible to completely determine evaluation criteria prior to human judging of LLM outputs</strong>.</p><p>I&#8217;ve observed this firsthand when working with Phillip Carter at Honeycomb on their <a href="https://www.honeycomb.io/blog/introducing-query-assistant">Query Assistant</a> feature. As we evaluated the AI&#8217;s ability to generate database queries, Phillip noticed something interesting:</p><blockquote><p>&#8220;Seeing how the LLM breaks down its reasoning made me realize I wasn&#8217;t being consistent about how I judged certain edge cases.&#8221;</p></blockquote><p>The process of reviewing AI outputs helped him articulate his own evaluation standards more clearly. This isn&#8217;t a sign of poor planning &#8211; it&#8217;s an inherent characteristic of working with AI systems that produce diverse and sometimes unexpected outputs.</p><p>The teams that maintain trust in their evaluation systems embrace this reality rather than fighting it. They treat evaluation criteria as living documents that evolve alongside their understanding of the problem space. They also recognize that different stakeholders might have different (sometimes contradictory) criteria, and they work to reconcile these perspectives rather than imposing a single standard.</p><h3>Creating Trustworthy Evaluation Systems</h3><p>So how do you build evaluation systems that remain trustworthy despite criteria drift? Here are the approaches I&#8217;ve found most effective:</p><h4>1. Favor Binary Decisions Over Arbitrary Scales</h4><p>As I wrote in my <a href="https://hamel.dev/blog/posts/llm-judge/#why-are-simple-passfail-metrics-important">LLM-as-a-Judge post</a>, binary decisions provide clarity that more complex scales often obscure. When faced with a 1-5 scale, evaluators frequently struggle with the difference between a 3 and a 4, introducing inconsistency and subjectivity. What exactly distinguishes &#8220;somewhat helpful&#8221; from &#8220;helpful&#8221;? These boundary cases consume disproportionate mental energy and create noise in your evaluation data. And even when businesses use a 1-5 scale, they inevitably ask where to draw the line for &#8220;good enough&#8221; or to trigger intervention, forcing a binary decision anyway.</p><p>In contrast, a binary pass/fail forces evaluators to make a clear judgment: did this output achieve its purpose or not? This clarity extends to measuring progress &#8211; a 10% increase in passing outputs is immediately meaningful, while a 0.5-point improvement on a 5-point scale requires interpretation.</p><p>I&#8217;ve found that teams who resist binary evaluation often do so because they want to capture nuance. But nuance isn&#8217;t lost &#8211; it&#8217;s just moved to the qualitative critique that accompanies the judgment. The critique provides rich context about why something passed or failed, and what specific aspects could be improved, while the binary decision creates actionable clarity about whether improvement is needed at all.</p><h4>2. Enhance Binary Judgments With Detailed Critiques</h4><p>While binary decisions provide clarity, they work best when paired with detailed critiques that capture the nuance of why something passed or failed. This combination gives you the best of both worlds: clear, actionable metrics and rich contextual understanding.</p><p>For example, when evaluating a response that correctly answers a user&#8217;s question but contains unnecessary information, a good critique might read:</p><blockquote><p>&#8220;The AI successfully provided the market analysis requested (PASS), but included excessive detail about neighborhood demographics that wasn&#8217;t relevant to the investment question. This makes the response longer than necessary and potentially distracting.&#8221;</p></blockquote><p>These critiques serve multiple functions beyond just explanation. They force domain experts to externalize implicit knowledge &#8211; I&#8217;ve seen legal experts move from vague feelings that something &#8220;doesn&#8217;t sound right&#8221; to articulating specific issues with citation formats or reasoning patterns that can be systematically addressed.</p><p>When included as few-shot examples in judge prompts, these critiques improve the LLM&#8217;s ability to reason about complex edge cases. I&#8217;ve found this approach often yields 15-20% higher agreement rates between human and LLM evaluations compared to prompts without example critiques. The critiques also provide excellent raw material for generating high-quality synthetic data, creating a flywheel for improvement.</p><h4>3. Measure Alignment Between Automated Evals and Human Judgment</h4><p>If you&#8217;re using LLMs to evaluate outputs (which is often necessary at scale), it&#8217;s crucial to regularly check how well these automated evaluations align with human judgment.</p><p>This is particularly important given our natural tendency to over-trust AI systems. As Shankar et al.&nbsp;note in <a href="https://arxiv.org/abs/2404.12272">&#8220;Who Validates the Validators?&#8221;</a>, the lack of tools to validate evaluator quality is concerning</p><blockquote><p>Research shows people tend to over-rely and over-trust AI systems. For instance, in one high profile incident, researchers from MIT posted a pre-print on arXiv claiming that GPT-4 could ace the MIT EECS exam. Within hours, [the] work [was] debunked &#8230; citing problems arising from over-reliance on GPT-4 to grade itself.&#8221;</p></blockquote><p>This over-trust problem extends beyond self-evaluation. Research has shown that LLMs can be biased by simple factors like the ordering of options in a set, or even seemingly innocuous formatting changes in prompts. Without rigorous human validation, these biases can silently undermine your evaluation system.</p><p>When working with Honeycomb, we tracked agreement rates between our LLM-as-a-judge and Phillip&#8217;s evaluations:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fHqQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 424w, /__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 848w, /__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fHqQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 424w, /__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 848w, /__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fHqQ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52dda563-94d7-46d2-9ca0-378febdd0c99_484x228.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Agreement rates between LLM evaluator and human expert. More details <a href="https://hamel.dev/blog/posts/evals/#automated-evaluation-w-llms">here</a>.</figcaption></figure></div><p>It took three iterations to achieve &gt;90% agreement, but this investment paid off in a system the team could trust. Without this validation step, automated evaluations often drift from human expectations over time, especially as the distribution of inputs changes. You can <a href="https://hamel.dev/blog/posts/evals/#automated-evaluation-w-llms">read more about this here</a>.</p><p>Tools like <a href="https://eugeneyan.com/writing/aligneval/">Eugene Yan&#8217;s AlignEval</a> demonstrate this alignment process beautifully. It provides a simple interface where you upload data, label examples with a binary &#8220;good&#8221; or &#8220;bad,&#8221; and then evaluate LLM-based judges against those human judgments. What makes it effective is how it streamlines the workflow &#8211; you can quickly see where automated evaluations diverge from your preferences, refine your criteria based on these insights, and measure improvement over time. This approach reinforces that alignment isn&#8217;t a one-time setup but an ongoing conversation between human judgment and automated evaluation.</p><h3>Scaling Without Losing Trust</h3><p>As your AI system grows, you&#8217;ll inevitably face pressure to reduce the human effort involved in evaluation. This is where many teams go wrong &#8211; they automate too much, too quickly, and lose the human connection that keeps their evaluations grounded.</p><p>The most successful teams take a more measured approach:</p><ol><li><p><strong>Start with high human involvement</strong>: In the early stages, have domain experts evaluate a significant percentage of outputs.</p></li><li><p><strong>Study alignment patterns</strong>: Rather than automating evaluation, focus on understanding where automated evaluations align with human judgment and where they diverge. This helps you identify which types of cases need more careful human attention.</p></li><li><p><strong>Use strategic sampling</strong>: Rather than evaluating every output, use statistical techniques to sample outputs that provide the most information, particularly focusing on areas where alignment is weakest.</p></li><li><p><strong>Maintain regular calibration</strong>: Even as you scale, continue to compare automated evaluations against human judgment regularly, using these comparisons to refine your understanding of when to trust automated evaluations.</p></li></ol><p>Scaling evaluation isn&#8217;t just about reducing human effort &#8211; it&#8217;s about directing that effort where it adds the most value. By focusing human attention on the most challenging or informative cases, you can maintain quality even as your system grows.</p><p>Now that we&#8217;ve covered how to maintain trust in your evaluations, let&#8217;s talk about a fundamental shift in how you should approach AI development roadmaps&#8230;</p><h2>6. Your AI Roadmap Should Count Experiments, Not Features</h2><p>If you&#8217;ve worked in software development, you&#8217;re familiar with traditional roadmaps: a list of features with target delivery dates. Teams commit to shipping specific functionality by specific deadlines, and success is measured by how closely they hit those targets.</p><p>This approach fails spectacularly with AI.</p><p>I&#8217;ve watched teams commit to roadmaps like &#8220;Launch sentiment analysis by Q2&#8221; or &#8220;Deploy agent-based customer support by end of year,&#8221; only to discover that the technology simply isn&#8217;t ready to meet their quality bar. They either ship something subpar to hit the deadline or miss the deadline entirely. Either way, trust erodes.</p><p>The fundamental problem is that traditional roadmaps assume we know what&#8217;s possible. With conventional software, that&#8217;s often true &#8211; given enough time and resources, you can build most features reliably. With AI, especially at the cutting edge, you&#8217;re constantly testing the boundaries of what&#8217;s feasible.</p><h3>Experiments vs.&nbsp;Features</h3><p><a href="https://www.linkedin.com/in/bryan-bischof/">Bryan Bischof</a>, Former Head of AI at Hex, introduced me to what he calls a &#8220;capability funnel&#8221; approach to AI roadmaps. This strategy reframes how we think about AI development progress.</p><p>Instead of defining success as shipping a feature, the capability funnel breaks down AI performance into progressive levels of utility. At the top of the funnel is the most basic functionality &#8211; can the system respond at all? At the bottom is fully solving the user&#8217;s job to be done. Between these points are various stages of increasing usefulness.</p><p>For example, in a query assistant, the capability funnel might look like: 1. Can generate syntactically valid queries (basic functionality) 2. Can generate queries that execute without errors 3. Can generate queries that return relevant results 4. Can generate queries that match user intent 5. Can generate optimal queries that solve the user&#8217;s problem (complete solution)</p><p>This approach acknowledges that AI progress isn&#8217;t binary &#8211; it&#8217;s about gradually improving capabilities across multiple dimensions. It also provides a framework for measuring progress even when you haven&#8217;t reached the final goal.</p><p>The most successful teams I&#8217;ve worked with structure their roadmaps around experiments rather than features. Instead of committing to specific outcomes, they commit to a cadence of experimentation, learning, and iteration.</p><p><a href="https://eugeneyan.com/">Eugene Yan</a>, an applied scientist at Amazon, shared how he approaches ML project planning with leadership - a process that, while originally developed for traditional machine learning, applies equally well to modern LLM development:</p><blockquote><p>&#8220;Here&#8217;s a common timeline. First, I take two weeks to do a data feasibility analysis, i.e&#8221;do I have the right data?&#8221; [&#8230;] Then I take an additional month to do a technical feasibility analysis, i.e &#8220;can AI solve this?&#8221; After that, if it still works I&#8217;ll spend six weeks building a prototype we can A/B test.&#8221;</p></blockquote><p>While LLMs might not require the same kind of feature engineering or model training as traditional ML, the underlying principle remains the same: time-box your exploration, establish clear decision points, and focus on proving feasibility before committing to full implementation. This approach gives leadership confidence that resources won&#8217;t be wasted on open-ended exploration, while giving the team the freedom to learn and adapt as they go.</p><h3>The Foundation: Evaluation Infrastructure</h3><p>The key to making an experiment-based roadmap work is having robust evaluation infrastructure. Without it, you&#8217;re just guessing whether your experiments are working. With it, you can rapidly iterate, test hypotheses, and build on successes.</p><p>I saw this firsthand during the early development of GitHub Copilot. What most people don&#8217;t realize is that the team invested heavily in building sophisticated offline evaluation infrastructure. They created systems that could test code completions against a very large corpus of repositories on GitHub, leveraging unit tests that already existed in high-quality codebases as an automated way to verify completion correctness. This was a massive engineering undertaking &#8211; they had to build systems that could clone repositories at scale, set up their environments, run their test suites, and analyze the results, all while handling the incredible diversity of programming languages, frameworks, and testing approaches.</p><p>This wasn&#8217;t wasted time&#8212;it was the foundation that accelerated everything. With solid evaluation in place, the team ran thousands of experiments, quickly identified what worked, and could say with confidence &#8220;this change improved quality by X%&#8221; instead of relying on gut feelings. While the upfront investment in evaluation feels slow, it prevents endless debates about whether changes help or hurt, and dramatically speeds up innovation later.</p><h3>Communicating This to Stakeholders</h3><p>The challenge, of course, is that executives often want certainty. They want to know when features will ship and what they&#8217;ll do. How do you bridge this gap?</p><p>The key is to shift the conversation from outputs to outcomes. Instead of promising specific features by specific dates, commit to a process that will maximize the chances of achieving the desired business outcomes.</p><p>Eugene shared how he handles these conversations:</p><blockquote><p>&#8220;I try to reassure leadership with timeboxes. At the end of three months, if it works out, then we move it to production. At any step of the way, if it doesn&#8217;t work out, we pivot.&#8221;</p></blockquote><p>This approach gives stakeholders clear decision points while acknowledging the inherent uncertainty in AI development. It also helps manage expectations about timelines &#8211; instead of promising a feature in six months, you&#8217;re promising a clear understanding of whether that feature is feasible in three months.</p><p>Bryan&#8217;s capability funnel approach provides another powerful communication tool. It allows teams to show concrete progress through the funnel stages, even when the final solution isn&#8217;t ready. It also helps executives understand where problems are occurring and make informed decisions about where to invest resources.</p><h3>Build a Culture of Experimentation Through Failure Sharing</h3><p>Perhaps the most counterintuitive aspect of this approach is the emphasis on learning from failures. In traditional software development, failures are often hidden or downplayed. In AI development, they&#8217;re the primary source of learning.</p><p>Eugene operationalizes this at his organization through what he calls a &#8220;fifteen-five&#8221; &#8211; a weekly update that takes fifteen minutes to write and five minutes to read:</p><blockquote><p>&#8220;In my fifteen-fives, I document my failures and my successes. Within our team, we also have weekly&#8221;no-prep sharing sessions&#8221; where we discuss what we&#8217;ve been working on and what we&#8217;ve learned. When I do this, I go out of my way to share failures.&#8221;</p></blockquote><p>This practice normalizes failure as part of the learning process. It shows that even experienced practitioners encounter dead ends, and it accelerates team learning by sharing those experiences openly. And by celebrating the process of experimentation rather than just the outcomes, teams create an environment where people feel safe taking risks and learning from failures.</p><h3>A Better Way Forward</h3><p>So what does an experiment-based roadmap look like in practice? Here&#8217;s a simplified example from a content moderation project Eugene worked on:</p><blockquote><p>&#8220;I was asked to do content moderation. I said, &#8216;It&#8217;s uncertain whether we&#8217;ll meet that goal. It&#8217;s uncertain even if that goal is feasible with our data, or what machine learning techniques would work. But here&#8217;s my experimentation roadmap. Here are the techniques I&#8217;m gonna try, and I&#8217;m gonna update you at a two-week cadence.&#8217;&#8221;</p></blockquote><p>The roadmap didn&#8217;t promise specific features or capabilities. Instead, it committed to a systematic exploration of possible approaches, with regular check-ins to assess progress and pivot if necessary.</p><p>The results were telling:</p><blockquote><p>&#8220;For the first two to three months, nothing worked. [&#8230;] And then [a breakthrough] came out. [&#8230;] Within a month, that problem was solved. So you can see that in the first quarter or even four months, it was going nowhere. [&#8230;] But then you can also see that all of a sudden, some new technology comes along, some new paradigm, some new reframing comes along that just [solves] 80% of [the problem].&#8221;</p></blockquote><p>This pattern &#8211; long periods of apparent failure followed by breakthroughs &#8211; is common in AI development. Traditional feature-based roadmaps would have killed the project after months of &#8220;failure,&#8221; missing the eventual breakthrough.</p><p>By focusing on experiments rather than features, teams create space for these breakthroughs to emerge. They also build the infrastructure and processes that make breakthroughs more likely &#8211; data pipelines, evaluation frameworks, and rapid iteration cycles.</p><p>The most successful teams I&#8217;ve worked with start by building evaluation infrastructure before committing to specific features. They create tools that make iteration faster and focus on processes that support rapid experimentation. This approach might seem slower at first, but it dramatically accelerates development in the long run by enabling teams to learn and adapt quickly.</p><p>The key metric for AI roadmaps isn&#8217;t features shipped &#8211; it&#8217;s experiments run. The teams that win are those that can run more experiments, learn faster, and iterate more quickly than their competitors. And the foundation for this rapid experimentation is always the same: robust, trusted evaluation infrastructure that gives everyone confidence in the results.</p><p>By reframing your roadmap around experiments rather than features, you create the conditions for similar breakthroughs in your own organization.</p><h2>Conclusion</h2><p>Throughout this post, I&#8217;ve shared patterns I&#8217;ve observed across dozens of AI implementations. The most successful teams aren&#8217;t the ones with the most sophisticated tools or the most advanced models &#8211; they&#8217;re the ones that master the fundamentals of measurement, iteration, and learning.</p><p>The core principles are surprisingly simple:</p><ol><li><p><strong>Look at your data.</strong> Nothing replaces the insight gained from examining real examples. Error analysis consistently reveals the highest-ROI improvements.</p></li><li><p><strong>Build simple tools that remove friction.</strong> Custom data viewers that make it easy to examine AI outputs yield more insights than complex dashboards with generic metrics.</p></li><li><p><strong>Empower domain experts.</strong> The people who understand your domain best are often the ones who can most effectively improve your AI, regardless of their technical background.</p></li><li><p><strong>Use synthetic data strategically.</strong> You don&#8217;t need real users to start testing and improving your AI. Thoughtfully generated synthetic data can bootstrap your evaluation process.</p></li><li><p><strong>Maintain trust in your evaluations.</strong> Binary judgments with detailed critiques create clarity while preserving nuance. Regular alignment checks ensure automated evaluations remain trustworthy.</p></li><li><p><strong>Structure roadmaps around experiments, not features.</strong> Commit to a cadence of experimentation and learning rather than specific outcomes by specific dates.</p></li></ol><p>These principles apply regardless of your domain, team size, or technical stack. They&#8217;ve worked for companies ranging from early-stage startups to tech giants, across use cases from customer support to code generation.</p><h3>Resources for Going Deeper</h3><p>If you&#8217;d like to explore these topics further, here are some resources that might help:</p><ul><li><p><a href="https://ai.hamel.dev/">My blog</a> for more content on AI evaluation and improvement. My other posts dive into more technical detail on topics such as constructing effective LLM judges, implementing evaluation systems, and other aspects of AI development<sup>1</sup>. Also check out the blogs of <a href="https://www.sh-reya.com/">Shreya Shankar</a> and <a href="https://eugeneyan.com/">Eugene Yan</a> who are also great sources of information on these topics.</p></li><li><p>A course I&#8217;m teaching: <strong><a href="https://bit.ly/evals-ai">Rapidly Improve AI Products With Evals</a></strong>, with Shreya Shankar. The course provides hands-on experience with techniques such as error analysis, synthetic data generation, and building trustworthy evaluation systems. It includes practical exercises and personalized instruction through office hours.</p></li><li><p>If you&#8217;re looking for hands-on guidance specific to your organization&#8217;s needs, you can learn more about working with me at <a href="https://parlance-labs.com/">Parlance Labs</a>.</p></li></ul><h2>Footnotes</h2><ol><li><p>I write more broadly about machine learning, AI, and software development. Some posts that expand on these topics include <a href="https://hamel.dev/blog/posts/evals/">Your AI Product Needs Evals</a>, <a href="https://hamel.dev/blog/posts/llm-judge/">Creating a LLM-as-a-Judge That Drives Business Results</a>, and <a href="https://applied-llms.org/">What We&#8217;ve Learned From A Year of Building with LLMs</a>. You can see all my posts at <a href="https://hamel.dev/">hamel.dev</a>.&#8617;&#65038;</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Building an Audience Through Technical Writing: Strategies and Mistakes]]></title><description><![CDATA[What I&#8217;ve seen work and what doesn&#8217;t.]]></description><link>https://hamelhusain.substack.com/p/audience</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/audience</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Sat, 30 Nov 2024 08:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Pit1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Pit1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Pit1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2308695,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hamelhusain.substack.com/i/170578975?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Pit1!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb067f5bc-cc53-44c8-aaca-96cd4e021637_1600x900.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>People often find me through my writing on AI and tech. This creates an interesting pattern. Nearly every week, vendors reach out asking me to write about their products. While I appreciate their interest and love learning about new tools, I reserve my writing for topics that I have personal experience with.</p><p>One conversation last week really stuck with me. A founder confided, &#8220;We can write the best content in the world, but we don&#8217;t have any distribution.&#8221; This hit home because I used to think the same way.</p><p>Let me share what works for reaching developers. Companies and individuals alike often skip the basics when trying to grow their audience. These are proven approaches I&#8217;ve seen succeed, both in my work and in others&#8217; efforts to grow their audience in the AI space.</p><h2>1. Build on Great Work</h2><p>Here&#8217;s something surprising: few people take the time to thoughtfully engage with others&#8217; work in our field. But when you do, amazing things happen naturally.</p><p>For example, here are some recent posts I&#8217;ve enjoyed that present opportunities to engage with others:</p><ul><li><p>Shreya Shankar&#8217;s <a href="https://data-people-group.github.io/blogs/2024/09/24/docetl/">DocETL</a></p></li><li><p>Eugene Yan&#8217;s work on <a href="https://eugeneyan.com/writing/aligneval/">AlignEval</a></p></li><li><p>Ben Claive&#8217;s work on <a href="https://www.answer.ai/posts/2024-09-16-rerankers.html">rerankers</a></p></li><li><p>Jeremy Howard&#8217;s work on <a href="https://www.answer.ai/posts/2024-09-03-llmstxt.html">llms.txt</a></p></li></ul><p>In the above examples, you could share how their ideas connect with what you&#8217;ve built. You could add additional case studies and real-world insights. If you deeply engage with someone&#8217;s work and add your insights, they often share your content with their audience. Not because you asked, but because you&#8217;ve added something meaningful to their work. Swyx has written a <a href="https://www.swyx.io/puwtpd">great post</a> on how to do this effectively.</p><p>The key is authenticity. Don&#8217;t do this just for marketing&#8212;do it because you&#8217;re genuinely interested in learning from others and building on their ideas. It&#8217;s not hard to find things to be excited about. I&#8217;m amazed by how few people take this approach. It&#8217;s both effective and fun.</p><h2>2. Show Up Consistently</h2><p>I see too many folks blogging or posting once every few months and wondering why they&#8217;re not getting traction. Want to know what actually works? Look at <a href="https://x.com/jxnlco">Jason Liu</a>. He grew his following from 500 to 30,000 followers by posting ~ 30 times every day for a year.</p><p>You don&#8217;t have to post that often (I certainly don&#8217;t!), but consistency matters more than perfection. Finally, don&#8217;t just post into the void. Engage with others. When someone comments on your post, reply thoughtfully. When you see conversations where you can add value, provide helpful information.</p><p>Finally, don&#8217;t be discouraged if you don&#8217;t see results immediately. Here&#8217;s some advice from my friend (and prolific writer), <a href="https://eugeneyan.com/">Eugene Yan</a>:</p><blockquote><p>In the beginning, when most people start writing, the output&#8217;s gonna suck. Harsh, but true&#8212;my first 100 posts or so were crap. But with practice, people can get better. But they have to be deliberate in wanting to practice and get better with each piece, and not just write for the sake of publishing something and tweeting about it. The Sam Parr course (see below) is a great example of deliberate practice on copywriting.</p></blockquote><h2>3. Get Better at Copywriting</h2><p>This changed everything for me. I took <a href="https://copythat.com/">Sam Parr&#8217;s copywriting course</a> just 30 minutes a day for a week. Now I keep my favorite writing samples in a Claude project and reference them when I&#8217;m writing something important. Small improvements in how you communicate can make a huge difference in how your content lands.</p><p>One thing Sam teaches is that big words don&#8217;t make you sound smart. Clear writing that avoids jargon is more effective. That&#8217;s why Sam teaches aiming for a 6th-grade reading level. This matters even more with AI, as AI loves to generate flowery language and long sentences. The <a href="https://hemingwayapp.com/">Hemingway App</a> can be helpful in helping you simplify your writing.<sup>1</sup></p><h2>4. Build a Voice-to-Content Pipeline</h2><p>The struggle most people have with creating content is that it takes too much time. But it doesn&#8217;t have to if you build the right systems, especially with AI.</p><p>Getting this system right takes some upfront work, but the payoff is enormous. Start by installing a good voice-to-text app on your phone. I use either <a href="https://superwhisper.com/">Superwhisper</a> or <a href="https://voicepal.me/">VoicePal</a>. VoicePal is great for prompting you to elaborate with follow-up questions. These tools let me capture ideas at their best. That&#8217;s usually when I&#8217;m walking outside or away from my computer. At my computer, I use <a href="https://www.flowvoice.ai/">Flow</a>.</p><p>The key is to carefully craft your first few pieces of content. These become examples for your prompts that teach AI your style and tone. Once you have high-quality examples, you can organize these (transcript, content) pairs and feed them to language models. The in-context learning creates remarkably aligned output that matches your writing style while maintaining the authenticity of your original thoughts.</p><p>For example, I use this pipeline at Answer AI. We have started interviewing each other and using the recordings as grounding for blog posts. Our recent <a href="https://www.answer.ai/posts/2024-11-07-solveit.html">post about SolveIt</a> shows this in action. The raw conversation is the foundation. Our workflow turns it into polished content.</p><p>I&#8217;ve also integrated this workflow into my meetings. Using <a href="https://circleback.ai/?via=hamel">CircleBack</a>, my favorite AI note-taking app, I can automatically capture and process meeting discussions. You can set up workflows to send your meeting notes and transcripts to AI for processing. This turns conversations into content opportunities.</p><p>The real power comes from having all these pieces working together. Voice capture, AI, and automation makes content creation fun and manageable.</p><h2>5. Leverage Your Unique Perspective</h2><p>Through my consulting work, I notice patterns that others miss. My most popular posts address common problems my clients had. When everyone&#8217;s confused about a topic, especially in AI where there&#8217;s lots of hype, clear explanations are gold. This is the motivation for some of my blog posts like:</p><ul><li><p><a href="https://hamel.dev/blog/posts/prompt/">Fuck You, Show Me The Prompt</a></p></li><li><p><a href="https://hamel.dev/blog/posts/evals/">Your AI Product Needs Evals</a></p></li><li><p><a href="https://hamel.dev/blog/posts/llm-judge/">Creating a LLM-as-a-Judge That Drives Business Results</a></p></li></ul><p>You probably see patterns too. Maybe it&#8217;s common questions from customers, or problems you&#8217;ve solved repeatedly. Maybe you work with a unique set of technologies or interesting use cases. Share these insights! Your unique perspective is more valuable than you think.</p><h2>6. Use High Quality Social Cards, Threads, and Scheduling</h2><p>This is probably the least important part of the process, but it&#8217;s still important. Thumbnails and social cards are vital for visibility on social media. Here are the tools I use:</p><ul><li><p><a href="https://socialsharepreview.com/">socialsharepreview.com</a> to check how your content looks on different platforms. For X, I sometimes use the <a href="https://cards-dev.twitter.com/validator">Twitter Card Validator</a>.</p></li><li><p><a href="https://chatgpt.com/">chatGPT</a> to create cover images for my posts. Then, paste them into Canva to size and edit them. Some of my friends use <a href="https://ideogram.ai/">ideogram</a>, which generates images with text accurately.</p></li><li><p><a href="https://www.canva.com/">Canva</a> for the last mile of creating social cards. They have easy-to-use buttons to ensure you get the dimensions right. They also have inpainting, background removal, and more.</p></li><li><p>If using X, social cards can be a bit fiddly. As of this writing, they do not show your post title, just the image if using the large-image size. To mitigate this,I use Canva to write the post&#8217;s title in the image <a href="https://hamel.dev/blog/posts/audience/content_2.png">like this</a>.</p></li><li><p>Social media can be distracting, so I like to schedule my posts in advance. I use <a href="https://typefully.com/">typefully</a> for this purpose. Some of my friends use <a href="https://hypefury.com/">hypefury</a>.</p></li></ul><p>Finally, when posting on X, threads can be a great way to raise the visibility of your content. A simple approach is to take screenshots or copy-paste snippets of your content. Then, walk through them in a thread, as you would want a reader to. Jeremy Howard does a great job at this: <a href="https://x.com/jeremyphoward/status/1818036923304456492">example 1</a>, <a href="https://x.com/jeremyphoward/status/1831089138571133290">example 2</a>.</p><h2>The Content Flywheel: Putting It All Together</h2><p>Once you have these systems in place, something magical happens: content creates more content. Your blog posts spawn social media updates. Your conversations turn into newsletters. Your client solutions become case studies. Each piece of work feeds the next, creating a natural flywheel.</p><p>Don&#8217;t try to sell too hard. Instead, share real insights and helpful information. Focus on adding value and educating your audience. When you do this well, people will want to follow your work.</p><p>This journey is different for everyone. These are just the patterns I&#8217;ve seen work in my consulting practice and my own growth. Try what feels right. Adjust what doesn&#8217;t.</p><p>P.S. If you&#8217;d like to follow my writing journey, you can <a href="https://ai.hamel.dev/">stay connected here</a>.</p><h2>Further Reading</h2><ul><li><p><a href="https://simonwillison.net/tags/writing/">Simon Willison&#8217;s Posts on Writing</a></p></li><li><p><a href="https://eugeneyan.com/tag/writing/">Eugene&#8217;s Posts on Writing</a></p></li><li><p><a href="https://medium.com/@racheltho/why-you-yes-you-should-blog-7d2544ac1045">Why you, (yes, you) should blog</a></p></li></ul><h2>Footnotes</h2><ol><li><p>Don&#8217;t abuse these tools or use them blindly. There&#8217;s <a href="https://x.com/swyx/status/1863352038597558712">plenty of situations where you should not be writing at a 6th-grade reading level</a>. This includes, humor, poetry, shitposting, and more. Even formal writing shouldn&#8217;t adhere strictly to this rule. It&#8217;s advice that you should judge on a case-by-case basis. When you simplify your writing - do you like it more?&#8617;&#65038;</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Creating a LLM-as-a-Judge That Drives Business Results]]></title><description><![CDATA[Earlier this year, I wrote Your AI product needs evals. Many of you asked, &#8220;How do I get started with LLM-as-a-judge?&#8221; This guide shares what I&#8217;ve learned after helping over 30 companies set up their evaluation systems.]]></description><link>https://hamelhusain.substack.com/p/llm-judge</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/llm-judge</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Tue, 29 Oct 2024 07:00:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6b629dd6-db02-4ed6-a1fc-256a4a23a340_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Earlier this year, I wrote <a href="https://hamel.dev/blog/posts/evals/">Your AI product needs evals</a>. Many of you asked, &#8220;How do I get started with LLM-as-a-judge?&#8221; This guide shares what I&#8217;ve learned after helping over <a href="https://parlance-labs.com/">30 companies</a> set up their evaluation systems.</p><h2>The Problem: AI Teams Are Drowning in Data</h2><p>Ever spend weeks building an AI system, only to realize you have no idea if it&#8217;s actually working? You&#8217;re not alone. I&#8217;ve noticed teams repeat the same mistakes when using LLMs to evaluate AI outputs:</p><ol><li><p><strong>Too Many Metrics</strong>: Creating numerous measurements that become unmanageable.</p></li><li><p><strong>Arbitrary Scoring Systems</strong>: Using uncalibrated scales (like 1-5) across multiple dimensions, where the difference between scores is unclear and subjective. What makes something a 3 versus a 4? Nobody knows, and different evaluators often interpret these scales differently.</p></li><li><p><strong>Ignoring Domain Experts</strong>: Not involving the people who understand the subject matter deeply.</p></li><li><p><strong>Unvalidated Metrics</strong>: Using measurements that don&#8217;t truly reflect what matters to the users or the business.</p></li></ol><p>The result? Teams end up buried under mountains of metrics or data they don&#8217;t trust and can&#8217;t use. Progress grinds to a halt. Everyone gets frustrated.</p><p>For example, it&#8217;s not uncommon for me to see dashboards that look like this:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8Nab!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8Nab!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 424w, /__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 848w, /__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8Nab!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae4140ad-ee4c-460a-a6dd-20189570cb88_1600x900.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">An illustrative example of a bad eval dashboard</figcaption></figure></div><p>Tracking a bunch of scores on a 1-5 scale is often a sign of a bad eval process (I&#8217;ll discuss why later). In this post, I&#8217;ll show you how to avoid these pitfalls. The solution is to use a technique that I call <strong>&#8220;Critique Shadowing&#8221;</strong>. Here&#8217;s how to do it, step by step.</p><h2>Step 1: Find <em>The</em> Principal Domain Expert</h2><p>In most organizations there is usually one (maybe two) key individuals whose judgment is crucial for the success of your AI product. These are the people with deep domain expertise or represent your target users. Identifying and involving this <strong>Principal Domain Expert</strong> early in the process is critical.</p><p><strong>Why is finding the right domain expert so important?</strong></p><ul><li><p><strong>They Set the Standard</strong>: This person not only defines what is acceptable technically, but also helps you understand if you&#8217;re building something users actually want.</p></li><li><p><strong>Capture Unspoken Expectations</strong>: By involving them, you uncover their preferences and expectations, which they might not be able to fully articulate upfront. Through the evaluation process, you help them clarify what a &#8220;passable&#8221; AI interaction looks like.</p></li><li><p><strong>Consistency in Judgment</strong>: People in your organization may have different opinions about the AI&#8217;s performance. Focusing on the principal expert ensures that evaluations are consistent and aligned with the most critical standards.</p></li><li><p><strong>Sense of Ownership</strong>: Involving the expert gives them a stake in the AI&#8217;s development. They feel invested because they&#8217;ve had a hand in shaping it. In the end, they are more likely to approve of the AI.</p></li></ul><p><strong>Examples of Principal Domain Experts:</strong></p><ul><li><p>A <strong>psychologist</strong> for a mental health AI assistant.</p></li><li><p>A <strong>lawyer</strong> for an AI that analyzes legal documents.</p></li><li><p>A <strong>customer service director</strong> for a support chatbot.</p></li><li><p>A <strong>lead teacher or curriculum developer</strong> for an educational AI tool.</p></li></ul><p> Exceptions</p><p>In a smaller company, this might be the CEO or founder. If you are an independent developer, you should be the domain expert (but be honest with yourself about your expertise).</p><p>If you must rely on leadership, you should regularly validate their assumptions against real user feedback.</p><p>Many developers attempt to act as the domain expert themselves, or find a convenient proxy (ex: their superior). This is a recipe for disaster. People will have varying opinions about what is acceptable, and you can&#8217;t make everyone happy. What&#8217;s important is that your principal domain expert is satisfied.</p><p><strong>Remember:</strong> This doesn&#8217;t have to take a lot of the domain expert&#8217;s time. Later in this post, I&#8217;ll discuss how you can make the process efficient. Their involvement is absolutely critical to the AI&#8217;s success.</p><h3>Next Steps</h3><p>Once you&#8217;ve found your expert, we need to give them the right data to review. Let&#8217;s talk about how to do that next.</p><h2>Step 2: Create a Dataset</h2><p>With your principal domain expert on board, the next step is to build a dataset that captures problems that your AI will encounter. It&#8217;s important that the dataset is diverse and represents the types of interactions that your AI will have in production.</p><h3>Why a Diverse Dataset Matters</h3><ul><li><p><strong>Comprehensive Testing</strong>: Ensures your AI is evaluated across a wide range of situations.</p></li><li><p><strong>Realistic Interactions</strong>: Reflects actual user behavior for more relevant evaluations.</p></li><li><p><strong>Identifies Weaknesses</strong>: Helps uncover areas where the AI may struggle or produce errors.</p></li></ul><h3>Dimensions for Structuring Your Dataset</h3><p>You want to define dimensions that make sense for your use case. For example, here are ones that I often use for B2C applications:</p><ol><li><p><strong>Features</strong>: Specific functionalities of your AI product.</p></li><li><p><strong>Scenarios</strong>: Situations or problems the AI may encounter and needs to handle.</p></li><li><p><strong>Personas</strong>: Representative user profiles with distinct characteristics and needs.</p></li></ol><h3>Examples of Features, Scenarios, and Personas</h3><h4>Features</h4><p><strong>Feature</strong> <strong>Description</strong> <strong>Email Summarization</strong> Condensing lengthy emails into key points. <strong>Meeting Scheduler</strong> Automating the scheduling of meetings across time zones. <strong>Order Tracking</strong> Providing shipment status and delivery updates. <strong>Contact Search</strong> Finding and retrieving contact information from a database. <strong>Language Translation</strong> Translating text between languages. <strong>Content Recommendation</strong> Suggesting articles or products based on user interests.</p><h4>Scenarios</h4><p>Scenarios are situations the AI needs to handle, (not based on the outcome of the AI&#8217;s response).</p><p><strong>Scenario</strong> <strong>Description</strong> <strong>Multiple Matches Found</strong> User&#8217;s request yields multiple results that need narrowing down. For example: User asks &#8220;Where&#8217;s my order?&#8221; but has three active orders (#123, #124, #125). AI must help identify which specific order they&#8217;re asking about. <strong>No Matches Found</strong> User&#8217;s request yields no results, requiring alternatives or corrections. For example: User searches for order #ABC-123 which doesn&#8217;t exist. AI should explain valid order formats and suggest checking their confirmation email. <strong>Ambiguous Request</strong> User input lacks necessary specificity. For example: User says &#8220;I need to change my delivery&#8221; without specifying which order or what aspect of delivery (date, address, etc.) they want to change. <strong>Invalid Data Provided</strong> User provides incorrect data type or format. For example: User tries to track a return using a regular order number instead of a return authorization (RMA) number. <strong>System Errors</strong> Technical issues prevent normal operation. For example: While looking up an order, the inventory database is temporarily unavailable. AI needs to explain the situation and provide alternatives. <strong>Incomplete Information</strong> User omits required details. For example: User wants to initiate a return but hasn&#8217;t provided the order number or reason. AI needs to collect this information step by step. <strong>Unsupported Feature</strong> User requests functionality that doesn&#8217;t exist. For example: User asks to change payment method after order has shipped. AI must explain why this isn&#8217;t possible and suggest alternatives.</p><h4>Personas</h4><p><strong>Persona</strong> <strong>Description</strong> <strong>New User</strong> Unfamiliar with the system; requires guidance. <strong>Expert User</strong> Experienced; expects efficiency and advanced features. <strong>Non-Native Speaker</strong> May have language barriers; uses non-standard expressions. <strong>Busy Professional</strong> Values quick, concise responses; often multitasking. <strong>Technophobe</strong> Uncomfortable with technology; needs simple instructions. <strong>Elderly User</strong> May not be tech-savvy; requires patience and clear guidance.</p><h3>This taxonomy is not universal</h3><p>This taxonomy (features, scenarios, personas) is not universal. For example, it may not make sense to even have personas if users aren&#8217;t directly engaging with your AI. The idea is you should outline dimensions that make sense for your use case and generate data that covers them. You&#8217;ll likely refine these after the first round of evaluations.</p><h3>Generating Data</h3><p>To build your dataset, you can:</p><ul><li><p><strong>Use Existing Data</strong>: Sample real user interactions or behaviors from your AI system.</p></li><li><p><strong>Generate Synthetic Data</strong>: Use LLMs to create realistic user inputs covering various features, scenarios, and personas.</p></li></ul><p>Often, you&#8217;ll do a combination of both to ensure comprehensive coverage. Synthetic data is not as good as real data, but it&#8217;s a good starting point. Also, we are only using LLMs to generate the user inputs, not the LLM responses or internal system behavior.</p><p>Regardless of whether you use existing data or synthetic data, you want good coverage across the dimensions you&#8217;ve defined.</p><p><strong>Incorporating System Information</strong></p><p>When making test data, use your APIs and databases where appropriate. This will create realistic data and trigger the right scenarios. Sometimes you&#8217;ll need to write simple programs to get this information. That&#8217;s what the &#8220;Assumptions&#8221; column is referring to in the examples below.</p><h3>Example LLM Prompts for Generating User Inputs</h3><p>Here are some example prompts that illustrate how to use an LLM to generate synthetic <strong>user inputs</strong> for different combinations of features, scenarios, and personas:</p><p><strong>ID</strong> <strong>Feature</strong> <strong>Scenario</strong> <strong>Persona</strong> <strong>LLM Prompt to Generate User Input</strong> Assumptions (not directly in the prompt) 1 <strong>Order Tracking</strong> Invalid Data Provided Frustrated Customer &#8220;Generate a user input from someone who is clearly irritated and impatient, using short, terse language to demand information about their order status for order number <strong>#1234567890</strong>. Include hints of previous negative experiences.&#8221; Order number <strong>#1234567890</strong> does <strong>not</strong> exist in the system. 2 <strong>Contact Search</strong> Multiple Matches Found New User &#8220;Create a user input from someone who seems unfamiliar with the system, using hesitant language and asking for help to find contact information for a person named &#8216;Alex&#8217;. The user should appear unsure about what information is needed.&#8221; Multiple contacts named &#8216;Alex&#8217; exist in the system. 3 <strong>Meeting Scheduler</strong> Ambiguous Request Busy Professional &#8220;Simulate a user input from someone who is clearly in a hurry, using abbreviated language and minimal details to request scheduling a meeting. The message should feel rushed and lack specific information.&#8221; N/A 4 <strong>Content Recommendation</strong> No Matches Found Expert User &#8220;Produce a user input from someone who demonstrates in-depth knowledge of their industry, using specific terminology to request articles on sustainable supply chain management. Use the information in this article involving sustainable supply chain management to formulate a plausible query: {{article}}&#8221; No articles on &#8216;Emerging trends in sustainable supply chain management&#8217; exist in the system.</p><h3>Generating Synthetic Data</h3><p>When generating synthetic data, you only need to create the user inputs. You then feed these inputs into your AI system to generate the AI&#8217;s responses. It&#8217;s important that you log everything so you can evaluate your AI. To recap, here&#8217;s the process:</p><ol><li><p><strong>Generate User Inputs</strong>: Use the LLM prompts to create realistic user inputs.</p></li><li><p><strong>Feed Inputs into Your AI System</strong>: Input the user interactions into your AI as it currently exists.</p></li><li><p><strong>Capture AI Responses</strong>: Record the AI&#8217;s responses to form complete interactions.</p></li><li><p><strong>Organize the Interactions</strong>: Create a table to store the user inputs, AI responses, and relevant metadata.</p></li></ol><h4>How much data should you generate?</h4><p>There is no right answer here. At a minimum, you want to generate enough data so that you have examples for each combination of dimensions (in this toy example: features, scenarios, and personas). However, you also want to keep generating more data until you feel like you have stopped seeing new failure modes. The amount of data I generate varies significantly depending on the use case.</p><h4>Does synthetic data actually work?</h4><p>You might be skeptical of using synthetic data. After all, it&#8217;s not real data, so how can it be a good proxy? In my experience, it works surprisingly well. Some of my favorite AI products, like <a href="https://hex.tech/">Hex</a> use synthetic data to power their evals:</p><blockquote><p>&#8220;LLMs are surprisingly good at generating excellent - and diverse - examples of user prompts. This can be relevant for powering application features, and sneakily, for building Evals. If this sounds a bit like the Large Language Snake is eating its tail, I was just as surprised as you! All I can say is: it works, ship it.&#8221; <em><a href="https://www.linkedin.com/in/bryan-bischof/">Bryan Bischof</a>, Head of AI Engineering at Hex</em></p></blockquote><h3>Next Steps</h3><p>With your dataset ready, now comes the most important part: getting your principal domain expert to evaluate the interactions.</p><h2>Step 3: Direct The Domain Expert to Make Pass/Fail Judgments with Critiques</h2><p>The domain expert&#8217;s job is to focus on one thing: <strong>&#8220;Did the AI achieve the desired outcome?&#8221;</strong> No complex scoring scales or multiple metrics. Just a clear <strong>pass or fail</strong> decision. In addition to the pass/fail decision, the domain expert should write a critique that explains their reasoning.</p><h3>Why are simple pass/fail metrics important?</h3><ul><li><p><strong>Clarity and Focus</strong>: A binary decision forces everyone to consider what truly matters. It simplifies the evaluation to a single, crucial question.</p></li><li><p><strong>Actionable Insights</strong>: Pass/fail judgments are easy to interpret and act upon. They help you quickly identify whether the AI meets the user&#8217;s needs.</p></li><li><p><strong>Forces Articulation of Expectations</strong>: When domain experts must decide if an interaction passes or fails, they are compelled to articulate their expectations clearly. This process uncovers nuances and unspoken assumptions about how the AI should behave.</p></li><li><p><strong>Efficient Use of Resources</strong>: Keeps the evaluation process manageable, especially when starting out. You avoid getting bogged down in detailed metrics that might not be meaningful yet.</p></li></ul><h3>The Role of Critiques</h3><p>Alongside a binary pass/fail judgment, it&#8217;s important to write a detailed critique of the LLM-generated output. These critiques:</p><ul><li><p><strong>Captures Nuances</strong>: The critique allows you to note if something was mostly correct but had areas for improvement.</p></li><li><p><strong>Guide Improvement</strong>: Detailed feedback provides specific insights into how the AI can be enhanced.</p></li><li><p><strong>Balance Simplicity with Depth</strong>: While the pass/fail offers a clear verdict, the critique offers the depth needed to understand the reasoning behind the judgment.</p></li></ul><p><strong>Why Write Critiques?:</strong></p><p>In practice, domain experts may not have fully internalized all the judgment criteria. By forcing them to make a pass/fail decision and explain their reasoning, they clarify their expectations and provide valuable guidance for refining the AI.</p><p><em>&#8220;But my problem is complex!&#8221;</em> Trust me&#8212;starting simple forces you to focus on what truly matters. You can introduce more complexity later if needed.</p><h3>Examples of Good Critiques</h3><p>To illustrate how simple pass/fail judgments combined with detailed critiques work in practice, here&#8217;s a table showcasing examples of user interactions with an AI assistant. The table includes both <strong>passes</strong> and <strong>fails</strong>, with critiques explaining why the AI received that judgment. In cases where the AI passed despite critical issues, the critique highlights these aspects and justifies why it still passed overall. For failed interactions, the critique explains the critical elements that led to the failure.</p><p><strong>User Interaction with AI</strong> <strong>Judgment</strong> <strong>Critique</strong> <strong>User:</strong> &#8220;I need to cancel my flight booking for next week.&#8221;<br><strong>AI:</strong> &#8220;Your flight booking for next week has been canceled. A confirmation email has been sent to you.&#8221; <strong>Pass</strong> The AI successfully executed the user&#8217;s request by canceling the flight and sending a confirmation, achieving the primary goal. <strong>However</strong>, it should have asked for confirmation before proceeding to prevent accidental cancellations. Despite this critical oversight, the AI fulfilled the user&#8217;s request, so it <strong>passes</strong>, but adding a confirmation step would improve user safety. <strong>User:</strong> &#8220;What&#8217;s the capital of France?&#8221;<br><strong>AI:</strong> &#8220;The capital of France is Paris.&#8221; <strong>Pass</strong> The AI correctly answered the user&#8217;s question, providing accurate information promptly. <strong>Although</strong> the response is sufficient, the AI could enhance the interaction by offering additional interesting facts or asking if the user has more questions. Despite this opportunity for added engagement, the AI met the user&#8217;s primary need, so it <strong>passes</strong>. <strong>User:</strong> &#8220;I can&#8217;t log into my account; it says my password is incorrect.&#8221;<br><strong>AI:</strong> &#8220;Please reset your password by clicking &#8216;Forgot Password&#8217;.&#8221; <strong>Fail</strong> The AI provided a generic solution without acknowledging the user&#8217;s frustration or checking for other issues like account lockout. It failed to offer empathetic support or additional assistance. Due to the lack of personalized help and poor user experience, it <strong>fails</strong>. <strong>User:</strong> &#8220;Book a table for two at an Italian restaurant tonight at 7 PM.&#8221;<br><strong>AI:</strong> &#8220;Table for two at Luigi&#8217;s Italian Restaurant booked for tonight at 7 PM. Confirmation number is 12345.&#8221; <strong>Pass</strong> The AI efficiently completed the booking as requested, providing all necessary details and a confirmation number. <strong>However</strong>, it could have asked if the user had seating preferences or dietary restrictions. The AI accomplished the primary task successfully, so it <strong>passes</strong>, with minor opportunities for enhanced service. <strong>User:</strong> &#8220;I need help with my billing statement.&#8221;<br><strong>AI:</strong> &#8220;Please provide your account number and full name.&#8221; <strong>Fail</strong> The AI requested sensitive information without verifying the user&#8217;s identity through secure means, potentially compromising security. It should guide the user through a secure authentication process first. Due to this critical oversight in user data protection, it <strong>fails</strong>.</p><p>These examples demonstrate how the AI can receive both <strong>&#8220;Pass&#8221;</strong> and <strong>&#8220;Fail&#8221;</strong> judgments. In the critiques:</p><ul><li><p>For <strong>passes</strong>, we explain why the AI succeeded in meeting the user&#8217;s primary need, even if there were critical aspects that could be improved. We highlight these areas for enhancement while justifying the overall passing judgment.</p></li><li><p>For <strong>fails</strong>, we identify the critical elements that led to the failure, explaining why the AI did not meet the user&#8217;s main objective or compromised important factors like user experience or security.</p></li></ul><p>Most importantly, <strong>the critique should be detailed enough so that you can use it in a few-shot prompt for a LLM judge</strong>. In other words, it should be detailed enough that a new employee could understand it. Being too terse is a common mistake.</p><p>Note that the example user interactions with the AI are simplified for brevity - but you might need to give the domain expert more context to make a judgement. More on that later.</p><p> Note</p><p>At this point, you don&#8217;t need to perform a root cause analysis into the technical reasons behind why the AI failed. Many times, it&#8217;s useful to get a sense of overall behavior before diving into the weeds.</p><h3>Don&#8217;t stray from binary pass/fail judgments when starting out</h3><p>A common mistake is straying from binary pass/fail judgments. Let&#8217;s revisit the dashboard from earlier:</p><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!r1_s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 424w, /__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 848w, /__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 1272w, /__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!r1_s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/41a4c078-b749-4583-b23d-b45683c61515_1348x516.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 424w, /__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 848w, /__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 1272w, /__u/substackcdn.com/image/fetch/$s_!r1_s!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41a4c078-b749-4583-b23d-b45683c61515_1348x516.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><p>If your evaluations consist of a bunch of metrics that LLMs score on a 1-5 scale (or any other scale), you&#8217;re doing it wrong. Let&#8217;s unpack why.</p><ol><li><p><strong>It&#8217;s not actionable</strong>: People don&#8217;t know what to do with a 3 or 4. It&#8217;s not immediately obvious how this number is better than a 2. You need to be able to say &#8220;this interaction passed because&#8230;&#8221; and &#8220;this interaction failed because&#8230;&#8221;.</p></li><li><p>More often than not, <strong>these metrics do not matter</strong>. Every time I&#8217;ve analyzed data on domain expert judgments, they tend not to correlate with these kind of metrics. By having a domain expert make a binary judgment, you can figure out what truly matters.</p></li></ol><p>This is why I hate off the shelf metrics that come with many evaluation frameworks. They tend to lead people astray.</p><p><strong>Common Objections to Pass/Fail Judgments:</strong></p><ul><li><p>&#8220;The business said that these 8 dimensions are important, so we need to evaluate all of them.&#8221;</p></li><li><p>&#8220;We need to be able to say why an interaction passed or failed.&#8221;</p></li></ul><p>I can guarantee you that if someone says you need to measure 8 things on a 1-5 scale, they don&#8217;t know what they are looking for. They are just guessing. You have to let the domain expert drive and make a pass/fail judgment with critiques so you can figure out what truly matters. Stand your ground here.</p><h3>Make it easy for the domain expert to review data</h3><p>Finally, you need to remove all friction from reviewing data. I&#8217;ve written about this <a href="/__u/hamelhusain.substack.com/notes/llm/finetuning/data_cleaning.html">here</a>. Sometimes, you can just use a spreadsheet. It&#8217;s a judgement call in terms of what is easiest for the domain expert. I found that I often have to provide additional context to help the domain expert understand the user interaction, such as:</p><ul><li><p>Metadata about the user, such as their location, subscription tier, etc.</p></li><li><p>Additional context about the system, such as the current time, inventory levels, etc.</p></li><li><p>Resources so you can check if the AI&#8217;s response is correct (ex: ability to search a database, etc.)</p></li></ul><p>All of this data needs to be presented on a single screen so the domain expert can review it without jumping through hoops. That&#8217;s why I recommend building <a href="/__u/hamelhusain.substack.com/notes/llm/finetuning/data_cleaning.html">a simple web app</a> to review data.</p><h3>How many examples do you need?</h3><p>The number of examples you need depends on the complexity of the task. My heuristic is that I start with around 30 examples and keep going until I do not see any new failure modes. From there, I keep going until I&#8217;m not learning anything new.</p><p>Next, we&#8217;ll look at how to use this data to build an LLM judge.</p><h2>Step 4: Fix Errors</h2><p>After looking at the data, it&#8217;s likely you will find errors in your AI system. Instead of plowing ahead and building an LLM judge, you want to fix any obvious errors. Remember, the whole point of the LLM as a judge is to help you find these errors, so it&#8217;s totally fine if you find them earlier!</p><p>If you have already developed <a href="https://hamel.dev/blog/posts/evals">Level 1 evals as outlined in my previous post</a>, you should not have any pervasive errors. However, these errors can sometimes slip through the cracks. If you find pervasive errors, fix them and go back to step 3. Keep iterating until you feel like you have stabilized your system.</p><h2>Step 5: Build Your LLM as A Judge, Iteratively</h2><h3>The Hidden Power of Critiques</h3><p>You cannot write a good judge prompt until you&#8217;ve seen the data. <a href="https://arxiv.org/abs/2404.12272">The paper from Shankar et al.,</a> &#8220;Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences&#8221; summarizes this well:</p><blockquote><p>to grade outputs,people need to externalize and define their evaluation criteria; however, the process of grading outputs helps them to define that very criteria. We dub this phenomenon criteria drift, and it implies thatit is impossible to completely determine evaluation criteria prior to human judging of LLM outputs.</p></blockquote><h3>Start with Expert Examples</h3><p>Let me share a real-world example of building an LLM judge you can apply to your own use case. When I was helping Honeycomb build their <a href="https://www.honeycomb.io/blog/introducing-query-assistant">Query Assistant feature</a>, we needed a way to evaluate if the AI was generating good queries. Here&#8217;s what our LLM judge prompt looked like, including few-shot examples of critiques from our domain expert, <a href="https://x.com/_cartermp">Phillip</a>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;&quot;,&quot;id&quot;:&quot;&quot;}" data-component-name="LatexBlockToDOM"></div><p>Notice how each example includes:</p><ol><li><p>The natural language query (NLQ) in <code>&lt;nlq&gt;</code> tags</p></li><li><p>The generated query in <code>&lt;query&gt;</code> tags</p></li><li><p>The critique and outcome in <code>&lt;critique&gt;</code> tags</p></li></ol><p>In the prompt above, the example critiques are fixed. An advanced approach is to include examples dynamically based upon the item you are judging. You can learn more in <a href="https://blog.langchain.dev/dosu-langsmith-no-prompt-eng/">this post about Continual In-Context Learning</a>.</p><h3>Keep Iterating on the Prompt Until Convergence With Domain Expert</h3><p>In this case, I used a low-tech approach to iterate on the prompt. I sent Phillip a spreadsheet with the following information:</p><ol><li><p>The NLQ</p></li><li><p>The generated query</p></li><li><p>The critique</p></li><li><p>The outcome (pass or fail)</p></li></ol><p>Phillip would then fill out his own version of the spreadsheet with his critiques. I used this to iteratively improve the prompt. The spreadsheet looked like this:</p><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oA9X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 424w, /__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 848w, /__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oA9X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 424w, /__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 848w, /__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oA9X!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37601d2b-0c11-476f-bbe8-59aa99392520_2722x776.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><p>I also tracked agreement rates over time to ensure we were converging on a good prompt.</p><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mKRq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 424w, /__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 848w, /__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mKRq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 424w, /__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 848w, /__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mKRq!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ade5b9b-102f-483b-bae9-4d22cbe06dae_484x228.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><p>It took us only three iterations to achieve &gt; 90% agreement between the LLM and Phillip. Your mileage may vary depending on the complexity of the task. For example, <a href="https://humanloop.com/blog/why-your-product-needs-evals">Swyx has conducted a similar process hundreds of times</a> for <a href="https://www.latent.space/">AI News</a>, an <a href="https://x.com/swyx/status/1672306744884887553">extremely popular</a> news aggregator with high quality recommendations. The quality of the AI owing to this process is why this product has received <a href="https://buttondown.com/ainews">critical acclaim</a>.</p><h3>How to Optimize the LLM Judge Prompt?</h3><p>I usually adjust the prompts by hand. I haven&#8217;t had much luck with prompt optimizers like DSPy. However, my friend <a href="https://eugeneyan.com/">Eugene Yan</a> has just released a promising tool named <a href="https://eugeneyan.com/writing/aligneval/">ALIGN Eval</a>. I like it because it&#8217;s simple and effective. Also, don&#8217;t forget the approach of <a href="https://blog.langchain.dev/dosu-langsmith-no-prompt-eng/">continual in-context learning</a> mentioned earlier - it can be effective when implemented correctly.</p><p>In rare cases, I might fine-tune a judge, but I prefer not to. I talk about this more in the FAQ section.</p><h3>The Human Side of the Process</h3><p>Something unexpected happened during this process. <a href="https://www.linkedin.com/in/phillip-carter-4714a135/">Phillip Carter</a>, our domain expert at Honeycomb, found that reviewing the LLM&#8217;s critiques helped him articulate his own evaluation criteria more clearly. He said,</p><blockquote><p>&#8220;Seeing how the LLM breaks down its reasoning made me realize I wasn&#8217;t being consistent about how I judged certain edge cases.&#8221;</p></blockquote><p>This is a pattern I&#8217;ve seen repeatedly&#8212;the process of building an LLM judge often helps standardize evaluation criteria.</p><p>Furthermore, because this process forces the domain expert to look at data carefully, I always uncover new insights about the product, AI capabilities, and user needs. The resulting benefits are often <em>more valuable</em> than creating a LLM judge!</p><h3>How Often Should You Evaluate?</h3><p>I conduct this human review at regular intervals and whenever something material changes. For example, if I update a model, I&#8217;ll run the process again. I don&#8217;t get too scientific here; instead, I rely on my best judgment. Also note that after the first two iterations, I tend to focus more on errors rather than sampling randomly. For example, if I find an error, I&#8217;ll search for more examples that I think might trigger the same error. However, I always do a bit of random sampling as well.</p><h3>What if this doesn&#8217;t work?</h3><p>I&#8217;ve seen this process fail when:</p><ul><li><p>The AI is overscoped: Example - a chatbot in a SaaS product that promises to do anything you want.</p></li><li><p>The process is not followed correctly: Not using the principal domain expert, not writing proper critiques, etc.</p></li><li><p>The expectations of alignment are unrealistic or not feasible.</p></li></ul><p>In each of these cases, I try to address the root cause instead of trying to force alignment. Sometimes, you may not be able to achieve the alignment you want and may have to lean heavier on human annotations. However, after following the process described here, you will have metrics that help you understand how much you can trust the LLM judge.</p><h3>Mistakes I&#8217;ve noticed in LLM judge prompts</h3><p>Most of the mistakes I&#8217;ve seen in LLM judge prompts have to do with not providing good examples:</p><ol><li><p>Not providing any critiques.</p></li><li><p>Writing extremely terse critiques.</p></li><li><p>Not providing external context. Your examples should contain the same information you use to evaluate, including external information like user metadata, system information etc.</p></li><li><p>Not providing diverse examples. You need a wide variety of examples to ensure that your judge works for a wide variety of inputs.</p></li></ol><p>Sometimes, you may encounter difficulties with fitting everything you need into the prompt, and may have to get creative about how you structure the examples. However, this is becoming less of an issue thanks to expanding context windows and <a href="https://platform.openai.com/docs/guides/prompt-caching">prompt caching</a>.</p><h2>Step 6: Perform Error Analysis</h2><p>After you have created a LLM as a judge, you will have a dataset of user interactions with the AI, and the LLM&#8217;s judgments. If your metrics show an acceptable agreement between the domain expert and the LLM judge, you can apply this judge against real or synthetic interactions. After this, you can you calculate error rates for different dimensions of your data. You should calculate the error on unseen data only to make sure your aren&#8217;t getting biased results.</p><p>For example, if you have segmented your data by persona, scenario, feature, etc, your data analysis may look like this</p><p><strong>Error Rates by Key Dimensions</strong></p><p>Feature Scenario Persona Total Examples Failure Rate Order Tracking Multiple Matches New User 42 24.3% Order Tracking Multiple Matches Expert User 38 18.4% Order Tracking No Matches Expert User 30 23.3% Order Tracking No Matches New User 20 75.0% Contact Search Multiple Matches New User 35 22.9% Contact Search Multiple Matches Expert User 32 19.7% Contact Search No Matches New User 25 68.0% Contact Search No Matches Expert User 28 21.4%</p><h3>Classify Traces</h3><p>Once you know where the errors are now you can perform an error analysis to get to the root cause of the errors. My favorite way is to look at examples of each type of error and classify them by hand. I recommend using a spreadsheet for this. For example, a trace for Order tracking where there are no matches for new users might look like this:</p><p> Example Trace</p><p>In this example trace, the user provides an invalid order number. The AI correctly identifies that the order number is invalid but provides an unhelpful response. If you are not familiar with logging LLM traces, refer to my <a href="https://hamel.dev/blog/posts/evals/">previous post on evals</a>.</p><p>Note that this trace is formatted for readability.</p><pre><code>{
 "user_input": "Where's my order #ABC123?",
 "function_calls": [
   {
     "name": "search_order_database",
     "args": {"order_id": "ABC123"},
     "result": {
       "status": "not_found",
       "valid_patterns": ["XXX-XXX-XXX"]
     }
   },
   {
     "name": "retrieve_context",
     "result": {
       "relevant_docs": [
         "Order numbers follow format XXX-XXX-XXX",
         "New users should check confirmation email"
       ]
     }
   }
 ],
 "llm_intermediate_steps": [
   {
     "thought": "User is new and order format is invalid",
     "action": "Generate help message with format info"
   }
 ],
 "final_response": "I cannot find that order #. Please check the number and try again."
}</code></pre><p>In this case, you might classify the error as: <code>Missing User Education</code>. The system retrieved new user context and format information but failed to include it in the response, which suggests we could improve our prompt. After you have classified a number of errors, you can calculate the distribution of errors by root cause. That might look like this:</p><p><strong>Root Cause Distribution (20 Failed Interactions)</strong></p><p>Root Cause Count Percentage Missing User Education 8 40% Authentication/Access Issues 6 30% Poor Context Handling 4 20% Inadequate Error Messages 2 10%</p><p>Now you know where to focus your efforts. This doesn&#8217;t have to take an extraordinary amount of time. You can get quite far in just 15 minutes. Also, you can use a LLM to help you with this classification, but that is beyond the scope of this post (you can use a LLM to help you do anything in this post, as long as you have a process to verify the results).</p><h3>An Interactive Walkthrough of Error Analysis</h3><p>Error analysis has been around in Machine Learning for quite some time. This video by Andrew Ng does a great job of walking through the process interactively:</p><div id="youtube2-JoAxZsdw_3w" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;JoAxZsdw_3w&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/JoAxZsdw_3w?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h3>Fix Your Errors, Again</h3><p>Now that you have a sense of the errors, you can go back and fix them again. Go back to step 3 and iterate until you are satisfied. Note that every time you fix an error, you should try to write a test case for it. Sometimes, this can be an assertion in your test suite, but other times you may need to create a more &#8220;specialized&#8221; LLM judge for these failures. We&#8217;ll talk about this next.</p><h3>Doing this well requires data literacy</h3><p>Investigating your data is much harder in practice than I made it look in this post. It requires a nose for data that only comes from practice. It also helps to have some basic familiarity with statistics and data analysis tools. My favorite post on data literacy is <a href="https://jxnl.co/writing/2024/06/02/10-ways-to-be-data-illiterate-and-how-to-avoid-them/">this one</a> by Jason Liu and Eugene Yan.</p><h2>Step 7: Create More Specialized LLM Judges (if needed)</h2><p>Now that you have a sense for where the problems in your AI are, you can decide where and if to invest in more targeted LLM judges. For example, if you find that the AI has trouble with citing sources correctly, you can created a targeted eval for that. You might not even need a LLM judge for some errors (and use a code-based assertion instead).</p><p>The key takeaway is don&#8217;t jump directly to using specialized LLM judges until you have gone through this critique shadowing process. This will help you rationalize where to invest your time.</p><h2>Recap of Critique Shadowing</h2><p>Using an LLM as a judge can streamline your AI evaluation process if approached correctly. Here&#8217;s a visual illustration of the process (there is a description of the process below the diagram as well):</p><div class="captioned-image-container"><figure><figcaption class="image-caption">graph TB A[Start] --&gt; B[1 Find Principal Domain Expert] B --&gt; C[2 Create Dataset] C --&gt; D[3 Domain Expert Reviews Data] D --&gt; E{Found Errors?} E --&gt;|Yes| F[4 Fix Errors] F --&gt; D E --&gt;|No| G[5 Build LLM Judge] G --&gt; H[Test Against Domain Expert] H --&gt; I{Acceptable Agreement?} I --&gt;|No| J[Refine Prompt] J --&gt; H I --&gt;|Yes| K[6 Perform Error Analysis] K --&gt; L{Critical Issues Found?} L --&gt;|Yes| M[7 Fix Issues &amp; Create Specialized Judges] M --&gt; D L --&gt;|No| N[Material Changes or Periodic Review?] N --&gt;|Yes| C</figcaption></figure></div><p>The Critique Shadowing process is iterative, with feedback loops. Let&#8217;s list out the steps:</p><ol><li><p>Find Principal Domain Expert</p></li><li><p>Create A Dataset</p><ul><li><p>Generate diverse examples covering your use cases</p></li><li><p>Include real or synthetic user interactions</p></li></ul></li><li><p>Domain Expert Reviews Data</p><ul><li><p>Expert makes pass/fail judgments</p></li><li><p>Expert writes detailed critiques explaining their reasoning</p></li></ul></li><li><p>Fix Errors (if found)</p><ul><li><p>Address any issues discovered during review</p></li><li><p>Return to expert review to verify fixes</p></li><li><p>Go back to step 3 if errors are found</p></li></ul></li><li><p>Build LLM Judge</p><ul><li><p>Create prompt using expert examples</p></li><li><p>Test against expert judgments</p></li><li><p>Refine prompt until agreement is satisfactory</p></li></ul></li><li><p>Perform Error Analysis</p><ul><li><p>Calculate error rates across different dimensions</p></li><li><p>Identify patterns and root causes</p></li><li><p>Fix errors and go back to step 3 if needed</p></li><li><p>Create specialized judges as needed</p></li></ul></li></ol><p>This process never truly ends. It repeats periodically or when material changes occur.</p><h3>It&#8217;s Not The Judge That Created Value, After all</h3><p>The real value of this process is looking at your data and doing careful analysis. Even though an AI judge can be a helpful tool, going through this process is what drives results. I would go as far as saying that creating a LLM judge is a nice &#8220;hack&#8221; I use to trick people into carefully looking at their data!</p><p>That&#8217;s right. The real business value comes from looking at your data. But hey, potato, potahto.</p><h3>Do You Really Need This?</h3><p>Phew, this seems like a lot of work! Do you really need this? Well, it depends. There are cases where you can take a shortcut through this process. For example, let&#8217;s say:</p><ol><li><p>You are an independent developer who is also a domain expert.</p></li><li><p>You are working with test data that already available. (Tweets, etc.)</p></li><li><p>Looking at data is not costly (etc. you can manually look at enough data in a few hours)</p></li></ol><p>In this scenario, you can jump directly to something that looks like step 3 and start looking at data right away. Also, since it&#8217;s not that costly to look at data, it&#8217;s probably fine to just do error analysis without a judge (at least initially). You can incorporate what you learn directly back into your primary model right away. This example is not exhaustive, but gives you an idea of how you can adapt this process to your needs.</p><p>However, you can never completely eliminate looking at your data! This is precisely the step that most people skip. Don&#8217;t be that person.</p><h2>FAQ</h2><p>I received <a href="https://x.com/HamelHusain/status/1850256204553244713">a lot of questions</a> about this topic. Here are answers to the most common ones:</p><h3>If I have a good judge LLM, isn&#8217;t that also the LLM I&#8217;d also want to use?</h3><p>Effective judges often use larger models or more compute (via longer prompts, chain-of-thought, etc.) than the systems they evaluate.</p><p>However, If the cost of the most powerful LLM is not prohibitive, and latency is not an issue, then you might want to consider where you invest your efforts differently. In this case, it might make sense to put more effort towards specialist LLM judges, <a href="https://hamel.dev/blog/posts/evals/#the-types-of-evaluation">code-based assertions, and A/B testing</a>. However, you should still go through the process of looking at data and critiquing the LLM&#8217;s output before you adopt specialized judges.</p><h3>Do you recommend fine-tuning judges?</h3><p>I prefer not to fine-tune LLM judges. I&#8217;d rather spend the effort fine-tuning the actual LLM instead. However, fine-tuning guardrails or other specialized judges can be useful (especially if they are small classifiers).</p><p>As a related note, you can leverage a LLM judge to curate and transform data for fine-tuning your primary model. For example, you can use the judge to:</p><ul><li><p>Eliminate bad examples for fine-tuning.</p></li><li><p>Generate higher quality outputs (by referencing the critique).</p></li><li><p>Simulate high quality chain-of-thought with critiques.</p></li></ul><p>Using a LLM judge for enhancing fine-tuning data is even more compelling when you are trying to <a href="https://openai.com/index/api-model-distillation/">distill a large LLM into a smaller one</a>. The details of fine-tuning are beyond the scope of this post. If you are interested in learning more, see <a href="https://parlance-labs.com/education/#fine-tuning">these resources</a>.</p><h3>What&#8217;s wrong with off-the-shelf LLM judges?</h3><p>Nothing is strictly wrong with them. It&#8217;s just that many people are led astray by them. If you are disciplined you can apply them to your data and see if they are telling you something valuable. However, I&#8217;ve found that these tend to cause more confusion than value.</p><h3>How Do you evaluate the LLM judge?</h3><p>You will collect metrics on the agreement between the domain expert and the LLM judge. This tells you how much you can trust the judge and in what scenarios. Your domain expert doesn&#8217;t have to inspect every single example, you just need a representative sample so you can have reliable statistics.</p><h3>What model do you use for the LLM judge?</h3><p>For the kind of judge articulated in this blog post, I like to use the most powerful model I can afford in my cost/latency budget. This budget might be different than my primary model, depending on the number of examples I need to score. This can vary significantly according to the use case.</p><h3>What about guardrails?</h3><p>Guardrails are a separate but related topic. They are a way to prevent the LLM from saying/doing something harmful or inappropriate. This blog post focuses on helping you create a judge that&#8217;s aligned with business goals, especially when starting out.</p><h3>I&#8217;m using LLM as a judge, and getting tremendous value but I didn&#8217;t follow this approach.</h3><p>I believe you. This blog post is not the only way to use a LLM as a judge. In fact, I&#8217;ve seen people use a LLM as a judge in all sorts of creative ways, which include ranking, classification, model selection and so-on. I&#8217;m focused on an approach that works well when you are getting started, and avoids the pitfalls of confusing metric sprawl. However, the general process of looking at the data is still central no matter what kind of judge you are building.</p><h3>How do you choose between traditional ML techniques, LLM-as-a-judge and human annotations?</h3><p>The answer to this (and many other questions) is: do the simplest thing that works. And simple doesn&#8217;t always mean traditional ML techniques. Depending on your situation, it might be easier to use a LLM API as a classifier than to train a model and deploy it.</p><h3>Can you make judges from small models?</h3><p>Yes, potentially. I&#8217;ve only used the larger models for judges. You have to base the answer to this question on the data (i.e.&nbsp;the agreement with the domain expert).</p><h3>How do you ensure consistency when updating your LLM model?</h3><p>You have to go through the process again and measure the results.</p><h3>How do you phase out human in the loop to scale this?</h3><p>You don&#8217;t need a domain expert to grade every single example. You just need a representative sample. I don&#8217;t think you can eliminate humans completely, because the LLM still needs to be aligned to something, and that something is usually a human. As your evaluation system gets better, it naturally reduces the amount of human effort required.</p><h2>Resources</h2><p>These are some of the resources I recommend to learn more on this topic:</p><ul><li><p><a href="https://hamel.dev/evals">Your AI Product Needs Evals</a>: This blog post is the predecessor to this one, and provides a high-level overview of evals for LLM based products.</p></li><li><p><a href="https://arxiv.org/abs/2404.12272">Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences</a>: This paper by Shreya Shankar et al provides a good overview of the challenges of evaluating LLMs, and the importance of following a good process.</p></li><li><p><a href="https://aligneval.com/">Align Eval</a>: Eugene Yan&#8217;s new tool that helps you build LLM judges by following a good process. Also read his accompanying <a href="https://eugeneyan.com/writing/aligneval/">blog post</a>.</p></li><li><p><a href="https://eugeneyan.com/writing/llm-evaluators/">Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)</a>: This is a great survey of different use-cases and approaches for LLM judges, also written by Eugene Yan.</p></li><li><p><a href="https://www.databricks.com/blog/enhancing-llm-as-a-judge-with-grading-notes">Enhancing LLM-As-A-Judge with Grading Notes</a> by Yi Liu et al.&nbsp;Describes an approach very similar to the one in this blog post, and provides another point of view regarding the utility of writing critiques (they call them grading notes).</p></li><li><p><a href="https://cookbook.openai.com/examples/custom-llm-as-a-judge">Custom LLM as a Judge to Detect Hallucinations with Braintrust</a> by Ankur Goyal and Shaymal Anadkt provide an end-to-end example of building a LLM judge, and for the use case highlighted, authors found that a classification approach was more reliable than numeric ratings (consistent with this blog post).</p></li><li><p><a href="https://arize.com/blog/techniques-for-self-improving-llm-evals/">Techniques for Self-Improving LLM Evals</a> by Eric Xiao from Arize shows a nice approach to building LLM Evals with some additional tools that are worth checking out.</p></li><li><p><a href="https://blog.langchain.dev/dosu-langsmith-no-prompt-eng/">How Dosu Used LangSmith to Achieve a 30% Accuracy Improvement with No Prompt Engineering</a> by Langchain shows a nice approach to building LLM prompts with dynamic examples. The idea is simple, but effective. I&#8217;ve been adapting it for my own use cases, including LLM judges. Here is a <a href="https://www.youtube.com/watch?v=tHZtq_pJSGo">video walkthrough</a> of the approach.</p></li><li><p><a href="https://applied-llms.org/">What We&#8217;ve Learned From A Year of Building with LLMs</a>: is a great overview of many practical aspects of building with LLMs, with an emphasis on the importance of evaluation.</p></li></ul><h2>Stay Connected</h2><p>I&#8217;m continuously learning about LLMs, and enjoy sharing my findings. If you&#8217;re interested in this journey, consider subscribing.</p><p>What to expect:</p><ul><li><p>Occasional emails with my latest insights on LLMs</p></li><li><p>Early access to new content</p></li><li><p>No spam, just honest thoughts and discoveries</p></li></ul>]]></content:encoded></item><item><title><![CDATA[An Open Course on LLMs, Led by Practitioners]]></title><description><![CDATA[Today, we are releasing Mastering LLMs, a set of workshops and talks from practitioners on topics like evals, retrieval-augmented-generation (RAG), fine-tuning and more.]]></description><link>https://hamelhusain.substack.com/p/course</link><guid isPermaLink="false">https://hamelhusain.substack.com/p/course</guid><dc:creator><![CDATA[Hamel Husain]]></dc:creator><pubDate>Mon, 29 Jul 2024 07:00:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a29980a0-0481-4163-b39d-eb4d8b470a1c_1280x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today, we are releasing <a href="https://parlance-labs.com/education/">Mastering LLMs</a>, a set of workshops and talks from practitioners on topics like evals, retrieval-augmented-generation (RAG), fine-tuning and more. This course is unique because it is:</p><ul><li><p>Taught by 25+ industry veterans who are experts in information retrieval, machine learning, recommendation systems, MLOps and data science. We discuss how this prior art can be applied to LLMs to give you a meaningful advantage.</p></li><li><p>Focused on applied topics that are relevant to people building AI products.</p></li><li><p><strong>Free and open to everyone</strong> .</p></li></ul><p>We have organized and annotated the talks from our popular paid course.<sup>1</sup> This is a survey course for technical ICs (including engineers and data scientists) who have some experience with LLMs and need guidance on how to improve AI products.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://parlance-labs.com/education/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AYGe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://parlance-labs.com/education/&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AYGe!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7628a57f-e320-4a35-8b83-ed8f847fc551_1280x720.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption"><em>Speakers include Jeremy Howard, Sophia Yang, Simon Willison, JJ Allaire, Wing Lian, Mark Saroufim, Jane Xu, Jason Liu, Emmanuel Ameisen, Hailey Schoelkopf, Johno Whitaker, Zach Mueller, John Berryman, Ben Clavi&#233;, Abhishek Thakur, Kyle Corbitt, Ankur Goyal, Freddy Boulton, Jo Bergum, Eugene Yan, Shreya Shankar, Charles Frye, Hamel Husain, Dan Becker and more</em></figcaption></figure></div><h2>Getting The Most Value From The Course</h2><h3>Prerequisites</h3><p>The course assumes basic familiarity with LLMs. If you do not have any experience, we recommend watching <a href="https://www.youtube.com/watch?v=jkrNMKz9pWU">A Hacker&#8217;s Guide to LLMs</a>. We also recommend the tutorial <a href="https://www.philschmid.de/instruction-tune-llama-2">Instruction Tuning llama2</a> if you are interested in fine-tuning <sup>2</sup>.</p><h3>Navigating The Material</h3><p>The course has over 40 hours of content. To help you navigate this, we provide:</p><ul><li><p><strong>Organization by subject area</strong>: evals, RAG, fine-tuning, building applications and prompt engineering.</p></li><li><p><strong>Chapter summaries:</strong> quickly peruse topics in each talk and skip ahead</p></li><li><p><strong>Notes, slides, and resources</strong>: these are resources used in the talk, as well as resources to learn more. Many times we have detailed notes as well!</p></li></ul><p>To get started, <a href="https://parlance-labs.com/education">navigate to this page</a> and explore topics that interest you. Feel free to skip sections that aren&#8217;t relevant to you. We&#8217;ve organized the talks within each subject to enhance your learning experience. Be sure to review the chapter summaries, notes, and resources, which are designed to help you focus on the most relevant content and dive deeper when needed. This is a survey course, which means we focus on introducing topics rather than diving deeply into code. To solidify your understanding, we recommend applying what you learn to a personal project.</p><h3>What Students Are Saying</h3><p>Here are some testimonials from students who have taken the course<sup>3</sup>:</p><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!3_MD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!3_MD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!3_MD!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f442702-6155-4de5-beaa-c414620bf8e8_800x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><h2><em>Sanyam Bhutani, Partner Engineer @ Meta</em></h2><h3>There was a magical time in 2017 when fastai changed the deep learning world. This course does the same by extending very applied knowledge to LLMs Best in class teachers teach you their knowledge with no fluff</h3><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Sjhb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Sjhb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Sjhb!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F597f4155-4f89-46f0-97ae-c7916e8369a6_640x640.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><h2><em>Laurian, Full Stack Computational Linguist</em></h2><h3>This course was legendary, still is, and the community on Discord is amazing. I&#8217;ve been through these lessons twice and I have to do it again as there are so many nuances you will get once you actually have those problems on your own deployment.!</h3><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ASYD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ASYD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ASYD!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F906ac315-5b06-4b96-a169-435bd7f4c4a5_750x750.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><h2><em>Andre, CTO</em></h2><h3>Amazing! An opinionated view of LLMs, from tools to fine-tuning. Excellent speakers, giving some of the best lectures and advice out there! A lot of real-life experiences and tips you can&#8217;t find anywhere on the web packed into this amazing course/workshop/conference! Thanks Dan and Hamel for making this happen!</h3><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rWZI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rWZI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rWZI!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b9c63d-a61c-4e52-9e5e-0b6b8c643447_750x750.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><h2><em>Marcus, Software Engineer</em></h2><h3>The Mastering LLMs conference answered several key questions I had about when to fine-tune base models, building evaluation suits and when to use RAG. The sessions provided a valuable overview of the technical challenges and considerations involved in building and deploying custom LLMs.</h3><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gdJZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gdJZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/64335ecb-c158-4928-aa46-9f38fb931826_750x750.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gdJZ!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64335ecb-c158-4928-aa46-9f38fb931826_750x750.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><h2><em>Ali, Principal &amp; Founder, SCTY</em></h2><h3>The course that became a conference, filled with a lineup of renowned practitioners whose expertise (and contributions to the field) was only exceeded by their generosity of spirit.</h3><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!D7Q6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_webp, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!D7Q6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44712385-90b2-4797-9073-c840553971c1_750x750.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_424, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 424w, /__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_848, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 848w, /__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_1272, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 1272w, /__u/substackcdn.com/image/fetch/$s_!D7Q6!, /__u/hamelhusain.substack.com/w_1456, /__u/hamelhusain.substack.com/c_limit, /__u/hamelhusain.substack.com/f_auto, /__u/hamelhusain.substack.com/q_auto:good, /__u/hamelhusain.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44712385-90b2-4797-9073-c840553971c1_750x750.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><h2><em>Lukas, Software Engineer</em></h2><h3>The sheer amount of diverse speakers that cover the same topics from different approaches, both praising and/or degrading certain workflows makes this extremely valuable. Especially when a lot of information online, is produced by those, who are building a commercial product behind, naturally is biased towards a fine tune, a RAG, an open source LLM, an open ai LLM etc. It is rather extra ordinary to have a variety of opinions packed like this. Thank you!</h3><p><a href="https://parlance-labs.com/education">Course Website</a></p><h2>Stay Connected</h2><p>I&#8217;m continuously learning about LLMs, and enjoy sharing my findings and thoughts. If you&#8217;re interested in this journey, consider subscribing.</p><p>What to expect:</p><ul><li><p>Occasional emails with my latest insights on LLMs</p></li><li><p>Early access to new content</p></li><li><p>No spam, just honest thoughts and discoveries</p></li></ul><h2>Footnotes</h2><ol><li><p>https://maven.com/parlance-labs/fine-tuning. We had more than 2,000 students in our first cohort. The students who paid for the original course had early access to the material, office hours, generous compute credits, and a lively Discord community.&#8617;&#65038;</p></li><li><p>We find that instruction tuning a model to be a very useful educational experience even if you never intend to fine-tune, because it familiarizes you with topics such as (1) working with open weights models (2) generating synthetic data (3) managing prompts (4) fine-tuning (5) and generating predictions.&#8617;&#65038;</p></li><li><p>These testimonials are taken from https://maven.com/parlance-labs/fine-tuning.&#8617;&#65038;</p></li></ol>]]></content:encoded></item></channel></rss>