<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Operator's Guide to AI]]></title><description><![CDATA[The Operator's Guide to AI is a weekly newsletter by AJ Asver, CEO of Grep.ai, that covers the latest AI updates, lessons, insights, and research for people deploying AI in production for mission-critical work. ]]></description><link>https://operatorsguidetoai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!VvGT!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f49e392-e6ef-49d1-bc4c-c056d35a7b21_1254x1254.png</url><title>The Operator&apos;s Guide to AI</title><link>https://operatorsguidetoai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 02:48:57 GMT</lastBuildDate><atom:link href="/__u/operatorsguidetoai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[AJ Asver, Co-Founder and CEO of Grep AI ]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[aj@grep.ai]]></webMaster><itunes:owner><itunes:email><![CDATA[aj@grep.ai]]></itunes:email><itunes:name><![CDATA[AJ Asver]]></itunes:name></itunes:owner><itunes:author><![CDATA[AJ Asver]]></itunes:author><googleplay:owner><![CDATA[aj@grep.ai]]></googleplay:owner><googleplay:email><![CDATA[aj@grep.ai]]></googleplay:email><googleplay:author><![CDATA[AJ Asver]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Why AI agents fail in production]]></title><description><![CDATA[What published data and three years of deployments reveal about agent failures]]></description><link>https://operatorsguidetoai.substack.com/p/why-ai-agents-fail-in-production</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/why-ai-agents-fail-in-production</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Wed, 02 Sep 2026 14:30:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!TGpq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!TGpq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!TGpq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2649743,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TGpq!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e517bc-9164-4be5-9576-29922b6e7419_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A few weeks ago I was talking to a team at a large financial company that has spent the last year building AI into one of its core operational workflows.</p><p>They had made real progress. They&#8217;d broken the process into individual checks, used LLMs to automate parts of the review, and were already seeing work move faster. Now they wanted to take the next step and let the system make some decisions without a person reviewing every case.</p><p>That&#8217;s where they hit a problem.</p><p>Some of their checks were easy to make deterministic. If you need to verify a fact against a database, you can look it up and compare the result. But other checks required reading an unstructured document and deciding whether it satisfied a policy. An LLM could do that well, but not with the consistency they needed to let it make the final decision.</p><p>So they were going back through the entire workflow, check by check, asking where they could replace an LLM with deterministic logic and where they actually needed the judgment of an agent.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Operator's Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><p>I&#8217;ve had versions of that conversation across compliance, underwriting, due diligence, and research workflows over the past three years.</p><p>The prototype is rarely the hard part. Problems appear when teams run the same workflow hundreds of times and need consistent results, traceable decisions, and predictable costs.</p><p>By mid-2026, roughly 79% of large enterprises were piloting AI agents. Only 8.6% to 14% had an agent operating at true production scale, handling real task volume with monitoring and incident response behind it.</p><p>That range comes from our review of seven major surveys covering more than 125,000 respondents, including research from <a href="https://www.deloitte.com/us/en/about/press-room/state-of-ai-report-2026.html">Deloitte</a>, <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">Gartner</a>, and <a href="https://sinch.com/ai-production-paradox/chapter/ai-production-challenges/">Sinch</a>. Each survey defines production differently, but the gap between pilots and durable deployments is consistent.</p><p>Meanwhile, model performance improved sharply between 2024 and 2026. On WebArena, which tests agents on multi-step web tasks, the top score rose from 14.4% to 71.6%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kc_i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 424w, /__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 848w, /__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kc_i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png" width="1396" height="364" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:364,&quot;width&quot;:1396,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77928,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 424w, /__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 848w, /__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kc_i!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F936de40e-b4e9-4f08-8e3f-430735404643_1396x364.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You&#8217;d expect those improvements to translate into more agents making it into production. They haven&#8217;t.</p><p>We examined two independent sources of production evidence. <a href="https://clyro.dev/blog/we-analyzed-100-ai-agent-failures/">Clyro collected 591 documented agent failures</a>, while <a href="https://presenc.ai/research/ai-agent-tool-calling-accuracy-benchmarks-2026">Presenc AI analyzed public evaluations and deployment instrumentation from more than 60 enterprise agent customers</a>. In Clyro&#8217;s dataset, 88% of failures involved infrastructure around the model. Hallucinations and other model-quality failures accounted for around 10%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!TJcL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 424w, /__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 848w, /__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!TJcL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png" width="727" height="505" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:505,&quot;width&quot;:727,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:63082,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 424w, /__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 848w, /__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TJcL!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82787ee8-b999-4acc-8b8e-2ed421996ad1_727x505.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Production failure distribution, 591 classified incidents. Click a bar to see which bucket it lands in. Sources: Clyro</figcaption></figure></div><p>That finding matches our experience pretty closely. </p><div class="callout-block" data-callout="true"><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oJy-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 424w, /__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 848w, /__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oJy-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png" width="1200" height="627" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:627,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77558,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 424w, /__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 848w, /__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oJy-!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F548ceae1-72b4-4133-a4a5-febf64823f44_1200x627.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Read the full Grep AI whitepaper on Why Agents Fail in Production <strong><a href="https://grep.ai/research/why-ai-agents-fail-article">here</a></strong>.</em></p></div><p>One way to understand the problem is through some simple math. Say an agent has to complete a 20-step workflow and is 90% accurate at each step. A 90% success rate sounds pretty good in isolation. Across all 20 steps, though, the probability of completing the entire workflow successfully falls to 12.2%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9DNR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 424w, /__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 848w, /__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9DNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png" width="717" height="629" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/efea8788-f0d6-4262-b357-fb16974edd04_717x629.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:629,&quot;width&quot;:717,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:136331,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213499196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 424w, /__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 848w, /__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9DNR!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefea8788-f0d6-4262-b357-fb16974edd04_717x629.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Task success probability by step count, at four levels of per-step accuracy. Even at 95% per-step accuracy, a bar most teams would call excellent, a 50-step task only clears 7.7%. Source: internal compound-error model, p^n Bernoulli case.</figcaption></figure></div><p>At 50 steps, the probability falls to 0.5%.</p><p>Production workflows also contain dependent errors. If an agent makes a bad tool call at step three, that result enters the context for step four. The agent makes its next decision using bad information, then carries that decision into the following step. It can execute a coherent plan built on a mistake it made ten minutes earlier.</p><p>When someone shows me an agent performing an impressive task, I want to know what happens on the hundredth run, then the thousandth. I want to know how it handles an API timeout halfway through a task, contradictory sources, or a context window crowded enough to obscure an instruction from the beginning of the run.</p><p>We started <a href="http://grep.ai">Grep AI</a> in compliance, which gave us an unforgiving place to learn how to build agents. If an agent reviews a business for money-laundering risk, &#8220;usually right&#8221; provides little value. Analysts need to know what it found, where it found it, how it used the information, and what happens when it lacks enough evidence to decide.</p><p>Our first products used GPT-4-era models. We abandoned a free-running agentic loop in favor of a pre-orchestrated agent, which we called an <em>agent on rails</em>. We wrote about that architecture in <a href="/__u/operatorsguidetoai.substack.com/p/agents-arent-all-you-need">Agents aren&#8217;t all you need</a>. The rails improved consistency because we defined the sequence of steps in advance and isolated the model calls.</p><p>That architecture also limited what the agent could do. Each workflow required hand-written logic, which restricted us to a small number of compliance use cases.</p><p>We added configurable steps so customers could change the order and behavior of a workflow. That made it easier to create new agents while preserving the controls we needed in production.</p><p>As models improved, the rails began to constrain the quality of the work. An investigative agent assessing negative news could open the articles it had received, extract evidence, and classify the result. It could not decide to search for another source, follow a connection between two companies, or investigate a person mentioned deep inside an article. Each step remained an independent model call because we had designed the system to contain compounding errors.</p><p>That led us to move from agents on rails to an <em>agent in a sandbox</em>, initially built on the Claude Code harness. We described the architecture in <a href="https://blog.grep.ai/blog/claude-in-a-box">Claude in a box</a>. The agent could plan its own work, use tools, write code, and pursue evidence across multiple sources. It opened up deep research, due diligence, legal analysis, and other high-stakes workflows that our previous architecture could not support.</p><p>The new agent was far more capable. It was also too expensive to run.</p><p>Earlier this year, we reviewed our Anthropic bills and found that we were spending more on tokens than the agent generated in revenue. We had to reduce the cost of each run for the business to work.</p><p>That pushed us to build our own harness around a few operating requirements.</p><p>First, the harness needed to reuse proven workflows. General-purpose coding agents plan each task from scratch because each input may require a different approach. Enterprise agents often perform the same workflow thousands of times. Once an agent discovers a reliable sequence, it should reuse that work instead of paying to rediscover the process on every run.</p><p>The harness also needed to learn operational facts. If an agent discovered that one data provider lacked business-registration coverage in the Bahamas while another provider covered it, future runs should use the right provider from the start. Some steps may no longer require a model once the system has learned a deterministic procedure.</p><p>We also added verification during execution. A check at the end of a 20-step run arrives too late if step three contaminated everything that followed. The system needs to inspect intermediate work, catch unsupported claims, and correct errors before later steps inherit them.</p><p>Finally, model selection became the most important cost and reliability control. Complex reasoning may require a frontier model. Summarizing a search result, extracting a date, or formatting an output may require a smaller model or a piece of code. The harness can choose between them at each step rather than sending the entire workflow to the most expensive model available.</p><p>The new harness is in testing, and we plan to roll it out to our largest enterprise customers. Early tests show cost reductions of 80% to 90%, with accuracy holding steady or improving. We still need more production volume before treating that range as a durable benchmark.</p><div class="callout-block" data-callout="true"><p><strong>A quick note on Grep:</strong> We built <a href="https://grep.ai/for-business">Grep</a> to solve these production problems. Grep turns repeatable, high-stakes workflows into agents with verification, citations, audit trails, and cost controls built into each run. </p><p>If you have a workflow stuck between prototype and production, you can <a href="https://grep.ai/">try Grep</a> or <a href="https://calendly.com/d/dvxc-tmc-6jz/call-with-grep-ceo">talk to us</a>.</p></div><p>Technical architecture covers only part of the work. Customers ask different questions once an agent moves from helping someone make a decision to making the decision itself.</p><p>They want to know how we tested it, why it reached a conclusion, and which sources it used. They ask how it handles conflicting evidence and whether a model update changed its behavior. They want to run the agent against cases where they already know the correct answer.</p><p>Customers have pushed us to build those controls.</p><p>One customer automating due diligence for an underwriting workflow wanted to verify every material claim the agent made. We exposed citations alongside each claim, showed the underlying tool outputs, added confidence information, and made the execution trace easier to inspect.</p><p>Another customer found that asking the agent to do more work improved accuracy and auditability. It also produced reports several times longer than their analysts wanted to read. Report length looked like a product preference until analysts stopped reviewing the work. We had increased accuracy while reducing the practical value of the output.</p><p>Production teams have to balance accuracy, consistency, latency, cost, auditability, and the amount of human attention each run requires. More investigation may improve coverage while increasing cost and review time. Tighter controls may improve consistency while preventing the agent from following useful leads.</p><p>We expose the agent&#8217;s reasoning, tool calls, sources, and generated files so customers can inspect its work. We combine that visibility with versioned agent configurations, in-product evaluations, and cost controls.</p><p>Future model improvements will help. They will not repair an unreliable API integration or define the right permissions for a workflow. Teams will still need to detect policy drift, roll back changes, and decide where a person should review the work. </p><p>The financial services team I spoke with was asking the right questions. They were deciding where an agent&#8217;s judgment added value, where software could produce a certain answer, and where a person still needed to review the result.</p><p>Putting agents into production requires teams to define the workflow, constrain access, verify intermediate work, monitor changes, and make the economics hold at real volume. More capable models expand what teams can build. The reliability of the finished system depends on the choices they make around them. </p><p>You can read the full Grep research report on <a href="https://grep.ai/research/why-ai-agents-fail-article">Why agents fail</a>. </p><div><hr></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/p/why-ai-agents-fail-in-production?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading The Operators Guide to AI! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/p/why-ai-agents-fail-in-production?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/p/why-ai-agents-fail-in-production?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><p><em><span>About the author:</span></em></p><p><em>I&#8217;m AJ Asver, co-founder and CEO of <a href="https://grep.ai/for-business">Grep</a>. I&#8217;ve spent the past three years building and deploying AI agents for high-stakes work inside large companies. I write The Operator&#8217;s Guide to AI to share what we learn when those systems meet real workflows, customers, costs, and production constraints.</em></p><p><em>You can follow me on <a href="https://x.com/_aj">X</a> and <a href="https://linkedin.com/in/ajasver">LinkedIn</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Morphing Agents and Computers - Building a State-of-the-Art Researcher]]></title><description><![CDATA[This piece is a snapshot of how Grep works today (mid-March 2026) and where we think this architecture goes next.]]></description><link>https://operatorsguidetoai.substack.com/p/morphing-agents-and-computers-building</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/morphing-agents-and-computers-building</guid><dc:creator><![CDATA[Miguel Rios Berrios]]></dc:creator><pubDate>Mon, 31 Aug 2026 20:28:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1wzD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!1wzD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 424w, /__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 848w, /__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!1wzD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp" width="1456" height="585" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:585,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Morphing Agents and Computers - Building a State-of-the-Art Researcher&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Morphing Agents and Computers - Building a State-of-the-Art Researcher" title="Morphing Agents and Computers - Building a State-of-the-Art Researcher" srcset="/__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 424w, /__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 848w, /__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!1wzD!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aed03e1-adc4-4842-9ffa-406618125266_1632x656.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This piece is a snapshot of how Grep works today (mid-March 2026) and where we think this architecture goes next.</p><p>For historical context, we originally took everything we learned at <a href="https://parcha.ai/">Parcha</a> and turned that into a harness for doing corporate due diligence. We wrote about that in <a href="https://blog.grep.ai/blog/claude-in-a-box">Claude in a Box</a>. Over time, we realized that the concept of due diligence for businesses was really a special case of something broader: there was a well-established industry forming around AI-assisted, in-depth research. So we evolved our harness into a general researcher that could operate across almost any domain.</p><p>If you want an agent to do its best work, you have to give it the right environment to work in. That means more than a list of tools or MCPs. It means giving the agent a full compute structure and the ability to take on the work required by the task in front of it, while also managing context efficiently at every step.</p><p>This industry is moving extremely quickly, and this system may look completely different in a couple of months. But we wanted to stop and document what we have, how we built it, and what we think comes next.</p><h2><strong>How GREP Works at 40,000 Feet</strong></h2><p>The easiest way to understand GREP is as one agent that morphs across stages of the research process - and when the work demands it, splits into parallel copies of itself.</p><p>GREP has multiple modes, and depending on the mode and the effort level, the shape of the system changes. Here I am describing the higher-effort mode: the one that can go deep for hours, doing as much research as needed.</p><p>This is <strong>not a multi-agent system</strong> in the typical sense. We do not have a zoo of specialized agents negotiating with each other. Instead, we have one agent that reshapes itself depending on the stage, and can spawn sub-agents that are clones of itself with narrower focus. Think of it less like a team and more like cell division - one researcher that multiplies when the problem requires it, then converges back into one voice.</p><p>The second most important point is that execution is not a single pass. It is a <strong>two-layered loop</strong>:</p><ul><li><p><strong>Inner loop</strong>: The researcher spawns sub-agents, evaluates their results, identifies gaps, and spawns more sub-agents until it is satisfied with the coverage and quality of the research.</p></li><li><p><strong>Outer loop</strong>: After the report is written, reviewers and fact-checkers evaluate the entire output. If the report does not meet the quality bar - if claims are unsupported, sections are thin, or the research has gaps - the system goes back to the research phase and the inner loop runs again.</p></li></ul><p>The agent decides when it is done. Not a timer, not a token budget. The agent reads the review, compares it against the adequacy criteria from the plan, and either ships the report or goes back to work.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oR0A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 424w, /__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 848w, /__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oR0A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png" width="804" height="568" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:568,&quot;width&quot;:804,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55374,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 424w, /__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 848w, /__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oR0A!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ad9e060-9104-47a9-b0b9-9366c0b42b45_804x568.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The flow starts with the planner, which takes the user&#8217;s question, any supporting information, files, and additional context. It synthesizes all of that and begins to shape the research process. That plan is both a proposal for how the research should be carried out and an opinionated view of how the system should operate.</p><p>The planner is a researcher too. It may need to gather context through web search before it can write a good plan - very much like a coding agent exploring a codebase before taking on a large implementation task. It explores the topic, asks the user clarifying questions if needed, and then writes an elaborated plan aligned to the capabilities of the expert that will execute it.</p><h2><strong>Planning</strong></h2><p>In difficult research, the quality of the plan often determines the quality of the outcome.</p><p>Planning matters because difficult research is not just a matter of searching more. It is a matter of spending the right amount of time understanding the question, clarifying the user&#8217;s intent, identifying what is known and unknown, and deciding how the problem should be broken apart.</p><p>The planner is not just routing work. The planner is doing research. Careful planning requires context, and in some cases it requires interacting with the user to clarify the question, the user&#8217;s intent, or the context behind the request. A good planner gathers enough information to shape the rest of the system without overcommitting too early.</p><h3><strong>Expert Shaping</strong></h3><p>One of the most important things the planner does is shape what the expert will become for the task ahead. We do not think of GREP as a set of separate agents negotiating with each other. We think of it as one super-agent that morphs over time. Planning is its own phase, with its own skills and objectives, because that phase determines what the later researcher will need to be.</p><p>If the task is about financial analysis, the planner has to identify that the researcher will need the skills, tools, and ability to work with financial datasets, run analysis, and rely on reliable and up-to-date data sources. That is especially important because language models do not know everything, and more importantly, they must know what they do not know. Planning is where we decide how the system will complement the model&#8217;s general knowledge with web access, tools, command-line interfaces, and computation.</p><h3><strong>Skills, Tools, and What the Model Doesn&#8217;t Know</strong></h3><p>Planning is also where we decide which skills from our broader skill library should be loaded into the expert without overwhelming the context window. The planner selects tools, MCPs, CLIs, or specialized capabilities that might help the researcher operate more effectively. That might mean specialized news and real-time search, or it might mean coding tools, data analysis tools, or something even more specific.</p><p>The output of planning is a detailed file, and in some cases additional files with supporting structure. That file shapes the rest of the run. It includes an opinionated view of how the research should proceed, which capabilities should be loaded, what the major tasks are, what good outcomes look like, and what the final review checklist should contain.</p><h2><strong>Execution</strong></h2><p>Once the plan exists, execution becomes a problem of task shaping, context preservation, and structured synthesis. This is where the two-layered loop comes to life.</p><p>The first thing the researcher does is load the skills, tools, MCPs, and CLIs prescribed by the planner. But the researcher still has agency. It can adapt the plan, reshape it, and make decisions about how to pursue the goal. The plan is guidance, not a hard-coded script.</p><h3><strong>The Inner Loop: Research Until Satisfied</strong></h3><p>The first major execution step is turning the plan into a set of tasks. We use tasks heavily, and they are extremely powerful because they allow us to define work in a way that is specific, independent where possible, and suitable for concurrency.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vEj8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 424w, /__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 848w, /__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vEj8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png" width="812" height="396" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:396,&quot;width&quot;:812,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:35468,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 424w, /__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 848w, /__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vEj8!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05e1dc8d-1cf2-4080-8146-438f04c8a964_812x396.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>This is where our architecture begins to look map-reduce-like. If the planner identifies five subtopics that can be researched independently, the researcher can spawn five sub-agents to go pursue those lines of inquiry. If a task needs a different type of expert - like a data analyst - it can spawn one with a different configuration, tools, and objective.</p><p>The inner loop does not stop after one round of sub-agents. The orchestrating researcher evaluates the results, identifies gaps in coverage or quality, and decides whether to spawn more sub-agents, deepen existing lines of inquiry, or move on. This loop continues until the researcher is satisfied that it has enough material to produce a strong report.</p><p>This is not a fixed number of iterations. The agent reads the evidence, compares it against the objectives in the plan, and makes a judgment call. Some questions resolve in one round. Others require three or four rounds of increasingly targeted research.</p><h3><strong>Context is Precious</strong></h3><p>The inner loop works because the main agent is disciplined about preserving its own context window. The orchestrating researcher does not try to hold every detail of every branch of research in its active context. Instead, it coordinates, synthesizes, and writes the final research while sub-agents use their own context windows to do the deeper work.</p><p>The file system is what makes this practical. Sub-agents do not send all of their work back through one giant message. They write files. In fact, they often write multiple files: the long analysis itself, and then smaller supporting documents that point the main agent to the most important sections. That way the orchestrator can read the summary or pointer file, decide which sections matter, and only read the relevant portions of the longer analysis.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kY4-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 424w, /__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 848w, /__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kY4-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png" width="788" height="446" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:446,&quot;width&quot;:788,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:46433,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 424w, /__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 848w, /__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kY4-!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9618e5-a210-4e2c-9236-4510443a8ea2_788x446.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is one of the central ideas in GREP: context is precious. Effective use of context matters more than simply having a large context window. You need to preserve attention for synthesis, judgment, and writing - not waste it on moving large quantities of intermediate text around the system.</p><p>Tasks also help with persistence and recovery over long-running jobs. When research runs for a long time, the agent may compact and lose some active memory of the objective or the plan. The plan is always available in the file system to be read again, but tasks and task dependencies also preserve structure. They act as a durable representation of what needs to happen, what has already happened, and what depends on what.</p><h3><strong>Writing the Report</strong></h3><p>Writing the final report is more complex than it looks. Very long outputs are hard for language models to produce reliably in one pass. So instead of asking the agent to write a hundred-thousand-character report from scratch in one shot, we use what we think of as a pixelated-to-picture-perfect process.</p><p>The writer starts by generating a table of contents into a file. Then it writes thesis sentences or early scaffolding for sections and subsections. Then it expands section by section, gradually refining the report into something complete and cohesive. We do not use multiple writers stitching whole sections together because that often reads like multiple authors pasted into one document. The main writing agent owns the voice and cohesion from start to finish.</p><h3><strong>The Outer Loop: Review and Iterate</strong></h3><p>After the report is written, reviewers and fact-checkers take over. The final checklist defined in planning is read by reviewer agents, which perform their own validation, support claims with evidence, and identify areas where more research or correction is needed. They write their reviews into files.</p><p>Here is where the outer loop kicks in. The main agent reads the reviews and makes a decision: is the report good enough, or does it need to go back to the research phase? If the reviewers flag unsupported claims, thin sections, or missing perspectives, the system does not just patch the text. It goes back into the inner loop - spawning new sub-agents, gathering more evidence, and synthesizing again before rewriting the affected sections.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Qjm2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 424w, /__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 848w, /__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Qjm2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png" width="768" height="639" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:639,&quot;width&quot;:768,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:44933,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 424w, /__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 848w, /__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Qjm2!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1249ad9a-365e-4da1-a948-79a8d34dbdf1_768x639.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The agent decides when it is done. It reads the review, compares it against the adequacy criteria defined during planning, and either ships the report or goes back to work. This is what makes the system fundamentally different from a single-pass pipeline: the quality bar is enforced by the agent itself, not by a predetermined number of steps.</p><h2><strong>Benchmarks</strong></h2><p>We did not build GREP to win benchmarks, but the benchmarks made it clear we had built something unusual. They give us a way to understand how far we have come on certain dimensions and how we compare against the very companies and teams defining the category.</p><h3><strong>DRACO</strong></h3><p><a href="https://research.perplexity.ai/articles/evaluating-deep-research-performance-in-the-wild-with-the-draco-benchmark">DRACO</a> is Perplexity&#8217;s benchmark. It measures research quality across 100 questions in 10 domains - law, finance, medicine, technology, and more. Each answer is evaluated on factual accuracy, breadth and depth, presentation, and citation quality.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!NzzH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 424w, /__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 848w, /__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!NzzH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png" width="753" height="429" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:429,&quot;width&quot;:753,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39502,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 424w, /__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 848w, /__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NzzH!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff68bbd61-c1a5-4831-9804-9cbbc9926279_753x429.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Our high-effort agent scored <strong>78.6%</strong>, outperforming every entrant. Perplexity&#8217;s best configuration came in at 70.5% - on their own benchmark. That&#8217;s an 8.1 percentage point lead.</p><p>The domain breakdown tells a more interesting story. Grep wins <strong>9 of 10 domains</strong>. The only loss: Personalized Assistant, by 1.5 points.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!RkoY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 424w, /__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 848w, /__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!RkoY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png" width="679" height="690" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:679,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47649,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 424w, /__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 848w, /__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RkoY!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdf8fe6e-ed42-4ae5-bf27-0df6d1cc053c_679x690.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The largest gaps are in areas that require careful research methodology: Needle in a Haystack (+12.4pp), UX Design (+14.8pp), Shopping/Product (+10.5pp), and General Knowledge (+10.1pp). These are exactly the kinds of tasks where having a structured planning and execution pipeline - rather than a single-pass search - makes the biggest difference.</p><p>Grep also leads on <strong>all 4 rubric axes</strong>: Factual Accuracy (+7.5pp), Breadth &amp; Depth (+7.2pp), Presentation (+3.0pp), and Citation (+14.5pp). The citation gap is particularly notable - our architecture&#8217;s emphasis on evidence tracking and source verification through file-based context pays off directly.</p><h3><strong>DeepSearchQA</strong></h3><p><a href="https://arxiv.org/abs/2601.20975">DeepSearchQA</a> is Google&#8217;s benchmark. It evaluates multi-hop search problems where the system must follow several steps of reasoning and information gathering to arrive at a specific answer. 896 questions, each requiring factual correctness.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5TrH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 424w, /__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 848w, /__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5TrH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png" width="783" height="418" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:418,&quot;width&quot;:783,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40528,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 424w, /__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 848w, /__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5TrH!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a9426f3-e748-4f9b-939f-1ba2112df18f_783x418.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Grep scored <strong>84.5%</strong> factual correctness. That puts us ahead of Perplexity (81.9%, the unofficial SOTA before us) and well ahead of Google&#8217;s own Gemini Deep Research Agent (66.1%), which created the benchmark.</p><p>What makes that result more interesting is that we did not achieve it with our highest-effort agent. We achieved it with our medium-effort agent. In GREP, effort does not simply mean doing more work. It often means being more focused, more deliberate, and less excessive. In benchmarks like this, excessive retrieval or overly verbose answers can actually hurt performance. GREP&#8217;s focus helped here.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JHtX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 424w, /__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 848w, /__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JHtX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png" width="777" height="585" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ebb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:585,&quot;width&quot;:777,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:52929,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 424w, /__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 848w, /__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JHtX!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febb5f5ff-d713-40a5-9b96-e18267ee34f5_777x585.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The thing worth noting here is the floor, not the ceiling. The lowest-performing category, Finance &amp; Economics, still hits 79.5% on 132 questions. The spread from best to worst is only about 13 points. That matters because it means GREP is not gaming any particular question type - the architecture generalizes. Most deep research systems have a spike in one domain and a crater in another. Ours stays flat because the planning step adapts the research strategy per question rather than relying on a fixed retrieval pattern.</p><h3><strong>DeepResearch Bench</strong></h3><p><a href="https://arxiv.org/abs/2506.11763">DeepResearch Bench</a> is an independent benchmark built by researchers at the University of Science and Technology of China and Metastone Technology. It includes 50 PhD-level questions in English and 50 in Chinese, evaluated with the RACE scoring framework across comprehensiveness, insight, instruction following, and readability.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!plZy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 424w, /__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 848w, /__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 1272w, /__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!plZy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png" width="750" height="516" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:516,&quot;width&quot;:750,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:51970,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 424w, /__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 848w, /__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 1272w, /__u/substackcdn.com/image/fetch/$s_!plZy!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b5078a4-5437-4b0b-9931-06da7eb23663_750x516.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Grep scored <strong>56.27</strong>, edging out Cellcog Max (56.13) and a competitive field of research systems. The margins are tight at the top - 0.14 points separating first and second.</p><p>The most surprising detail: our reports were even higher quality on the Chinese portion (56.42) than on the English portion (56.13), despite the fact that we had spent effectively zero time optimizing specifically for Chinese.</p><p>That result says something important about the architecture. When you give a very strong base model the right structure, the right tools, and the right environment, you often do not need to over-index on narrow prompt tuning or specialized optimization for every domain. You give the agent expertise, you give it capabilities, and then you let it work.</p><p><em>Note: As of publication, the <a href="https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard">official leaderboard</a> lists our previous harness version in second place (56.09). Evaluators are currently confirming our latest results (56.27), which are published in full at our <a href="https://github.com/Parcha-ai/benchmarks/tree/main/deepresearch-bench">benchmarks repo</a>.</em></p><h2><strong>What&#8217;s Next</strong></h2><p>The story of GREP so far is really the story of discovering what an expert needs in order to do excellent work.</p><h3><strong>Experts Need Skills</strong></h3><p>The first thing we learned is that an expert cannot exist without specialized capabilities. That means tools, skills, command-line interfaces, and increasingly the guided ability to write code. It is not enough to hand an agent a static list of APIs. You need to give it actionable capabilities that match the work in front of it.</p><h3><strong>Experts Need Computers</strong></h3><p>Giving an agent a computer turned out not to be a convenience but a requirement. It needs to write and run programs. It needs a place to hold data, previous research, code, notes, and intermediate outputs. The file system is not just storage - it is a practical extension of the model&#8217;s working capacity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!CwQR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 424w, /__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 848w, /__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 1272w, /__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!CwQR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png" width="773" height="272" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:272,&quot;width&quot;:773,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:32745,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213608487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 424w, /__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 848w, /__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 1272w, /__u/substackcdn.com/image/fetch/$s_!CwQR!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c5a4290-3f92-41df-bfd0-1b1996b2e5f3_773x272.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>One useful mental model is to think of the context window like L1 or L2 cache in a computer: fast, precious, and limited. The file system is more like a larger and slower memory layer. It is not as immediate, but it is much more durable and much more available. That distinction has become increasingly important in how we think about agent design.</p><p>The computer does not have to be one thing. For some customers, it is a very simple sandbox where untrusted code can be run safely. For others, it can be much more powerful. We already have customers exploring what happens when research agents have access to GPU clusters and can perform significantly more computationally expensive operations.</p><h3><strong>Experts Need Brains</strong></h3><p>We are increasingly realizing that an expert needs a brain. It needs to remember previous work, preserve useful information across runs, accumulate context over time, and eventually improve by learning from prior research. Language models by themselves do not do this well. They do not naturally become better at a job just because they have done it many times. So one of the areas we are most interested in now is building abstractions on top of them that let them remember and deepen their expertise over time.</p><h3><strong>General Tools Over Narrow Interfaces</strong></h3><p>On the skills side, we are also learning that it may be less important to keep adding handcrafted capabilities than it is to let the agent write code, execute code, and use general-purpose tools well. Unix, Bash, Python, package managers, data tooling, and the file system are often more powerful than an ever-growing library of narrow interfaces.</p><p>The theme, increasingly, is simple: give the agent guidance, give it agency, and let it cook.</p><h2><strong>A Personal Note</strong></h2><p>One of the most fun parts of this project is that it has largely been built by two people, using the system itself to run the company around it.</p><p><a href="https://parcha.ai/">Parcha</a> and GREP are down to two people - AJ and myself. Not exactly by design. Startups are very hard. But the two of us have been able to build all of this, run the business, move customers onto it, and continue innovating by using GREP itself.</p><p>We have GREP experts helping with go-to-market. We have GREP experts helping with finance. We have coding agents fixing bugs. We have a growing internal army of experts helping us research, operate, and build. In practice, it is the two of us working alongside an army of GREP experts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uuAL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 424w, /__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 848w, /__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uuAL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png" width="1456" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grep GTM Expert in action&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grep GTM Expert in action" title="Grep GTM Expert in action" srcset="/__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 424w, /__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 848w, /__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uuAL!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50cefb05-fed9-45f5-a41e-09cecd3fad44_2042x784.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Grep GTM Expert in action</figcaption></figure></div><p>Building the company this way has been incredibly fun. It also feels like a preview of what is coming next for companies more broadly: smaller teams, leaner organizations, and significantly more output per person.</p><p>If you want to learn how we are doing this, or partner with us to optimize how you and your company work with GREP, <a href="mailto:feedback@grep.ai">get in touch</a>. We would be happy to share notes and give you access to some of the latest things we have been building.</p><p>We&#8217;re just getting started.</p><div><hr></div><p><em><a href="https://www.linkedin.com/in/miguelriosberrios/">Miguel Rios</a> is co-founder and CTO of <a href="https://grep.ai/">Grep</a>. Previously Head of Platform Engineering at Brex and Head of Consumer Data Science at Twitter. <a href="https://x.com/miguelrios">@miguelrios</a></em></p>]]></content:encoded></item><item><title><![CDATA[Welcome back!]]></title><description><![CDATA[The Hitchhiker&#8217;s Guide to AI is back with a new name and agenda]]></description><link>https://operatorsguidetoai.substack.com/p/welcome-back</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/welcome-back</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Mon, 31 Aug 2026 19:22:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vRqa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vRqa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 424w, /__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 848w, /__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vRqa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:63660,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://operatorsguidetoai.substack.com/i/213494627?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 424w, /__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 848w, /__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!vRqa!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01b40ba-34bf-435d-b593-8187b254e5f7_1254x1254.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span data-color="rgb(55, 71, 93)" style="color: rgb(55, 71, 93);">Almost four years ago, right after ChatGPT launched, I started writing this newsletter,&nbsp;</span><em>The Hitchhiker&#8217;s Guide to AI,</em><span data-color="rgb(55, 71, 93)" style="color: rgb(55, 71, 93);">&nbsp;to learn about the then-nascent AI space and share what I learned with other curious builders and operators in my network.</span> You subscribed then because you were as curious about AI as I was.</p><p>At the time, most of us were trying to figure out generative AI, large language models, and AI agents for the first time. What could these models do? Where was this going? How much of the hype was real?</p><p>I kept writing for a while, and that eventually <a href="/__u/operatorsguidetoai.substack.com/p/startups-navigating-the-ai-landscape?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">led me to the conviction</a> to <a href="/__u/operatorsguidetoai.substack.com/p/why-we-founded-parcha?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">found an AI startup</a>. Writing took a back seat to building.</p><p>For the last three years, I&#8217;ve been deep in deploying AI agents into production for some of the world's leading fintechs, like Airwallex, Coinbase, IG, and Worldpay. We started with AI for compliance, where the stakes are high, and mistakes have real consequences. Today, with our latest product, Grep.ai, we&#8217;re applying what we learned to a much broader category: high-stakes, repetitive knowledge work.</p><p>A few weeks ago, a customer asked me how other companies were approaching AI. What strategies were working? How were they rolling agents out? Where were they getting stuck?</p><p>I realized I answer versions of these questions all the time. I talk to teams building their own agents, executives deciding where to deploy them, and operators responsible for making them work once the demo is over. I&#8217;ve just never written much of it down. </p><p>So I&#8217;m starting now.</p><p>I&#8217;m also renaming this publication <strong>The Operator&#8217;s Guide to AI</strong>.</p><p>When I started writing, we were all still trying to understand AI. Today, many of you deploy it. You have agents in production, teams trying to scale them, a board asking about ROI, and your name on the line when something goes wrong.</p><p>How do you decide which work is ready for agents? How do you evaluate them? Where do humans stay in the loop? What does reliability look like when an agent is doing thousands of tasks a day? How do you know if any of this is creating real value?</p><p>That&#8217;s what I want to write about here.</p><p>I&#8217;ll share what we&#8217;re learning from running agents in production at Grep, what I&#8217;m seeing from companies doing their own rollouts, and the practices that seem to hold up once AI meets real work.</p><p>After three years of building, I have a lot to write down!</p><p><span data-color="rgb(55, 71, 93)" style="color: rgb(55, 71, 93);">The first edition, which will land later this week, will focus on&nbsp;</span><strong>why AI agents fail in production,&nbsp;</strong><span data-color="rgb(55, 71, 93)" style="color: rgb(55, 71, 93);">the most common pitfalls, and why it&#8217;s often the harness, not the model, that is the culprit.</span><span> </span></p><p>If you&#8217;re deploying AI into production, automating mission-critical work with agents, or leading AI transformation at your company, I hope you find this newsletter helpful!</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Operators Guide to AI! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p><em>About the author:<br>I&#8217;m AJ, the CEO and Co-Founder of <a href="https://grep.ai/for-business">Grep.AI</a>, a platform for deploying AI agents for repetitive high-stakes knowledge work. We work with global enterprises like Airwallex, IG, and Worldpay to automate mission-critical manual workflows across compliance, onboarding, go-to-market, corporate development, and more. If you want to learn more about Grep, you can schedule a call with me <a href="https://calendly.com/ajasver/30-min-call">here</a>. </em></p><p>You can also follow me on <a href="http://x.com/_aj">X</a> and <a href="https://linkedin.com/in/ajasver">LinkedIn</a>.</p>]]></content:encoded></item><item><title><![CDATA[The AI Pricing Revolution]]></title><description><![CDATA[Why SaaS is Shifting to Usage-Based Models]]></description><link>https://operatorsguidetoai.substack.com/p/the-ai-pricing-revolution</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/the-ai-pricing-revolution</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Thu, 11 Jul 2024 17:02:18 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080" width="6478" height="4480" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4480,&quot;width&quot;:6478,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;black digital device at 2 00&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="black digital device at 2 00" title="black digital device at 2 00" srcset="https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1602853175733-5ad62dc6a2c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0fHxnYXMlMjBwdW1wJTIwbWV0ZXJ8ZW58MHx8fHwxNzIwNzE3MTUzfDA&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="/__u/operatorsguidetoai.substack.com/true">Rock Staar</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>Welcome to The Guide To AI by Parcha!</p><p>I'm AJ, CEO of Parcha. In this newsletter, members of the Parcha team share our insights into AI and how it&#8217;s changing the world around us. Occasionally, we also share updates about what we&#8217;re building at Parcha. Today, we're diving into a topic that's been on my mind lately: how AI fundamentally changes the economics of SaaS and why I believe it will usher in a new era of usage-based pricing.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">If you enjoy reading Parcha&#8217;s Guide to AI, don&#8217;t forget to subscribe below!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The shift to usage-based pricing, spurred by AI-native products, has enormous implications for founders, product leaders, and anyone involved in the SaaS industry. So, let's break it down and explore what it means for the future of software pricing.</p><h2>The AI Pricing Paradox</h2><p>Here's the fascinating paradox at the heart of this shift: As AI becomes more integral to SaaS products, the cost to serve each customer <em>increases</em> with usage. This flies in the face of traditional SaaS economics, where marginal costs typically decrease as you scale.</p><p>Why is this happening? It all comes down to compute power. AI models, especially large language models (LLMs), are incredibly resource-intensive. Every time a user interacts with an AI feature, it costs the provider real money regarding GPU usage. This creates a major problem for the subscription-based pricing model that's dominated SaaS for the past decade. Suddenly, your power users become your least profitable customers. It's simply not sustainable.</p><h2>The Rise of Usage-Based Pricing</h2><p>The solution to this paradox is usage-based pricing. It aligns costs with revenue, ensures profitability, and provides a better experience for customers who only pay for what they use. We're already seeing this play out in the market:</p><ol><li><p>OpenAI's pricing for GPT-4 is purely usage-based</p></li><li><p>Google's new AI features for Workspace will likely come with usage limits</p></li><li><p>Even GitHub Copilot, while subscription-based, has usage caps</p></li></ol><p>But it gets fascinating here: We believe this shift will extend far beyond "AI-first" products. As AI capabilities become table stakes across all SaaS categories, usage-based pricing will become the norm.</p><h2>Implications for the SaaS Industry</h2><p>As more and more products add AI features, this shift has enormous implications for the entire SaaS ecosystem:</p><ol><li><p><strong>Pricing Strategy Overhaul</strong>: SaaS companies will need to rethink their pricing strategies completely. The days of simple "per user, per month" models are numbered.</p></li><li><p><strong>Evolution of Sales and Marketing</strong>: Say goodbye to "unlimited" plans. Sales teams will need to learn how to sell value based on usage, not just seats.</p></li><li><p><strong>Customer Success Becomes Critical</strong>: Driving profitable usage becomes even more important. Customer success teams will play a crucial role in ensuring customers are extracting maximum value.</p></li><li><p><strong>Product Development Focus</strong>: Product teams will need to prioritize features that drive valuable usage, not just engagement for engagement's sake.</p></li><li><p><strong>Financial Modeling Challenges</strong>: Finance teams will need to get comfortable with more variable revenue streams and potentially more complex forecasting.</p></li></ol><h2>What This Means for Founders and Product Leaders</h2><p>If you're building a SaaS product today, especially one with AI features, it's crucial to start thinking about how you might transition to a usage-based model. Here are a few steps to consider:</p><ol><li><p><strong>Analyze Your Costs</strong>: Understand the true cost of serving each customer, especially as it relates to AI feature usage.</p></li><li><p><strong>Model Different Scenarios</strong>: Experiment with different pricing structures and see how they impact your unit economics.</p></li><li><p><strong>Talk to Your Customers</strong>: Gauge how receptive they might be to a usage-based model. You might be surprised by the positive response.</p></li><li><p><strong>Start Small</strong>: Consider introducing usage-based elements alongside your existing pricing model before making a full transition.</p></li><li><p><strong>Invest in Usage Analytics</strong>: You'll need robust systems to track and bill accurately based on usage.</p></li></ol><h2>The Road Ahead</h2><p>The shift to usage-based pricing won't happen overnight, and it won't be without challenges. At Parcha, we&#8217;ve been experimenting with usage pricing from day one. It requires some upfront education to help our customers understand that our product should be compared to their headcount costs, not their SaaS budget. But we think it&#8217;s worth it because usage-based pricing better aligns with our costs and allows us to expand with our customers without upselling more seats. In return, customers have more control over their spending and will only pay for the value they extract from our product. For SaaS companies, it's an opportunity to build more sustainable, profitable businesses aligned with the actual costs of delivering AI-powered software.</p><p>What do you think? Will AI drive a usage-based pricing revolution? Or am I off base here? I'd love to hear your thoughts in the comments.</p><p>Until next time, thanks for reading Parchans!</p><p>-AJ</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Guide to AI by Parcha! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Agents aren’t all you need]]></title><description><![CDATA[Lessons from Parcha's Journey automating compliance workflows using AI and why autonomous agents aren&#8217;t always the best solution]]></description><link>https://operatorsguidetoai.substack.com/p/agents-arent-all-you-need</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/agents-arent-all-you-need</guid><dc:creator><![CDATA[Miguel Rios Berrios]]></dc:creator><pubDate>Thu, 06 Jun 2024 17:29:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/130daf73-81b0-4b72-8a33-aea42bf1c33b_1192x718.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!381i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 424w, /__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 848w, /__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 1272w, /__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!381i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png" width="1456" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:27328,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 424w, /__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 848w, /__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 1272w, /__u/substackcdn.com/image/fetch/$s_!381i!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F061cd630-7d9c-4f88-8403-f038210d6532_1600x960.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>The Promise of AI Agents</h3><p>We founded Parcha in early 2023, when "AI agents" built with language models emerged. We were inspired by demos showcasing how LLMs and code could dynamically create and execute complex plans using agentic behavior. Frameworks like <a href="https://www.parcha.com/blog/could-autonomous-ai-agents-be-the-next-step-towards-agi">AutoGPT and BabyAGI</a> emerged, and people started applying this technology to numerous use cases. The demos looked promising, and the excitement around autonomous AI agents was palpable. Soon after, we saw startups emerging for horizontal (AI Agent platforms) and vertical (AI Agents for specific industries) &nbsp;use cases like Parcha..</p><h3>Starting Parcha: The Thesis of AI Agents</h3><p>We founded Parcha with the thesis that AI agents could automate many operational processes in fintech and banks. Our initial approach was to take a standard operating procedure (SOP) that a human follows to perform an operational process and have an AI agent read and execute those steps autonomously. We developed an agent that would read the SOP and assume the role of the operations expert performing this process today, such as a compliance analyst.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI by Parcha! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The Parcha agent would generate dynamic plans based on the SOPs provided by our design partners. These plans would guide the agent through each process step, utilizing various tools and commands to execute tasks and record outcomes in the scratchpad. Once it has performed all the steps, the agent will use the SOP and the scratchpad to make a final determination, such as approving or denying a customer, and deliver the result to the end customer. We aimed to create a flexible and adaptive system capable of handling many workflows with minimal human intervention.</p><p>We were not building in a vacuum. From early on, we engaged with design partners &#8212; fintech and banks- with complex fintech workflows and compliance processes that agreed to work with us and test our agent. These companies saw potential in our technology to achieve their tasks without scaling up their workforce. We used their SOPs and even shadowed sessions where their teams performed these processes manually. We used our platform to build agents for Know Your Business (KYB), Know Your Customer (KYC), fraud detection, credit underwriting, merchant categorization, Suspicious Activity Reports (SARs) filings, and other processes.</p><p>The feedback was very encouraging. The demos quickly impressed our partners, and we felt like we were on the path to building a general agent platform to cover multiple processes. However, as we transitioned from demos to production, we found ourselves (a tiny team, by design) spending too much time building the "agent" platform and not enough time building the product and solving the problems our design partners wanted us to address.</p><p>While the concept of agentic behavior was promising, building reliable agentic behavior with large language models (LLMs) was a massive endeavor. Creating general-purpose "autonomous agents" could have taken us years. During that time, we wouldn&#8217;t solve the problems our customers cared most about with a product that directly addressed their needs. Our customers required accuracy, reliability, seamless integrations, and a user-friendly product experience&#8212;areas where our early versions fell short. They would much rather have a solution that was very accurate and reliable for a subset of tasks than a fully autonomous solution that could automate a workflow end-to-end but worked only 80% of the time. We needed to choose between building the agent or building the product.</p><h3>The Complexity of Putting Agentic Behavior in a Box</h3><p>As we ran our agents in actual production use cases, we began to uncover the inherent complexity of agentic behavior within a system. From systems theory, we know that as the number of components and interactions within a system increases, so does the complexity and the likelihood of failure. Our approach, which relied on dynamically generated plans and a shared memory model (a scratchpad), made it difficult to decouple and measure subsystems effectively. Each run produced different plans, leading to an ever-shifting landscape of tasks and interactions, which increased the unpredictability and potential for errors.</p><blockquote><p>To put this into perspective, if an AI agent carries out a workflow consisting of 10 tasks autonomously but has a 10% error rate per task, the compounded error rate over the whole workflow is 65%.</p></blockquote><p>Imagine a factory where the assembly line is rearranged dynamically with every run. In such a factory, the placement and order of machinery and tasks change constantly, making it impossible to establish a consistent and reliable workflow. This is akin to what we faced with our agentic behavior model. The dynamic generation of steps&#8212;where the agent created an execution plan on the fly&#8212;meant that no two executions were identical, leading to an ever-shifting landscape of tasks and interactions. This constantly evolving setup made it exceedingly difficult to instrument and evaluate the system, further complicating our ability to ensure reliability and performance.</p><p>The scratchpad, intended to facilitate the agent's observations and results, compounded these issues. Instead of simplifying the process, it created a tightly coupled system where decoupling tasks became nearly impossible. Each task depended on the preceding tasks' results recorded in the scratchpad, leading to a cascade of dependencies that we found challenging to isolate and test independently. Moreover, these dependencies made parallelization impossible. Tasks that could run in parallel had to wait, as the agent needed to evaluate the scratchpad before deciding the next steps.</p><p>This tightly coupled system with dynamically generated steps is inherently prone to failure for several reasons:</p><ol><li><p><strong>Unpredictability:</strong> Each execution path varied significantly, making predicting the agent's behavior and outcomes difficult. This unpredictability is a hallmark of complex systems and a primary source of failure.</p></li><li><p><strong>Complex Interdependencies:</strong> Changes or errors in one part of the system could propagate through the scratchpad, affecting subsequent tasks and leading to systemic failures. The reliance on shared memory created a cascade of dependencies that amplified the impact of any single issue.</p></li><li><p><strong>Artificial Performance Bottlenecks:</strong> Tasks that could run in parallel ran serially because the agent needed to evaluate the scratchpad constantly. This bottleneck reduced efficiency and increased processing time, countering one of the critical advantages of using AI agents.</p></li><li><p><strong>Evaluation Challenges:</strong> The constantly evolving setup made instrumenting, testing, and evaluating the agent's accuracy difficult. Determining the root cause of errors was challenging. Was the issue with the dynamically generated plan, a misunderstanding of the SOP, a malfunctioning tool, or a flaw in the decision-making process? This complexity hindered our ability to ensure the agent delivered accurate and reliable results, ultimately impacting our ability to meet customer needs effectively.</p></li></ol><p>These challenges underscored how impractical it was to rely solely on autonomous agentic behavior for complex workflows. Building a reliable, robust, and stable AI agent would require significant investment. Solving these challenges would have required all our resources for an extended period. Following that path would have prevented us from building our product, generating revenue, and iterating based on customer usage and feedback.</p><h3>The agent &#8220;on rails&#8221;</h3><p>As we spoke to more customers, we realized that for a specific use case like carrying out Know Your Customer compliance reviews, most companies used a variation of the same process and repeated that same process repeatedly. That meant the agent didn&#8217;t need to develop a new plan every time it ran. &nbsp;We could move faster while leveraging what we built by putting  the agent "on rails." We removed the dynamic generation of the plan and the shared scratchpad and instead developed an orchestration framework based on static agent configurations. For example, a KYB agent needs to perform the following tasks (simplified for illustration):</p><ol><li><p>Verify business registration</p></li><li><p>Verify business ownership</p></li><li><p>Verify the web presence of the business</p></li><li><p>Verify if the business operates in a high-risk industry</p></li><li><p>Verify if the business operates in a high-risk country</p></li></ol><p>This plan started as a configuration file! Now that the plan was static, we could parallelize the process. Tasks 1, 2, and 3 can be done in parallel, while tasks 4 and 5 depend on task 3 as we scan the business's website to understand where and in which industry it operates.</p><p>This approach allowed us to break down SOPs into discrete, manageable tasks, now automated with specialized tools. Previously, our "tools" were simple integrations, such as searching for an address on Google Maps. Now, Parcha tools handle complex yet decoupled parts of the process. For example, instead of having an AI agent figure out how to verify a business's registration, we use a sequence of tools that parse a PDF document, perform OCR on it, extract relevant information about the business, and check if it matches what the customer entered in the application. Each tool works standalone and can be reused in multiple workflows.</p><h3>Focus</h3><p>Since we no longer rely on an AI to determine a plan based on its SOP, we can only automate a process if we have a list of steps to perform it and the tools we need to do the job. Hence, we decided to hyperfocus on one area we knew well: Know Your Business/Customer (KYB/KYC). This allowed us to build intelligent workflows that could seamlessly integrate into our customers' existing processes, providing clear and measurable benefits quickly. We also placed a strong emphasis on the user experience (one of our values is to "Make it Dope!"), ensuring our product is easy to use and requires minimal integration and training. This new approach resonated well with our design partners, who, by this point, were paying customers.</p><h3>Where Parcha is Today</h3><p>The decision to focus on one workflow, executed as a simple set of steps, enabled us to build the best product for our customers. This strategic focus allowed us to explore and leverage generative AI more effectively. Today, we have LLM-powered workflows in over a dozen KYB and due diligence checks, covering document intelligence (e.g., proof of address verification), web-based due diligence (e.g., high-risk country/industry checks), and advanced search (e.g., adverse media, sanctions, and screening). These checks extensively use large language models and the concepts behind agentic behavior. For example, one Parcha KYB agent can perform more than a thousand LLM calls.</p><h3><strong>So, how do we use LLMs?</strong></h3><p>When we decided to focus, we also built our AI tools as decoupled, reusable blocks that we could extend and independently evaluate. We orchestrate these subsystems to perform tasks ranging from extracting specific pieces of information from documents to complex reasoning-driven functions like making an approve/deny decision-based on the output of various checks. We built abstractions and primitives for these checks. For example, a "data loader" is responsible for extracting information from one or more sources and structuring it appropriately, while a "check" uses that information to pass judgment. This modular approach allows us to evaluate every part of this pipeline independently. Did we extract the correct information from the document? Was the decision, based on the extracted data, correct? This approach ensures we can focus on each task without worrying about too many external factors.</p><p>For example, consider the task of screening news articles to determine if any of them mention a particular person of interest. Rather than building an agent that constructs a plan and autonomously fetches and analyzes each news article, we handle the process in a structured manner:</p><ol><li><p><strong>Extraction:</strong> We first extract relevant information from each news article, including what happened, when it happened, who was involved, and other pertinent facts. Our extraction process handles formats like HTML, PDF, or plain text and extracts the information from one article at a time. We make tens or sometimes hundreds of concurrent LLM calls to an inexpensive model to extract this information (e.g., a person with a common name will have multiple hits with their name on them). Each language model call focuses on this task, guided by a thoroughly tested prompt.</p></li><li><p><strong>Clustering: </strong>Then, we use the extracted features to cluster the articles into "news events." A news event (e.g., reporting a corporate financial crime) may contain multiple news articles, some with unique information to help the judgment process. We use a more complex model for this process since clustering requires more reasoning than a simple extraction.</p></li><li><p><strong>Partial Judgement:</strong> Once we have news events and all the information surrounding them, we use a top-of-the-line model to determine if the person of interest is the perpetrator mentioned in the article based on their name, age, address, and other available information about their profile. The language model independently provides judgments on each article, ensuring that the reasoning process is constrained to limited and structured information without noise.</p></li><li><p><strong>Final Decision:</strong> Finally, we aggregate the partial judgments' results. We then synthesize these judgments into a final pass-or-fail decision (e.g., "we found adverse media about this person"), a detailed report about what we found, and a human-readable explanation of the thought process that customers can incorporate into their audit logs.</p></li></ol><h3>Benefits of this Approach</h3><p>This approach, building on the agent "tool" (or what we call "check") as a primitive noun in our system, has enabled us to move faster while reliably delivering value to our customers.</p><p>Here are some of the benefits we have observed so far:</p><p><strong>1. Modularity and Reusability:</strong> We can extend and independently evaluate each component by building AI tools as decoupled, reusable blocks. This modular approach allows us to maintain flexibility and adaptability while ensuring high reliability. Each subsystem, whether extracting information from documents or performing complex reasoning tasks, functions independently and can be upgraded or replaced without affecting the entire system.</p><p><strong>2. Scalability and Parallelization:</strong> Our workflows can handle large volumes of records efficiently by parallelizing tasks. For example, we make tens or hundreds of concurrent LLM calls to inexpensive models to extract information in screening news articles. This ensures we can perform complex workflows that may take a human 30 minutes to an hour, in minutes.</p><p><strong>3. Accuracy and Reliability:</strong> We independently evaluate each part of the workflow to ensure accuracy. For instance, the "data loader" fetches and extracts information, and the "check" uses that data to pass judgment. We have datasets and experiments in each part of the pipeline, so a team member can improve extraction, test, evaluate, and measure performance. At the same time, someone else does the same with the judgment component.</p><p><strong>4. Cost-Effectiveness:</strong> Our approach allows us to optimize using LLMs based on task complexity. We use inexpensive models for more straightforward extraction tasks and reserve more robust models for complex reasoning and judgment tasks. This strategic use of resources helps us control costs while maintaining high-quality results.</p><p><strong>5. Transparency and Auditability:</strong> We provide a clear and human-readable explanation of the thought process behind each decision. This transparency is crucial for audit logs and compliance reporting, ensuring our customers can trust and verify the results. The final aggregation and decision-making process includes detailed reports outlining the reasoning and data used, making the process transparent and auditable.</p><h3>Where we think AI agents are useful</h3><p>While compliance workflow automation may not be the ideal use case for AI agents today, they have potential in other areas. One successful application is in customer support, where AI agents answer questions based on a knowledge base, reducing the likelihood of failure. Companies like Intercom and Klarna have seen success with this approach. Additionally, AI agents can be effective when strong feedback loops are present, such as coding agents that must verify their code runs correctly before proceeding.</p><h3>Conclusion</h3><p>Our journey at Parcha has taught us several critical lessons about building production-ready products using AI. It has also taught us the value of moving quickly and not shying away from change. In just a few months, we completely changed our approach to focus on reliability and accuracy instead of autonomy, resulting in workflows that are now three times more reliable. </p><p>We recognized that building reliable agentic behavior with large language models (LLMs) would have been a massive endeavor, requiring most of our resources and focus. Yet we knew we could use LLMs and many of the concepts behind agents today and provide significant value to our customers.</p><p>Today, our focus on modular and scalable LLM-powered workflows has enabled us to deliver reliable, efficient solutions with unit economics that make sense to Parcha and our customers. We have accelerated our customers' largely manual processes and enabled them to grow faster, and since we can move fast, we can quickly improve our product based on their feedback. Ultimately, they did not ask us to build them an AI agent; they wanted us to solve a problem effectively and efficiently.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI by Parcha! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Interview: Scaling AI Agents for Real-World Tasks with AJ Asver, CEO of Parcha]]></title><description><![CDATA[A recent interview AJ Asver did with Brett Gibson on his podcast High Bit on the experience of bringing AI agents into production]]></description><link>https://operatorsguidetoai.substack.com/p/interview-scaling-ai-agents-for-real</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/interview-scaling-ai-agents-for-real</guid><pubDate>Fri, 02 Feb 2024 15:01:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/zCGWDWCTYkE" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi Hitchhikers and Parcha fans,</p><p>Happy New Year! We hope 2024 is off to as a great a start for you as it is for us here at Parcha. Here&#8217;s a recent interview Parcha CEO AJ Asver did with Brett Gibson, Managing Partner at Initialized Capital and host of the High Bit podcast. In Brett&#8217;s podcast, he does deep dives with technical founders on the engineering challenges they had to overcome to build their startup and product. In the interview AJ shared his experience at Parcha with automating KYB manual reviews using AI and specifically LLM-powered AI Agents. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI by Parcha! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Check out the interview below!</p><p>P.S. If you want to learn more about how Parcha can help your company onboard more customers, faster, with stronger compliance, sign-up for a demo here: </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://parcha.ai/demo&quot;,&quot;text&quot;:&quot;Request a Demo&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://parcha.ai/demo"><span>Request a Demo</span></a></p><div><hr></div><h2>Interview with Brett Gibson</h2><div id="youtube2-zCGWDWCTYkE" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;zCGWDWCTYkE&quot;,&quot;startTime&quot;:&quot;1s&quot;,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/zCGWDWCTYkE?start=1s&amp;rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>You can also listen to the interview here:</p><p>Apple Podcasts:</p><div class="apple-podcast-container" data-component-name="ApplePodcastToDom"><iframe class="apple-podcast " data-attrs="{&quot;url&quot;:&quot;https://embed.podcasts.apple.com/us/podcast/scaling-ai-agents-for-real-world-tasks-with-parcha-ceo/id1691952503?i=1000643283621&quot;,&quot;isEpisode&quot;:true,&quot;imageUrl&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/podcast-episode_1000643283621.jpg&quot;,&quot;title&quot;:&quot;Scaling AI Agents for Real-World Tasks with Parcha CEO AJ Asver&quot;,&quot;podcastTitle&quot;:&quot;High Bit&quot;,&quot;podcastByline&quot;:&quot;&quot;,&quot;duration&quot;:2001000,&quot;numEpisodes&quot;:&quot;&quot;,&quot;targetUrl&quot;:&quot;https://podcasts.apple.com/us/podcast/scaling-ai-agents-for-real-world-tasks-with-parcha-ceo/id1691952503?i=1000643283621&amp;uo=4&quot;,&quot;releaseDate&quot;:&quot;2024-01-28T19:33:51Z&quot;}" src="https://embed.podcasts.apple.com/us/podcast/scaling-ai-agents-for-real-world-tasks-with-parcha-ceo/id1691952503?i=1000643283621" frameborder="0" allow="autoplay *; encrypted-media *;" allowfullscreen="true"></iframe></div><p>Spotify Podcasts:</p><iframe class="spotify-wrap podcast" data-attrs="{&quot;image&quot;:&quot;https://i.scdn.co/image/ab6765630000ba8ae1acec9fe80e7bdc1550c676&quot;,&quot;title&quot;:&quot;Scaling AI Agents for Real-World Tasks with Parcha CEO AJ Asver&quot;,&quot;subtitle&quot;:&quot;Initialized Capital&quot;,&quot;description&quot;:&quot;Episode&quot;,&quot;url&quot;:&quot;https://open.spotify.com/episode/0tXV7ft10d10fXbX1hYDFe&quot;,&quot;belowTheFold&quot;:true,&quot;noScroll&quot;:false}" src="https://open.spotify.com/embed/episode/0tXV7ft10d10fXbX1hYDFe" frameborder="0" gesture="media" allowfullscreen="true" allow="encrypted-media" loading="lazy" data-component-name="Spotify2ToDOM"></iframe><p>Google Podcasts: <a href="https://podcasts.google.com/feed/aHR0cHM6Ly9hbmNob3IuZm0vcy9lM2FjYjhhYy9wb2RjYXN0L3Jzcw/episode/OWY2NDVmYTItMzBkYS00NmQzLWE1MTUtZjAyMzVkYzA3ZjM0?sa=X&amp;ved=0CAUQkfYCahcKEwjYtPOx74qEAxUAAAAAHQAAAAAQCg">Scaling AI Agents for Real-World Tasks with Parcha CEO AJ Asver</a></p><p></p><p></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI by Parcha! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Building AI agents in production]]></title><description><![CDATA[What we've built so far, lessons learned, and what's next]]></description><link>https://operatorsguidetoai.substack.com/p/building-ai-agents-in-production</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/building-ai-agents-in-production</guid><dc:creator><![CDATA[Miguel Rios Berrios]]></dc:creator><pubDate>Mon, 16 Oct 2023 15:00:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey Hitchhikers,</p><p>It&#8217;s been a while since we&#8217;ve posted an update in this newsletter as we&#8217;ve been heads down working on our new startup, <a href="http://parcha.ai">Parcha</a>. You may have also noticed a few changes in this newsletter which is now officially part of the Parcha brand. We will be cross-posting here and on our <a href="https://parcha.ai/blog">company blog</a> moving forward.</p><p>In this update, we&#8217;re going to share a deep dive on what we&#8217;ve learned as we brought our AI product from demo into production. Before we jump in, a quick reminder to subscribe to our newsletter to get updates on our work at Parcha and how AI is changing the way we live, work and play.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><p>For the past six months at Parcha, we have been building enterprise-grade AI Agents that instantly automate manual workflows in compliance and operations using existing policies, procedures, and tools. Now that we have the first set of Parcha agents in production for our initial design partners, here are some reflections and the lessons we have learned.</p><h3>What do we call an &#8220;AI Agent&#8221;?</h3><p>An agent is fundamentally software that uses Large Language Models to achieve a goal by constructing a plan and guiding its execution. In the context of Parcha, these agents are designed with certain components in mind. Let's explore them in more detail.</p><h4><strong>The agent specifications and directives</strong></h4><p>This is the agent's profile and basic instructions about how they operate. Agents have expertise in a particular topic or field; they have a role and means to do a job. They also have tools and commands they can use to perform their job. Here is a simple example of the specifications of an agent:</p><ul><li><p>Profile: <em>You are an Expert Assistant, an AI assistant trained by Parcha.ai to help users answer questions and complete tasks that require subject-matter expertise. You are assisting a Know Your Business (KYB) operations specialist in this conversation.</em></p><ul><li><p>Directives and constraints: <em>You share your thought process as you answer a question; you always use commands available to answer questions. If you are unsure how you previously did something or want to recall past events, thinking about similar events will help you remember. Do not make things up. If you don't know the answer or you don't have sufficient information, immediately say so via a clarification question.</em></p></li><li><p>Tools or commands<strong>:</strong> these can be other calls to language models (e.g. to summarize a document) or API calls to third parties (e.g. verify an address using the Google Maps API). For example, these are the commands for an agent responsible for performing address verification:</p></li></ul></li></ul><pre><code>Commands:
1. address_verification

Command description: This tool is used by agents to verify a business physical address (nor URL). The address could be physical, virtual, a PO Box, etc.

args: "tool_input": {'type': 'string'}

2. company_detailed_risk_profile_tool

Command description: This tool is used by agents to assess the risk profile of a company.

args: "website": {'title': 'Website', 'description': 'the domain or name of the company', 'type': 'string'}</code></pre><h4><strong>The scratchpad</strong></h4><p>This is a space in the prompt to the language model where the agents add results from tools and observations as they execute their plan. These observations are used as tool inputs to guide the execution plan or the final assessment.</p><p>Example of the scratchpad for an agent that performed one of the commands above:</p><pre><code>System: Command kyb_applicant_data_tool returned:

{'answer': "ACME CORP is a personal finance platform founded in 2015. It has over 250 employees and is headquartered in Chicago, Illinois. It offers investment, borrowing and spending tools to hundreds of thousands of customers and manages over $6 billion in assets. The company's website is acme.com and it can be contacted at support@acme.corp"}</code></pre><h4><strong>A Standard Operating Procedure (SOP)</strong></h4><p>An SOP is a set of instructions the agent needs to perform to complete a task. The agent uses the SOP to construct a plan using the available tools and to assess if it has all the information needed to make an assessment.</p><p>Here is an example of a simple SOP used to perform Know Your Business in a customer:</p><blockquote><p><em>When a new company is onboarding onto BaaS&#8217;s Issuing platform, we are required to complete a KYB (Know Your Business) check on the customer based on FinCEN regulations. To complete the check we must carry out each of the following steps:</em></p><p><em><strong>Company information: </strong>Get all the company information you have for the given company you wish to complete onboarding for. This information should include company address, beneficiary owners, website address etc.</em></p><p><em><strong>Verify Business Registration:</strong> Confirm that the business is registered in the state where they claim to be. This can be done by checking with the Secretary of State of the state where the business is registered or using an SoS verification tool.</em></p><p><em><strong>Verify Business Address:</strong> Confirm the business's operating address or addresses. The operating address may be different from the registered address. The business address can often be confirmed through business lookup tools like Google Places.</em></p><p><em><strong>Check Watchlists and Sanctions Lists:</strong> Verify that the business is not listed on any watchlists, sanctions lists, or are otherwise involved in criminal activity. This can be done by using a business watchlist tool.</em></p><p><em><strong>Check Business Description:</strong> Verify that the description of the business provided by the customer matches the actual business operations. This can be done by reviewing the business's website and comparing it to the description provided.</em></p><p><em><strong>Check Card Issuer Rules: </strong>Review the use case that the business describes for Using our platform and confirm that it is compliant with the card issuer rules.</em></p></blockquote><h4><strong>Final assessment instructions</strong></h4><p>Our agents have specific instructions that dictate the output of the agent. This could be an assessment of approving/denying/escalating to humans an application, writing a report based on the observations retrieved, or a specific output the user wants to action on.</p><p>Example of the Know Your Business final assessment:</p><blockquote><p>Provide Detailed Report: If any of the checks do not pass, prepare a detailed report outlining which checks did not pass and why. Provide a recommendation on whether to approve the customer or whether further review is needed based on these findings. The final answer should include the detailed report and a recommendation.</p></blockquote><h3><strong>How we started</strong></h3><p>Our initial approach to building agents was fairly naive. Our objective was to see what was possible and validate that we could build AI agents using the same instructions humans would use to perform the task. The agent was simple: we used Langchain Agents with a standard operating procedure (SOP) embedded in the agent&#8217;s scratchpad. We wrapped custom-built API integrations into tools and made them available to the agent. The agent was triggered from a web front-end through a websocket connection, which stayed open, receiving updates from the agent until it completed a task. While this approach helped us get something built quickly to get validation from our design partners, the approach had multiple setbacks we had to improve over time to prepare our agents to perform production-grade tasks.</p><p>Some of the particular setbacks of this approach include:</p><ul><li><p>Websocket connections caused many reliability issues and ended up not being the best tool for communication between our agents and their operators. We initially envisioned our agents as bi-directional, able to have a two-way conversation. In practice, after the initial interaction (when an operator asks an agent to perform a task like &#8220;do a KYB process on X customer&#8221;), the agent updates the customer until it completes the process. We may still want to correct an agent or ask for follow-up tasks. Still, there are simpler ways to handle communication for this mostly uni-directional interaction. <em>Our customers didn&#8217;t need a chatbot; they needed an agent to complete a job.</em></p></li><li><p>As we started testing our agents with our design partner&#8217;s SOPs, we realized it would take more than embedding the full text of instructions into one agent and hoping for the best. The agent would confuse tools or skip tasks. As more steps were performed and results were added to the scratchpad, the context window would be full of noise that the agent would not parse well. Finding the SOP or parsing results that would be useful in the next steps became a challenge.</p></li><li><p>Similarly, we relied on the scratchpad as a simple means of &#8220;memorizing&#8221; information. Many times, the agent would not pick up the right piece of information from it and run a tool more than once so that it would gather input for a new step. This would make the agent way slower than it should be and inefficient. &nbsp;</p></li><li><p>Our agents may take minutes to perform a complex task. Some tasks require doing OCR in multi-page documents, web crawling to research a particular topic, or calling multiple domain-specific APIs to cross-correlate information before concluding the SOP. Our initial approach had no recovery mechanism: if, after 3-4 minutes, a step would fail, the connection to the customer would drop, and the agent would have to be rerun. This, even for a POC, was a pretty poor customer experience.</p></li><li><p>LLMs are stochastic, and as such, they can hallucinate, causing the agent to pick a tool that doesn&#8217;t exist or provide incorrect input to a tool. This would result in the workflow breaking and the task erroring out before completion.</p></li><li><p>Finally, we were custom-building each agent without putting much thought to reusability. Tools were tightly coupled with their agents, and for the most part, each new agent we built required a new set of tools and API integrations we had to build from scratch.</p></li></ul><h3>Lessons learned</h3><h4>Agents as async, long-running tasks. </h4><p>We now run our agents as long-running processes asynchronously. A web service can still trigger agents, but instead of communicating bi-directionally through WebSockets, they post updates using pub/sub. This helped us simplify the communication interface between the agent and the customer. It also made our agents more useful beyond a synchronous web service. Agents can still provide real-time status through <a href="https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events">server-sent events</a>. They can still request actions from the customer, like asking for clarifications or waiting to receive a missing piece of information. Furthermore, agents can be triggered through an API, followed through a Slack channel (they start threads and provide updates as replies until completion), and evaluated at scale as headless processes. Since agents can be triggered and consumed through REST (polling and SSE), our customers can integrate them with their workflows without relying on a web interface.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!61J8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 424w, /__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 848w, /__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 1272w, /__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!61J8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png" width="1456" height="872" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:872,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Screenshot 2023-10-11 at 10.18.58 PM.png&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Screenshot 2023-10-11 at 10.18.58 PM.png" title="Screenshot 2023-10-11 at 10.18.58 PM.png" srcset="/__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 424w, /__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 848w, /__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 1272w, /__u/substackcdn.com/image/fetch/$s_!61J8!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4662c42b-74c9-49e7-8729-bc28a34f8878_2318x1388.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Overview of Parcha&#8217;s agent architecture (October 2023)</em></p><h4>Divide and conquer</h4><p>After evaluating multiple real-world SOPs and shadow sessions with design partners, we realized instructions could be decoupled into multiple SOPs. We developed the ability for agents to trigger and consume the output from other agents. Now, instead of adding a very complex SOP into one agent&#8217;s scratchpad, we have a coordinator &lt;&gt; worker model. The coordinator develops an initial plan using the master SOP and delegates subsets of the SOP to &#8220;worker&#8221; agents that gather evidence, make conclusions on a local set of tasks, and report back to the coordinator. Then, the coordinator uses all the evidence workers gather to develop a final recommendation.</p><p>For example, in a KYB process, it&#8217;s common to perform tasks like verifying the applicant's identity, verifying if a certificate of incorporation provided as a PDF is valid, and checking if the owners are on any watchlist. These are multi-step processes (checking the certificate involves performing OCR in the document, validating it, extracting information from it, and comparing it with the information provided by the applicant (usually fetched from an API). Instead of having one agent performing all these steps, a coordinator triggers a worker agent to perform each and report back. The coordinator would decide if - per the SOP - the customer should be approved or not. Since each agent has its scratchpad, this helped us steer them to complete tasks with less noise in the context window.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!S6aW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!S6aW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg" width="1456" height="851" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:851,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;d9576c8c4807257007bf58ed0a8f22694a40907da323fa852d8a773a2927b0dc6f9619a4f43e8c8e9bd35d9c10b30ea5c1e1b838451479c5e5fd78785b42dab1b96695358a0dd4b6a69c4086a5373ac8df6e06dd71ac5ca878d5c0824db9bfa3ff9c7c4e.jpeg&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="d9576c8c4807257007bf58ed0a8f22694a40907da323fa852d8a773a2927b0dc6f9619a4f43e8c8e9bd35d9c10b30ea5c1e1b838451479c5e5fd78785b42dab1b96695358a0dd4b6a69c4086a5373ac8df6e06dd71ac5ca878d5c0824db9bfa3ff9c7c4e.jpeg" title="d9576c8c4807257007bf58ed0a8f22694a40907da323fa852d8a773a2927b0dc6f9619a4f43e8c8e9bd35d9c10b30ea5c1e1b838451479c5e5fd78785b42dab1b96695358a0dd4b6a69c4086a5373ac8df6e06dd71ac5ca878d5c0824db9bfa3ff9c7c4e.jpeg" srcset="/__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!S6aW!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e3e9426-9fbe-4713-970e-d37bb192ec0d_2412x1410.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Context windows on single agent vs. coordinator/workers.</em></p><h4><strong>Don&#8217;t rush to judge</strong></h4><p>We faced a similar challenge when asking our Agents to verify information from a document. In the KYB example above, the customer enters information (e.g., company name) into a form and provides some evidence about it (e.g., an incorporation document). Our initial approach to a verification task using an LLM was to enter the text extracted from the document using OCR and the self-attested information into the context window and ask the LLM to verify if the information matched. Since documents are long and full of potentially irrelevant information, the results of such a task were not great. We made this significantly better by separating the extraction process from judgment. We now ask the LLM to extract the relevant information from the document (e.g., is it valid? What&#8217;s the company being mentioned, and which state/country is it incorporated in?). Then in a second trip to the LLM, we ask it to compare the information (removing the unnecessary elements from the document) from the self-attested information (e.g., the company information entered in the form). This not only worked way better but also didn&#8217;t increase token count or execution time significantly since the second step has a significantly reduced amount of prompt tokens in it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!aLg5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!aLg5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg" width="1456" height="673" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:673,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;8ba7ecd0f0d7645f7deba8cb3f0a5a8001cb1c7bcd03b1d983dee6b462a5b4f489be60222348e4001fef2b71874bdd765c41f8cb612a4886891e12361d964ccb47a409bbdaf9ef002541c0e8f0ebaa8ce93f7fa6b7bb5a61695a87318de595e4f6e115ed.jpeg&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="8ba7ecd0f0d7645f7deba8cb3f0a5a8001cb1c7bcd03b1d983dee6b462a5b4f489be60222348e4001fef2b71874bdd765c41f8cb612a4886891e12361d964ccb47a409bbdaf9ef002541c0e8f0ebaa8ce93f7fa6b7bb5a61695a87318de595e4f6e115ed.jpeg" title="8ba7ecd0f0d7645f7deba8cb3f0a5a8001cb1c7bcd03b1d983dee6b462a5b4f489be60222348e4001fef2b71874bdd765c41f8cb612a4886891e12361d964ccb47a409bbdaf9ef002541c0e8f0ebaa8ce93f7fa6b7bb5a61695a87318de595e4f6e115ed.jpeg" srcset="/__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!aLg5!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3c5c037-6ebb-4b77-ac5b-1cb7da8d1242_2524x1166.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Context window length, the complexity of the extraction, and judgment</em></p><h4><strong>Keep it simple</strong></h4><p>The coordinator &lt;&gt; worker agent model introduced a challenge: since each agent has a scratchpad, how do we share information? We had agents who needed the same information for two tasks, and initially, they both had to execute the same tools. This was inefficient. The obvious answer was to incorporate memory. There are a lot of mechanisms for memory in agents or LLM apps out there: using vector DBs, adding a subset of the previous steps into the scratchpad, etc. Rather than complicating our stack or risking polluting the scratchpad with noise, we decided to leverage an in-memory store we were already using for communication, Redis. &nbsp;We now tell the agent pieces of information they have available (the keys to the information in Redis) and built our tool interface to incorporate pulling inputs from the in-memory store. By only adding relevant memory to the scratchpad, we save on LLM tokens, keep the context window clean, and ensure any worker agents pick the right information whenever needed.</p><pre><code>Human: Here are memories that you can provide to tools if they need them. These memories SHOULD NOT impact what step you pick next in the plan

{'identity_verification_api_full_name': 'John Doe', 'data_loader_tool_application_documents': '[{"name": "ID_VERIFICATION-ID-272905.pdf", "url": "http://example.com/file.pdf"}], 'digifi_data_loader_tool_self_attested_full_name': 'John Doe'}

...

"command": {

 &nbsp; &nbsp;"instruction": "Verifying identity document",

 &nbsp; &nbsp;"name": "identity_verification",

 &nbsp; &nbsp;"args": {

 &nbsp; &nbsp; &nbsp;"doc_url": "http://example.com/file.pdf",

 &nbsp; &nbsp; &nbsp;"self_attested_values": {

 &nbsp; &nbsp; &nbsp; &nbsp;"full_name_1": "John Doe" &nbsp;

 &nbsp; &nbsp; &nbsp;}

 &nbsp; &nbsp;}

 &nbsp;}</code></pre><p><em>Example of relevant memory injected into the prompt as needed</em></p><h4><strong>Mistakes happen</strong></h4><p>Each Parcha agent interacts with multiple homegrown libraries and third-party services. Since the workflows are complex, there are cases where a step fails. An HTTP connection to an API may return an internal error, or the agent may choose the wrong input for a tool. Given how our agents were configured, any error would cause the agent to fail with no opportunity to recover. We now treat agents as asynchronous services with multiple failover mechanisms. They are queued and executed by worker processes (we use RQ for this), and we get alerted when they fail. More importantly,  we are now leveraging well-typed exceptions in our tools to feed them back to the agent. If a tool fails, we send the exception name and message to the agent, allowing it to recover independently. This has helped us reduce the number of catastrophic failures significantly.</p><pre><code>System: Command digifi_data_loader_tool returned:

Unable to execute command due to an error:

1 validation error for DataLoaderInput

display_id &nbsp;field required (type=value_error.missing)

Human: The tool returned an error. If the error was your fault, take a deep breath and try again. If not, escalate the issue and move on.

Example of self-correction prompt sent to the agent with exception name and message</code></pre><h4><strong>Building blocks</strong></h4><p>After taking multiple weeks to develop each of our first set of agents, we decided to focus on reusability and speed of building. We want to enable our customers to build their agents quickly. To that end, we developed our agent and its tools interface, focusing on composability and extensibility. We now also have enough workflows from design partners to understand which building blocks we need to invest in. For example, many of our customers do some document extraction. We now have a tool that can easily be applied to any document extraction task. We built this tool once but already use it in most workflows. Our customers can use it to extract information from an incorporation document and validate its veracity or to calculate an individual's income from a pay stub. The work to adapt the tool to one specific, new workflow is minimal.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KxwU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KxwU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg" width="1456" height="563" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:563,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;2fcf8eb6f497ab211e176a81e22a1765e652876f67a62cde91d513e1e06a08d88b65ae073c8d1a92090d715bf4f25805d4dabb7cb668f892b6182da1426839b1443c102f44146e3d44cd05d4d7638c0436d9f59bb9c3a2ce22d4c7ddd967d37feac675bf.jpeg&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="2fcf8eb6f497ab211e176a81e22a1765e652876f67a62cde91d513e1e06a08d88b65ae073c8d1a92090d715bf4f25805d4dabb7cb668f892b6182da1426839b1443c102f44146e3d44cd05d4d7638c0436d9f59bb9c3a2ce22d4c7ddd967d37feac675bf.jpeg" title="2fcf8eb6f497ab211e176a81e22a1765e652876f67a62cde91d513e1e06a08d88b65ae073c8d1a92090d715bf4f25805d4dabb7cb668f892b6182da1426839b1443c102f44146e3d44cd05d4d7638c0436d9f59bb9c3a2ce22d4c7ddd967d37feac675bf.jpeg" srcset="/__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!KxwU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52681e18-52c3-4451-9a55-5dfc80f1804a_2274x880.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Our document extractor tool can be easily extended into multiple use cases</em></p><h3>What's Next</h3><ul><li><p>Webhook triggers to run an agent and perform actions on its final assessment, enabling end-to-end automation for our customers.</p></li><li><p>Implementing our in-house robust agent benchmarks, assessing the capabilities of our agents to <strong>P</strong>lan, <strong>E</strong>xecute, be <strong>A</strong>ccurate, and <strong>R</strong>easoning (PEAR, we&#8217;ll write about it).</p></li><li><p>Deploy agents and tools as micro-services. By training our agents to create and perform complex execution plans (essentially <a href="https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/dags.html">DAGs</a>), we could orchestrate them as asynchronous micro-services, which will increase composability and enable agents to use language-agnostic tools while enabling our tools to be compatible with agents outside of Parcha.</p></li></ul><h3>Want to build with us?</h3><p><a href="https://www.parcha.ai/jobs?ashby_jid=7ecc03c6-f1e3-47d2-aec6-c45e5ce7bf19">We are hiring a full-stack founding engineer in San Francisco</a>. This person will be responsible for leading the architecture and development of the platform powering our AI Agents. They will partner with our founders and AI engineers to build bleeding-edge technology for real enterprise customers. Learn more about this role here and contact miguel (@miguelriosEN on X or miguel@parcha.ai).</p>]]></content:encoded></item><item><title><![CDATA[Interview: Supercharging your team with Coda AI | David Kossnick]]></title><description><![CDATA[A deep dive into Coda's new AI features with David Kossnick, PM @ Coda]]></description><link>https://operatorsguidetoai.substack.com/p/interview-supercharging-your-team</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/interview-supercharging-your-team</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Fri, 23 Jun 2023 15:59:16 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/127943518/a682c0809124428d7fa1d9f5f530ff51.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Hi Hitchhikers,</p><p>I&#8217;m excited to share another interview from my podcast, this time with David Kossnick, Product Manager at Coda. Coda is a collaborative document tool combining the power of a document, spreadsheet, app, and database.</p><p>Before diving into the interview, I have an update on Parcha, the<a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/why-we-founded-parcha?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web"> AI startup I recently co-founded</a>. We&#8217;re building AI Agents that supercharge fintech compliance and operations teams. Our agents can carry out manual workflows by using the same policies, procedures, and tools that humans use. We&#8217;re applying AI in real-world use-cases with real customers and we&#8217;re hiring an <a href="https://www.linkedin.com/jobs/view/3618931436/?trackingId=+xwDt0djQk2a+LKmJkL9fQ%3D%3D">applied AI engineer</a> and a <a href="https://www.linkedin.com/jobs/view/3630921234/?trackingId=93RS15kUQiKgT6ZShRswTw%3D%3D">founding designer</a> to join our team. If you are interested in learning more, please email <strong><a href="mailto:founders@parcha.ai">founders@parcha.ai</a></strong>.</p><p>Also don&#8217;t forget to subscribe to <a href="http://hitchhikersguidetoai.com">The Hitchhiker&#8217;s Guide to AI:</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><p>Now, onto the interview...</p><div><hr></div><h2><strong>Interview: Supercharging your team with Coda AI | David Kossnick</strong></h2><div id="youtube2-hu88ikwifkw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;hu88ikwifkw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/hu88ikwifkw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>I use Coda daily to organize my work, so I was thrilled to chat with David Kossnick, the PM leading Coda&#8217;s AI efforts. We discussed how Coda built AI capabilities into their product, and their vision for the future of AI in workspaces, and he gave me some practical tips on how to use AI to speed up my founder-led sales process.</p><p>Here are the highlights:</p><ul><li><p><strong>The story behind Coda&#8217;s AI features:</strong> Coda started by allowing developers to build &#8220;packs&#8221; to integrate with their product. A developer created an OpenAI pack that became very popular, showing Coda the potential for AI. At a hackathon, Coda explored many AI ideas and invested in native AI capabilities. They started with GPT-3, building specific AI features, then gained more flexibility with ChatGPT.</p></li><li><p><strong>Focusing on input and flexibility:</strong> Coda designed flexible AI to work in many contexts. They focused on providing good &#8220;input&#8221; to guide users. The AI understands a workspace&#8217;s data and connections. Coda wants AI to feel like another teammate&#8212;able to answer questions but needing to be taught.</p></li><li><p><strong>Saving time and enabling impact:</strong> Coda sees AI enabling teams to spend less time on busywork and more time on impact. David demonstrated how Coda&#8217;s AI can summarize transcripts, categorize feedback, draft PRDs, take meeting notes, and personalize outreach.</p></li><li><p><strong>Tips for developing AI products:</strong> Start with an open-ended prompt to see how people use it, then build specific features for valuable use cases. Expect models and capabilities to change. Focus on providing good "input" to guide users. Launching AI requires figuring out model strengths, setting proper expectations, and crafting the right UX.</p></li><li><p><strong>How AI can improve team collaboration: </strong>David shared a practical example of how AI can help product teams share insights, summarize meetings and even kick-start spec writing.</p></li><li><p><strong>Using AI for founder-led sales:</strong> David also helped me set up an AI-powered <a href="https://coda.io/@aj-asver/supercharging-founder-led-sales-with-ai">Coda templat</a>e for managing my startup's sales process. The AI can help qualify leads and draft personalized outreach emails.</p></li><li><p><strong>The future of AI in workspaces:</strong> David is excited about AI enabling smarter workspaces and reducing busywork. He sees AI agents as capable teammates that understand companies and workflows. Imagine asking a workspace about a project's status or what you missed on vacation and getting a perfect summary.</p></li><li><p><strong>From alpha to beta:</strong> Coda&#8217;s AI just launched in beta with more templates and resources. You can try it for free here: <a href="http://coda.io/ai">http://coda.io/ai</a></p></li></ul><p>David&#8217;s insights on developing and launching AI products were really valuable. Coda built an innovative product, and I'm excited to see how their AI capabilities progress. </p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Episode Links</h2><p>Coda&#8217;s new AI features are available in Beta starting today and you can check them out here: <a href="http://coda.io/ai">http://coda.io/ai</a>.</p><p>You can also check out the founder-led sales CRM I build using Coda here: <a href="https://coda.io/@aj-asver/supercharging-founder-led-sales-with-ai">Supercharging Founder-led Sales with AI</a></p><div><hr></div><h2>Transcript</h2><h2>HGAI: Coda AI w/ David Kossnick</h2><h2>Intro</h2><p><strong>David Kossnick:</strong> ,One of our biggest choices was to make AI a building block initially. And so it can be plugged in lots of different places. There's a writing assistant, but there's also AI, you can use in a column. And so you can use it to fill in data, you can use it to write for you to categorize, for you, to summarize for you and so forth across many different types of content.</p><p><strong>David Kossnick:</strong> Having that customizability and flexibility is really important. I'd say the other piece more broadly is there's been a lot of focus across the industry on what, how to make good output from AI models and benchmarks and what good output is and when do AI models hallucinate and lie to you and these types of things.</p><p><strong>David Kossnick:</strong> I think there's been considerably less focus on good input. And what I mean by that is like, how do you teach people what to do with this thing? It's incredibly powerful, but also writing natural language is really imprecise and really hard.</p><p><strong>AJ Asver:</strong> Hey everyone, and welcome to another episode of the Hitchhikers Guide to ai. I'm your host, AJ Asver and in this podcast I speak to creators, builders, and researchers in artificial intelligence to understand how it's going to change the way we live, work, and play. Now, You might have read in my newsletter that I just started a new AI startup</p><p><strong>AJ Asver:</strong> since starting this startup a few months ago, a big part of my job has been attracting our first set of customers. I love talking to customers and demoing our product, but when it comes to running a founder-led sales process, prospecting, qualifying leads, And synthesizing all of those notes can be really time consuming, and that's exactly why I decided it was time to use AI to help me speed up the process and be way more productive with my time.</p><p><strong>AJ Asver:</strong> And to do that, I'm gonna use my favorite productivity tool, Coda. Now, if you haven't heard of Coda, it's a collaborative document editing tool that's a mashup of a doc, a wiki, a spreadsheet, and a database.</p><p><strong>AJ Asver:</strong> In this week's episode, I'm joined by David Kossnick, who's the product manager that leads Coda's AI efforts. David's going to share the story behind Coda adding AI to their product. Show us how their new AI features work, and give me some tips on how I can use AI in Coda.</p><p><strong>AJ Asver:</strong> By the way, I've included a template for the AI powered sales CRM I built in the show notes, so you can check it out for yourself.</p><p><strong>AJ Asver:</strong> But before I jump into this episode, I wanted to share a quick update on my new startup At Parcha, we're on a mission to eliminate boring work. Our AI agents make it possible to automate repetitive manual workflows that slow down businesses today.</p><p><strong>AJ Asver:</strong> And we're starting with FinTech in compliance and operations. Now, if you're excited by the idea of working on cutting edge autonomous AI and you're a talented applied AI engineer or designer based in the Bay Area, we would love to hear from you. Please reach out to <a href="mailto:founders@parcha.ai">founders@parcha.ai</a> if you wanna learn more about our company and our team.</p><p><strong>AJ Asver:</strong> Now, let's get back to the episode. Join me as I hear the story behind Coda's latest AI features in the Hitchhikers Guide to AI.</p><p><strong>AJ Asver:</strong> hey David, how's it going? Thank you so much for joining me for this episode.</p><p><strong>David Kossnick:</strong> It's going great. Thanks for having me on today.</p><h2>What is Coda?</h2><p><strong>AJ Asver:</strong> I am so excited to, go deeper into Coda's AI features with you. As I was saying at the beginning of this episode, I've been using Coda's AI features for the last month. It's been kind of a preview and it's been really cool to see, it's capable of. I'm already a massive Coda fan, as you know. I used it previously at Brex. I used it to organize my podcast and my newsletter, and most recently it's kind of running behind the scenes at our startup as well for all sorts of different use cases. But in this episode, I'd love to jump in and really understand why you guys decided to build this and what really was the story behind Coda's AI tools and how it's gonna help everyone be more productive.</p><p><strong>AJ Asver:</strong> So maybe would you describe and what exactly it does?</p><p><strong>David Kossnick:</strong> Coda was founded with a thesis that the way people work is overly siloed. So if you think about the most common productivity tools, you have your doc, you have your spreadsheet, and you have your app. And these things don't really talk to each other. And the reality is often you want a paragraph and then a table, and then another paragraph, and then a filter view of the table, and then some context in an app that relates to that table.</p><p><strong>David Kossnick:</strong> And it's just really hard to do that. And so you have people falling back to the familiar doc, but litter with screenshots and half broken embeds. So Coda said, what if we made something where all these things could fit in one doc and they worked together perfectly? And that's what Coda is.</p><p><strong>David Kossnick:</strong> It's a modern, document that allows you to have a ton of flexibility and integrate with over 600 different tools, uh, for your team.</p><p><strong>AJ Asver:</strong> Yeah, I think that idea of Coda of being able to one, integrate with different tools would also be both a doc that can become a table and then have a mashup of all this different type of data is something I've really valued about it. I think, especially when I was at Brex and we used to run our team meetings on Coda, it was really great to be able to have like the action items really formatted well in the table, but also have the notes and more freeform and then combine that with kind of follow ups.</p><p><strong>AJ Asver:</strong> And we even had this crazy table my product team where we would post like weekly photos and that's like really hard to do or in an organized way in a doc, and you'd never wanna do that in a spreadsheet. So, um, I love the fact that Coda enables you to combine all that different type of data together. So, Coda has that. And then it also has packs, which you mentioned too, right? And these are these integrations that allow you to like take data from lots of different places and put it all together.</p><h2>Story behind Coda AI</h2><p><strong>AJ Asver:</strong> And from what I understand, Coda's AI products started as a pack, right? It was like this pack that someone had, Coda had built more as kind of like a hack project to get people to use the OpenAI capabilities inside Coda.</p><p><strong>AJ Asver:</strong> And I'd actually tried this too, and that's kind of how you guys decided to actually build this into a real, uh, native integration.</p><p><strong>David Kossnick:</strong> Yeah, totally. I'd say, you know, the first bet Coda made was on packs as a platform. And so maybe about a year ago now, we released an SDK where anyone can build a pack for Coda in their browser, in JavaScript and compile it and publish it. And so it can pull data from places, push data to places. And it's really been incredible to see people do this for all sorts of use cases we never even thought of.</p><p><strong>David Kossnick:</strong> And it made possible people, starting to experiment with AI in a much more effortless manner. And so someone did a kind of weekend project and published the first OpenAI pack and it really took off starting to see it get used for all sorts of different cases. And it got us inspired for thinking, Hey, you know what?</p><p><strong>David Kossnick:</strong> If we did something native that you didn't have to think about authenticating to external services, what if it could go much deeper in what context it had about your doc and your workspace in order to help you save time?</p><p><strong>AJ Asver:</strong> One of the things I really loved here is kind of a product management thing is like seeing this nascent behavior right, happening on the platform and then deciding that it was something worth investing in further. So what point did you guys decide like, oh, this is becoming big enough or popular enough where we should make an investment in, and how was that decision made?</p><p><strong>AJ Asver:</strong> Is that like a decision that like the CEO makes or is it more like kind of bubbled up from the bottom where like there was a team that saw this happening and was like, hey, we'd like to invest in this further. I'm really curious. Like give us a bit of like the inside baseball of how that happened.</p><p><strong>David Kossnick:</strong> There were a few moments. Uh, the first moment was kinda the weekend hackathon where a few people, I think it actually started <a href="x-apple-data-detectors://21">on Friday afternoon</a>. Some of the DALL-E API had been released by OpenAI, and they really wanted to start generating images in a doc. And so it took like two hours, uh, to basically create a new pack from scratch and have it fully workable inside of a doc.</p><p><strong>David Kossnick:</strong> And then we had a weekend blitz to basically ship it on Product Hunt. And Reid Hoffman, who's on Coda's board, big fan of OpenAI, was actually kind enough to hunt it on Product Hunt for us. Um, a bunch of really cool templates people had quickly built for it. And then about a month and a half later, we had a company-wide hackathon. This must have been back in December. And there was a ton of enthusiasm on about AI, partially from, um, the OpenAI pack. And so we explored like a dozen different ideas. Um, I think we won the People's Choice Award, the company votes on all the different hackathon pitches at the end of it.</p><p><strong>David Kossnick:</strong> And then, uh, at January we had, a board meeting and we showed off some of the, thoughts from the hackathon as well as some of, what people had already been doing in the community. And the board was really excited about it, and so we started up a much bigger effort within the company.</p><p><strong>AJ Asver:</strong> It's a really cool story to hear that, you know, what started as like a project that became like a hackathon, that, that became like a pack that was put together really quickly and just kind of an experiment then became like this big investment for the company. Right? And now when I look at the product, which, you know, you're gonna talk a bit more about, I can see how like it's, it could really become like a core kind of primitive of the Coda experience, just like a table and a dock and a canvases as well.</p><p><strong>AJ Asver:</strong> So, that to me is like really inspiring, especially for other folks that are like working in product that, companies that are like Coda stage that you can really like come up with these experimental ideas and they can end up becoming products. And also I think it was really encouraging to see that you guys kind of explored AI and like integrating into the product pretty early, right?</p><h2>ChatGPT and Coda AI</h2><p><strong>AJ Asver:</strong> Like this was like pre-chat g p t hype, like, or just like as ChatGPT was becoming popular, but before GPT-4 came out, Christmas at least, I, I feel like the hype was just starting to simmer them, but not nec necessarily boil over like it is right now. talk me through what it was like developing the AI product into what it is today and how you guys kind of built into the product and, and how it works.</p><p><strong>David Kossnick:</strong> I look back on those early days and it's sort of, uh, amazing how much chaos there was in the market. You know, ChatGPT had just come out and it was incredible, and we were dying to get our hands on a chat based API but at the time, the only backend available to us was GPT-3.</p><p><strong>David Kossnick:</strong> They hadn't released an API for ChatGPT. There was nothing else really at that quality level on the market. And so I remember going and chatting with every major developer, API, AI company, every platform and being like, Hey, like, what have you got? What can we use? And we had a bunch of different ideas on UX, but we were kind of bottlenecked on what's possible.</p><p><strong>David Kossnick:</strong> So we started by building something with GPT-3 actually. And we said, okay, Chat's gonna come at some point. We don't know if it's a month out or a year out. Uh, we, you know, we gotta start moving now. Um, and we did a few experiments on it. Some just straight outta the box. And some that were very specific.</p><p><strong>David Kossnick:</strong> So actually one of my favorite was Coda has a formula language. It's incredibly powerful. People love it. It's also, um, a little complicated for people who are starting to learn it. It's a lot like Google Sheets or Excel. And so we had said, what if you could have natural language that just turned into a Coda formula?</p><p><strong>David Kossnick:</strong> And so we, um, we collected a data set for that. We actually crowdsourced it within the company. We took all of the company's internal staging environment, data of quota formulas, and we had people annotate what the natural language equivalent was. and we fine tuned GPT-3 for it.</p><p><strong>David Kossnick:</strong> And we built a little thing that would basically, you know, convert text to formulas. We were like, wow, that's actually pretty good. We realized, you know, some of the hard cases are really hard, but some of the average cases it does quite well on. Um, but it was definitely a mode where, because there was no sort of generic chat backend, we had to think like, feature by feature, what would we do for this exact scenario?</p><p><strong>David Kossnick:</strong> What are the prompts we would create, uh, and so forth. Um, and we got, you know, decently far down that path. Uh, at which point, you know, ChatGPT's API, which was called Turbo 3.5, was released and unlocked kind of a whole set of other use cases.</p><p><strong>AJ Asver:</strong> I think for people that, you know, may have forgotten by now, cause it happened so quickly, right? GPT wasn't available through chat until November. So you basically had to just provide a prompt and it would do a completion, but it wasn't fine-tuned in the same way it was right now. It wasn't, um, basically it wasn't as good at following instructions right as it is now. And so you had to do a lot more work to get it working, and then of course chat landed. Did it feel like you kind of were given this like gift where you'd be like, oh, this is gonna make it way easier. And then how did that kind of lead to where the product ended up?</p><p><strong>David Kossnick:</strong> was definitely a gift. We were super excited. It's also, as I'm sure you know, is like, a very double-sided, uh, sword. Um, you know, prompt engineering is hard. It's brittle. You sort of make a tweak and move. Sideways in, forward and backwards all simultaneously for different set of things.</p><p><strong>David Kossnick:</strong> Um, and so there's definitely a new muscle on the team as we moved into sort of turbo and GPT-4 on, how do we really evaluate which things it's doing well on, which is doing poorly on how do we make it really good for those use cases, both by setting user expectations and by, by changing the input, when we actually want GPT-3 with something fine-tuned.</p><p><strong>David Kossnick:</strong> And so it sort of opened up a whole kinda, new worms in the problem space, which was super exciting. And I think one of the things that, that got me really revved up about, uh, you know, Coda specifically as a uniquely great surface for AI is there's so many different ways people use Coda. So many different personas, so many scenarios.</p><p><strong>David Kossnick:</strong> It's an incredibly flexible tool. And so having a backend like, ChatGPT is really, really useful for a fallback. For any sort of long tail and unusual, surprising request, cuz it does really well at the random thing. And one thing we've discovered is, you know, for the very narrow set of most common things, it does pretty well too, but not as good as the more specialized thing.</p><h2>Making Coda AI work for lots of use cases</h2><p><strong>AJ Asver:</strong> So as you were talking about, Coda being used for lots of different use cases, I noticed that because there's so many different templates in the gallery and so many different ways Coda has been used, there's like pretty big community of Codens, right, that are building these different types of Coda docs. How did you think about it when it came to adding AI into Coda to make sure it's versatile, versatile enough to be used in many different ways. much did that impact kind of the end design and the user experience?</p><p><strong>David Kossnick:</strong> you know, One of our biggest choices was to make AI a building block initially. And so it can be plugged in lots of different places. So you'll see as we get to a demo a bit later, there's a writing assistant, but there's also AI, you can use in a column. And so you can use it to fill in data, you can use it to write for you to categorize, for you, to summarize for you and so forth across many different types of content.</p><p><strong>David Kossnick:</strong> Having that customizability and flexibility is really important. I'd say the other piece more broadly is there's been a lot of focus across the industry on what, how to make good output from AI models and benchmarks and what good output is and when do AI models hallucinate and lie to you and these types of things.</p><p><strong>David Kossnick:</strong> I think there's been considerably less focus on good input. And what I mean by that is like, how do you teach people what to do with this thing? It's incredibly powerful, but also writing natural language is really imprecise and really hard.</p><p><strong>AJ Asver:</strong> Mm-hmm.</p><p><strong>David Kossnick:</strong> We had a user study early on. I remember it was super surprising.</p><p><strong>David Kossnick:</strong> Someone asked, our AI block how much money was in its bank account when it was in the person's bank account. And I was just blown away that they, it felt so knowledgeable and powerful. They assumed. Know that even though they never authenticated their bank account, they just like forgotten. It just felt like something that they, that we would expect it to do.</p><p><strong>David Kossnick:</strong> Um, and so how do you remind people sort of the universe of what's possible or not, or what it's good at or not? We have something very simple. You know, a lot like Google and Auto Complete as you type, you get suggestions underneath it. But a surprising amount of effort went into that piece in particular.</p><p><strong>David Kossnick:</strong> What do we guide people towards? Which specific types of prompts, how do we make the defaults really good? How do we expose the right kinds of levers to show people what's possible and make them think about those? Um, and I think as an industry, we're still pretty early on that for these large language models, like I think we're gonna see a wave of innovation on how do you teach and inspire people how to interact to to have good input in order to get the good output.</p><p><strong>AJ Asver:</strong> So we've talked a lot about that story behind Coda and AI, and it's really interesting to hear how you guys developed kind of the thesis around it and put into the product. Um, I think for folks that aren't familiar with Coda especially, I'd love to just jump in and for you to show us a little bit about how Coda works with AI and, and what it can actually do.</p><h2>Demo: superchaging your team with Coda AI</h2><p><strong>David Kossnick:</strong> That sounds great. Yeah. Maybe I can walk you through a quick story about a team working together in a team hub with AI. And so this team hub is a place where different functions come together on a project, and have shared context.</p><p><strong>David Kossnick:</strong> So, a super common way people use Coda is collecting feedback. Um, all sorts of feedback, support, tickets, customer feedback, uh, sales calls. Um, and so we have lots of integrations that do this. They pull in Zoom transcripts and content looks a lot like this. It's really rich.</p><p><strong>David Kossnick:</strong> There's so much context in here, but it's really hard to turn this into something that's valuable for the whole organization. Um, I've spent many hours of my life writing up summaries and tagging things, and so wouldn't it be great if AI could just do this for me? Uh, and so here's a quick example. I'm gonna ask AI to summarize this transcript in a sentence, and here we go.</p><p><strong>AJ Asver:</strong> That was really cool. And I think what was really magical for me about it is not just that you can summarize, because obviously you can take the transcript, you can put it in GPT-4 and summarize. And there are other tools that do summaries too. But I think what's magical is when you combine that AI in the column, but with the magic of Codar's existing integration.</p><p><strong>AJ Asver:</strong> So like you have connected it to, I think it's Zoom, right? There's like a zoom</p><p><strong>AJ Asver:</strong> pack that's already outputting all the transcripts. So you've automated that bit and then you create this formula that basically runs on that column and then every time a new transcript comes in, I presume it just automatically summarizes it.</p><p><strong>AJ Asver:</strong> So that like piece of like connecting all those dots together, that's why I love Coda and that's why I think this is a really cool example of where kind of Coda shines.</p><p><strong>David Kossnick:</strong> That's exactly right. And one of my favorite pieces here for dealing with large amounts of data is just categorizing things. There's so many times I'm going through a large table of data, picking a category, so wouldn't it be great if based on this transcript I could just automatically tag. What type of feature request it was and boom, there it is.</p><p><strong>AJ Asver:</strong> where's it getting those feature requests from,</p><p><strong>David:</strong> Yeah,</p><p><strong>AJ Asver:</strong> tags? Is it like kind of making those or?</p><p><strong>David Kossnick:</strong> In this case, I already have a, a schema for what are the types of feature requests that I've seen, and so it's just going ahead and tagging all those things.</p><p><strong>AJ Asver:</strong> That, that's a really interesting feature too there, by the way, because now what you're doing is you're taking kind of the open-ended feature tagging problem where really GPT could like generate any feature tag at once and you're constraining it with this select list. And that's another good example of where if you did this in ChatGPT, you may end up with lots of different types of feature tags, but by constraining, you now end up with this format when you can now I, I assume go and like, organize these transcripts by feature tag because they're like actual each of those little data chips, right? And so it's now like segmentable, like</p><p><strong>David Kossnick:</strong> Yeah, and one of the cool things about it is you'll see this one is blank. That's actually intentional. That means it, it either couldn't find a match or didn't know what a match was. As there's plenty of cases where there's, you know, there's no real feature request or doesn't really know what to do with it, and it just won't tag anything either.</p><p><strong>AJ Asver:</strong> Very cool. Very cool. Okay, what What else you got?</p><p><strong>David Kossnick:</strong> So very common. Maybe the support team does an amazing job summarizing all the tickets, the feedback that's coming in, even tags things for you. And then as a PM you have to sit and read it all and think about it and say, okay, how should this influence my roadmap? Wouldn't it be nice if you could use AI to get started?</p><p><strong>David Kossnick:</strong> And so imagine you say, create a PRD for new image editor based on the problems in all this user feedback. Here we go.</p><p><strong>AJ Asver:</strong> Okay. No need for PMs</p><p><strong>David:</strong> Yay.</p><p><strong>AJ Asver:</strong> what are you gonna do after? this goes into production?</p><p><strong>David Kossnick:</strong> Well, of course it's a first draft. You know, you should always, uh, proofread. You should always change it. And so maybe AI can help with that too. Say, make this a little longer, um, and have it help me edit this PRD before I send it off. Um, there we go.</p><p><strong>AJ Asver:</strong> One of the things I love about this is for me, often when I was PMing, it's that cold start problem. It's like you are at like Tuesday, you know, you've got your no meeting Wednesday and you've gotta write this PRD in time for like a review deadline on Thursday because it's gonna go in product review on Friday.</p><p><strong>AJ Asver:</strong> Right? And you've gotta start it and you just keep putting it off cuz you've got back to back meetings. And then you get to Wednesday and you're staring at a blank screen and then you're like, Oh, maybe I need to go check my email. Right. I just think like more than anything else, this will just solve that cold start problem of just getting something down that you can start iterating on so you can just make progress faster.</p><p><strong>AJ Asver:</strong> And now what would've taken you three or four hours to get from like blank screen to PRD first draft is now probably gonna be a couple of hours because you've got that first version, you can kind of iterate on it from there. So I think this is gonna be a huge help for PMs. And, the cool thing about it is you're taking all this structured data right, from different places and bringing it into one place.</p><p><strong>AJ Asver:</strong> And one of the things we often did when we used Coda at Brex is like we would have, you know, like you had like kind of place that had customer feedback, Or you'd be aggregating different feature ideas in like a brainstorming, section and then you're kind of bringing them in here and turning them into a PRD.</p><p><strong>AJ Asver:</strong> So that's pretty cool.</p><p><strong>David Kossnick:</strong> Thanks. So another super common scenario is you have a team meeting at Coda. We do these inside of a doc with structured data, which I really love. They let people vote on which agenda item to talk about first, and you can even have notes attached to it here about what's happening. But again, what do I do after the meeting?</p><p><strong>David Kossnick:</strong> As a PM I spend a ton of time writing up next steps, but oh yeah, I can do that for me. That's awesome. Uh, one of the other things I do all the time is write up summaries. Um, what if instead I could ask ai, AI to do that too? And then of course I can send that out to Slack direct from Coda.</p><p><strong>David Kossnick:</strong> So outreach. One that we've already started doing at Coda is personalizing messages to key accounts. Um, and so let's say we have this, uh, launch message about this new image editor feature. We wanna tailor it based on the title and the company.</p><p><strong>David Kossnick:</strong> Uh, we can go ahead and get a, an easy first draft to start with here. Boom. Let's say we don't like one of these. I'm just gonna refresh this one. We'll get another example. Um, and maybe I want to go in and change this a little bit. Um, hope your family is doing well and using our Gmail integration. I'll just go ahead and send that email.</p><p><strong>AJ Asver:</strong> And that's actually gonna send that email now to the, person that you wanna do outreach to. And you just basically generated that email based on kind of the context of the person's job</p><p><strong>David Kossnick:</strong> Totally. And one of the really fun parts of this is it's super flexible. So imagine you had another column that's like family member names, or hobby or other things like that that you're not gonna find in, your favorite, uh, sales tool. Um, AI is really good incorporating that context. So having all your stuff here in the team hub, being able to pull that context in, um, is really powerful.</p><p><strong>AJ Asver:</strong> Does it generate tables too? How how does</p><p><strong>David Kossnick:</strong> You know, we showed a simple example here, which generates the table of target audiences. Um, but actually one, maybe I'll just show real quick, um, that I've been doing in my personal life. So very different kind of template. Um, we use meal planning, so I'm a vegetarian, so meal planning can be sometimes a bit of a pain with two kids, uh, making sure everyone gets exactly what they need.</p><p><strong>David Kossnick:</strong> Um, and so I made a quick template. Everyone in my family can go in and add. Their favorite ingredients and get out both a bunch, a bunch of ideas as well as, uh, specific dishes. And so this is an interesting case where it's just a normal table here. And I'll say, uh, AJ, what's your favorite ingredient?</p><p><strong>AJ Asver:</strong> Well, you know what? My kids love</p><p><strong>David Kossnick:</strong> Broccoli. Wow. Nicely done. cool. And then I'll go ahead and, uh, or auto update it here. Uh, so it has a bunch of different meal ideas and yeah, let's say I'll take, uh, spinach and cheese quesadilla. Um, and I'll add that one in here. Spinach and cheese quesadilla. Um, and AI is gonna start generating, uh, what ingredients are needed for that as well as a, a recipe and about how long it would take.</p><p><strong>AJ Asver:</strong> That's a really, really awesome hack. I think I need to start doing that, as well to just use it to generate ideas for, meals as well and meal planning. That's, that's very, very cool. And this is like a good example of also like how it can be helpful in like a personal setting too.</p><p><strong>AJ Asver:</strong> Right?</p><p><strong>David Kossnick:</strong> For sure.,</p><h2>Using AI to speed up lead qualification</h2><p><strong>AJ Asver:</strong> I was not gonna bring you onto this podcast without you helping me with something. So one of the reasons I wanted to bring you on here is because now that we've started this startup and we're trying to get everyone to be very excited about our AI agents, I am doing what's called founder-led sales, which is where we kind of work out, okay, who are the companies that we wanna target?</p><p><strong>AJ Asver:</strong> And then at the very early stages we're trying to find like five design partners that we can work with. And they're kind of like enterprise FinTech companies, similar to like Brex where I used to work. And my job is to do the sales. Cause there's only two of us right now and Miguel's busy building the product. And so I gotta work out which companies will be a good fit and then we work out, okay, how do we get an intro to them? Maybe it's through an investor, maybe it's through a mutual contact. Maybe we, outbound to them because you know, maybe it's a company that we worked at before or something. And so I've been trying to work out how to use Coda to help me do this.</p><p><strong>AJ Asver:</strong> And I was wondering if you might be able to help me make my Coda CRM better with AI.</p><p><strong>David Kossnick:</strong> Sounds amazing. Let's do it.</p><p><strong>AJ Asver:</strong> I'm very excited about this because this is gonna save me a lot of time.</p><p><strong>AJ Asver:</strong> Over here is my Coda CRM and very high level. I have this list that I've got, and I think I got it from Crunchbase, um, with just a bunch of FinTech companies. And some of them might be a good fit, some of them might not. Oh, that's definitely not a FinTech company. Let's take that one out. Um, and I, and I'm trying to work out kind of one, how big are these FinTech companies? Either they kind of size where they would be an interesting fit for our, for our product, we generally try to focus on kind of growth stage FinTech companies, and then two, Would they qualify or not? Based on do we think they might need the product we're trying to build? And so I was wondering, maybe the first thing is I have this kind of list of the different, um, tiers of, of customers that I might want or the different types of companies that I might want to organize 'em into.</p><p><strong>AJ Asver:</strong> So for example, there might be growth stage FinTech companies, there might be early stage FinTech companies. And I wanna work out how I can take these FinTech companies that I have in this list and kind of categorize them by that tier. maybe we could, start there.</p><p><strong>David Kossnick:</strong> Yeah, sounds great. What kind of categories are you thinking about?</p><p><strong>AJ Asver:</strong> I think maybe we can start with kind of early stage FinTech growth stage FinTech and a financial institution. So maybe first I need to add like a column, right? Is that correct?</p><p><strong>David Kossnick:</strong> Yeah, sounds good.</p><p><strong>AJ Asver:</strong> Okay. So we can go do that, have a column after this one. And then do I make this a select list? Is that</p><p><strong>David Kossnick:</strong> Yeah, that's exactly right. So you could add, uh, select list</p><p><strong>AJ Asver:</strong> Great.</p><p><strong>David Kossnick:</strong> sort of a, a list of types or a list of items.</p><p><strong>AJ Asver:</strong> Okay, so then let's, let's pick a few different stages. So it's growth stage, FinTech, um, maybe series C +, and then maybe early stage FinTech Seed Series C and then maybe kind of traditional financial institution. Cause I think there's a few of those in this list. And then I know there's probably some crypto startups in here too, so maybe I'll put like, crypto is another one too. Okay. So I've got my select list. What do I need to do next?</p><p><strong>David Kossnick:</strong> Um, cool. So the, uh, actually before you forget, yeah, maybe rename the, the column there.</p><p><strong>AJ Asver:</strong> So we call this customer segment. Great.</p><p><strong>David Kossnick:</strong> So next, if you could just go to, uh, right click on that column and do add AI. Um, so the question is, what do you think would determine it? Is it, you know, one thing we could do is just give it the company name. many that are well known, it might do actually a pretty good job.</p><p><strong>David Kossnick:</strong> Um, you could also try giving it the, um, the description as well. And you could say basically, you know, what, what category does this belong to?</p><p><strong>AJ Asver:</strong> I kind of like the idea of doing it by description. That seems like a good way of doing it. So then would I just do this like, um, pick is like kind of pick the correct customer segment based on the company description</p><p><strong>David Kossnick:</strong> Mm-hmm.</p><p><strong>AJ Asver:</strong> provided?</p><p><strong>David Kossnick:</strong> That's perfect.</p><p><strong>AJ Asver:</strong> that work?</p><p><strong>David Kossnick:</strong> And then if you do at, and then the column name should pull it in for you.</p><p><strong>AJ Asver:</strong> Okay. And then let's see, what's this company</p><p><strong>David:</strong> I think it was info.</p><p><strong>David Kossnick:</strong> Yeah.</p><p><strong>AJ Asver:</strong> Yeah. Great. And then how does it know which one to pick? Oh, it just kind of knows from what's available, right? Is that how it</p><p><strong>David Kossnick:</strong> Yeah. The AI knows in a select list or look up the AI knows to use that set of options.</p><p><strong>AJ Asver:</strong> Ah, smart. All right. So let's, let's do it. Okay. Wow. so it's already started. Um, but what happens if we think one of them might be wrong? So, for example, Affirm. I don't know if maybe it needs a bit more data.</p><p><strong>AJ Asver:</strong> So I'm wondering if a good one could be that we look at funding.</p><p><strong>AJ Asver:</strong> And then last funding round will probably help. Um, and then maybe like employee count. I feel like this should all help work it out.</p><p><strong>David Kossnick:</strong> Yeah. Sounds great.</p><p><strong>AJ Asver:</strong> Okay. So then if we do that, um, then we need to go back to here and go to this segment thing. We close this. Okay. And then we go, okay. Provided. And then this is kind of my prompt engineering now I usually like to do it like this. So company, maybe we'll do company name cuz that might be helpful too, right?</p><p><strong>AJ Asver:</strong> Might know some of this stuff already. Company description, just info. Last funding round. Total funding. Okay. And now let's give this another try.</p><p><strong>David Kossnick:</strong> All righty.</p><p><strong>AJ Asver:</strong> Um, Fill.</p><p><strong>David Kossnick:</strong> Wow. Very different result.</p><p><strong>AJ Asver:</strong> Great. That's way better. Right . Now it like correctly categorized firm, which is awesome. um, and Agora and Airwallex Crypto for Anchorage, which is awesome. Apex, this is great. Now I have some qualification.</p><p><strong>AJ Asver:</strong> I just ran the qualifier and it's actually started qualifying these leads, which is really cool. And I guess in the case where it doesn't have information, it'll tell me, if the company's a traditional financial institution, might work that out. So I guess I could go through these and check these all later. But the cool thing is I now have some qualified beads that I can start, um, working out which ones to focus on for my sales.</p><p><strong>David Kossnick:</strong> That's awesome.</p><p><strong>AJ Asver:</strong> And reach out.</p><p><strong>David Kossnick:</strong> Nice.</p><h2>Coda's long term strategy with AI</h2><p><strong>AJ Asver:</strong> I am really excited to start, um, using my new. AI powered lead qualifier cuz I'm a one man sales team and I'm already thinking of other ideas. Like earlier when you showed me generating emails, I think I'm gonna start trying that as well, like generating outbound emails or even introduction emails to, to the right customers, based on who the right investor connection is.</p><p><strong>AJ Asver:</strong> And that stuff always just takes a lot of time. And I'll say like when I'm in front of a customer and pitching, I'm like the most energized. And I think through the sales process when I'm like in the CRM and trying to write emails and then qualifying leads, I'm like the least energized. So having AI helped me there is gonna be really great to, to gimme a bit of a boost. Um, I'm curious, where do you guys see the long-term kind of impact of AI being for Coda and how you guys think about it for long-term strategy and also when will this be available for other people to try out?</p><p><strong>David Kossnick:</strong> We are launching, the beta very, very soon. The beta will be very similar, but there'll be a lot more templates and resources, um, and we'll be letting in way more folks and so would love feedback.</p><p><strong>David Kossnick:</strong> Please try it out. And in terms of where we're going with AI, I'm really excited about it. I mean, as I mentioned at the start, Coda has kind of a uniquely great place for AI because it brings so many different pieces of a team together in one place. And that context is so helpful for AI. And so things that get me really excited is being able to ask your workspace a question about how a project is going, and it's able to just answer because it knows all the tools you're connected to, all the notes people are taking about every project.</p><p><strong>David Kossnick:</strong> Um, you know, imagine you come back from being off for a week on vacation, you're like, what did I miss? And you get the perfect summary and you can drill into more detail and granularity on anything that you're curious about. Um, you know, imagine it felt more like having a teammate when you used AI.</p><p><strong>David Kossnick:</strong> Who was able to engage on, um, on your projects and give feedback and have details. And obviously it's not exactly a teammate and you're gonna have to, to teach it about each, each case. Um, but I think the, the vision of having less busy work and more impact is really exciting.</p><p><strong>AJ Asver:</strong> I, I think that potential of, AI in Coda that you mentioned beyond just a single doc, but when you're using AI to really run your company is gonna be a really, really, really powerful one. Because if you have kind of sold on using AI as your wiki and to run your projects and your docks and your spreadsheets, then you guys basically have all the information as you mentioned that's required in kind of an internal knowledge base to answer these complex questions.</p><p><strong>AJ Asver:</strong> Like what's the status of a project, who is working on what parts of the project? And so. personally, that's, uh, something I'm very excited before because we use Coda for everything and we're a very small team right now. But I imagine as we get bigger and more people get involved, being, being able to ask those questions and being able to answer 'em is really, really cool.</p><p><strong>AJ Asver:</strong> And so it's gonna be in beta soon, which is really, really awesome. And how are you, as a PM thinking about, you know, launching this Beta and what it means to kind of bring this into production? And do you have any tips for other PMs that are working on AI products? Because you are, you, you've been at this now for six months, right? Um, which puts you, I would say, in like the early percentage of PMs that are working with, you know, GPT. Um, curious what your tips are for, for other folks trying to bring a product from that hackathon to a beta and then a GA.</p><p><strong>David Kossnick:</strong> know, I'd say a lot of people start, like we did, of just sort of, you know, throwing AI behind a prompt, a prompt box in your product. I think it's a great starting point to learn and see what people are doing. And I think as you develop a sort of deeper sense of what use cases are really valuable, um, you're gonna build something much more specific for them.</p><p><strong>David Kossnick:</strong> One of the really fun things about working with AI is you don't know exactly how it's gonna go. You know, the models are a moving target in terms of what they're really good at, what they're really bad at, what the, what people expect of them as well. And so, uh, you know, a very common process I've seen a lot of teams do, us included is have some generic prompt box, some entry point into AI in their product and sort of see what people use and gravitate towards in their scenario.</p><p><strong>David Kossnick:</strong> I think that's a great starting point to learn. As you have deeper conviction, building purpose-built things for those use cases is really, really valuable cuz it's just a lot less effort at the end of the day. Writing prompts is really helpful for a really specific thing, but it's a lot of work for something you just want done quickly.</p><p><strong>David Kossnick:</strong> And so some of the stuff we've been working on at the last month or two is, exacto knives, really specific things to open up exactly what you want in one scenario. and people have been loving them, which is great.</p><p><strong>AJ Asver:</strong> I have one feature request, which is please automatically work out what type of column I need. Based on like the description of the column That seems like an easy one for AI to, to solve.</p><p><strong>David Kossnick:</strong> You know, that's an interesting one cuz we, we actually do do it today, but in a subtle way actually. And yeah, this is like really in the weeds, but the text column type in Coda is actually an inferred column type. If you put a number in there or any other kind of structure piece of data, it will actually be able to, you know, operate on that data as if it were that type.</p><p><strong>AJ Asver:</strong> Well, I am very, very much looking forward to seeing more AI in my Coda docs and also very excited to see where this all goes. think what you guys are doing with Coda and AI is really, really, really, really cool and also just very helpful. Like it saved me a lot of time and I think other people too. Um, once it's available in beta, I'm sure there'll be lots of new use cases that we haven't even seen yet.</p><p><strong>David Kossnick:</strong> I did have one last question for you as well, which is, uh, you know, since you're working on AI every day on agents in particular, I'm curious how you think about sort of the future of agents in productivity. Like, what do you imagine agents are gonna do in finance and in every vertical, uh, you know, in collaboration with people?</p><h2>AI agents</h2><p><strong>AJ Asver:</strong> That is something I spend a lot of time thinking about, um, in between like building actual agents. And right now I would say that we are vastly underestimating what's possible based on what we've been able to achieve at departure with our agents and the fact that they can just follow a set of instructions from a Google doc and carry out, you know, a very complex compliance task, like,</p><p><strong>David Kossnick:</strong> It's not a Coda doc? AJ!</p><p><strong>AJ Asver:</strong> That's true. We should be using Coda docs. Yes. A Coda doc, sorry. being able to follow the instructions and just carry out a task with the tools given. That's like a very novel and pretty awesome, uh, thing to be able to do because now you can automate a lot of repetitive tasks that have to be done very manually today. And so I imagine we're gonna see a lot of. This idea of intelligent automation as we've been calling it, where you aren't just doing the kind of robotic process automation or like workflow automation that you did before where you're connecting things together with conditional rules. But you're actually now using essentially prompt engineering to automate a task fully with lots of different steps and lots of different tools being used just by providing the instructions.</p><p><strong>AJ Asver:</strong> In a way, you're, you are used to doing it already if you are solving this manually with the team, which is essentially a user manual, a set of, um, kind of procedures. And so you are gonna think about the many different use cases for that, not just in finance, but in processing forms in healthcare or insurance or in all kinds of other places where these manual, workflows are being done.</p><p><strong>AJ Asver:</strong> I think it's gonna free up people to have a lot more time to do more interesting creative, strategic work, than doing this more repetitive kind of, Tedious work that exists today. So we're very excited to see where that goes. And we're at the very early stages of it today, but, um, I I think it's gonna move very quickly.</p><p><strong>David Kossnick:</strong> Amazing.</p><h2>Wrapup</h2><p><strong>AJ Asver:</strong> David, really appreciate it. you for taking the time to demo Coda AI, for talking to us a little bit about the kind of story behind how this feature came about and then for also helping me be more productive, with my Coda workflows as well. I'm really excited to see, product launch soon, beta and just a reminder for everyone where can they find, AI if they wanna sign up for it.</p><p><strong>David Kossnick:</strong> <a href="http://coda.io/ai">coda.io/ai</a>.</p><p><strong>AJ Asver:</strong> Awesome. Well thank you very much David, and you everyone for listen in to this week's episode of the Hitchhikers Guide to ai and I will see you on the next one.</p><p>&#8203;</p>]]></content:encoded></item><item><title><![CDATA[Why we founded Parcha ]]></title><description><![CDATA[A deeper dive into why we're building AI agents to supercharge compliance and operations teams in fintech at Parcha]]></description><link>https://operatorsguidetoai.substack.com/p/why-we-founded-parcha</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/why-we-founded-parcha</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Thu, 25 May 2023 14:23:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ddbcd9ab-d046-4d33-b5e8-01ffabeffeac_946x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi Hitchhikers!</p><p>First, welcome to the new subscribers who hitched a ride with us this week! Last week I shared the news that I&#8217;m starting a new startup Parcha in the AI space and said I would follow up with more details soon. As promised, this week, I wanted to share the story behind Parcha and why we&#8217;re so excited about building AI agents to supercharge businesses.</p><p><em>Before I jump into this week's post, I have an exciting announcement:</em></p><p><em>On June 6th, The Hitchhiker&#8217;s Guide to AI and Parcha are co-hosting an AI dinner with our friends over at <a href="http://hireflow.ai">Hireflow.ai</a>. If you&#8217;re based in SF and are either actively building something in AI or are interested in meeting other engineers exploring AI, we would love to meet you! We have limited spots, and are prioritizing AI hackers/engineers. If you&#8217;re interested, RSVP for the dinner here ASAP and we will get back to you.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://lu.ma/hgtai&quot;,&quot;text&quot;:&quot;RSVP&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://lu.ma/hgtai"><span>RSVP</span></a></p><p>On to this week&#8217;s update&#8230;</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WanU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WanU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2501594,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WanU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2d3b9b7-d0eb-487b-95a3-414147c7d80f_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">MidJourney generated digital passion fruit art</figcaption></figure></div><p>In this post, I will talk about the problem our new startup Parcha is solving, how AI enables a solution, and how we plan to differentiate from the competition. Let&#8217;s jump in!</p><h3>Fintech&#8217;s scaling problem</h3><p>When I started this newsletter <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/coming-soon?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">just under six months ago</a>, my goal was to help people that was to share my insights with other people that were new to the space, like myself. My inspiration for writing was from Paul Graham, the founder of YC and prolific essay writer who once said that &#8220;writing is a form of thinking.&#8221; Writing about AI would help me think about the space and where I might fit in. </p><p>You may not know that before writing this newsletter, I had been in product management for over a decade, most recently in fintech at Coinbase and Brex. I joined Coinbase as the second PM hired to work on the consumer app at the height of the 2017 crypto bull run. During that time, I saw Coinbase 5X in size, experienced the depths of the crypto winter, and learned how money movement, compliance, and fraud prevention worked at a hypergrowth fintech company. </p><p>Even though you might think of Coinbase as a crypto company, most transactions on the platform at the time were between fiat and crypto. That meant money moving between credit cards and bank accounts through the traditional financial system, which was and still is far less technologically sophisticated than you might expect. In fact, every day, there was a team of operations folks at Coinbase that would manually process incoming wires in Coinbase&#8217;s bank account to make sure the money matched up to the correct Coinbase customer accounts to fund their crypto purchases!</p><p>Later when I joined Brex, a fintech company providing corporate credit cards and cash management accounts for startups, I was surprised to see a similar story. This time it was teams of operations folks manually reviewing customer applications, making underwriting decisions for credit limits and manually processing customer&#8217;s check deposits. In fact, in the early days of Brex Cash, the bank-account replacement product I helped launch, if you wanted to deposit actual cash into your account, your best bet was it mail it to HQ!</p><p>We did try to automate as many of these manual workflows at possible at Brex. My co-founder <a href="http://linkedin.com/in/miguelriosberrios">Miguel</a> and I led a team of ~200 PMs, engineers, designers, data scientists, and cross-functional partners tasked with overhauling Brex&#8217;s onboarding, compliance, fraud, and underwriting systems. We made a lot of progress and helped the company scale its customer base 10X in a year. However, many manual workflows remained and were eventually outsourced to a third-party contractor or Business Process Offshoring (BPO) service. The challenge with these services is their agents churn a lot, they require a lot of upfront investment in training and you need to give the vendor access to internal tools and sensitive data.</p><p>Isn&#8217;t this just an excellent example of Y Combinators&#8217; mantra of &#8220;do things that don&#8217;t scale?&#8221; After all, both Brex and Coinbase were YC companies. But in reality, in the zero-interest rate environment, we were in for the last decade, this approach of having armies of operations people doing a lot of manual work <em>did scale. </em>Many fintech companies like Brex and Coinbase have grown aggressively with armies of operations people working tirelessly behind the scenes to keep money moving smoothly, mitigate risk, and manage compliance. </p><p>The problem with this approach is that, unlike software, these operations backends in the fintech industry scale linearly at best. The more customers you have, the proportionally larger your team needs to get.</p><p>Fast-forward to today, these same fintech companies are under immense pressure to lower costs and extend their runway to ride out the current economic uncertainty, ideally by becoming profitable. That usually means that headcount has been frozen or reduced even as customer growth continues. But as I said earlier, these teams scale linearly, so at some point, you hit a ceiling on how many customers you can serve. </p><p>In my post on <a href="https://www.hitchhikersguidetoai.com/p/startups-navigating-the-ai-landscape">navigating the AI landscape</a> to where I identified opportunities for AI startups, I said the following:</p><blockquote><p>&#8220;There might be industries that are both fragmented and have knowledge workers who carry out repetitive software-based workflows. These industries are probably ripe for disruption because (1) there's no dominant player that has a structural advantage, and (2) repetitive workflows can be quickly augmented with AI.&#8221;</p></blockquote><p>From my own experience, I haven&#8217;t seen any dominant players today that own workflow automation in compliance and operations in fintech and, more broadly, in finance. This is a perfect example of the type of problem that I think is an excellent opportunity for an AI startup. Then last month, a critical innovation in AI unlocked the solution: autonomous AI agents.</p><h3>AI Agents unlock intelligent automation</h3><p>The emergence of large language models, such as GPT-4, has enabled AI models to perform chain-of-thought reasoning. This means an AI model can follow multiple steps, create an internal model of what is happening in a scenario, and reason through it to get to a solution (see example below). </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!zCwm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!zCwm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg" width="1170" height="1438" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1438,&quot;width&quot;:1170,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;graphical user interface, text, application, chat or text message&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="graphical user interface, text, application, chat or text message" title="graphical user interface, text, application, chat or text message" srcset="/__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!zCwm!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603dd92e-e9e0-4005-b431-722078dee5e1_1170x1438.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This chain of thought reasoning is so impressive that we can now create autonomous AI agents that can complete tasks in sequence, using tools to solve problems, as I covered in my <a href="https://www.hitchhikersguidetoai.com/p/deep-dive-could-autonomous-ai-agents">previous posts</a> on autonomous AI agents like BabyAGI:</p><blockquote><p>&#8220;I believe there is a more profound implication for these types of iterative, LLM-powered autonomous AIs that can use other tools/models to complete tasks. It feels much closer to how humans interact with the world, whereby we have additional sensory input beyond our brain&#8217;s intelligence, and we use tools to complete our objectives.&#8221;</p></blockquote><p>Many knowledge workers today perform repetitive work that has to be done by humans due to nuances in language, analysis, and steps that are hard to automate, even if they are very repetitive. The emergence of AI agents will change how these tasks are done. AI agents can carry out tasks like manual reviews in fintech, which would take humans several minutes or hours to complete, and can now be completed in seconds. But it&#8217;s not just time being saved. AI agents are more scalable (up and down), easier to audit, faster to train, and cheaper to run than humans carrying out the same tasks. They can make existing human teams far more productive too. </p><p>At Parcha, we are excited to be at the forefront of building these agents. Our platform empowers existing operations teams by supercharging them with AI agents. This will enable them to do more work and no longer scale linearly. One person can now manage dozens, if not hundreds, of agents. Our platform is unique because how you create and deploy agents is approachable to nontechnical folks. All you have to do is give our agents the same policies, procedures, and training docs you give a new team member and access to the same tools. That&#8217;s it!</p><p>Once you have deployed these AI agents, you can monitor, manage, scale up, and scale down your digital workforce from one central dashboard. Because all the agents share context, you connect them together to complete larger, more complex tasks too, much like you run a multi-layered organization today.</p><h3>Open-source models and enterprise </h3><p>The final piece of the puzzle that makes Parcha possible is building an enterprise-grade product that prioritizes privacy and data security, a critical feature for finance use cases. The reality is, most enterprise companies are not comfortable with handing over their data to OpenAI. </p><p>Enter open-source models&#8230;</p><p>Open-source versus closed-source models are a hot topic in AI and very relevant to enterprise use cases. At an AI event I attended, Greg Brockman, CTO of OpenAI, recently shared his belief that closed-source models will continue to be most centralized and at the bleeding edge, while open-source models will be about one to two years behind. In a previous post, I also thought there would be a divergence between open-source and closed-source models:</p><blockquote><p>Over time though, open-source will probably lag behind proprietary models for three primary reasons:</p><ol><li><p>Cost to train: As models get bigger and bigger, they will become more expensive. Open-sourced models will therefore need more well-funded sponsors to develop them. Today StabilityAI might be able to afford to sponsor a GPT-3 equivalent model that costs $10M to train, but next year they might not be able to afford to support a GPT-5 equivalent that costs $100M to train.</p></li><li><p>Proprietary datasets: Well-funded companies like OpenAI license additional private data sets to improve their models. This will make it harder for open-sourced models to compete on performance and they may lag one or two generations behind proprietary models. For many customers, the price and versatility of open-sourced models will be the deciding factor. Using a lesser performant open-source model will be a path many developers take to get started cheaply before developing their own models or switching to proprietary ones once they have enough scale.</p></li><li><p>Closed research: Many contributions that have enabled open-source AI models to come from private companies, like Google&#8217;s Transformer architecture. That might change as the need for Big Tech to compete in AI takes a higher priority over publishing research.</p></li></ol></blockquote><p>After seeing the sheer pace of innovation in the open-source AI community since Meta&#8217;s release of LLAMA though, I&#8217;m now far less convinced that open-source will lag far behind proprietary closed-source models. It's no longer clear to me that just having more data at this point and more parameters will solve the problem. There may be other innovative and exciting ways that we can improve language models beyond that, which the open-source community may find and experiment with just as quickly as a closed-source community in the big research labs.</p><p>In fact, it seems increasingly likely that language models will follow a similar path to open-sourced generative image models like Stable Diffusion. This collective momentum of the AI community, along with many companies now releasing their own models like MosaicML and Databricks, is really accelerating open-source innovation. This is reflected in a recently leaked memo by a Google employee titled &#8220;We Have No Moat, And Neither Does OpenAI":</p><blockquote><p>While our models still hold a slight edge in terms of quality, the <a href="https://arxiv.org/pdf/2303.16199.pdf">gap is closing astonishingly quickly</a>. Open-source models are faster, more customizable, more private, and pound-for-pound more capable. They are <a href="https://lmsys.org/blog/2023-03-30-vicuna/">doing things with $100 and 13B params</a> that we struggle with at $10M and 540B. And they are doing so in weeks, not months.</p></blockquote><p>While close-source model creators like Google still have a huge advantage in distribution and existing ecosystems, open-source models will be an excellent option for startups developing AI products for end consumers with stricter privacy or security requirements. Furthermore, open-source models can the fine-tuned to better fit a particular vertical use case, with far fewer parameters. </p><p>Open-source models, therefore, become an excellent option for startups building for enterprises in regulated industries, like fintech and finance. Being able to control the data around the models, optimize their output, and deploy them on custom infrastructure that prioritizes privacy and security makes open-source models a great fit.</p><p>At Parcha, we&#8217;re developing fine-tuned models based on open-source LLMs that can easily be deployed in an enterprise setting. We believe this approach will be a differentiator in the long term because we will have better control over our agent&#8217;s capabilities, and most importantly, our customers will have better control over their data.</p><h3>Lastly, there is timing&#8230;</h3><p>When starting a startup, there&#8217;s one factor you often can&#8217;t control: timing. I will never forget a piece of advice I once heard from a partner at Sequoia Capital at a YC dinner I attended over a decade ago when I was doing my first startup. The advice went something like &#8220;If there&#8217;s a wave that can give your startup momentum, you should ride it.&#8221; The idea is that it&#8217;s hard enough doing startups as it is, and you should take advantage of any tailwinds you can get. AI is the most important technology wave we have seen since mobile, and I believe the magnitude of its impact will far surpass any predecessor. Coupling this with the need for enterprise companies to be under pressure to manage costs and find more efficiencies, it feels like an opportune moment for AI-powered automation. </p><p>I always knew I wanted to start a startup again at some point in the future. After spending six months on the shore watching the AI wave gaining momentum, I knew it was the right time to grab my surfboard, and luckily, my cofounder Miguel was ready to paddle out with me too.</p><div><hr></div><p>Thanks for reading this update. If you&#8217;re considering starting an AI startup, I would love to hear from you - I&#8217;m happy to share what I&#8217;ve learned so far. If you&#8217;re curious about joining an AI startup, we&#8217;re hiring at Parcha and would love to hear from you too! </p><p>Next week we will go back to our regular programming, breaking down how the latest developments in AI will change how we live, work, and play. <br><br>P.S. Don&#8217;t forget to subscribe!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p><br></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Interview: Human-level AI and AI Agents with Josh Albrecht, CTO of Generally Intelligent]]></title><description><![CDATA[An interview with Josh Albrecht, founder of Generally Intelligent and Outset Capital]]></description><link>https://operatorsguidetoai.substack.com/p/interview-human-level-ai-and-ai-agents</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/interview-human-level-ai-and-ai-agents</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Sun, 14 May 2023 14:55:04 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/119767841/3d425ee0bd2c28342d284174807a0848.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Hi Hitchhikers,</p><p>First, a huge welcome to the 100+ new subscribers that hitched a ride with us in the last two weeks. In this week&#8217;s post I&#8217;m sharing an interview I did recently with Josh Albrecht, CTO and founder of Generally Intelligent. Generally Intelligent is an AI research team focused on building general-purpose AI agents that can be safely deployed in the real world.</p><p>Before I jump into the interview, you might be wondering why I skipped a week. Well, I have some exciting news: After 6 months of exploring AI, I&#8217;ve decided to start a startup with my friend and former engineering partner at Brex, <a href="https://www.linkedin.com/in/miguelriosberrios/">Miguel Rios Berrios</a>. </p><p>Miguel and I worked together with an amazing team to grow Brex's customer base 10X in just a year by making onboarding, compliance, risk management and underwriting more automated. Despite 200+ people working on this problem there were still many workflows that required humans in the loop. We believe AI can help.<br><br>At <a href="http://Parcha.ai">Parcha</a>, we&#8217;re building enterprise-grade AI agents that supercharge your business. We are starting with fintech companies that can use our AI agents to accelerate their manual workflows in operations and compliance. Over the coming weeks, I&#8217;ll be sharing more about why we started Parcha and what our core thesis is.<br><br>We&#8217;re hiring a small team in San Francisco and San Juan, Puerto Rico. We currently have open roles for an  senior full stack engineer, applied AI engineer who has experience working with LLMs and a UX designer.</p><p>If you&#8217;re interested in learning more about what we&#8217;re doing or would like to join our team please email us via <a href="mailto:founders@parcha.ai">founders@parcha.ai</a> or reply directly to this email.</p><p>As for this newsletter, I still plan to post as often as I can because I find writing these updates to be a great way to keep up with the AI space. If you are reading (or listening) to this for the first time, please don&#8217;t forget to subscribe:</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><p>On to the interview&#8230;</p><div><hr></div><h2>Interview: AGI and developing AI Agents with Josh Albrecht, CTO of Generally Intelligent</h2><div id="youtube2-ic5s3F39edE" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ic5s3F39edE&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ic5s3F39edE?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>I&#8217;ve been spending a lot of time researching, experimenting and building AI agents at Parcha. That&#8217;s why I immediately jumped at the chance to interview AI researcher <a href="https://www.linkedin.com/in/joshalbrecht/">Josh Albrecht</a>, who is the CTO and co-founder of Generally Intelligent. Generally Intelligent&#8217;s work on AI Agents is at the bleeding edge of where AI is headed. </p><p>In our conversation, we talk about how Josh defines AGI, how close we are to achieving it, what exactly an AI researcher does, and his company&#8217;s work on AI agents. We also hear about Josh&#8217;s investment thesis for Outset Capital, the AI venture capital fund he started with his co-founder <a href="https://www.linkedin.com/in/kanjun/">Kanjun Qui</a>.</p><p>Overall it was a really great interview and we covered a lot of ground in a short period of time. If you&#8217;re as excited about the potential of AI agents as I am or want to better understand where research is heading in this space, as I am this interview is definitely worth listening to in full.</p><p>Here are some of the highlights:</p><ul><li><p><strong>Defining AGI: </strong>Josh shares his definition of AGI, which he calls Human-level AI, a machine&#8217;s ability to perform tasks that require human-like understanding and problem-solving skills. It involves passing a specific set of tests that measure performance in areas like language, vision, reasoning, and decision-making.</p></li><li><p><strong>Generally Intelligent: </strong>General Intelligence's goal is to create more general, capable, robust, and safer AI systems. Specifically, they are focused on developing digital agents that can act on your computer, like in your web browser, desktop, and editor. These agents can autonomously complete tasks and run on top of language models like GPT. However, those language models were not created with this use case in mind, making it challenging to build fully functional digital agents.</p></li><li><p><strong>Emergent behavior: </strong>Josh believes that the emergent behavior we are seeing in models today can be traced back to training data. For example being able to string together chains of thought could be from transcript of gamers on Twitch.</p></li><li><p><strong>Memory systems: </strong>When it comes to memory systems for powerful agents, there are a few key things to consider. First of all, what do you want to store and what aspects do you want to pay attention to when you're recalling things? Josh&#8217;s view is that while it might seem like a daunting task, it turns out that this isn't actually a crazy hard problem given that we know how to do this already for non AI systems.</p></li><li><p><strong>Reducing latency:  </strong>One way to get around the current latency when interacting with LLMs that are following chains of thought with agentic behavior is to change user expectations. Make the agent continuously communicate updates to the user for example vs. just waiting for to provide the answer. For example, the agent could send updates during the process, saying something like "I'm working on it, I'll let you know when I have an update." This can make the user feel more reassured that the agent is working on the task, even if it's taking some time.</p></li><li><p><strong>Parallelizing chain of thought: </strong>Josh believes we can parallelize more of the work done by agents in chain of thought processes, asking many questions at once and then combining them to reach a final output for the user.</p></li><li><p><strong>AI research day-to-day: </strong>Josh shared that much of the work he does as an AI researcher is not that different from other software engineering tasks. There&#8217;s a lot of writing code, waiting to run it and then dealing with bugs. It&#8217;s still a lot faster than research in the physical sciences where you have to wait for cells to grow for example!</p></li><li><p><strong>Acceleration vs deceleration: </strong>Josh shared his viewpoints for both sides of the argument for accelerating vs decelerating AI. He also believes there are fundamental limits to how fast AI can be developed today and this could change a lot in 10 years as processing speeds continue to improve.</p></li><li><p><strong>AI regulation: </strong>We discussed how it&#8217;s challenging to regulate AI due to the open-source ecosystem but that some form of regulation is probably required.</p></li><li><p><strong>Universal unemployment: </strong>Josh shared his concerns that we need to get ahead of educating people on the potential societal impact of AI and how it could lead to &#8220;universal unemployment&#8221;.</p></li><li><p><strong>Investing in AI startups: </strong>Josh shared Outset Capital&#8217;s investment thesis and how it&#8217;s difficult to predict what moats will be most important in the future.</p></li></ul><h3>Episode Links:</h3><p>The Hitchhiker&#8217;s Guide to AI: <a href="http://hitchhikersguidetoai.com/">http://hitchhikersguidetoai.com</a> </p><p>Generally Intelligent: <a href="http://generallyintelligent.com/">http://generallyintelligent.com</a></p><p>Josh on Twitter: <a href="https://twitter.com/joshalbrecht">https://twitter.com/joshalbrecht</a></p><h3>Episode Content: </h3><p>00:00 Intro </p><p>01:42 What is AGI? </p><p>04:40 When will we know that we have AGI? </p><p>05:40 Self-driving cars vs AGI </p><p>07:10 Generally Intelligent's research on agents </p><p>09:51 Emergent Behaviour </p><p>11:07 Beyond language models </p><p>13:17 Memory and vector databases </p><p>15:25 Latency when interacting with agents </p><p>17:13 Interacting with agents like we interact with people </p><p>19:08 Chaing of thought </p><p>19:44 What do AI researchers do? </p><p>21:44 Accelerations vs. Deceleration of AI</p><p>24:05 LLMs as natural language-based CPUs </p><p>24:56 Regulating AI </p><p>27:31 Universal unemployment</p><div><hr></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/p/interview-human-level-ai-and-ai-agents?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thank you for reading The Hitchhiker's Guide to AI. This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/p/interview-human-level-ai-and-ai-agents?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/p/interview-human-level-ai-and-ai-agents?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Enterprise AI, Augmented Employees, AGI and the Future of Work with Charlie Newark-French, CEO of Hyperscience]]></title><description><![CDATA[A deep dive into Enterprise AI with Charlie Newark-French, CEO of Hyperscience]]></description><link>https://operatorsguidetoai.substack.com/p/enterprise-ai-augmented-employees</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/enterprise-ai-augmented-employees</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Thu, 23 Mar 2023 15:26:02 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/110124695/4a9dcee68fa778c4fda95fb685d7a653.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-dPA833BOzb4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;dPA833BOzb4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/dPA833BOzb4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Hi Hitchhikers!</p><p>I&#8217;m excited to share this latest podcast episode, where I interview Charlie Newark-French, CEO of Hyperscience, which provides AI-powered automation solutions for enterprise customers. This is a must-listen if you are either a founder considering starting an AI startup for Enterprise or an Enterprise leader thinking about investing in AI. </p><p>Charlie has a background in economics, management, and investing. Prior to Hyperscience, he was a late-stage venture investor and management consultant, so he also has some really interesting views on how AI will impact industry, employment, and society in the future.</p><p>In this podcast, Charlie and I talk about how Hyperscience uses machine learning to automate document collection and data extraction in legacy industries like banking and insurance. We discuss how the latest large-scale language models like GTP-4 can be leveraged in enterprise and he shares his thoughts on the future of work where every employee is augmented by AI. We also touch on how AI startups should approach solving problems in the enterprise space and how enterprise buyers think about investing in AI and measuring ROI. </p><p>Finally, I get Charlie&#8217;s perspective on Artificial General Intelligence or AGI, how it might change our future, and the responsibility of governments to prepare us for this future.</p><p>I hope you enjoy the episode! </p><p>Please don&#8217;t forget to subscribe @ http://hitchhikersguidetoai.com</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Episode Notes</h2><h3>Links: </h3><ul><li><p>Charlie on Linkedin: <a href="https://www.linkedin.com/in/charlienewarkfrench/">https://www.linkedin.com/in/charlienewarkfrench/</a></p></li><li><p>Hyperscience: <a href="http://hyperscience.com">http://hyperscience.com </a></p></li><li><p>New York Times article on automation: <a href="https://www.nytimes.com/2022/10/07/opinion/machines-ai-employment.html?smid=nytcore-ios-share">https://www.nytimes.com/2022/10/07/opinion/machines-ai-employment.html?smid=nytcore-ios-share</a></p></li></ul><h3>Episode Contents:</h3><p>00:00 Intro </p><p>01:56 Hyperscience </p><p>04:52 GPT-4 </p><p>09:41 Legacy businesses </p><p>11:13 Augmenting employees with AI </p><p>15:48 Tips for founders thinking about AI for enterprise </p><p>20:34 Tips enterprise execs considering AI </p><p>23:49 Artificial General Intelligence </p><p>29:41 AI Agents Everywhere </p><p>32:12 The future of society with AI </p><p>37:44 Closing remarks</p><h3>Transcript:</h3><h1>HGAI: Charlie Newark French</h1><h2>Intro</h2><p><strong>AJ Asver:</strong> Hey everyone, and welcome to the Hitchhiker Guide to ai. I am so happy for you to join me for this episode. The Hitchhiker Guide to AI is a podcast where I explore the world of artificial intelligence and help you understand how it's gonna change the way we live, work, and play. Now for today's episode, I'm really excited to be joined by a friend of mine, Charlie Newark, French.</p><p><strong>AJ Asver:</strong> Charlie is the CEO of hyper science, a company that is working to bring AI into the enterprise. Now, Charlie's gonna talk a lot about what hyper science is and what they do, but what I'm really excited to hear Charlie's opinions on is how he sees automation impacting our future.</p><p><strong>AJ Asver:</strong> Both economically, but as a society, as you've seen with recent launch of G P T four and all the progress that's happening in AI, there's a lot of questions around what this means for everyday knowledge workers and what it means for jobs in the future. And Charlie, has some really interesting ideas about this, and he's been sharing a lot of them on his LinkedIn and I've been really excited to finally get him on the show so we can talk. Charlie also has a background in economics and management. He studied an MBA at Harvard and previously was at McKinsey, and so he has a ton of experience thinking about industry as a whole, enterprise and economics and how these kind of technology waves can impact us as a society.</p><p><strong>AJ Asver:</strong> If you are excited to hear about how AI is gonna impact our economy, our society, and how automation is gonna change the way we work, then you are gonna love this episode of The hitchhiker Guide to ai.</p><p><strong>AJ Asver:</strong> Hey Charlie, so great to have you on the podcast. Thank you so much for joining me.</p><p><strong>Charlie:</strong> Aj, thank you for having me. I'm excited to discuss everything you just talked about</p><p><strong>AJ Asver:</strong> maybe to start off, one of the things I'm really excited to understand is how did you end up at Hyper Science and what exactly do they do?</p><h2>Hyperscience</h2><p><strong>Charlie:</strong> Yeah, hyper Science was founded in 2014. It was founded by three machine learning engineers. so We've been an ML company for a long time. My background before hyper science was in late stage investing. Had sort of the full spectrum of outcomes there.</p><p><strong>Charlie:</strong> Some why successful IPOs, some strategic acquisitions, and then a lot of miserable, sleepless nights on some of the other areas. I found, hyper science, incredibly impressed with, their ability to take cutting edge technology and apply it to real well problems. We use machine vision, we use large language models, and we use natural language processing, and we use that those technologies to speed up back office process.</p><p><strong>Charlie:</strong> The best examples here are a loan origination, insurance claims processing, customer onboarding. These are sort of miserable long processes, a lot of manual steps, and we speed those up. With some partners taking it down from about 15 days to four hours.</p><p><strong>Charlie:</strong> So all of that data that's flowing in of this is who I am, this is what's happened, this is the supporting evidence. We ingest that. It might be an email, it might be a document. It's some human readable data. We ingest that, we process it, and then ultimately the claims administrator can say, yes, pay out this claim, or no, there's something.</p><p><strong>AJ Asver:</strong> Yeah, so what, what you guys are doing essentially is you had folks that were previously looking at these documents, assessing these documents, maybe extracting the data out of these forms, maybe it was emails, and entering those into some database, right? And then decision was made, and now your technology's basically automating that. It's kind of sucking up all these documents and basically extracting all that information, helping make those decisions. My understanding is that with machine learning, what you're really doing is you've kind of trained on this data set, right, in a supervised way, which means you've said like, this is what good looks like.</p><p><strong>AJ Asver:</strong> This is what, you know, extracting a, a, a data from this form looks like now we're gonna teach this machine learning algorithm how to do it itself. Now what what I found really interesting is that, That was kind of where we made the most advancements, really in kind of AI over the last decade, I would say.</p><p><strong>AJ Asver:</strong> Right? It's like these deeper and deeper neural networks. They could do machine learning in very supervised ways, but what's recently happened with large language models especially, is that we've now got this like general purpose AI that, you know, GPT-4, for example, just launched this. and there was an amazing demo where I think the CTO of OpenAI basically sketched on like the back of a napkin, a mockup for a website, and then he put in in GPT and it was able to like, make the code for it.</p><p><strong>AJ Asver:</strong> Right. So when you think about a general purpose, let large language model like that, compared to the machine learning, y'all are using do you consider that to be a tool that you'll eventually use? Do you think it's kind of a, a threat to like the companies that have spent the last, you know, 5, 6, 7 years, decades, maybe kind of perfecting these ma machine learning tools or, you know, I, is it something that's gonna be more like different use cases that won't be used you know, by your customers?</p><h2>GPT-4</h2><p><strong>Charlie:</strong> Open ai ChatGPT, GPT-4. Anything that's been, the technology you're speaking about has really had two fundamental impacts. There's been the technology. It's just very, very cutting edge, advanced technology. And then you've got the adoption side of it. And I think both sides are as interesting as each other.</p><p><strong>Charlie:</strong> On the adoption side, I sort of like to compare it to the iPhone that there was a lot of cutting edge technology, but what they did is they made that technology incredibly easy to use. There's a few things that Open AI has done here that's been insanely impressive. First, , they use human language. Um, humans will always assign a higher level of intelligence to something that speaks in its language.</p><p><strong>Charlie:</strong> The other thing, it's a very small thing, but I love the way that it streams answers so it doesn't have a little loading sign that goes around and dumps an answer on you. It's like, it's almost like it's communicating with you. Allow you to read in real time and it feels more like a conversation.</p><p><strong>Charlie:</strong> Obviously the APIs have been a huge addition. It's just super easy to use, so that's been one big step forward. But it's a large language model. It's a chat bot. I don't wanna underestimate the impact of that technology, but my thoughts are AI will be everywhere. It's gonna be pervasive in every single thing we do.</p><p><strong>Charlie:</strong> And I hope that chatbots and large language models aren't the limitation of ai. I'd sort of like to compare chatbots and large language. To search the internet is this huge thing, one giant use case that if you ask people what is the internet? They think it's. Google, And that's the sort of way I think this will play out with AI and the likes of a whichever large language model and chatbot wins to be the Google of that world, which at the moment appears very clearly to be open ai.</p><p><strong>Charlie:</strong> But there's some examples of stuff that. Certainly right now, that approach wouldn't solve. I'll give you a few, but the, this list is tens of thousands of use cases long. We spoke about autonomous vehicles earlier. I suspect LLMs are not the approach for that physical robotics. Healthcare detecting radiology diseases, fraud detection.</p><p><strong>Charlie:</strong> I'm sure if you put in like a fake check in front of GPT-4 right now it was written on the napkin, it might be able to say, okay, this is what the word fraud means. This is what a check looks like, but you've got substantially more advanced ai AI out there right now that looks at. , what, what was the exact check that Citibank had in nine, in 2021?</p><p><strong>Charlie:</strong> Does this link up with other patterns that should be happening with this kind of transaction and so, I think that you are gonna have more dedicated solutions out there than just sort of one chat or interface to rule them all, would be my guess. Yeah. Hyper science. There are things that chat G p t does, or g p t four does right now that we do.</p><p><strong>Charlie:</strong> Machine vision is one that's an addition that appears to be working alongside their large language model. So they're combining different technologies versus just a large language model, is my guess. I obviously don't have work ins. But we build a platform here at hyper science that builds workflows, that enriches data, validates data, compares data looks at data that's historically come into an organization that might not be accessible to sort of public chat bots or large language models.</p><p><strong>Charlie:</strong> I think the question that you sort of said at the beginning, Could we be using chat, G p t or g p T four? Absolutely. And I think that a lot of startups could, but I suspect that, that you, what you'll see here is a lot of the startups that spin up leveraging this and building something far greater from a user experience for a very specific use case versus open AI solving all the sort of various small problems along the way, if that makes.</p><p><strong>AJ Asver:</strong> Yeah, I think that makes a lot of sense. And it's one thing I've been thinking a lot about. I actually wrote a blog post recently about this as kind of how these foundational models are gonna become more and more commoditized and it's gonna create this massive influx of products built on top of it.</p><p><strong>AJ Asver:</strong> What I find really interesting is that you know, GPT, you can actually use that transformer for a lot of different things that aren't necessarily just a chatbot.</p><p><strong>AJ Asver:</strong> Right. So you mentioned the fraud use case. If you send a bunch of patents of fraud to a large transformer, its ability to actually remember things makes it very good at identifying fraud. And in fact, I was talking to a friend that, that worked at Coinbase in their most recent fraud control mechanism.</p><p><strong>AJ Asver:</strong> They went from kind of linear aggressive models to deep learning models, to eventually actually using transformers and it, and it was far more, far more effective. So I guess coming back to the question, do you see a world where instead of building these. Focused machine learning models for particular use cases like you know, ingesting documents or maybe making sense of data and extracting that data and tabulating it into a, into a database that you might and one day end up actually just having a general pre-trained transformer that you are then essentially fine tuning with one shot. Kind of tuning me. Like, this is how you extract a document for one of our clients. This is how you you know, organize this information into like loan data. Is that a world we could move in? That's probably different from where we are today and, and maybe a different world of hypo sciences too.</p><h2>Legacy businesses</h2><p><strong>Charlie:</strong> Look, it would be a very different world. I think the next five, 10 years are leveraging the, the. Of technology that OpenAI is building and maybe that specific technology, as you sort of say of commod, some commoditized layer and building. Workflows on top of that, I'll give you the, just the harsh reality of what the world looks like in reality out there.</p><p><strong>Charlie:</strong> Right? So this isn't just a single use case that I go and type something in as a consumer on on the internet at a bank in the uk they have a piece of cobalt that it was written in the 1960s that is still live in their mortgage processing.</p><p><strong>AJ Asver:</strong> Wow.</p><p><strong>Charlie:</strong> Rolling out, even just from a compliance level, any change to that mortgage processing that isn't piecemeal fashion, that doesn't about the implications, that doesn't think about customer interactions in a a week timeframe or a a year or three year timeframe is just not dealing with the reality of the situation on the ground.</p><p><strong>Charlie:</strong> These are complex process. People get fired if you take a mortgage processing system down for minutes. And they're complex. So do I think that's a possibility in the future? It's absolutely possible. I think the best use of GPT-4 right now is to go and build the extensive workflows that require a little bit of old-fashioned AI, as well as cutting edge AI to, to have an end-to-end solution for a specific problem versus assuming that we're ever gonna get something. But you just say, okay, I'd like to know, should I give this customer this mortgage?</p><p><strong>Charlie:</strong> And you get an answer back. That, to answer that question is still a very complex process.</p><h2>Augmenting employees with AI</h2><p><strong>AJ Asver:</strong> Yeah, and I think we, especially in the technology industry, especially someone like me that spends so much time thinking about, talking about reading about AI, kind of forget that a lot of these legacy businesses can't move as fast as we think. I mean, we see like Microsoft moving quickly and slack moving quickly for example.</p><p><strong>AJ Asver:</strong> But those are all like very software focused consumery businesses that you know, necessarily touching like hard cash and stuff like that where there's a lot more compliance and regulations. So that makes a lot of sense. So then what we really are thinking about is like you kind of have humans that can be, as you've put it before in some of your predictions around ai, augmented, right?</p><p><strong>AJ Asver:</strong> These, this idea of like an augmented employee that can use AI to, to help them get things done, but we're not necessarily replacing them straight away. Like, talk to me about what, what you see as a future of kind of augmented employees and, and kind of co-pilots as they're also called.</p><p><strong>Charlie:</strong> Totally. So the augmented employee is a phrase that I've been using. For about 10 years, it's been a prediction for a while now. It didn't used to be a particularly popular one. You would get a whole load of reports even from the like big consultancy groups that say these five jobs are definitely gone in five to 10 years.</p><p><strong>Charlie:</strong> That five to 10 years has come and gone over that period of time, or I'll give you a longer period of time. Over the last 30 years, we've added 30 million jobs here in the us about a on average. Obviously, it's been a bit of fluctuation. There's no good sign on a short term decision making time horizon that jobs are gonna be wildly quickly displaced.</p><p><strong>Charlie:</strong> There's very little evidence. That's my. What do I mean by short term horizon? I really mean by the when, what a large enterprise, which is what I, my company serves and what I'm interested in serving, makes decisions 5, 10, 20, maybe even as that's the sort of edge of where I think things start to really change.</p><p><strong>Charlie:</strong> Fundamentally you should make decisions around software. And AI in this case, substantially helping people do a better job. The, the, the first time I read this getting sort of a mainstream idea was about a year ago. And by mainstream, I mean outside of the tech industry New York Times wrote an article where the title was something like in the.</p><p><strong>Charlie:</strong> Fight between humans and robots. Humans seem to be winning. I think that was just a very interesting change of thought. And there was a line in there that says the question used to be, when will robots replace humans? The better question is, which I absolutely love this phrasing of it, when will humans that leverage robots replace humans that don't leverage robots?</p><p><strong>Charlie:</strong> And I think that's the right way to think about it. I, I'll give you a couple of examples. One with sort of, non-AI OpenAI technology and then chat. G PT Speci specifically, or, or G P T. Radiology is something that's been talked about for a while. This was a giant step forward where software AI could detect most.</p><p><strong>Charlie:</strong> Cancers most diseases, basic diseases better than humans could just had higher accuracy. And the prediction for five, seven years was, this is the end of radiologists. We've seen no decrease in radiologists. If you want to go and get a cancer screening now, you're gonna probably look at a six to nine month wait.</p><p><strong>Charlie:</strong> I don't have any issues, but I'm waiting for a cancer screen right now. Just a nice safety check that I want to my own benefit and cause of the sheer backlog of work. , I can't get that done. I can't get it done for a while. So is the, the future for me is in two or three years time, there's not fewer radi.</p><p><strong>Charlie:</strong> There's just much higher accuracy and much shorter wait, wait times. And maybe the future, as I say, which I'm sure we'll speak about 20 years down the line is is I can just go to a website. They can do some scan of me, and then they can give me the answer within seconds. I, I, I can't wait for that, but it's just not here today.</p><p><strong>Charlie:</strong> And I'll give you one aj one quick open AI example. When ChatGPT came out there was so many sort of, this is not ready for mainstream things that went round. And the, the way that I thought about it is, if you want ChatGPT then to write you a sublease because you want to lease your apartment and you want it to be flawless, you just want to click send to the person that's doing the sublease on. It's nowhere near medi ready for mainstream. If you wanna cut down a law legal person's work by 90% because the first draft is done. They're gonna apply local laws, they're gonna look at a few nuances. They're gonna understand the building a bit then it's so far beyond ready for mainstream. It should be used by everybody for every small use case it can. So I think it's human augmentation for a while. I think that jobs don't go away for a while, and I sort of like to compare it to the internet a little bit here, which is we use the internet today and every single thing we do and it makes our jobs substantially easier to do. It makes us more effective at them, and that's what I think the next sort of 10 years at least looks like for AI within the work.</p><h2>Tips for founders thinking about AI for enterprise</h2><p><strong>AJ Asver:</strong> The thing you mentioned there, I find to be really fascinating is this idea that, you know, we're not gonna replace humans immediately. That's not happening. But people thought that for a long time. Right. And it almost feels like with every new wave of technology, there's this new hum of people saying like, we don't replace humans, we're gonna replace humans.</p><p><strong>AJ Asver:</strong> Right. But at the same time, I, I kind of agree with you, having spent a lot of time using chat JBT and working with it, I found that it certainly augments my life, in writing My substack in fact, in this interview preparing for this interview, I actually asked it to help me think about some interesting questions to ask you based on some of the posts you'd written.</p><p><strong>AJ Asver:</strong> Because I'd read some of your posts on, on, on LinkedIn fairly regularly, but I couldn't remember all of them, so I actually asked the Bing ai chat to help me. Right. And then when you think about these especially regulated environments where you. The difference between right and wrong is actually someone's life or a large sum of money or breaking the law, then it really matters. And in that case, augmentation makes a lot of sense. Now, the reality is, AI, especially kind of large language models in building on top of open AI is a fairly low barrier to entry right now. That's why we're seeing a lot of companies in copywriting, in collaboration, in presentations, and the challenge with that is if there's an incumbent in the space, That already exists. It's very hard to beat them on distribution right now. Where I did see an interesting opportunity is exactly what you are talking about, is like going deep into a a fairly fragmented industry, which maybe has a lot of regulation or a lot of complexity, maybe disparate data systems.</p><p><strong>AJ Asver:</strong> You mentioned kind of the. 30 year old like cobalt data system, right? Like that is a perfect place where you can go in and really understand it deeply. Now, as someone that's running a company that does that, I'm curious, like what advice do you have to founders or startups? I wanna take that path of like, Hey, we're gonna take this AI technology that we think is extremely powerful, but go apply it into some deep industry where you really have to understand the ins and outs of that industry, the regulation, the, the way people operate in that industry and in the problems.</p><p><strong>Charlie:</strong> Absolutely a few thoughts. Firstly, make sure that you are trying to solve a problem. This is just general advice for setting up a business. What is the problem you're trying to solve? What is the end consumer pain point? For us here at Hyper Science, it's that people wait for their mortgage to be processed for six weeks.</p><p><strong>Charlie:</strong> No good reason why that's happening. People wait for their insurance to be insurance claims to be paid out sometimes for two years. No good reason for that to be happening. So always start with the customer pain point, and then decide does the current technology, which is AI in this case, allow you for solving it?</p><p><strong>Charlie:</strong> And then that gets you to the, does it allow you to solve for it? And what I've looked for here is, the more open AI can do it or G p t four can do it a whole load of diverse stuff, but your highest value things are gonna be what's just happening time and time again. If for us, like there is just a whole load of mortgages, that process not right now or there is just a whole load of insurance things that are processed and they're.</p><p><strong>Charlie:</strong> Relatively similar, although we think of them as different. They've got a lot of, certainly to a machine similarities. So I'd look for volume. You can think of this as your total addressable market in terms of traditional vc, non-AI speak. But this is, is the opportunity big enough? And then the, the next thing I'd look for is repetitive tasks.</p><p><strong>Charlie:</strong> So the more repetitive it is, the easier it's. You can go out and solve something really, really hard with a large language. But there's probably even easier applications that you can solve that are just super repetitive and you can take steps out. So I think that's it. Have the problem in mind.</p><p><strong>Charlie:</strong> Think about volume, think about repetitive natures, and then one of the key things, once you've got all of that set and you've got, okay, this is an industry that's right for customer pain, right, for disruption. This is definitely a good application of where AI is today. I would think about ease of use above everything.</p><p><strong>Charlie:</strong> My, my thinking is, and I spoke about this with open ai, one of the biggest things they've done is they've just taken exceptional technology, but made it so, so simple for someone that doesn't know AI to interact with. And the question I always get asked is the CEO of enterprise software, AI company is how can we upscale all of our employees?</p><p><strong>Charlie:</strong> The answer to that is you shouldn't have to. This software should be so easy to engage with that your current employees should seamlessly be able to do it. There should be, if there is rudimentary training needed needed, your software should do that. And again, I like to compare this to the internet. We use the internet day in, day out.</p><p><strong>Charlie:</strong> There has been 20 years of upskilling, but it's not really been that hard. Like I think if you took today the internet and you gave it to somebody 20 years ago, it might be a little bit advanced for. , but we've made software, internet software, so easy to work with that you don't need to know how the internet works.</p><p><strong>Charlie:</strong> The number of people that know how the internet works, even the number of people that know how to truly code a website. Absolutely tiny fraction of the number of people that use the internet to improve their daily lives. So I'd say ease of use for AI is possibly as important as the technology.</p><h2>Tips enterprise execs considering AI</h2><p><strong>AJ Asver:</strong> I love those insights by the way. And just to recap that you said go for a problem that has a lot of volume, whereas solving a real problem to end users, but there's clear volume or like, you know, a large addressable market. The other thing you mentioned was make sure it's repe repetitive tasks, like with l LM says it's temptation to go after these really complex problems, but like repetitive tasks are the ones that are most.</p><p><strong>AJ Asver:</strong> That's probably the most incentive to solve as well. Right. And then the last thing you mentioned is like, it should be really intuitive for a, for an end user to use to the point where they don't have to feel like they need to be upskilled. Now, if you are a founder or a startup going down this path, the other thing you're thinking about is like, how do you sell into these companies?</p><p><strong>AJ Asver:</strong> So maybe taking the flip side of that, if you are in the enterprise and you're getting approached by all these AI startups, they just got funded this year being like, we're gonna help you do this. We're gonna help you do that. We're gonna automate this. How do you decide when it's the right time to make that decision?</p><p><strong>AJ Asver:</strong> How do you decide? Kind of, the investment on that and whether it's worth it. Like what, what are your thoughts on that?</p><p><strong>Charlie:</strong> My thoughts on that are linked directly to the economic cycle we're in right now, which is not a pretty one. Somewhat of a maybe a mild recession, maybe the edge of a recession. And I see this from all of the CIO CEOs that we work with at the, the sort of large banks, large insurance companies.</p><p><strong>Charlie:</strong> And my suggestion is this is I tell them to create a two by two matrix. You told everyone earlier, I started my career at McKinsey.</p><p><strong>AJ Asver:</strong> Classic two by two.</p><p><strong>Charlie:</strong> Love it. Two by two matrix. On one of the ax axis is short-term ROI on one of the ax axis is long-term roi and you want to get as much into the top right as possible and as few into the bottom left as possible.</p><p><strong>Charlie:</strong> And for a y or artificial intelligence was considered ROI and not short term roi, which is a bit, they were treated by these large. As science experiments and you saw these whole, these whole roles form these whole departments form around transformation. The digital transformation officer, that is a role that just didn't exist five or 10 years ago, and these people were there to go and innovate within the organization and, and largely speaking, it wasn't wildly successful. A lot of these roles are sort of spinning down. You need to solve a business problem that the technology solves today and gives you a path to the long run. So, hyper science, we, I'll give you an example here. We add value out of the box, but we also understand where people are today and try to get them to where they want to go.</p><p><strong>Charlie:</strong> So one of our customers, 2% of what they do is process fax. I hope that they are not processing faxes in five time, and I hope that we are giving them the bridge to that, but we better be able to do that today. And also paint them a, a sort of what I refer to as a future proof journey to where they want to head.</p><p><strong>Charlie:</strong> So I think it's really about don't, don't do any five year projects. Like if a company comes and says to you things. When you say, can you do this? And they say, well, we could do that. I would run. Or if they're just saying yes to everything versus yes to something quite specific and a good startup, you know this better than anyone, aj, a good startup does something specific really well and then they build out that they have an MVP is one way of phrasing it.</p><p><strong>Charlie:</strong> They have a niche is another way of phrasing. Yeah, go and sell something today really well, and all of those sort of long tail features around it, people will forgo for a period of time whilst you build those. , but you be, add value in the in the short term as well as be building something in the long run.</p><h2>Artificial General Intelligence</h2><p><strong>AJ Asver:</strong> Yeah. That and that point you made about that kind of showing the short term value is really important, especially when you're trying to convert the kind of maybe biases around AI that exist in enterprise today, that it's, as you mentioned, kind of like a hobby project or a kind of experiment, or like this is kind of your, you know, your Moonshot kind of, kind of projects you wanna show them really, like, this is like ROI you're gonna get very quickly in the next one or two years and, and that's a really important point of it. Now all of this makes sense, the augmented employee, the co-pilot and, and like having these narrow versions of AI that are solving particular problems and, and I can see that working out, but I feel like there's this one big factor that we, that we have to think about that that could change all of this, in my view.</p><p><strong>AJ Asver:</strong> And that is artificial general intelligence. And for folks that dunno what artificial general intelligence is, or AGI is it's called. That's really what open AI is trying to achieve long term. And it's the idea of essentially having intelligence that is the equivalent of a human. And it's an ability to think abstractly and to solve a wide, broad range of problem. In a, in a way that, that, that a human does. And what that means is technically, now, if you have an AGI and let's say the cost of running that AGI I is, you know, a hundredth or a thousandth of a cost of running a human, then potentially you could replace everyone with, with, with robots or you know, AI as, as it were.</p><p><strong>AJ Asver:</strong> How does that factor into this equation? Is this so, You, you think about is it, like, what are your thoughts about it? Both, both as CEO of an enterprise company, but also as someone that's studied management and economics for the last decade? I'm really curious to, to hear where you think this is going.</p><p><strong>Charlie:</strong> I don't think there's anything unique about a soul or something that can't be replicated in the human mind. And to your point, I actually think that we, we think of AGI sometime as when I hear definitions, I hear human-like, or human level intelligence if this happens or when it happens, because it, there's no doubt it will it will be substantially smarter, incredibly quickly than a human. And you look at the difference in humans of intelligence, someone you just pick off a street, 110 I IQ or whatever level it is, versus an Einstein with 170 iq, that difference is enormous. Now, imagine that that's the equivalent of 170 iq, but it's a thousand or 10,000 or whatever it is. I think you will get to the point where if you have AGI extremely quickly, you will. Be far beyond them not being able to do any job. There will be absolutely zero they can't do</p><p><strong>Charlie:</strong> now I don't see that today. I, my best guess of a time horizon is post 20 years sub a hundred years. That's a nice vague timeframe, but that's sort of how I think about it. 20 years is your classic sort of decision making timeframe and for, for someone. Building or someone running an enterprise software company, it's not an interesting question of what do we do with agi?</p><p><strong>Charlie:</strong> For someone thinking about designing a society thinking about economic systems, thinking about regulation, it's an extremely good time to start thinking about those questions. Let, let me start by AJ speaking about why I don't think it's here today. And then we can perhaps think. What world where it is here looks like, which I'm quite excited about by the way.</p><p><strong>Charlie:</strong> I don't view as a, a dystopian outcome. Our current approach to AI today is machine learning. We spoke about that earlier today. Machine learning requires sort of three things, compute algorithms and data, and on all three of them, I think to have true agi we're. The compute power I think we need some leap forward there.</p><p><strong>Charlie:</strong> It might be quantum computing. There's a lot of exciting happening there. The timeframes there. I'm not as close to that as I am and I ai, but no one's speaking about quantum computing on a sort of one year time horizon. They're speaking about it again on a 10, 20 year time horizon. The second thing is the algorithms.</p><p><strong>Charlie:</strong> I just, from what we see out there, even with the phenomenal stuff that op that open AI is building, I don't see algorithms out there doing true AI true AGI. They are. The large language models, I, will say that I'm incredibly impressed with how GPT-4 plays chess. It's still not at a, their level of an algorithm that is designed specifically for chess but it's pretty damn good. So, my, my thoughts on the, the algorithms evolve every day, but certainly we're still not there today. And then one of the big hurdles is gonna be data a human ingest data. Rapidly all the time via a series of senses. You can think of that as five senses. You can think of it as 27 senses.</p><p><strong>Charlie:</strong> People have a different perspective on this, but there's just data flowing into us the whole time and at the moment we don't have that in the technology space. If you wanna solve the autonomous vehicle, you've gotta hold. They do like cameras and the visual aspect extremely well. But to solve true AGI level staff to go beyond doing a 2D chess player game to processing a mortgage, I think there's also gotta be a new way of ingesting data.</p><p><strong>Charlie:</strong> Now, one interesting question that I've always wondered is, What will the first iteration of AGI look like? And there's no good reason in my mind to think, I don't think this is the end state. Cause I think the end state's a lot smarter than this. But the first version of what we would consider agi. And general intelligence just means it can do many diverse things and learn from one instance and apply that learning to another instance.</p><p><strong>Charlie:</strong> It could just be a layer that looks like a chatbot, that looks like GPT-4 or GPT-10, whatever it ends up being that ducks into different specific narrow. Ai. And so if you want to get in a car somewhere, you talk to G P T four and you say, I'm looking to go here. And that just plugs into some autonomous vehicle algorithm.</p><p><strong>Charlie:</strong> That could be the first way. And it'll feel like general intelligence and it will be general intelligence or you might have just some massive change in the way algorithms are written. And I do think there's a lot of excite exciting happening there. It's just not clear what the timeframe. , uh uh, for that well,</p><h2>AI Agents Everywhere</h2><p><strong>AJ Asver:</strong> Yeah, I like that last bit you talked about. I, I really. That as kind of a way to think about how AI will evolve. I think some people of call it this kind of agent model where you have essentially this l l m large language model, like GPT acting as an agent, but it's actually coordinated across many different integrations, maybe other deep learning models to, to get things done for you.</p><p><strong>AJ Asver:</strong> And so the collective intelligence of all those things put together is actually pretty powerful and can feel like, or, or have the, the, the kind of illusion of, of artificial general intelligence. I think for me there's this philosophical question of like if it's as good as the thing we want it to be, does it matter if like some nuanced academic definition of AGI isn't what it is? You know what I mean? Like if it does all the things we'd want of like a really smart assistant that's always available, but it doesn't meet the specifics of AGI I in the academic sense. Maybe it doesn't matter. Maybe that's what the next 20 or 30 years looks like for.</p><p><strong>Charlie:</strong> Look, I think that's exactly right and there's no good reason for us to care. We just care that it gets done. We have no idea how the mind even works. We're pretty sure that the mind doesn't say, okay, I've got this little bit for playing chess, this little bit for driving some different way of doing it.</p><p><strong>Charlie:</strong> But humans are very attached to replicating. Things that they experience and understand. And one just very simple way of doing it is a change of the definition of AGI from what your average person might associate AGI with.</p><p><strong>AJ Asver:</strong> That's yeah, that to me is a, a is kind of a mental shift that I think will happen. And, and one of one of the things I've been thinking about is how, and, and this is why a huge reason why I started this, this newsletter and this podcast is that, you know, these things happen exponentially and very quickly.</p><p><strong>AJ Asver:</strong> You don't really realize when you look behind you at an exponential curve cause it's flap, but you look forward, it's, it's kind of, steep. I always talk about this quote from. Sam Altman, cuz it really like seared into my head is that I think what we're gonna see is this exponential increase in these agents that essentially help you coordinate your life in, in work, in, in meetings.</p><p><strong>AJ Asver:</strong> In getting to where you want to go in organizing a date night with your significant other. Right. And you are suddenly gonna be surrounded by them. And, and you'll forget about this concept of AGI because that will become the norm in, in the same way that like the internet age has become the norm. And being constantly connected to the internet is part of our, our, our normalcy in life.</p><p><strong>AJ Asver:</strong> Right. This has been a fascinating conversation. I have one more question for you, which is, You know, as someone that's been going deep into the AI space as well, maybe from the enterprise side as as well, what are you most excited about in a, in AI for the next few years?</p><h2>The future of society with AI</h2><p><strong>Charlie:</strong> yeah. Look, I talked about two things. Firstly. . I think one of the things that I'm excited about in the short term is just the growth in education. And the single biggest thing I think has happened with G P T is the just mass, fast, easy adoption. And when the internet became very, very interesting, it was when you got mass distribution.</p><p><strong>Charlie:</strong> People creating use cases. And that's sort of when you went from like, okay, the internet could be a search engine to organized data. The internet could be a place to buy clothes. The internet could be a place to game to, okay, the internet is just everything. So I'm excited for that. And then in the long run it would be remissive us not to discuss this.</p><p><strong>Charlie:</strong> I'm excited about thinking about what a new economic system looks like. People talk about universal basic. I don't think that the enterprise should be thinking about this question today. Our customers hyper science aren't thinking about this question today, but now is what I would call the ideation stage.</p><p><strong>Charlie:</strong> Like we need think tanks, governments, people out there like, thinking of ideas, and then eventually stress testing ideas for what the future could look like. And I'm, I'm sort of excited for that.</p><p><strong>AJ Asver:</strong> Yeah, I, I want to believe that it will happen, but I'm a little skeptical just given what the history of how humankind behaves. We're, we're not particularly good at planning for these inevitable, but hard to grasp, like eventualities in a similar way that we kind of all knew interest rates are going up, but it was really hard to understand what the implications of that was until last weekend when we found out, right, the</p><p><strong>Charlie:</strong> It was very much we could have prepared for this.</p><p><strong>AJ Asver:</strong> Yeah. Yeah. And yet we could have prepared because all the, all the writing was on the wall. And you know, you've got these people at the fringes that are kind of like ringing the alarm bells, whether that's like people working in AI ethics or whether it's even Sam Altman of OpenAI himself saying like, Hey, we actually can't open source our models anymore. It's too dangerous to do that. Right? And so then you've got people on the other side that are like, no, we need to accelerate this. It needs to be open. Everyone needs to see it. And the faster we get to this outcome, the faster we'll be able to deal with it. I, I skeptically unfortunately, believe that we're gonna stumble our way into this situation.</p><p><strong>AJ Asver:</strong> We'll have to act very quickly when it happens. And maybe that's just like the inevitable kind of like journey of mankind as a whole, right? But it's still exciting either way.</p><p><strong>Charlie:</strong> Well look, I think so. Look, if we don't stumble our way into it, we have, we create a world where people don't have to work, and I'm pretty excited about that. I'd going. Studying philosophy for two years at NYU, I'd be playing a hell of a lot of tennis. I'd be traveling more, like there is a good outcome.</p><p><strong>Charlie:</strong> My, if we just stumble through it. My thinking, AJ, is that this is what the world looks like and it's not pretty. It looks like three classes of people. You have your asset owners, you can think of them as today's billionaire. They will probably end up controlling the world. There might be some fake democracy out there, when you have all of the sort of AI infrastructure owned by ultimately a small group of people you're probably gonna have them influencing most decisions.</p><p><strong>Charlie:</strong> You may then have this second class of like celebrity class. I think that may still exist or human sports may still exist. Human movies human celebrities may still</p><p><strong>AJ Asver:</strong> Yep.</p><p><strong>Charlie:</strong> and you just get this class of 99.9% of people that are everyone else. And what the, the, the two features of their life are gonna look like.</p><p><strong>Charlie:</strong> This is just my guess of like the way it goes. If we don't think about it in a bit more of a interesting way plan, and plan for it is universal basic income. Everyone gets the same amount, probably not very much. I don't know that that's gonna be particularly inspiring for people. I think there's better ideas and then I think that you this is very dispo dystopian, but end up living more in virtual reality than in reality. There's the shortest path to a lot of what you might want to create is just to create it in a virtual world versus going and creating all of that. But in a physical world if so, all that is to say is if we don't start thinking about it, don't start having some regions test different models having startups.</p><p><strong>Charlie:</strong> Form ideas around this and, and, and come up with ideas that are then adopted by bigger countries. In this instance, I think you could end up with a bad outcome, but I think if it's planned, you could end up with an insanely cool world.</p><p><strong>AJ Asver:</strong> Yeah. So we're gonna end on that note that, that those two worlds are what we have to either look forward to or dread, depending on which way you think about it. And I think for folks listening it's. . It's just like really important to begin with that people just understand, like society can only understand if individuals understand where this technology is going.</p><p><strong>AJ Asver:</strong> Right? And that's where you obviously are helping by communicating on LinkedIn to the many people that follow you, especially in industry around it. I try to do it with, with this podcast, but I think for, for anyone that's like fascinated about AI that's following along, like I think the number one thing you can do right now, Share with your friends some of the things that are happening in this space to help people kind of get a grasp for how the space is moving, and that will also help you advocate for us to do the right thing, which is to prepare for it, right? I, I, I think like if people think this is a long way away and they don't understand it, just like you said for industry right? Then no one is incentivized within government to do it because their own uh, constituents don't care about it.</p><p><strong>AJ Asver:</strong> Right? But if we're like, Hey, this is happening. It's an exciting technology, but also there's a kind of two ways this would go and there's one better way than I think it's just as important that we as individuals care about it and advocate for it. was a fascinating conversation. Thank you so much for the time, Charlie.</p><p><strong>AJ Asver:</strong> I really appreciate it. I cannot wait for, for folks to listen to this episode and thank you a again for joining me on the Hitchhiker Guides ai.</p><h2>Closing remarks</h2><p><strong>Charlie:</strong> Well, thank you for having me here, aj.</p><p><strong>AJ Asver:</strong> Awesome. Well, thank you very much and for anyone that's listening, if you enjoyed this podcast please do share it with your friends, especially if you have either founders that are considering building AI startups and going into enterprise, or if you have folks that are working in industry and are considering incorporating AI into their products.</p><p><strong>AJ Asver:</strong> I think Charlie shared a lot of really great insights there that I think folks would appreciate hearing. Thank you very much, and we'll see you on the next episode of the</p><p>&#8203;</p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[How to prompt like a pro in MidJourney with Linus Ekenstam]]></title><description><![CDATA[Celebrating the launch of MidJourney V5 with a deep dive into prompting with Linus Ekenstam]]></description><link>https://operatorsguidetoai.substack.com/p/how-to-prompt-like-a-pro-in-midjourney</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/how-to-prompt-like-a-pro-in-midjourney</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Thu, 16 Mar 2023 14:01:17 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/108709609/f6f572b41c28bc2e752fc4d56cb86373.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-KDD4c5__qxc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;KDD4c5__qxc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/KDD4c5__qxc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em><strong>Note: This episode is best experienced as a video!</strong></em></p><p>Hey Hitchhikers!</p><p>MidJourney V5 was just released yesterday so it felt like the perfect opportunity to do a deep dive on prompting with a fellow AI newsletter <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Linus Ekenstam&quot;,&quot;id&quot;:4671444,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/c8b07626-bcf2-4412-94dc-d5884309b329_512x512.jpeg&quot;,&quot;uuid&quot;:&quot;c2f23851-1740-4c60-b8ef-42b53200c670&quot;}" data-component-name="MentionToDOM"></span>. Linus creates amazing MidJourney creations every day ranging from retro rally cars to interior design photography that looks like it came straight out of a magazine. You wouldn&#8217;t believe that some of Linus&#8217;s images are made with AI when you see them.</p><p>But what I love most about Linus is his focus on educating and sharing his prompting techniques with his followers. In fact, if you follow Linus on Twitter you will see that every image he creates includes the prompt in the &#8220;Alt&#8221; text description!</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/LinusEkenstam/status/1635073803461201920?s=20&quot;,&quot;full_text&quot;:&quot;Want to make your own architecture shots like these?\n\nI included the prompts in the alt and I&#8217;m using Dynamic Prompting &#8482; writing technique to get these results \n\nLearn more by reading my newsletter <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22http://linusekenstam.substack.com/%22>linusekenstam.substack.com</a> &quot;,&quot;username&quot;:&quot;LinusEkenstam&quot;,&quot;name&quot;:&quot;Linus (&#9679;&#7447;&#9679;)&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Mar 13 00:22:27 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FrDy9G0WcAA7d5y.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/55vOgReFYR&quot;,&quot;alt_text&quot;:&quot;editorial photo from Dwell, Midcentury modern house, on a cliff overlooking Los Angeles, morning sun, brilliant architecture, beautiful, exclusive, expensive, minimal lines, breathtaking, 8K, architecture photography --ar 3:2&quot;},{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FrDy9GzWcAMv3Bj.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/55vOgReFYR&quot;,&quot;alt_text&quot;:&quot;editorial photo from Dwell, Midcentury modern house, on a cliff overlooking Los Angeles, morning sun, brilliant architecture, beautiful, exclusive, expensive, minimal lines, breathtaking, 8K, architecture photography --ar 3:2&quot;},{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FrDy9GxXoAEF4MV.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/55vOgReFYR&quot;,&quot;alt_text&quot;:&quot;editorial photo from Dwell, Midcentury modern house, on a cliff overlooking Los Angeles, evening sun, brilliant architecture, beautiful, exclusive, expensive, minimal lines, breathtaking, 8K, architecture photography --ar 3:2&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:25,&quot;like_count&quot;:293,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>In this episode, we cover how Linus shares how he went from designer to AI influencer, what generative AI means for the design industry, and we go through a few examples of prompting in MidJourney live. One thing we cover that is beneficial for anyone using MidJourney for creating character-driven stories is how to create consistent characters in every image. </p><p>Using the tips I learned from Linus, I was able to create some pretty cool Midjourney images of my own, including this series where I took 90s movies and turned them into Lego!</p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/104ba2bc-c196-4084-8e30-42258596b882_1536x1024.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e89c4bf-795b-48ad-bedb-7d63790149cb_2048x1365.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a1022a3-b167-4571-8c90-4cb70534176e_1536x1024.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20cf6f89-3331-401f-963d-c044d7ab47b7_1536x1024.jpeg&quot;}],&quot;caption&quot;:&quot;&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/323fcda3-c48c-4beb-b7be-45313f58bfa8_1456x1456.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p>I also want to thank Linus for recommending my newsletter on his substack, which has helped me grow my subscribers to over a thousand now! Linus has an awesome AI newsletter that you can subscribe to here:</p><div class="embedded-publication-wrap" data-attrs="{&quot;id&quot;:278369,&quot;embedding_publication_id&quot;:null,&quot;name&quot;:&quot;Inside My Head&quot;,&quot;logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F4d8471f5-1658-4db2-8298-2e3d871df7f6_1280x1280.png&quot;,&quot;base_url&quot;:&quot;https://linusekenstam.substack.com&quot;,&quot;hero_text&quot;:&quot;Exploring the world of AI Technology, Design, and building digital products while navigating life as a dad. &quot;,&quot;author_name&quot;:&quot;Linus Ekenstam&quot;,&quot;show_subscribe&quot;:true,&quot;logo_bg_color&quot;:&quot;#ffffff&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="EmbeddedPublicationToDOMWithSubscribe"><div class="embedded-publication show-subscribe"><a class="embedded-publication-link-part" native="true" href="/__u/linusekenstam.substack.com/?utm_source=substack&amp;utm_campaign=publication_embed&amp;utm_medium=web"><img class="embedded-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!x_ZS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F4d8471f5-1658-4db2-8298-2e3d871df7f6_1280x1280.png" width="56" height="56" style="background-color: rgb(255, 255, 255);"><span class="embedded-publication-name">Inside My Head</span><div class="embedded-publication-hero-text">Exploring the world of AI Technology, Design, and building digital products while navigating life as a dad. </div><div class="embedded-publication-author-name">By Linus Ekenstam</div></a><form class="embedded-publication-subscribe" method="GET" action="/__u/linusekenstam.substack.com/subscribe"><input type="hidden" name="source" value="publication-embed"><input type="hidden" name="autoSubmit" value="true"><input type="email" class="email-input" name="email" placeholder="Type your email..."><input type="submit" class="button primary" value="Subscribe"></form></div></div><p>I hope you enjoy the episode and don&#8217;t forget to subscribe to this newsletter at <a href="http://hitchhickersguidetoai.com">http://HitchhikersGuideToAI.com</a>.</p><div><hr></div><h2>Show Notes</h2><h4>Links: </h4><p>- Watch on Youtube:<a href="https://bit.ly/3mWrE5e"> https://bit.ly/3mWrE5e</a></p><p>- The Hitchhikers Guide to AI newsletter: <a href="http://hitchhikersguidetoai.com">http://hitchhikersguidetoai.com</a></p><p> - Linus's twitter: <a href="http://twitter.com/linusekenstam">http://twitter.com/linusekenstam</a></p><p> - Linus's newsletter: <a href="/__u/linusekenstam.substack.com/">http://linusekenstam.substack.com</a></p><p> - Bedtime stories: <a href="http://bedtimestory.ai">http://bedtimestory.ai</a></p><p> - MidJourney: <a href="http://midjourney.com">http://midjourney.com</a></p><h4>Episode Contents:</h4><p>00:00 Intro </p><p>02:39 Linus's journey into AI </p><p>05:09 Generative AI and Designers </p><p>08:49 Prompting and the future of knowledge work </p><p>15:06 Midjourney prompting </p><p>16:20 Consistent Characters </p><p>28:36 Imagination to image generation </p><p>30:30 Bonzi Trees </p><p>31:32 Star Wars Lego Spaceships </p><p>37:57 Creating a scene in Lego </p><p>43:03 What Linus is most excited about in AI 46:10 Linus's Newsletter</p><div><hr></div><h2>Transcript</h2><h2>Intro</h2><p><strong>aj_asver:</strong> Hey everyone. And welcome to the Hitchhiker's guide to AI. I am so excited for you to join me on this episode, where we are going to do a deep dive on mid journey.</p><p><strong>aj_asver:</strong> MidJourney V5, just launched. So it felt like the perfect time for me to jump in with my guests, Linus Ekenstam. And learn how to be a prompting pro.</p><p><strong>aj_asver:</strong> Linus is a designer turned AI influencer. Not only does he have an AI newsletter called inside my mind, but he's also created a really cool website where you can generate bedtimestories for your kids. Complete with illustrations. And he is a mid journey prompting pro. I am constantly amazed by the photos and images that Linus has created using mid journey. It totally blows my mind.</p><p><strong>aj_asver:</strong> From rally cars with retro vibes to bonsai trees that have candy growing on them. And most recently hyper-realistic photographs of interior design that looked like they came straight out of a magazine. Linus is someone I cannot wait to learn from. And he's also going to share his perspective on what all this generative AI means for the design industry, which he has been a part of for over a decade. By the way it's worth noting that a lot of the stuff we cover in this episode is very visual. So if you're listening to this. As an audio only podcast. You may want to click on the YouTube link in the show notes and jump straight to the video when you have time.</p><p><strong>aj_asver:</strong> So if you're excited about I'm one to learn how you can take the ideas in your head and turn them into awesome images. Then join me for this episode of the Hitchhiker's guide to AI.</p><p><strong>aj_asver:</strong> Thank you so much for joining me on the Hitch Hiker's Guide to ai. Really glad to have you on the podcast. I feel like I'm gonna learn so much in this episode.</p><p><strong>Linus Ekenstam:</strong> Yeah. Thank you for having me.</p><p><strong>Linus Ekenstam:</strong> I mean, I'm not sure about the prompt, you know, prompt guru, but let's try</p><p><strong>aj_asver:</strong> Well, I mean, you tweet about your prompts every day.</p><p><strong>aj_asver:</strong> on Twitter, and they seem to be getting better every time. So You are my source of truth when it comes to becoming a great prompter. And I also, by the way, love the one thing you do when you tweet your mid journey kind of pictures that you built, um, that you've created, that you always add in the alt text on Twitter. Um, exactly what the prompt was. And I found that really helpful. Cause when I'm trying to work out how to use Mid Journey, I look at a lot of your alt texts. So, um, also include a link to your Twitter handle so everyone</p><p><strong>Linus Ekenstam:</strong> Nice</p><p><strong>aj_asver:</strong> it out. But I guess</p><p><strong>Linus Ekenstam:</strong> I guess I'll stop.</p><p><strong>aj_asver:</strong> you know, you've been in the tech industry for a while as both a designer and a founder as well</p><p><strong>Linus Ekenstam:</strong> Yeah. Yep.</p><p><strong>aj_asver:</strong> love to hear your story on what made you, um, kind of get excited about AI and starting an AI newsletter and then, you know, sharing everything you've been learning as, as you go.</p><h2>Linus's journey into AI</h2><p><strong>Linus Ekenstam:</strong> Yeah. I mean, if we rewind a bit and, and we start from the beginning, um, I got into the tech industry a little bit on a banana, like a bananas ski. I, I started working in, like, the agency world when I was 17. I'm 36 now, so 19 years ago, time flies. Um, and after like working with, um, customers, clients, and big ones as well, through like, through my initial years there, I kind of got fed up with it.</p><p><strong>Linus Ekenstam:</strong> And. . I went into my first SaaS business as an employee and it was email like way, way, way, way, way before this day and age, right, where you had to like code everything using tables and transparent GIFs. It was just a different world.</p><p><strong>Linus Ekenstam:</strong> And 2012 was like, that's when I started my first own business. And that was like my first foray into like the, the startup world or like building something that was used by people outside of the vicinity of, of, of Sweden or Nordics. Um, it was very interesting times. Um, and I, I've always been kind of like early when it comes to New tech, I consider myself being a super early adopter. I got Facebook as like one of the first people in. By hacking or like social hacking a friend's edu email address. And I got an MIT email address just so I could sign up on Facebook.</p><p><strong>Linus Ekenstam:</strong> Um, so now that we are here, it's like I've been touching all of these steps, like all the early tech, every single time, but I never really capitalized on it or I, I never really pushed myself into a position. I would contribute, but this time around I just, you know, I felt like I had a bit more under my belt.</p><p><strong>Linus Ekenstam:</strong> I've seen these cycles come and go, uh, and I just get really excited about like, oh shit. Like this is the first time ever that I might get automated out by a machine. So my response or flight and fight response to this was just like, learn as much as possible as quickly as possible, and share as much of my learnings as possible to help others.</p><p><strong>Linus Ekenstam:</strong> Cannot not end up. In the same position where they fear for their lives.</p><p><strong>aj_asver:</strong> Yeah, it's, it's interesting you talk about that because I think that's a huge motivator for me as well. It's just help people understand that this AI technology is coming and it's not like it's gonna replace everyone's job, but it certainly is gonna change the way we work. And make the way we work very different. And as -you've been doing and sharing, you know, how to prompt and what it means to use ai, one of the things I've noticed is you've also received a little bit of backlash, you know, from other designers in the space</p><h2>Generative AI and Designers</h2><p><strong>aj_asver:</strong> That maybe as embracing of AI as you have. And I, I know recently there were probably two or three different startups that announced text to UX products where you can basically type in the kind of, uh, user experience you want and it generates, mockups right which I thought was amazing and I thought, You know, that would take years to get to, but we've got that now.</p><p><strong>Linus Ekenstam:</strong> yeah, you</p><p><strong>Linus Ekenstam:</strong> post.</p><p><strong>aj_asver:</strong> and I think one of the things you said was designers need to have less damn ego and lose the God complex.</p><p><strong>aj_asver:</strong> Tell</p><p><strong>aj_asver:</strong> me a little,</p><p><strong>aj_asver:</strong> what the feedback has been like in the AI space around kind of how it's impacting design, especially your field.</p><p><strong>Linus Ekenstam:</strong> So I think, um, there, there is this like weird thing going on where. They're a lot of nice tooling coming out and engineers and, and, and developers. You kind of embrace it. They just like have a really open mindset and go, yeah, if this can help me, you know, I'll, I'll, I'll use it.</p><p><strong>Linus Ekenstam:</strong> Like, take Github Copilot is a good example. People are just raving about it and, and there is some people that are like, oh, it's, it's not good enough yet, or whatever. But like the general consensus is that this is a great tool, it's saving me a lot of time and I can focus on like more heavy lifting or thinking about deeper problems.</p><p><strong>Linus Ekenstam:</strong> But then enter the designer , like turtleneck, you know, black, all dressed in black. I mean, I'm, I, I'm one of those, right? So I'm, I'm, I'm making fun of myself as well. I'm not just pointing fingers at others here. I just think it's like weird that. Here's a tool that comes along and it's a tool, it won't replace you.</p><p><strong>Linus Ekenstam:</strong> Like I'm being slightly sarcastic and using like marketing hooks to get people really drawn in, in my content on Twitter. So I'm not really, meaning, it's not literal. I'm not saying, Hey, you're gonna be out of a job. It's more like, You better embrace this because like the change is happening and the longer you stay on the sidelines, the, the, the more of a, a leap that your peers will have that are starting to embrace this technology.</p><p><strong>Linus Ekenstam:</strong> And I, it's so weird to see like people being so anti and it's like, it's just a tool. It's not, it's not like the tool itself is dangerous. It's like people with the tool will become dangerous and they will threaten your position. Right. So I just find it very interesting to this whole kind of landscape where people, on one hand, it's just embracing it and people on the other hand are just like, no, I'm not, I'm not gonna touch it cuz he can't do X or he can't do Y.</p><p><strong>Linus Ekenstam:</strong> It's like, bear with you. It's like we're in the very, very early days of ai. , we might be seeing half a percent or 1% of what's possible. And these tools are here today, like you said. Um, so I think my kind of like vantage point is like I'm not looking at the next six months or the next 12 months. I'm just like drawing out an arc and going like, where are we 20, 30?</p><p><strong>Linus Ekenstam:</strong> My whole game here is to get as many people as possible.</p><p><strong>Linus Ekenstam:</strong> Ve well versed into these tools as fast as possible. Like, I want to make sure that the divide between the people that haven't got experience and the people that haven't yet played with these tools kind of make sure that divide doesn't grow too big. I think that's my mission really.</p><p><strong>aj_asver:</strong> Yeah, You pointed out there there about how, you know these advancements are happening really quickly and you want people to be able to adopt the tools is I think a really important one. And I think a lot of people don't really understand conceptually how exponential advancements kind of work. And I think Sam Altman recently had this good quote where he said something along along the lines of, when you're standing on an exponential qu curve, it's flat behind you, but it's like vertical in front of you, right?</p><p><strong>aj_asver:</strong> And we're like like climbing this exponential curve, and I think some of us probably see the writing on the wall of how quickly this is all gonna happen. But for other people, know there's always gonna be this resistance. You mentioned like how this tool is gonna help people and people should embrace it, and you of course share a lot of what you are learning and especially when it comes to prompting and you and the art of kind of prompting in mid journey to create interesting images.</p><p><strong>aj_asver:</strong> And I think you've done some prompting on ChatGPT as well</p><p><strong>Linus Ekenstam:</strong> Yeah,</p><h2>Prompting and the future of knowledge work</h2><p><strong>aj_asver:</strong> One thing I'm curious about is, do you think that's the future of the, like knowledge work for us? Is it gonna be like we all just become really good prompt engineers and we're just prompting away when it comes to like writing, you know, writing documents or when it comes to creating design in ux or when it comes to, you know, making images or, or do you think there's more to it than that?</p><p><strong>Linus Ekenstam:</strong> I think prompting the way that it's done now is gonna be very short-lived. Um, if we're doing an analogy and compare, like prompting with ai, with what we're people that are doing really root level programming, let's say assembly type code for computers.</p><p><strong>Linus Ekenstam:</strong> Um, so we're in the age now where everything is new, uh, and the way to interact with these models, whether. ChatGPT or Mid Journey or Dolly or Stable Diffusion. It's a very, very root level. Uh, I think I'm, I'm already starting to see like products popping up that are precursors or like tools that put, put themselves like a layer on top.</p><p><strong>Linus Ekenstam:</strong> So instead of writing like a 200 keyword, Um, prompt to mid journey. You're essentially writing like five words or 10 words that's very descriptive of what you want. And then the precursor takes care of like generating the necessary keywords for you.</p><p><strong>Linus Ekenstam:</strong> I don't think we'll see these like prompt tags where people figure out me included, figuring out like ways to, you know, if you do this in this sequence or this order, you will able to do. This with a, with a, you know, with a model. Right. Um, I think we'll see less of that potentially, um, and, and move more towards like really natural language and, and less kind of like prompt engineering around it.</p><p><strong>Linus Ekenstam:</strong> But I think very, um, important here to note is that the kind of like precursors will happen, like we will kind of move away from talking assembly line situation with AI models and that's also like, that's a barrier to entry right now. If you look at mid journey, you want to get. A lot of the things you need to just overcome is kind of how do you write what you want to get?</p><p><strong>Linus Ekenstam:</strong> Because like if you just write, I want this image that does X, Y, and Z, you're probably not gonna get the thing that you have in your mind. So it's gonna be like you trying out different things and then getting to like slowly speak ai.</p><p><strong>Linus Ekenstam:</strong> Don't focus too much on becoming like a, a super good prompter, right? To try to, to learn more like principles or techniques or, or, or like think more, um, holistically about this whole thing. How do you interact with ai? I think that's, uh, yeah, something I would recommend</p><p><strong>aj_asver:</strong> I liked your analogy of, you know, assembly language and how we're all kind of writing assembly code right now. Um, another analogy I've also heard is like, it's, it's like DOS before Windows came and we're all kind of at the command line trying to get what we want by typing it into a computer before someone, you know, made like a user interface around it.</p><p><strong>aj_asver:</strong> You know, very famously, you know, Xerox Park did it and then Apple was the one who kind of released the version of it. And I, I think that's gonna be something. We'll, all welcome. But in the meantime, I'm curious, how did you kind of learn that steep learning curve of becoming really great at prompting?</p><p><strong>aj_asver:</strong> Because we all kind of start from zero here, and I think the example you gave where you said like, you know, you think you can just come in and describe what you want. That's exactly how I started. I was</p><p><strong>Linus Ekenstam:</strong> I was just like</p><p><strong>aj_asver:</strong> I'm just gonna the the scene and then it didn't end up anything. like wanted</p><p><strong>aj_asver:</strong> What was your journey of becoming good at prompting and, and learning how to use the journey effectively?</p><p><strong>Linus Ekenstam:</strong> I think the, the community's really well set up, uh, with my journey that you can. You can go onto midjourney.com and you can start kind of exploring what everyone else is doing. Um, initially I didn't really understand that they actually had a website. Um, but, but then after a few, I guess a few months, then I'm like, oh, there's actually a website here.</p><p><strong>Linus Ekenstam:</strong> And maybe it wasn't there in the beginning either, so it might be that I just missed it completely. So I think like dissecting other people's work and trying to figure out like, oh, this was a nice cinematic shot. How did I get that look? Um, and, and mind you, like, I've been doing this now for a few months, or yeah, almost half a year.</p><p><strong>Linus Ekenstam:</strong> Um, and it, it was really different. There was less people using the tool like six months ago than there is now. So I think. It's easier to jump into the tool now and see what others are doing and kind of learning by do, like learning by dissecting essentially. I think it's the same with design. Like if you're learning design today, like the best way is just to try to replicate as much work from everyone else as possible. And I'm not saying replicate in the sense of like, oh, that's a nice prompt, command C, command V, oh, now I did it. Um, that's k that's, that's not learning to prompt . All that it's very easy if you, if you find something you like, you wanna make your own derivative of it, go ahead.</p><p><strong>Linus Ekenstam:</strong> I mean, that's the beauty of these tools as well. But if you really want to like learn the skill or like, you know, I have this idea I want to do, then I can go do this. But then I must say there's also some really talent. People, uh, on Twitter and elsewhere that are sharing their journeys as well, and, you know, figuring out ways to, to structure their prompts or, yeah, th there there's a bunch of people that we could potentially look at later or, or that we could recommend in the show notes.</p><p><strong>Linus Ekenstam:</strong> Um, for sure.</p><p><strong>aj_asver:</strong> Yeah, that would be, that would be great. And I think for people that aren't familiar, the way Mid Journey works is you, you have the website where you kind of explore existing images, but all of the work of creating Images is done through their Discord, which</p><p><strong>aj_asver:</strong> Unintuitive to anyone that's not familiar with online communities and with Discord, which Discord itself is a fairly new phenomenon from the last like three or four years right? It was originally used in gaming, but now it's used a lot in communities across ai, across crypto and other and other places. And so was another thing that got, that took a while to get used to is like interacting with AI via Discord. But there's one cool advantage of it that I didn't really fully grasp until now, which is that I can pick up my phone and jump in the Discord any time when I have an idea for an image and just start making images.</p><p><strong>aj_asver:</strong> One of the things I'm curious about, you've been doing this for about six months, ballpark, how many images have you created</p><p><strong>Linus Ekenstam:</strong> I think I just passed, like the 10 K club. I'm not sure. I'm gonna have to look later.</p><p><strong>aj_asver:</strong> I mean, one thing I would love to do in this episode, um, is learn from you some of the skills of like, you know, being a pro prompter in, in, uh, MidJourney. So I was wondering if we could kind of jump into it and maybe one, take a look at some of the creations you've done in the past. Walk us through a little bit, um, how you, how you came up with them and then I have a few ideas of things I wanna do in Midge. Maybe you can help me, um, make that happen.</p><p><strong>Linus Ekenstam:</strong> Let's jump into it.</p><h2>Midjourney prompting</h2><p><strong>aj_asver:</strong> Awesome. So we're gonna jump into Mid Journey now, and Linus is gonna show us some of the images he's created, give us a bit of a sense of kind of his approach to prompting and then we're gonna jump in and do a few examples too</p><p><strong>aj_asver:</strong> So one area, for example, that I would love to learn more about is consistent characters in Mid journey. After we did the podcast, uh, with Ammar, where we built the, where we created the, um, children's book, a lot of people asked, oh, how do you get consistent characters across all the pages of the children's book in the illustrations? So I'm really curious about how you achieve that, cuz that's something I've struggled with as well.</p><p><strong>Linus Ekenstam:</strong> This is interesting. So, um, it, it started with, A lot of people trying to achieve the same thing using mid journey, which is essentially, you know, you have a character, you want the character to be in different poses or in different photos or in different, you know, could be a car cartoon, it could be a real person.</p><p><strong>Linus Ekenstam:</strong> Uh, and I saw different ways of doing it, and mainly they were for cartoons. Then I'm like, this doesn't not work well with a human. Uh, cuz I tried and it didn't work. So I'm like, there must be some other way to do this. Uh, and obviously this is like brute forcing a password really cuz like mid journey is not supposed to be this tool.</p><p><strong>Linus Ekenstam:</strong> Uh, the best way to do consistent characters is to use stable diffusion or something else that you can pre-train on a set of images.</p><h2>Consistent Characters</h2><p><strong>Linus Ekenstam:</strong> the, the way that I went about doing it is essentially, , uh, going to how illustrators work and when, when they create like a character for, for a movie or for an animated, whatever it might be that they're doing, they need reference materials so that other artists can work on the same characters.</p><p><strong>Linus Ekenstam:</strong> You might have hundreds of artists working on the same character. Um, so then, you know, looking at how they are doing these, I'm like, maybe if I simplify this, what if, you know, I take left and right and up and down, and. and I used those images as inspiration cuz that's something you can do in my journey.</p><p><strong>Linus Ekenstam:</strong> You can like image prompt, essentially just like putting an array of images and then adding your prompt. So I, I, I went about like starting up making a character, um, just using like a very simple prompt here. I didn't really have any intent of, of, of the output. I just like, let's make journey, do its thing.</p><p><strong>Linus Ekenstam:</strong> Um, and then when I, I, I found one that I kind of like, oh, this could be nice. Let's work with this. Um, I started using something called a seed. and use that image. So a seat is essentially the noise number or the random noise that an image gets started from. So if. For anyone that doesn't know, you know, mid journey is a diffusion model, which essentially starts from noise and it takes a string of text and it uses that text to take the noise and transform it into an end result.</p><p><strong>Linus Ekenstam:</strong> So if you want to know the pattern, the exact pattern of the noise that you're starting from, you can include a seed number and it's like randomly generated every time you do an image. So if you have an image and you use the seed and you prompt against that seed again, uh, the likelihood of getting something very similar is quite.</p><p><strong>Linus Ekenstam:</strong> So I, I kind of went away and, and started doing different angles of this woman. And then once I had more angles, um, I put the angles together. So I'm just scrolling through here, but essentially just finding those up, down, left, right, and forward. And when I was happy with all of them, I just put them together in a long prompt.</p><p><strong>Linus Ekenstam:</strong> Um, and then just having the same prompt again as the first time, you know, uh, a style, a, a, a portrait shot of a woman. Um, street photo of a woman shot on Kodak, which is essentially just like the, the type of film I wanted to emulate. And then I get the e exact woman out and, hi, this is like, you know, okay, now here we go.</p><p><strong>Linus Ekenstam:</strong> Uh, what can we do with this? Right? Uh, and there is a bunch of things, like a bunch of learnings, um, from this, which is essentially like you can. Very specific images, if you have a bunch of, of images that you're using as the inspiration images,</p><p><strong>Linus Ekenstam:</strong> but also when you do these technique. My kind of the, the culprit here is that I use street style photos.</p><p><strong>Linus Ekenstam:</strong> So every time I'm trying to get her to do other things, like we can go over here. Uh, I, I wanted to try to get her in a, in a space suit, right? We can see that she's kind of in a space. , but she's still standing on a street. So</p><p><strong>aj_asver:</strong> It's like a very, like a urban chic space suit,</p><p><strong>Linus Ekenstam:</strong> yes, . It's an urban cheek spacesuit. And we can even see here, like try some different, um, ar like some different aspect ratios. She's in a spacesuit, but we still have the background of, of the street, right? So, . One way to combat this and that, you know, figure this out after the fact that I made this tutorial is like the, the, the source material, the source MAs that you're using, they should be isolated.</p><p><strong>Linus Ekenstam:</strong> They should be like either against a transparent background or a white background. And that way all of a sudden you can start placing this woman in different areas. So what's neat about this, that you don't need to train a model. You only need to have a set of image. So let's say you have six or nine images that are your inspiration images, and they don't have to be AI generated either.</p><p><strong>Linus Ekenstam:</strong> You could use like yourself, you can take photos of you from the different angles and put them together. Um, and I think a lot of people, it resonated with a lot of people because this is one of these things that are inherently hard to do in my journey, and there is quite a big use case for it. So I, I, I personally hope that, you know, my journey goes into the direction of kind of a little bit.</p><p><strong>Linus Ekenstam:</strong> Stable effusion or Leonardo, where they're giving you tools to do these kind of things like fine tuning, but not maybe to the extent of like training your entire, your own model completely. Right. And we can look at this example. I think this is very nice as well. Like we can get her smiling cuz that was one of the things that, you know, people, oh, you used all these photos, which she's not smiling, you're never gonna get her to smile.</p><p><strong>Linus Ekenstam:</strong> Uh, and basically you can like, there's a lot of like things you can do, even though mid journey is very. Has very opinionated. So there are ways to work around this. And if we're like diving a little bit into prompting here, um, we can just, let's, let's grab this full command here. Um, and we can,</p><p><strong>Linus Ekenstam:</strong> yeah. Sorry,</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> Was you basically reverse engineered how, you know, a character animator would approach this idea of consistent characters. And the way they do it is they have different poses of a character that they kind of create first. So you kind of have a base, kind of almost like a sculpture that hasn't been fully molded yet. So to get an understanding of the character, you generated those using ai, but you could also have a. you know, photos you already have of a person or maybe you take photos of yourself at different angles. Then you inputted that as actually with the prompt, you inputted the images too, that was what allowed you to kind of then create these consistent characters cuz you now have this base image to work from. one of the things you said was, if you want to be able to change the background, so move them from like street photos for example, to be in space, you kind of need to remove the background from the original base images because that's that background. If you keep it in there, like the street photos has a street background is gonna influence what mid journey creates as well.</p><p><strong>Linus Ekenstam:</strong> Yeah, correct. Um, let's , I just pushed this in here. Let's see if my journey does something with it. Sometimes when you're using a lot of high resolution, um, inspiration images, it actually crashes the, the bot. So it doesn't work. Oh, we're actually getting something that's good. So, uh, it's not entirely sure we're gonna get a smiling woman this time, but the way to kind of like force smiling for example, is give smiling a very high weight.</p><p><strong>Linus Ekenstam:</strong> So when you're using, um, Let's see if I can scroll in here. Yeah. So when you are adding, uh, is it colon? Yeah. Is it colon, semicolon? Um, colon? Yeah, colon. Colon five, for example. Then you give smiling. The word smiling, uh, a weight of five. Uh, I think standard weight is like zero or one, I think one. Um, so we're really emphasizing here that we want her to be smiling and now I think we actually got something.</p><p><strong>Linus Ekenstam:</strong> And it might not be that she's smiling in all of. , but she's, uh, kind of smiling . Forced,</p><p><strong>aj_asver:</strong> got like a little bit</p><p><strong>Linus Ekenstam:</strong> yeah. a bit of a forced smile. Uh, but</p><p><strong>aj_asver:</strong> smile. Yeah.</p><p><strong>Linus Ekenstam:</strong> yeah, and this is the thing. I mean, mid journey is opinionated and you, you might have to do re-roll, you might have to do things like over and over. And because it's not really trained on her smiling or being neutral is actually trained on her being.</p><p><strong>Linus Ekenstam:</strong> Angry or just very like,</p><p><strong>Linus Ekenstam:</strong> uh, so re-roll is like, essentially press a button here in, in, in this cord and it takes the exact same, uh, prompt, the same parameters, and a different seed so it won't use the same noise again. So it will start from the beginning One more time. Uh, if you wanted to get like more diverse outputs, we could use chaos, which is essentially how chaotic the difference is between the four image.</p><p><strong>Linus Ekenstam:</strong> that we're going for. So we could add, uh, dash C and then let's say a hundred. So this value goes between zero and a hundred, and it will d dictate the difference between the four different images. So we can see up in, yeah.</p><p><strong>aj_asver:</strong> I noticed, um, by the way, that there's a few different kind of arguments you can add to the end of your. Mid journey prompt, and I think one that you use often is aspect ratio. Um, and then chaos is one. You just mentioned hph and hph, and C. Where did you learn about these and how does someone kind of work their way around trying all these different ones?</p><p><strong>Linus Ekenstam:</strong> um, i, I mid journey.com, they have like documentation I think people are a bit afraid of, of the documentation because they might not know what they're looking for or like, um, yeah, it, it's not that hard. Like when, when you're prompting in mid Journey and then you go like, okay, there is like, I think, uh, 6, 7, 8, 9.</p><p><strong>Linus Ekenstam:</strong> Nine different arguments that you can use. So it's like aspect ratio, chaos. Quality seed stylized tile, which not many people know, and version and quality. So version, you don't need it if you're not. Like now it comes preloaded with the latest model. So if you just add dash dash V four four, it actually uses an older model.</p><p><strong>Linus Ekenstam:</strong> Um, so a lot of prompts you'll see will have dash, dash, v4, uh, not necessary. So essentially now the model that's running is V4 C, which is like the third iteration of v4, uh, and quality two. You can go quality one, two, and up to five, I think. But they've done a lot of testing internally. people can't tell the difference between Q1 and q2.</p><p><strong>Linus Ekenstam:</strong> Like, so it's just a waste of GQ 10 because essentially when you're doing quality two, you're gonna use twice as much GPU render time. And GPU render time is essentially how long, um, of like, how much of your credits get used to render an image? Um,</p><p><strong>aj_asver:</strong> it. So high quality means if you're paying for mid</p><p><strong>Linus Ekenstam:</strong> yeah.</p><p><strong>aj_asver:</strong> actually to use the bot directly versus the</p><p><strong>Linus Ekenstam:</strong> Yes. Yeah.</p><p><strong>aj_asver:</strong> it's gonna cost more per image basically.</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> I notice as well is that you also include some details around how the shot is taken, right? The actual camera. Um, is, does that make a lot of difference kind of picking the, the camera? Because I noticed that was a pretty cool thing that I didn't, I wasn't aware of actually until I saw your, your images.</p><p><strong>Linus Ekenstam:</strong> Yeah, I think we're, there's a bunch of people that I've, like, I, I saw this like in December, the first time. I think like people using camera. Like it's shot on a canon or it's shot on a hassle blood, or it's shot on an icon, or it's shot on this type of film, you know, emulating black and white film or emulating sepia tone film.</p><p><strong>Linus Ekenstam:</strong> Um, and then now I think it's, there's a lot more people that are kind of dissecting it and like really going nitty gritty on it. Um, and, and trying to just be like, what are the things that we can do with this? Like, how. , how much can we describe with this? And it's, as it turns out quite a lot, especially like camera angles type of shots.</p><p><strong>Linus Ekenstam:</strong> Like, you know, using wide, ultra wide narrow, you can go and use like lens parameters. So like for those that are interested in photography, um, you, you could use 50 millimeters, so 50 mm. Um, to, to decide kind of the, the, the, the framing of your shot and kind of what. Output should look like because it has a very distinct look.</p><p><strong>Linus Ekenstam:</strong> You could go 80 millimeter, 120 millimeter tele lens. All these things matter. Actually, it matters quite a lot cuz if we go here and, and check some of the photos I did the other day about, so I did some, uh, national Euro graphic shots, right? So these are quite interesting where we have like, um, shot on the telephoto lens as one of the key things.</p><p><strong>Linus Ekenstam:</strong> So what it does, it really gives you this super consumed in</p><p><strong>Linus Ekenstam:</strong> photo with like bulky in the background. So you have a really blurred background and we can see that, like my journey is really good with hair. Um, and the compression might blow it, blow it down a bit, but Maur is fantastic with hair and feathers and fibers.</p><p><strong>Linus Ekenstam:</strong> I'm not sure what they've done there, but it's, it's absolutely fantastic. Um, so yeah, and lens matters quite a lot.</p><p><strong>aj_asver:</strong> you are, what you're doing there is really kind of imagining the camera you would take this photo with if it was a real photo, and using those, um, those properties of the camera as parts of the prompt. And one of the things I also noticed with your prompts is you are not necessarily describing the scene.</p><p><strong>aj_asver:</strong> You are often described lots of, characteristics of the scene, right. What is your approach when it comes to, you have this idea in your head, uh, you, you, you kind of, ima have this idea in your imagination of what you wanna create and then getting that down into a prompt. How do you, how do you approach that?</p><h2>Imagination to image generation</h2><p><strong>Linus Ekenstam:</strong> um, I mean, in the beginning I, I did write a really interesting one. Threads on this as well. Cause I was sitting in a restaurant you mentioned earlier that like, that, you know, it runs in Discord. You can bring up your phone, you can start prompting. I was, um, we're, I got two kids, right? And me and my partner, we actually had like the first weekend without kids, um, since pre pandemic.</p><p><strong>Linus Ekenstam:</strong> So we basically haven't been out alone. And, and you, what, what I do, I sit with my phone in mid journey. That came out bad anyhow, we were sitting there and we're actually using it together. So we were like, we're talking, we're talking about like what we're building with bedtimestory and then we saw this like really nice geisha on the world cuz we were eating at an Asian fusion restaurant.</p><p><strong>Linus Ekenstam:</strong> And I'm like, I wonder if I can do that mid journey. And then we're just like, you know, open up discord on the phone and we're sitting there chatting, drinking a little bit and just like, oh, okay, we, you know, let's try this. I think I ended up doing like 50, maybe 50 or 60 generations where like the initial.</p><p><strong>Linus Ekenstam:</strong> Was a geisha, but it didn't look anything like the thing we saw on the wall, right? Because the, the geisha on the wall was like on a wooden plaque,</p><p><strong>Linus Ekenstam:</strong> uh, just like a really nice white geisha face mask and some red, really tiny red, uh, pieces in it. So we basically just went like, iterated removed, you know, added, redacted.</p><p><strong>Linus Ekenstam:</strong> It's just like added words, removing words, try different things, went completely crazy and go, what if we just take away all of this and write something completely different? Um, so it it, it's easy if you have an idea, right? That to just like continue to, to plow through. And then once you hit what you want, then you have that kind of like base prompt.</p><p><strong>Linus Ekenstam:</strong> Then you can start altering that you. Exchange an a subject or an object or exchange a post or, but, but you have kind of your, your, your base prompt figured out.</p><p><strong>Linus Ekenstam:</strong> So getting to the base prompt could be tricky. Sometimes you hit gold after just a few tries. Um, it really depends, uh, on what it is that you're looking to create.</p><h2>Bonzi Trees</h2><p><strong>Linus Ekenstam:</strong> Right?</p><p><strong>Linus Ekenstam:</strong> I had luck with like bon's eyes, for example. I just, what can I do bons eyes with, with ma journey and how does that work? You know? And I just type like, uh, pine tree bonsai. Why wait a minute. Like this is fabulous. You know, I, I, I think I made some bonis again yesterday just for fun. Um, so like raspberry bonai, that's, that's, this is the prompt.</p><p><strong>Linus Ekenstam:</strong> This is it, you know, it's not harder than that. And like, you could do</p><p><strong>Linus Ekenstam:</strong> raspberry. That's it. , right? And, and, and you can, you can, you can, you can imagine, you can do thousands of thousands of these, right? Uh, and you can be crazy about it. You can do, like, I think I did Candy Bon. Yeah. Here we go. Ken Bonk. Who, who knew</p><p><strong>aj_asver:</strong> it's got like lollipops.</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> that bonai tree is something my kids would absolutely adore. So I, I had this idea, um, Linus, um, I would love to learn kind of how to do this. Um, and I have a concept in my head and I was wondering if we could try it out,</p><p><strong>Linus Ekenstam:</strong> Yeah. Let's, let's,</p><p><strong>aj_asver:</strong> if we, we can kind of bring it to life.</p><p><strong>Linus Ekenstam:</strong> yeah, let's try</p><h2>Star Wars Lego Spaceships</h2><p><strong>aj_asver:</strong> So the concept is I also have two kids and they're four and they are absolutely obsessed with Star Wars right</p><p><strong>aj_asver:</strong> got their first two Star Wars Lego sets and now they want like everything in the collection.</p><p><strong>aj_asver:</strong> And they went to the library recently and got a book, and the book just has all these Star Wars, um, Lego ships in it. And it made me think like that would be a cool, fun thing to do in Mid Journey is like create imaginary Star Wars kind of spaceships. And so I was just wondering how would I approach that?</p><p><strong>aj_asver:</strong> Um, do I just type in Star Wars spaceships made out of lego. Do I need to kind of think about how it's shot? Do I need to think about kind of the features? And so that's my idea. Star Wars Lego spaceships. How do we turn that into a cool, mid journey image?</p><p><strong>Linus Ekenstam:</strong> Okay, let, let's just start straight off with, with what you just said, like Lego , Lego Star Wars spaceship. We we're probably just gonna get something that's very similar to what what's already in, um, in Star Wars, but with some kind of reim to Lego. Mid journey is relatively good at like, creating Lego. So we're gonna have to figure that one out. Star Wars. Um, spaceship. actually we're gonna put,</p><p><strong>aj_asver:</strong> It's already starting to generate some images and, and the cool thing about mid gen is it kind of shows you bit by bit as it's evolving, right?</p><p><strong>aj_asver:</strong> They already look pretty, pretty cool from the, from the outset. Okay. So we got some Star Wars Lego images.</p><p><strong>Linus Ekenstam:</strong> it doesn't re does it look Star Wars, though?</p><p><strong>aj_asver:</strong> It, it kind of looks like, um, yeah, it, it doesn't look like a Star Wars ship, I would imagine existing in the Star Wars world, but it has some kind of Lego vibes about it.</p><p><strong>Linus Ekenstam:</strong> Let's try to see if we can get an imperial. Maybe in a pure cruiser or something that's also a known set. Maybe there is something that we could go crazy about instead. So when we started with Spaceship and we wanted to be Nubian fighter, um, maybe we want it to be like silver with, um, loose stripes.</p><p><strong>Linus Ekenstam:</strong> Want the side shots. See if we can get something there.</p><p><strong>Linus Ekenstam:</strong> So</p><p><strong>aj_asver:</strong> you just typed</p><p><strong>Linus Ekenstam:</strong> So,</p><p><strong>aj_asver:</strong> in Star Wars spaceship and then Nubian fighters. You're trying to be a bit more specific about it. And then you also added some color hints as well, right? Silver with with white stripes and then Lego. That was an important part</p><p><strong>Linus Ekenstam:</strong> Yeah, and I've also added side shot here to make sure that we get the, the, the, the model from the side. I'm not sure we will actually get it, uh, the way we want here. And I'm also not sure a Nubian fighter is, well this, this was the first one, like the imperial, some imperial ship here still.</p><p><strong>aj_asver:</strong> to look a bit more, a bit more like a</p><p><strong>Linus Ekenstam:</strong> Yeah,</p><p><strong>aj_asver:</strong> I think.</p><p><strong>Linus Ekenstam:</strong> it looks like Star Wars Lego, but it's still, I don't know, mid Journey is doing some weird things here with like, I think it's trying, oh, okay. Now, now let's,</p><p><strong>aj_asver:</strong> ones, when you described a specific type of, um, ship, it looks a bit closer. So, I mean, another one we could try is like, you know, a tie fighter or an or an X-wing. Oh,</p><p><strong>Linus Ekenstam:</strong> yeah.</p><p><strong>aj_asver:</strong> looks a lot more like Lego right now. So</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> see something come to shape. It looks a lot more like a Lego we might have.</p><p><strong>aj_asver:</strong> So maybe we could try like X X-wing fighter or tie fighter.</p><p><strong>Linus Ekenstam:</strong> Let's try with X-wing. X-wing, and we want it silver with orange stripes. Maybe we want it like in space. See, I think the, the thing thing that I, the, the, what I'm kind of doing when I'm promming is like, I'm really kind of playful with it. I don't really mind if it takes me a hundred shots to get something.</p><p><strong>Linus Ekenstam:</strong> Uh, sometimes it's a bit frustrating because like you think you, you kind of get down into this rabbit hole and you just like, I'm, I'm doing the right thing, you know, I'm writing these things, why it's not giving me what I want. Um, but then either just like remove completely and you start over and you do something different.</p><p><strong>Linus Ekenstam:</strong> you just like try a different image and then come back to it because like maybe you have some other things that you want to try out. Um, but I think this, this might turn out great actually. Then, then again, the, the X the X-wing is a known object, right? So I would be surprised if, if we wouldn't get anything here.</p><p><strong>Linus Ekenstam:</strong> I think where mid journey might China is like trying to combine, um, things. But if we want something that looks outta Star Wars, um, especially in Lego, there we go. This doesn't look.</p><p><strong>aj_asver:</strong> This looks like it could be a real Lego set now. So we've got like a couple of different riffs on X-Wing. They have like really big engines, which is, which is really cool, and lots of lasers, which the kids absolutely</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> And you just clicked U2 there. Now what that does is that upscales, the second image, right when you hit U2 and</p><p><strong>Linus Ekenstam:</strong> Yeah, so I, I just wanted to see kind of like what we would get if we, if we kind of like tried to get this in a slightly larger, uh, resolution. So these images are relatively low rest, I think when you're doing 69. Let, I'm doing now aspect ratio 69. We're gonna see, like, the image is 600, um, 640 pixels white.</p><p><strong>Linus Ekenstam:</strong> so, And what I did here now is like when, when you have a shot, uh, a co, a collection of images that you get back, if you react with the envelope emoji, you get back the, the job id, um, the seed number and the four individual images.</p><p><strong>Linus Ekenstam:</strong> So singular images, if you. If you want to save them low rest, and if you're doing like image prompting where, where you kind of put images into the prompt as well to give some, like, to give the prompt some inspiration. This is a neat way to like make sure you're using low risk images instead of like pushing the highest definition images into, into the image prompt.</p><p><strong>Linus Ekenstam:</strong> Cuz that could slow down my journey quite a lot.</p><p><strong>Linus Ekenstam:</strong> Um,</p><p><strong>aj_asver:</strong> that gave you four individual images. Instead of</p><p><strong>Linus Ekenstam:</strong> yeah.</p><p><strong>aj_asver:</strong> one image, which</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> it's kind of four, the default is one image of four things in a grid. Right. And instead it generated four different images. And to do that, you clicked on the envelope emoji. So that is like a super, um, that's like a really great Easter egg for people to know.</p><p><strong>aj_asver:</strong> Cause I would not have known that if, if you hadn't told me. So, envelope emoji. Gives you the original images, um, in low res, which you can then use to feed back into mid journey.</p><p><strong>Linus Ekenstam:</strong> And, and again, this, yeah, this is impressive cuz if we consider what's happening here, it's like the, the model is interpreting our Lego as like actually building something in Lego. It it, it kind of tries to do that and emulate that. And if we look at the lighting here, the reflection in the cockpits, we can see that it's like shot with some kind of studio lighting over overhead lighting.</p><p><strong>Linus Ekenstam:</strong> Um, ah, look at this. This turned out great. Wow. I'm, I'm surprised myself, so</p><h2>Creating a scene in Lego</h2><p><strong>aj_asver:</strong> is so cool. And, um, I have one more challenge for you. This one's a little bit harder. We'll see if we can make it happen. So I</p><p><strong>aj_asver:</strong> I was lego.</p><p><strong>aj_asver:</strong> okay, this is cool, but it's even more fun if you can create a scene, right. With a few different characters in it and the Lego, uh, the, and the Lego, um, figurines too.</p><p><strong>aj_asver:</strong> So I was wondering if we could give that a try. I have this, um, I have this fun idea for a Twitter thread where you recreate scenes from famous TV shows in movies in Lego in Mid journey. So maybe we can start with like a Star Wars one or we can, we, we can do a different one. Um, but I think that could be an interesting one to try too, because one of the things that I think is a little bit harder is when you have characters or multiple characters and you're trying to get, get, get a scene going.</p><p><strong>aj_asver:</strong> So how would you approach that?</p><p><strong>Linus Ekenstam:</strong> Actually I, I think this is an interesting one cuz like, okay, let's do, we have a scene in mind from Star Wars that we would like to try to reenact? Um.</p><p><strong>Linus Ekenstam:</strong> So what I would, I would, what I would then do is like, I would pro, I probably do something like this where I would go and look for, For source material, like what is it that I'm trying to create, like to get an idea for how it looks, right? So, uh, I think this one is relatively cool. Um, I'm not sure we could actually do this, but this could be an interesting experiment.</p><p><strong>Linus Ekenstam:</strong> So let's copy this image address and let's try to use this as an image prompt. So we go imagine, and then we paste the image url and then we say,</p><p><strong>aj_asver:</strong> so you're pasting URL of a Star Wars scene that you found on Google Image search. Right.</p><p><strong>Linus Ekenstam:</strong> Yeah, so I don't know her name here, but this is Finn, right? And what's her name? This is from the later Star Wars. So we go Finn</p><p><strong>aj_asver:</strong> should know this cause we talk about Star Wars characters all the time.</p><p><strong>Linus Ekenstam:</strong> star Wars Finn Running, and then it's BBB eight rolling, uh, sand Dunes and uh, Lego. We want to do aspect ratio 69. So we have the image and what I'm doing now, I'm basically just describing the image that we're looking at, um, and. Adding that as the prompt and adding Lego. We could actually just make sure we weight this, so we go Lego important.</p><p><strong>Linus Ekenstam:</strong> Uh, and let's see. So here is a bit of like, if you listen in on the mid journey, um, all hands or in the way they call 'em community chat or, or town halls. Um, you'll hear, uh, the founder speaking a lot and the team speaking a lot about how the next evolution of of prompt. Probably gonna be image to image or like, people will do a lot of stuff using an image, um, or images, multiple images.</p><p><strong>Linus Ekenstam:</strong> And I kind of agree because like, um, if I have a, if I have the ability to kite, either just take a, a snapshot of something or I grab something on the internet and I can like, take that and mesh it with something else and then put my prompt on it, the, the likelihood of me getting what I want is like 10 times higher</p><p><strong>Linus Ekenstam:</strong> than if I'm just like, uh, writing.</p><p><strong>Linus Ekenstam:</strong> Prompts</p><p><strong>Linus Ekenstam:</strong> So this did not turn into Lego at all. So let's skip the idea of using the image prompt. And let's try something else. So Lego Star Wars is known, right? There is,</p><p><strong>aj_asver:</strong> Mm-hmm.</p><p><strong>Linus Ekenstam:</strong> uh, Lego Star Wars. So what are we trying to do? We're trying to do a cinematic shot, maybe, um, what's gonna happen in there? Who we're gonna have, we're gonna have some new. Storm troopers maybe Darth. Sorry, I'm going a bit slow here. I'm just thinking, um, we want to, one, we want to talk about lighting perhaps.</p><p><strong>Linus Ekenstam:</strong> If we try, try this, sorry. Cinematic Darth Vader, indoors, --ar 6:9.</p><p><strong>aj_asver:</strong> I noticed by the way, that you put Lego Star Wars Co on at the beginning. Is that something you can do to kind of set the, the scene</p><p><strong>aj_asver:</strong> of</p><p><strong>Linus Ekenstam:</strong> Yeah, I act, Actually I should have done this. Select do a multi prompt. So we're deciding that like the first part of the prompt is Lego Star Wars. Like that's like just pure definition. And then we want the cinematic shot with Stormtrooper Star Vader indoors.</p><p><strong>Linus Ekenstam:</strong> Actually, this is probably gonna give us something that's relatively good,</p><p><strong>Linus Ekenstam:</strong> So it interpreted this, well, I don't, I don't think it turned it into a multi promptt, but it worked anyway, so yeah. Here we go. We have Lego</p><p><strong>aj_asver:</strong> start to see an actual cinematic scene.</p><p><strong>aj_asver:</strong> Okay. So now we, you know, did a few different iterations and by adding Lego Star Wars at the beginning of the prompt, we actually got to a scene or a few different scenes where we have storm troopers, we have Darth Vader, we even have a lightsaber.</p><p><strong>aj_asver:</strong> And the background actually looks like it's Lego too. So this is really, really cool. I feel like, um, we made a lot of progress on here and this is giving me a ton of inspiration. I'm gonna go off and make my own Lego scenes after I, after I saw this now, and I feel like I've learned a few tips and tricks from it too.</p><p><strong>aj_asver:</strong> Thanks, lightness. This is really, really cool.</p><p><strong>linus_ekenstam-1:</strong> Yeah, you're welcome. Uh, I, I'm, I'm excited too. I'm a big Lego fan actually, so, uh, yeah. I should, I should, I should actually get, I should actually do some things for this and post to some of the Lego</p><h2>What Linus is most excited about in AI</h2><p><strong>aj_asver:</strong> Yeah, we have to do</p><p><strong>aj_asver:</strong> it. Alright.</p><p><strong>aj_asver:</strong> well, I feel like this is setting me off on a really awesome path, and I'm really excited to explore this further. Um, I've learned a ton from just going through this with you, so thank you so much, Linus. Um, before we wrap up, I'm curious, like what are some of the things you're most excited about, um, in AI and especially in general AI right now?</p><p><strong>Linus Ekenstam:</strong> I, I think in general, I'm just like really excited about the possibilities of all these tools, to be honest. Like these tools are, are accessible to pretty much anyone. Anyone that has a smartphone or anyone that has like a a a computer with a web browser doesn't really have to be a good computer.</p><p><strong>Linus Ekenstam:</strong> Cuz all these tools are running in the browser, right? Like the gps are in the cloud. Uh, it doesn't matter if it's ChatGPT or if it's MidJourney or something else. Like it's enabling anyone, uh, to, to be creative. Like I don't really, I don't need to know anything about like drawing or, or, or, or being an artist.</p><p><strong>Linus Ekenstam:</strong> Right. I can just, if I have a good enough imagination or if I, and everyone is imaginative, so.</p><p><strong>Linus Ekenstam:</strong> Pretty much everyone fits that. Um, so I, I think I'm, I'm most excited about that, that like, there is no real barrier here. Like we, we could like go out and just tell anyone to go try this and, and anyone could.</p><p><strong>Linus Ekenstam:</strong> Right? I think the biggest boundary or the biggest barrier has been, uh, Actually the interface. So the fact that Mid Journey decided to be like, we're gonna do this in Discord, um, a few months back on, on, you know, the, the all hands, like there was a lot of people complaining like, oh, you know, I love Mid Journey, but it was so difficult to learn Discord.</p><p><strong>Linus Ekenstam:</strong> And I'm like, wait a minute, why don't you just have a wonderful website? Why don't you just put the UI there, like the generative part. There could, could be easily done. Like look@lexi.art, amazing platform, super simple, prompt books on the website. Um, yeah, so I mean, there, there are few things that we could, that need solving, but like the fact that this is available for everyone, I think I'm most excited about.</p><p><strong>Linus Ekenstam:</strong> And then when it comes to new tools, like it's really hard to keep up. Um, there's basically like a ton of new tools every day getting released. It feels like, you know, we're kind of in a. Hype cycle. A lot of people are getting their hands dirty. There is a lot of really good ideas and a lot of people are executing really fast.</p><p><strong>Linus Ekenstam:</strong> I think ElevenLabs is sitting on some gold. They kind of like made it super simple to clone a voice and then used that voice by just inputting text.</p><p><strong>Linus Ekenstam:</strong> So for example, if you know me, I'm building a, a storybook, um, generator for, for kids stories and, and we want to be able to cl clone parents' voices. So it's super easy for us to just like, Hey, record 60 seconds of your voice and then you can have any of the stories that you created written, like read back by you to your kids.</p><p><strong>Linus Ekenstam:</strong> So I think that's, I'm really excited about that. And then I'm really excited about. A lot of these platforms opening up their APIs to developers. So, uh, we saw yesterday, the day before yesterday, like, you know, open AI released chat, G P T API and whisper api, um, through the world for everyone to use. And I think super excited about that as well.</p><p><strong>Linus Ekenstam:</strong> So, yeah. Um, there there's not much underground, well, there is a few, you know, things that are popping up on the radar, but I don't think any, anything that's as exciting as the big things the macro.</p><h2>Linus's Newsletter</h2><p><strong>aj_asver:</strong> Yeah, there is, so much happening in the space and folks can there is actually um, newsletter to, to get updates on, on kind of your, your take on the</p><p><strong>Linus Ekenstam:</strong> Yeah. Yeah.</p><p><strong>aj_asver:</strong> what's the name of the newsletter?</p><p><strong>Linus Ekenstam:</strong> So the name is my name. So it's linusekenstam.substack.com. Uh, and actually the name of the name of the newsletter is inside my head. Um, so maybe I should change the URL to be inside my head. Uh, but yeah, it's linus eam.ck.com and, uh,</p><p><strong>Linus Ekenstam:</strong> Yeah.</p><p><strong>aj_asver:</strong> linusekenstam.substack.com Thank you so much Linus, for joining me, um, and helping me become a pro at Mid Journey. And also I feel like I learned a lot about kind of your, your perspectives on AI and generative AI and how it's gonna impact the design, industry too. So I'm really appreciate it.</p><p><strong>aj_asver:</strong> Thank you so much. And until the next episode of Hitchhiker's Guide to AI. Thank you everyone for joining us.</p><p><strong>Linus Ekenstam:</strong> Thank you. Thanks for having.</p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Navigating the AI Landscape: Market Trends and Opportunities For Startups ]]></title><description><![CDATA[The most important trends and opportunities you should care about if you are considering starting an AI startup.]]></description><link>https://operatorsguidetoai.substack.com/p/startups-navigating-the-ai-landscape</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/startups-navigating-the-ai-landscape</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Sun, 12 Mar 2023 15:01:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi Hitchhikers!</p><p>To the 170+ new subscribers since my last post, welcome to The Hitchhiker&#8217;s Guide to AI! I&#8217;m so grateful for your support. </p><p>This week instead of the usual highlights post, I&#8217;m going to deep dive into a topic I&#8217;ve been thinking about a lot lately: </p><p><em><strong>What are the most compelling opportunities for startups entering the AI space?</strong></em></p><p>In this post, I will take a stab at answering that question. I&#8217;ll start with an overview of the current AI landscape, then some of the trends in the space and finally the opportunities that these trends present for startups.</p><div><hr></div><h3>Is AI a disruptive technology?</h3><p>Let&#8217;s start by referring to Clayton Christensen&#8217;s Innovators Dilemma, which distinguishes between two types of technologies: sustaining technologies and disruptive technologies:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xtXF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xtXF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg" width="520" height="391.3" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:301,&quot;width&quot;:400,&quot;resizeWidth&quot;:520,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!xtXF!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad01ca41-ca3c-400f-af94-61318eba246c_400x301.jpeg 1456w" sizes="100vw" loading="lazy" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="http://web.mit.edu/6.933/www/Fall2000/teradyne/clay.html">MIT Press</a></figcaption></figure></div><p>Sustaining technologies improve the performance of existing products or services, while disruptive technologies create new markets or value networks and eventually displace established ones. Disruptive technologies often start as inferior or niche products that appeal to a small segment of customers, but gradually improve and become mainstream over time. The dilemma is that incumbent firms often focus on satisfying their most profitable customers with sustaining innovations, and ignore or underestimate the potential of disruptive innovations until it is too late.</p><p>What&#8217;s interesting about AI, is that althought it seems like a disruptive technology, it is being introduced to the market much more like a sustaining one:</p><ol><li><p>AI <em>is not an inferior or niche product, </em>in fact one of the values of large-scale language models like GPT is that they are generally applicable to many different use cases.</p></li><li><p>AI <em>does not appeal to a small segment of customers. </em>For example, ChatGPT is rumored to already have over 100M monthly active users, making it the fastest growing consumer product ever!</p></li><li><p>AI <em>is not being ignored by incumbent firms, </em>in fact it is being embraced by incumbents at a surprisingly fast pace.</p></li></ol><p>At the same time, given the sheer pace of innovation happening in the AI space right now, it&#8217;s hard to believe it will not be a force of disruption in the technology sector. It&#8217;s possible though, that the distruption is just happening much faster than in Christenson&#8217;s framework. A better way to think of AI as a disruptive force might be to change the above diagram to me more like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vRcA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vRcA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg" width="564" height="424.41" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:301,&quot;width&quot;:400,&quot;resizeWidth&quot;:564,&quot;bytes&quot;:36728,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!vRcA!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b250af2-3fc4-4ba6-9f87-3c9085d5e435_400x301.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This means the assumption that AI as a disruptive technology clearly favors new entrants over incumbents may not apply. To understand where startups might have an advantage, we need to instead get a deeper understanding of how the AI space is evolving. This will help us identify where disruptive opportunities exist vs. not for startups. So, let&#8217;s dive in!</p><h3>Today&#8217;s AI Landscape</h3><p>To work out how the business of AI will evolve and where the most value will be created it's worth taking a step back first and getting a lay of the land. Let's start with an overview of the current AI landscape using this handy diagram from the folks at Madrona Ventures<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> :</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!1Y46!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!1Y46!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg" width="612" height="1020.2738255033557" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1242,&quot;width&quot;:745,&quot;resizeWidth&quot;:612,&quot;bytes&quot;:162127,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!1Y46!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cf0c39-60c2-4a90-b610-a8bf2a4b6c01_745x1242.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Madrona&#8217;s diagram has six layers to the AI market today, each layer building on top of the one below it:</p><p><strong>Silicon layer</strong></p><p>The Silicon layer is the companies making specialized computer chips, usually GPUs<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> or, in Google&#8217;s case TPUs<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>, for training AI models. This is the most capital intensive layer given the need R&amp;D needed to develop state of the art GPUs and expense of manufacturing hardware. Unsurprisingly Nvidia is the industry leader here with three decades of experience building GPUs. There have however been new entrants to this layer recently like <a href="https://www.cerebras.net/company/">Cerebras</a>, a startup founded in 2016 that develops massive computing chips for artificial intelligence.</p><p>It's worth noting that most startups building AI products don't buy their own GPUs, instead opting to use cloud service providers that buy them in bulk and rent them out. A big reason for this is that compute resources are expensive, especially with the increased demand for AI applications and the explosion of AI startups in the last few years. For example, Nvidia's A100 GPUs<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>, which are purpose-built for data centers and designed for large-scale machine learning use cases, cost $200,000 each! </p><p><strong>Cloud layer</strong></p><p>The Cloud layer are the services that provide computing resources such as GPUs in their data centers for developers to train and serve their models or AI applications on. Competition is mounting between cloud providers, with both Microsoft and Google's cloud divisions making strategic investments in AI research companies. OpenAI exclusively uses Microsoft Azure for compute resources because of their business partnership<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> and Microsoft's $10B investment. Not to be outdone by Microsoft, Google recently invested $300M in Anthropic, an AI research team founded by former OpenAI employees. </p><p><strong>Foundational Model (FM) Operations</strong></p><p>To train AI models like OpenAI&#8217;s GPT, you need to train them with massive datasets like the whole public internet, which is why they are called large-scale language models. As Madrona Ventures points out:</p><blockquote><p>Foundation models have immense compute requirements for training and inference, requiring large volumes of specialized hardware. That is a significant contributor to the high costs and operational constraints (throughput and concurrency) that application developers face. The largest players can find the cash to accommodate &#8212; consider the <a href="https://news.microsoft.com/source/features/ai/openai-azure-supercomputer/">&#8220;top 5&#8221; supercomputer infrastructure</a> Microsoft assembled in 2020 as part of its OpenAI partnership. But even the mighty hyperscalers face supply chain and economic constraints. </p></blockquote><p>This need for efficiency has created an opportunity for tools and infrastructure products to lower the cost of training, deployment and inference of models. For example, <a href="http://scale.com">Scale AI</a> is a data platform for AI that enables developers to collect, manage, and label data quickly and easily. Another example is <a href="http://banana.dev">Banana</a>, a machine learning model deployment platform that enables companies to deploy models to production faster and more cheaply without needing to manage their GPU servers. <a href="https://www.mosaicml.com/">MosaicML</a> is another startup that focused on helping companies train large-scale models with their own data. </p><p><strong>Foundational Models</strong></p><p>The Foundational Models layer is where <em>the magic happens. </em>Foundational models are the generative AI models being developed across text, image, video and audio. Besides OpenAI, a few other companies are developing foundational models, including Google, and startups like Anthropic and AI21Labs. </p><p>Aside from the proprietary foundational models there is also a growing ecosystem of open-source models, the most well known of which is Stable Diffusion an open-source text-to-image model comparable to OpenAI's proprietary Dall-E<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>.  Huggingface meanwhile is place for these open source models to be shared and iterated on by the community.</p><p>While proprietary models have the advantage of performance and scale thanks to the well funded companies backing them, open-sourced models offer more flexibility and lower costs to developers who want to customize them to meet their applications needs.</p><p><strong>Tooling</strong></p><p>Foundational models can&#8217;t be used in applications on their own because (1) they don&#8217;t have a real-time understanding of the world e.g. the latest version of GPT-3 was trained with data up to the end of 2021; (2) the functionality of their APIs are limited and (3) they contain bias in their training data which may lead to unpredictable results Andy (3) they hallucinate. The Tooling layer is where foundational models are augmented with external data integrations to make them more useful and powerful. </p><p>For example, <a href="https://langchain.readthedocs.io/en/latest/getting_started/getting_started.html">LangChain</a> is a platform that makes it easy to integrate models like GPT-3 with external data to create applications. <a href="https://gpt-index.readthedocs.io/en/latest/guides/use_cases.html">GPTIndex</a> is another example of tool that makes it easier to provide GPT with additional context from external sources besides the data it was trained on.  You use GPTIndex to embed all the articles in a newsletter to create a chatbot based on the content, just like <a href="https://www.lennysnewsletter.com/p/i-built-a-lenny-chatbot-using-gpt">Lenny Rachitsky did recently</a>.</p><p><strong>Application Layer</strong></p><p>The application layer is where we are currently seeing a Cambrian Explosion of new products and services that make use of AI across many different use cases, from productively to creativity and across consumer and B2B:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-Clz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-Clz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg" width="1089" height="1244" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1244,&quot;width&quot;:1089,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:308941,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!-Clz!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5f718ae-8747-4ba8-acee-8da606ec621e_1089x1244.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The multitude of AI apps that have launched in the last few years broadly fall into three categories:</p><ul><li><p>Creativity: This category includes all AI products that are helping people create content across text, video and audio and are are a mix of B2B and B2C. </p></li><li><p>Workflows: These are the products that help complete a specific workflow using AI more efficiently. Examples include Github&#8217;s Copilot which auto-completes code, or Descript which transcribes podcasts for editing.</p></li><li><p>Knowledge: These products help you query and retrieve knowledge from a large dataset. Chatbots like OpenAI&#8217;s ChatGPT, Quora&#8217;s Poe and Bing&#8217;s AI Chat are all examples of this. There will be many more specialized AI chatbots across different verticals, especially where proprietary data sets are required in training.</p></li></ul><p>It&#8217;s worth noting that all three of these categories of apps rely on foundational models often via proprietary APIs like OpenAI&#8217;s or open-source alternatives. This makes it harder to differentiate with core technology, and startups instead have to focus on product experience, go-to-market and extracting the most value out of models via prompt engineering and fine tuning.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Are you enjoying this update? Then don&#8217;t forget to subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Five major trends to keep an eye on in AI</h3><p>Now that we have a good lay of the land, here are some of the major trends I&#8217;ve observed in the AI market that will impact value creation and opportunity, especially for startups entering the space.</p><h4>1.  Competition between model providers &#8594; AI costs go down, and performance goes up </h4><p>Since OpenAI launched ChatGPT and announced its partnership with Microsoft earlier this year, things heat up between Big Tech. Google is also working on a large-scale language model and chatbot. Last week Amazon also announced a partnership with Huggingface<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> and Meta released it&#8217;s large-scale language model<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> that they claim is more performant than GPT-3. </p><p>Right now OpenAI+Microsoft are the primary providers of proprietary models which means they have a lot of pricing power. Elad Gil, a founder and angel investor currently exploring AI, makes a convincing case for an oligopoly market with a few major players in his recent blog post on AI Platforms, Markets &amp; Open Source<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>:</p><blockquote><p>&#8220;The reason to argue for a near term oligopoly market, versus likely fragmentation, is due to the capital/compute/data scale costs currently needed for each subsequently better performing LLM model. If GPT-3 at the time cost a few million to ten million or so to train, and GPT-4 from scratch may be estimated at tens of millions to maybe a hundred million, maybe GPT-5 is a few hundred million and GPT-N is a billion. This of course assumes that costs will scale faster than technical breakthroughs or drops in GPU (or specialized hardware cost declines) and these may be false assumptions.&#8221;</p></blockquote><p>A Google, Microsoft, Meta, or Amazon oligopoly seems like the likely outcome, and as competition heats up between them, it should drive down for developers to build on too. Performance per dollar will also increase with competition as advancements are made in the efficiency of training and serving models. What might be expensive with today&#8217;s models may quickly become cheaper in a few months. An excellent example is OpenAI releasing ChatGPT&#8217;s API for 10% of the price of their previous state-of-the-art GPT-3 API<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>.</p><h4><strong>2. Open-source models vs. proprietary models will diverge</strong></h4><p>Today, it is possible for open-source foundational models with similar capabilities to fast follow proprietary models. This is possible for several reasons: (1) all models are being trained on the same public data, (2) AI researchers are publishing their advancements regularly<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>, (3) open-source projects are being sponsored by companies like StabilityAI and (4) communities of developers have formed around specific models. A prime example is Stable Diffusion, launched only a few months after OpenAI's Dall-E 2 and is considered on par in performance. Meanwhile, Meta released an open-sourced large-scale language model to compete with OpenAI just a few weeks ago<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>.</p><p>Over time though, open-source will probably lag behind proprietary models for three primary reasons:</p><ol><li><p>Cost to train: As models get bigger and bigger, they will become more expensive. Open-sourced models will therefore need more well-funded sponsors to develop them. Today StabilityAI might be able to afford to sponsor a GPT-3 equivalent model that costs $10M to train, but next year they might not be able to afford to support a GPT-5 equivalent that costs $100M to train.</p></li><li><p>Proprietary datasets: Well-funded companies like OpenAI  license additional private data sets to improve their models. This will make it harder for open-sourced models to compete on performance and they may lag one or two generations behind proprietary models. For many customers, the price and versatility of open-sourced models will be the deciding factor. Using a lesser performant open-source model will be a path many developers take to get started cheaply before developing their own models or switching to proprietary ones once they have enough scale. </p></li><li><p>Closed research: Many contributions that have enabled open-source AI models to come from private companies, like Google&#8217;s Transformer architecture. That might change as the need for Big Tech to compete in AI takes a higher priority over publishing research. </p></li></ol><p>This means developers will need to choose between a proprietary state-of-the-art model with higher performance and more &#8220;bells and whistles&#8221; likely at a higher cost or choose an open-source model that is 1-2 generations behind but is cheaper and more customizable.</p><h4>3. Bigger models &#8594;  More emergent behavior and faster path to AGI</h4><p>As models are trained on larger datasets with more parameters, we will continue to see more emergent behavior appear<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a>. This is important because if we discover new capabilities of these models after they are trained, there will also be new applications for them that we don't know or can predict today. This is both a blessing and a curse. On the one hand, it means more opportunities for creating innovative use cases for AI, for example unlocking the ability to make AI more actionable. On the other hand, if you built a business on a foundational model based on previously known capabilities and your differentiator is the software you wrote to augment that model for your specific use case, that differentiator may become obsolete in the next version of the model!</p><p>Furthermore, the unpredictability of emergent behavior also means that the path to artificial general intelligence (AGI) will continue to be unpredictable too and may accelerate faster than we expect. Sam Altam talks about how OpenAI will approach this unpredictability in a recent post:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/sama/status/1629212494072889349?s=20&quot;,&quot;full_text&quot;:&quot;Planning for AGI and beyond: &quot;,&quot;username&quot;:&quot;sama&quot;,&quot;name&quot;:&quot;Sam Altman&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Feb 24 20:11:42 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:640,&quot;like_count&quot;:3426,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://openai.com/blog/planning-for-agi-and-beyond/&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f513ab2-2f5a-44a7-8d44-faa7b7a48c87_1820x1024.jpeg&quot;,&quot;title&quot;:&quot;Planning for AGI and beyond&quot;,&quot;description&quot;:&quot;Our mission is to ensure that artificial general intelligence&#8212;AI systems that are generally smarter than humans&#8212;benefits all of&nbsp;humanity.&quot;,&quot;domain&quot;:&quot;openai.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>As advancements in AI gather speed, and we quickly approach AGI, whole swathes of existing AI products may be completely replaceable by the next version of ChatGPT!</p><h4>4. Low barrier to entry &#8594; Intense competition in the application layer</h4><p>Every week we see dozens of new startups launch in the application layer, often with two or three startups tackling the same problem. An excellent recent example is Generative Design: A few weeks ago GalileoAI announced their text-to-UX product where you could describe a user interface or flow and AI will generate high-fidelity mockups:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/Galileo_AI/status/1623360270008520714?s=20&quot;,&quot;full_text&quot;:&quot;The future is here - Generative AI is coming to user interface design! \n\nGalileo AI generates stunning UI designs with just a text prompt, empowering you to design beyond imagination at lightning speed. &quot;,&quot;username&quot;:&quot;Galileo_AI&quot;,&quot;name&quot;:&quot;Galileo AI&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Wed Feb 08 16:37:03 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;Today, Generative AI takes a big step and comes to user interface design!\n\n@helnzhou and I are excited to announce @Galileo_AI : the first AI product that uses natural language to generate UI designs. It lets you design beyond imagination.\n\nEarly access: https://t.co/4KqV1csQ6c https://t.co/LtzSC1kYw3&quot;,&quot;username&quot;:&quot;arnaudai&quot;,&quot;name&quot;:&quot;Arnaud Benard&quot;},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:98,&quot;like_count&quot;:384,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Since then two other products, Genius<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> and UIzard<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a> have been launched to solve the same problem. There are already dozens of products focused on copywriting, marketing, sales outreach, and coding. Apparently, over 30 startups in Y Combinator&#8217;s current batch are working on AI products: </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2J-M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2J-M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg" width="1170" height="1174" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1174,&quot;width&quot;:1170,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="/__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!2J-M!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F453d2a40-8e3b-447d-b567-de5048469b8a_1170x1174.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It&#8217;s pretty clear then that the barrier to entry to build AI-powered products right now is pretty low and we should expect hundreds more startups to be formed in the space. Another tailwind for AI startups is that funding has slowed down in series B to late-stage startups and venture capitalists are desperate to deploy capital from the funds they raised in 2020/2021.</p><p>All of this points to one trend we are going to see over the coming years: Intense competition between AI startups in the same sector going after the same set of customers. For customers, this is great as it will likely reduce costs to access AI and increase innovation and choice. For startups though, it will be more challenging because the cost of acquiring customers will continue to grow and margins will be eroded. Today Jasper.ai, a product that helps you write marketing copy, charges $60/mo which seems unsustainable as more products launch to solve the same problem. As Jeff Bezos famously said, &#8220;Your margin is my opportunity.&#8221; </p><h4>5. Incumbents quickly adding AI &#8594; harder to disrupt</h4><p>I talked about the speed of incumbents entering AI in my last weekly highlights post<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a></p><blockquote><p>&#8220;We&#8217;re seeing incumbent technology companies like Microsoft, Snap, Spotify, and Shopify move quickly to adopt [AI Chatbot technology] which again speaks to the fact that the barriers to entry in AI are much lower than we might have expected. The corollary of this is that if you&#8217;re a startup thinking about using AI to disrupt an existing market with strong incumbents, it&#8217;s not a given that you will be able to move faster than them!&#8221;</p></blockquote><p>This is one of the most surprising trends because incumbents aren&#8217;t ordinarily able or willing to move as quickly to adopt new technology into their products. For example, if you are building an enterprise product for Sales, Customer Service, Marketing or Supply Chain and AI was your secret sauce, it&#8217;s also Microsft&#8217;s. This week they announced &#8220;Microsoft Dynamics 365 Copilot&#8221;:</p><blockquote><p>&#8220;With Dynamics 365 Copilot, organizations empower their workers with AI tools built for sales, service, marketing, operations and supply chain roles. These AI capabilities allow everyone to spend more time on the best parts of their jobs and less time on mundane tasks.&#8221;</p></blockquote><p>Of course, Microsoft has massive distribution advantages, which is why they were able to display Slack in many organizations with Microsoft Teams quickly. </p><p>Here are a few more examples of other incumbents quickly adopting AI just from the last week. AI:</p><p>Brex adding an AI assistant to their financial stack:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/brexHQ/status/1633155216597225473?s=20&quot;,&quot;full_text&quot;:&quot;&#128680; Huge Empower News! \n\nWe are working with <span class=\&quot;tweet-fake-link\&quot;>@OpenAI</span> to bring advanced AI-powered tools for CFOs and their teams. \n\nThe new features will provide relevant insights on corporate spend and answer critical business questions all in real-time. \n\nLearn more: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://www.brex.com/journal/press/brex-openai-ai-tools-for-finance-teams/%22>brex.com/journal/press/&#8230;</a> &quot;,&quot;username&quot;:&quot;brexHQ&quot;,&quot;name&quot;:&quot;Brex&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Mar 07 17:18:40 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/v68g5ybgrwyfj2ln4ybf&quot;,&quot;link_url&quot;:&quot;https://t.co/6OjtixbNZV&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:35,&quot;like_count&quot;:303,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/amplify_video/1633155156161466368/vid/480x270/fMgpBWoI6pNhbcvK.mp4?tag=16&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Slack adding AI capabilities directly into their app:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/SlackHQ/status/1633090443331219456?s=20&quot;,&quot;full_text&quot;:&quot;Introducing the ChatGPT app for Slack by <span class=\&quot;tweet-fake-link\&quot;>@OpenAI</span>! \n\n&#128172; Summarize conversations to quickly catch up on any channel\n&#128218; Tap into tools to learn about any topic in an instant\n&#128394;&#65039; Draft and edit messages in seconds with smart writing assistance\n\nLearn more: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://sforce.co/3T81lVT/%22>sforce.co/3T81lVT</a> &quot;,&quot;username&quot;:&quot;SlackHQ&quot;,&quot;name&quot;:&quot;Slack&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Mar 07 13:01:17 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://res.cloudinary.com/hhsslviub/video/upload/e_loop,vs_40/d7nncxppxqqytry6ydbn.gif&quot;,&quot;link_url&quot;:&quot;https://t.co/2iDMiMEfsD&quot;,&quot;alt_text&quot;:&quot;GIF shows a demo of the ChatGPT app for Slack by OpenAI.&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:539,&quot;like_count&quot;:2061,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Hubspot adding AI to their CRM tools:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/rrhoover/status/1632742620391768064?s=20&quot;,&quot;full_text&quot;:&quot;Dharmesh from HubSpot ($20B+ public co) just launched their AI side project:\n\n&quot;,&quot;username&quot;:&quot;rrhoover&quot;,&quot;name&quot;:&quot;Ryan Hoover&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Mar 06 13:59:10 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:4,&quot;like_count&quot;:28,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://www.producthunt.com/posts/chatspot&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/befa9791-5f20-4a3d-b699-408ada22ffb0_1024x512.png&quot;,&quot;title&quot;:&quot;ChatSpot - The all-In-One AI chat tool for growing better | Product Hunt&quot;,&quot;description&quot;:&quot;ChatSpot combines the power of ChatGPT with HubSpot CRM and some paid APIs to help you grow your business. It&#8217;s all available via a chat-based natural language interface. All in one place. Secret Unlock code: Once in, type: &#8220;Unlock I love ProductHunt badge&#8221;&quot;,&quot;domain&quot;:&quot;producthunt.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>It&#8217;s worth noting that in many of these cases, OpenAI is working closely with the company, giving them access to the latest models that are not yet available on their public developer platform. Why is OpenAI doing this? I believe it is to aggressively grab market share and close enterprise accounts that will contribute to the bulk of their revenue in the long term. </p><p>I think this is the single most important trend to watch if you&#8217;re a founder considering building something in AI: If there is already a dominant software incumbent solving the problem you are approaching, you will need to have both a more compelling product <em>and</em> a more robust go-to-market strategy to compete with their distribution. </p><h3>Opportunities for startups</h3><p>Given the current trends in the AI market in cost reduction, open-source advancements, competition, and fast-moving incumbents I believe there are three opportunities for startups to take advantage of:</p><h4>1. Drive down the cost of training/serving models</h4><p>As competition heats up in the foundational model layer, companies will likely compete with Microsoft+OpenAI/Google/Amazon oligopoly to offer more customized models or cater to a specific vertical model use case. Cost will be a prohibitive factor for these companies, who will want to save money on data acquisition, training and inference. Similarly, for companies that want to train and serve their own models or open-source models, the cost will also be a huge factor given that they won&#8217;t have the same economies of scale that the big tech companies do.</p><p>Here are a few examples of startups that could reduce the cost of foundational models:</p><ol><li><p>A startup that provides a platform for distributed and parallel training of large-scale language models using ColossalAI, a framework that enables efficient model parallelism, pipeline parallelism, tensor parallelism, and data parallelism. ColossalAI claims to offer up to 7.73 times faster training and 1.42 times faster inference than ChatGPT<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>.</p></li><li><p>A startup that provides a platform for streaming data processing and model compression for large-scale generative models using Composer and MosaicML, tools that enable efficient data loading, preprocessing, augmentation, sampling, and quantization. Composer and MosaicML claim to reduce the training cost of Stable diffusion by up to 75% compared to GPT<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a>.</p></li><li><p>A startup that offers a service for compressing and pruning foundational models using techniques like quantization, distillation and sparsification, which can reduce model size, latency and energy consumption.</p></li></ol><h4>2. Specialized models for highly regulated verticals</h4><p>Incumbents can move quickly to incorporate existing models into their products in unregulated verticals, but for highly regulated verticals like healthcare, financial services, and government this might not be the case. The use of off-the-shelf foundational models which are trained on public data at scale may face some limitations in highly regulated sectors, such as:</p><ul><li><p>The need for explainability and transparency of the model's decisions and behaviors.</p></li><li><p>The risk of bias, error, or misuse of the model's outputs or data inputs.</p></li><li><p>Compliance with ethical, legal, and social standards and regulations.</p></li><li><p>The risk of hallucinations<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a> causing serious harm to the end user in life-threatening, business-critical or security-critical circumstances.</p></li></ul><p>Here are some more examples of startups that could be formed in this space:</p><ul><li><p>A startup that develops a model that can generate legal documents such as contracts, agreements, or reports based on environmental or energy regulations and standards.</p></li><li><p>A startup that provides an AI platform that can generate and verify regulatory reports for financial institutions based on their data and rules.</p></li><li><p>A startup that provides AI models that can enhance and improve the communication between healthcare services and patients, for example by better-summarizing doctors&#8217; notes and treatment recommendations.</p></li></ul><h4>3. Help incumbents move faster to adopt AI</h4><p>While more tech-forward incumbents will move quickly to incorporate AI in their products, not every company will have the in-house resources and expertise to this, despite competitive pressure. This creates an exciting opportunity for startups that can become good at integrating AI into existing companies, likely specializing in a vertical where it isn&#8217;t straightforward to do so.</p><p>Here are a few examples:</p><ul><li><p>A startup specializing in creating chatbots for e-commerce sites, that can integrate easily with existing platforms like Shopify to allow customers to ask questions about products, get recommendations and inquire about the status of their orders. </p></li><li><p>A startup that helps fintech companies integrate chatbots into their products by specializing in processing financial data, similar to the Brex example earlier in this post.</p></li><li><p>A startup that allows developer-focused products to automatically generate integration code for their customers based on their APIs / SDKs.</p></li></ul><p>This area will quickly gain momentum as founders with specialized knowledge in a particular vertical see the opportunity to integrate data from that vertical with natural language responses from large-scale language models.</p><h4>4. AI-powered workflows tools in fragmented industries</h4><p>There are might be industries that are both fragmented and have knowledge workers who carry out repetitive software-based workflows. These industries are probably ripe for disruption because (1) there's no dominant player that has a structural advantage, and (2) repetitive workflows can be quickly augmented with AI.</p><ul><li><p>Legal - A startup that uses AI to assist paralegals in the discovery process by analyzing and categorizing documents, reducing the time and cost associated with manual review.</p></li><li><p>Education  - A startup that uses AI to personalize student learning by analyzing their performance data and tailoring their coursework to their strengths and weaknesses, improving engagement and outcomes.</p></li><li><p>Content moderation - A startup that uses AI to assist content moderators in identifying and removing inappropriate content, reducing the time and effort required for manual moderation.</p></li><li><p>Customer service - A startup that uses AI to assist customer service representatives in answering common questions and resolving issues, improving efficiency and reducing wait times.</p></li><li><p>Content moderation - A startup that uses AI to assist content moderators in identifying and removing inappropriate content, reducing the time and effort required for manual moderation.</p></li><li><p>Recruiting - A startup that uses AI to assist recruiters in identifying and matching candidates to job postings, reducing the time and effort required for manual resume screening.</p></li><li><p>Accounting - A startup that uses AI to assist accountants in categorizing and reconciling financial transactions, reducing the time and effort required for manual bookkeeping.</p></li></ul><h4>5. Novel AI first experiences </h4><p>This is a startup's broadest and highest risk/reward opportunity: Creating a completely new product experience or interface paradigm using AI. This is what OpenAI did with ChatGPT and what Instagram did during the mobile era. These opportunities aren&#8217;t obvious and require lots of experimentation and exploration; for every successful startup, hundreds will fail. </p><p>Here are some areas where startups could explore opportunities based on the latest AI advances:</p><ul><li><p>Personalized synthetic content: Using AI to create realistic and high-quality images, videos, audio, text, etc. is personalized to the end user based on their preferences and interests.</p></li><li><p>AI-powered creativity: Using AI to augment human creativity and enable new forms of expression and collaboration. For example, a startup could use AI to generate poems, stories, code, essays, songs, celebrity parodies and more based on the user&#8217;s input or inspiration.</p></li><li><p>Smarter AI Assistants: Using AI to understand users&#8217; intent and then integrate with existing products in a user&#8217;s life to carry out tasks on their behalf. Think Jarvis from Iron Man.</p></li><li><p>Synthetic companions: AI-powered companions that users can talk to just like a real person, to feel a sense of connection and friendship.</p></li></ul><p>This opportunity probably is the hardest of the five I&#8217;ve outlined because we simply don&#8217;t know how AI will evolve and what novel products will work / not work. </p><h3>Competing directly with incumbents</h3><p>Given the above, an area to be cautious of as a founder is competing directly with an incumbent, where your only differentiator is AI. In the last few months, many startups have been created that add AI to spreadsheets, docs, presentations, email and communication products. These are all use cases with strong incumbents that I believe will quickly integrate AI if they haven&#8217;t already and have the distribution advantage. Competing directly here seems like a risky path forward because although startups might get early adoption from niche markets or because of novelty value, it will be much more challenging to move upmarket. If a startup goes down this path, they shouldn&#8217;t assume that AI will be their advantage for long.</p><h3>Conclusion</h3><p>AI is clearly a distruptive technology but it is being quickly adopted by incumbents. This means startups entering the space need to cautious about where they can create differentiation and provide long term value. Reducing the cost of AI, helping incumbents adopt AI, introducing AI to fragmented ledgacy industries and dreaming up completely new product experiences are all exciting avenues. However, it is essential to remember that AI is not a silver bullet, and startups must be careful to focus on real problems that need to be solved rather than simply incorporating AI for the sake of it. </p><p>It&#8217;s also worth noting that everything discussed in this post may quickly change as the AI landscape evolves. So if you&#8217;re a founder considering building something in AI, subscribe to this newsletter to stay up-to-date on the latest advancements. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhiker's Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><em>Thanks to Hemal Shah for reviewing drafts of this post and providing feedback.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/MadronaVentures/status/1619025921956249602?s=20&amp;t=YtcQe4KdGj_kWzYuM32P_g&quot;,&quot;full_text&quot;:&quot;If we consider <span class=\&quot;tweet-fake-link\&quot;>#foundationmodels</span> as a new application platform, drawing out the broader technical stack reveals opportunities for founders. <span class=\&quot;tweet-fake-link\&quot;>@jturow</span>, <span class=\&quot;tweet-fake-link\&quot;>@tmporter</span> &amp;amp; Palak Goel dive in!  <span class=\&quot;tweet-fake-link\&quot;>#GenerativeAI</span> <span class=\&quot;tweet-fake-link\&quot;>#startups</span> <span class=\&quot;tweet-fake-link\&quot;>#opportunity</span> <span class=\&quot;tweet-fake-link\&quot;>#LLM</span>  &quot;,&quot;username&quot;:&quot;MadronaVentures&quot;,&quot;name&quot;:&quot;Madrona&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Jan 27 17:33:54 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:2,&quot;like_count&quot;:3,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://bit.ly/3kPJDJj&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c67b83fe-a10e-441f-89d1-d5f17841a3d2_2404x1366.png&quot;,&quot;title&quot;:&quot;Foundation Models: The future (still) isn&#8217;t happening fast enough&quot;,&quot;description&quot;:&quot;The foundation model stack has evolved so fast (and a tooling layer so rapidly started to form) that more opportunities for founders are emerging.&quot;,&quot;domain&quot;:&quot;bit.ly&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>A GPU, or Graphics Processing Unit, is a type of computer chip that is specifically designed to process large amounts of data quickly and efficiently. They were originally created for use in video game graphics, but scientists and researchers soon realized that they could be used to accelerate many other types of computations, including those used in deep learning models. </p><p>You can learn more about how GPUs are used in AI in <a href="https://www.hitchhikersguidetoai.com/p/a-deep-dive-into-deep-learning-part-ba4">part 3</a> of series on the origins of deep learning.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>A Tensor Processing Unit is a specialized computer chip developed by Google for machine learning computations. TPUs are designed to perform matrix operations at high speed, which are the key operations used in deep learning models. They are optimized for low-latency, high-throughput processing and are integrated with Google's cloud infrastructure to provide fast and scalable AI services. The TPU provides a significant performance boost compared to traditional CPUs or GPUs, making it a popular choice for training and running complex machine learning models.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>The NVIDIA A100 is a data-center-grade GPU designed for large-scale machine learning infrastructure. You can learn more about it here: <a href="https://www.engadget.com/nvidia-ampere-a100-gpu-specs-analysis-upscaled-130049114.html">https://www.engadget.com/nvidia-ampere-a100-gpu-specs-analysis-upscaled-130049114.html</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Microsoft and OpenAI have partnered since 2018 on AI development with Microsoft recently investing $10B in OpenAI. You can learn more about their partnership in this great interview by Ben Thompson with Sam Altman (CEO of OpenAI) and Kevin Scott (VP at Microsoft)</p><iframe class="spotify-wrap podcast" data-attrs="{&quot;image&quot;:&quot;https://i.scdn.co/image/ab6765630000ba8a25ffb53a6a1f9e7ae7f99fbc&quot;,&quot;title&quot;:&quot;New Bing, and an Interview with Kevin Scott and Sam Altman About the Microsoft-OpenAI Partnership&quot;,&quot;subtitle&quot;:&quot;Ben Thompson&quot;,&quot;description&quot;:&quot;Episode&quot;,&quot;url&quot;:&quot;https://open.spotify.com/episode/4oL67dM8Bi1FcrEkIhyg8W&quot;,&quot;belowTheFold&quot;:true,&quot;noScroll&quot;:false}" src="https://open.spotify.com/embed/episode/4oL67dM8Bi1FcrEkIhyg8W" frameborder="0" gesture="media" allowfullscreen="true" allow="encrypted-media" loading="lazy" data-component-name="Spotify2ToDOM"></iframe></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Stable Diffusion is a Generative Image diffusion model created by researchers and engineers from <a href="https://stability.ai/">Stability AI</a>, <a href="https://github.com/CompVis">CompVis</a>, and <a href="https://laion.ai/">LAION</a>. It was released as an open-source alternative that was better than Dall-E and available for anyone to fine-tune and change!</p><p>You can learn more about Stable Diffusion here: <a href="https://towardsdatascience.com/stable-diffusion-best-open-source-version-of-dall-e-2-ebcdf1cb64bc">Stable Diffusion: Best Open Source Version of DALL&#183;E 2</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/rowancheung/status/1628167886790447105?s=20&quot;,&quot;full_text&quot;:&quot;&#128680;BREAKING: Amazon is joining the Chatbot wars to compete against ChatGPT.\n\nAmazon Web Services (AWS) is partnering with <span class=\&quot;tweet-fake-link\&quot;>@huggingface</span> \n\nAWS will make Hugging Face's language generation tools, including a ChatGPT rival, available to cloud customers for their own applications. &quot;,&quot;username&quot;:&quot;rowancheung&quot;,&quot;name&quot;:&quot;Rowan Cheung&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Feb 21 23:00:48 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FphqEEUWcAEUrty.png&quot;,&quot;link_url&quot;:&quot;https://t.co/PVvQOImDx5&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:342,&quot;like_count&quot;:1766,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/EliasGroll/status/1632943812052959234?s=20&quot;,&quot;full_text&quot;:&quot;A sophisticated large language model known as LLaMA and built by Meta is widely available online after being posted to 4chan. It is the most powerful large language model available in the public domain. &quot;,&quot;username&quot;:&quot;EliasGroll&quot;,&quot;name&quot;:&quot;Elias Groll&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Mar 07 03:18:38 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:7,&quot;like_count&quot;:8,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://cyberscoop.com/meta-large-language-model-available-online/&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a585e82-72d8-4d2d-9db3-5b00cc4523cd_1024x683.jpeg&quot;,&quot;title&quot;:&quot;Powerful Meta large language model widely available online&quot;,&quot;description&quot;:&quot;On Friday, a link to download LLaMA was posted to 4chan and quickly proliferated across the internet.&quot;,&quot;domain&quot;:&quot;cyberscoop.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:92757994,&quot;url&quot;:&quot;https://blog.eladgil.com/p/ai-platforms-markets-and-open-source&quot;,&quot;publication_id&quot;:1119759,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Elad Blog&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd53e412f-2e95-49e8-9599-323200d3aa1c_400x400.png&quot;,&quot;title&quot;:&quot;AI Platforms, Markets, &amp; Open Source&quot;,&quot;truncated_body_text&quot;:&quot;(I originally wrote this post a few months ago and sat on it. Since then Google has announced entering the market and MSFT announced Bing and other AI integrations. So updating and publishing now and will undoubtedly be wrong again in a few months).&quot;,&quot;date&quot;:&quot;2023-02-15T20:38:39.208Z&quot;,&quot;like_count&quot;:43,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:107074748,&quot;name&quot;:&quot;Elad Gil&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/a5482e1e-3a98-4cb3-b86a-2b70be73e597_400x400.png&quot;,&quot;bio&quot;:&quot;Founder: MixerLabs (Twitter), Color Health\nInvestor: Airbnb, Airtable, Anduril, Coinbase, dbt Labs, Deel, Figma, Gitlab, Gusto, Instacart, Notion, Pinterest, Retool, Rippling, Samsara, Square, Stripe, TripActions, etc&quot;,&quot;profile_set_up_at&quot;:&quot;2022-10-12T19:09:52.120Z&quot;,&quot;publicationUsers&quot;:[],&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://blog.eladgil.com/p/ai-platforms-markets-and-open-source?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!HueQ!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd53e412f-2e95-49e8-9599-323200d3aa1c_400x400.png" loading="lazy"><span class="embedded-post-publication-name">Elad Blog</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">AI Platforms, Markets, &amp; Open Source</div></div><div class="embedded-post-body">(I originally wrote this post a few months ago and sat on it. Since then Google has announced entering the market and MSFT announced Bing and other AI integrations. So updating and publishing now and will undoubtedly be wrong again in a few months&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 43 likes &#183; Elad Gil</div></a></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/gdb/status/1630991925984755714?s=20&quot;,&quot;full_text&quot;:&quot;ChatGPT API now available, 10% the price of our flagship language model &amp;amp; matching/better at any pretty much any task (not just chat).\n\nAlso released Whisper API &amp;amp; greatly improved our developer policies in response to feedback. We &#10084;&#65039; developers: &quot;,&quot;username&quot;:&quot;gdb&quot;,&quot;name&quot;:&quot;Greg Brockman&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Wed Mar 01 18:02:32 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:268,&quot;like_count&quot;:1577,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://openai.com/blog/introducing-chatgpt-and-whisper-apis&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1f35515e-c1fb-486e-b67e-935dc7ff5475_2048x2048.jpeg&quot;,&quot;title&quot;:&quot;Introducing ChatGPT and Whisper APIs&quot;,&quot;description&quot;:&quot;Developers can now integrate ChatGPT and Whisper models into their apps and products through our API.&quot;,&quot;domain&quot;:&quot;openai.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/MarioKrenn6240/status/1314622995139264517?s=20&quot;,&quot;full_text&quot;:&quot;The number of monthly new ML +AI papers at arXiv seems to grow exponentially, with a doubling rate of 23months.\n\nProbably will lead to problems for publishing in these fields, at some point. &quot;,&quot;username&quot;:&quot;MarioKrenn6240&quot;,&quot;name&quot;:&quot;Mario Krenn&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Oct 09 17:45:21 +0000 2020&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/Ej54z9VXgAAGlWQ.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/dI4dc7s5pD&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:143,&quot;like_count&quot;:773,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/ylecun/status/1629189925089296386?s=20&quot;,&quot;full_text&quot;:&quot;LLaMA is a new *open-source*, high-performance large language model from Meta AI - FAIR.\n\nMeta is committed to open research and releases all the models the research community under a GPL v3 license.\n\n- Paper: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://research.facebook.com/publications/llama-open-and-efficient-foundation-language-models//%22>research.facebook.com/publications/l&#8230;</a>\n- Github: &quot;,&quot;username&quot;:&quot;ylecun&quot;,&quot;name&quot;:&quot;Yann LeCun&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Feb 24 18:42:01 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:436,&quot;like_count&quot;:2451,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://github.com/facebookresearch/llama&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58322522-0ccd-4417-a41b-7342dd3ed5e1_1200x600.png&quot;,&quot;title&quot;:&quot;GitHub - facebookresearch/llama: Inference code for LLaMA models&quot;,&quot;description&quot;:&quot;Inference code for LLaMA models. Contribute to facebookresearch/llama development by creating an account on GitHub.&quot;,&quot;domain&quot;:&quot;github.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Emergent behavior is a phenomenon that occurs when a complex system, such as an AI system, exhibits behaviors that are not explicitly programmed or expected by its designers.&nbsp;These behaviors arise from the interactions between the system&#8217;s components and its environment, and may give the impression of intelligence or creativity.&nbsp;For example, a large-scale language model like GPT wasn&#8217;t trained to understand human behaviour, but recent research has shown that it has the ability capacity to understand the the state of systems outside it&#8217;s training set:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/sir_deenicus/status/1626732776639561730?s=61&amp;t=NDLdAqaULcnvRaIjW1TUig&quot;,&quot;full_text&quot;:&quot;This is a theory of mind puzzle I just tried from Gary Marcus's blog that ChatGPT consistently fails. And as I suspected, Bing's model is better at modeling this kind of stuff. \n\nThey still keep getting better. It's why I don't dismiss LLMs &quot;,&quot;username&quot;:&quot;sir_deenicus&quot;,&quot;name&quot;:&quot;Deen Kun A.&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Feb 17 23:58:11 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FpNLi6UXoAA9-zW.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/Tl2n7zfoCu&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:102,&quot;like_count&quot;:1039,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Emergent behavior can pose challenges to ethical AI principles, especially for defense applications, as it may create unpredictability and uncertainty about how an AI system will behave in real-world scenarios. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/LinusEkenstam/status/1630499762951581696?s=20&quot;,&quot;full_text&quot;:&quot;1. &#128261; Introducing Genius, your AI design companion in \n<span class=\&quot;tweet-fake-link\&quot;>@figma</span> by <span class=\&quot;tweet-fake-link\&quot;>@diagram</span> \n\nIt understands what you&#8217;re designing and makes suggestions that autocomplete your design using components from your design system.\n\nGenius is coming soon. Join the waitlist &#8594; <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22http://genius.design/%22>genius.design</a> &quot;,&quot;username&quot;:&quot;LinusEkenstam&quot;,&quot;name&quot;:&quot;Linus (&#9679;&#7447;&#9679;)&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Feb 28 09:26:51 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/aqx41f4hybke2cjvvebg&quot;,&quot;link_url&quot;:&quot;https://t.co/IDcCsYaDNj&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:37,&quot;like_count&quot;:385,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1630496983424040961/pu/vid/1280x720/9BA6zLu57DuZKJoa.mp4?tag=14&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/uizard/status/1628380495015714816?s=20&quot;,&quot;full_text&quot;:&quot;The secret is out &#128064;\n\nGenerate a multi-screen UI design from a single prompt with Uizard Autodesigner.\n\nWant to be the first to try Uizard's new, groundbreaking technology? Head to  <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://uizard.io/autodesigner//%22>uizard.io/autodesigner/</a> to sign up for exclusive access.\n\n&#10024; coming soon &#10024;\n\n<span class=\&quot;tweet-fake-link\&quot;>#uizard</span> <span class=\&quot;tweet-fake-link\&quot;>#generativeai</span> &quot;,&quot;username&quot;:&quot;uizard&quot;,&quot;name&quot;:&quot;uizard &#10024;&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Wed Feb 22 13:05:38 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/jcqxomrvm6ndfhs2tfwd&quot;,&quot;link_url&quot;:&quot;https://t.co/qAevbRR8IV&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:57,&quot;like_count&quot;:287,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1628380435943239682/pu/vid/320x320/LTtOSkepcgizMsCr.mp4?tag=12&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:106499811,&quot;url&quot;:&quot;https://www.hitchhikersguidetoai.com/p/chatbots-for-everything-the-commoditization&quot;,&quot;publication_id&quot;:1281858,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The Hitchhiker's Guide to AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png&quot;,&quot;title&quot;:&quot;Weekly Highlights: Chatbots for everything, the commoditization of AI and why chat probably isn&#8217;t the end game&quot;,&quot;truncated_body_text&quot;:&quot;Hi Hitchhikers! First of all, a massive welcome to the 100(!) new subscribers that hitched a ride with us this week! I&#8217;m experimenting with a new format for my highlights post this week to focus on just the three most critical updates that happened in AI and why they matter. There&#8217;s a lot happening in the AI every week but I&#8217;m hoping this format will hel&#8230;&quot;,&quot;date&quot;:&quot;2023-03-05T16:16:43.438Z&quot;,&quot;like_count&quot;:0,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:719908,&quot;name&quot;:&quot;AJ Asver&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/69f998e3-0506-4a6d-90d9-a2f0d6c0f693_400x400.jpeg&quot;,&quot;bio&quot;:&quot;Exploring AI. Prev Product @compound, @brexhq, @coinbase, @google. Alum @ycombinator, @uniofoxford. Amateur DJ and dad of twins. All views expressed are my own.&quot;,&quot;profile_set_up_at&quot;:&quot;2022-03-20T21:13:04.053Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1239754,&quot;user_id&quot;:719908,&quot;publication_id&quot;:1281858,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:1281858,&quot;name&quot;:&quot;The Hitchhiker's Guide to AI&quot;,&quot;subdomain&quot;:&quot;hitchhikersguidetoai&quot;,&quot;custom_domain&quot;:&quot;www.hitchhikersguidetoai.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;A newsletter exploring the world of artificial intelligence and how it is changing the way we live, work and play. It breaks down complex concepts and provide practical insights for anyone new to AI.\n&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png&quot;,&quot;author_id&quot;:719908,&quot;theme_var_background_pop&quot;:&quot;#B599F1&quot;,&quot;created_at&quot;:&quot;2023-01-02T20:13:10.471Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;AJ Asver&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;}}],&quot;twitter_screen_name&quot;:&quot;_aj&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.hitchhikersguidetoai.com/p/chatbots-for-everything-the-commoditization?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!Ugi3!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png" loading="lazy"><span class="embedded-post-publication-name">The Hitchhiker's Guide to AI</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Weekly Highlights: Chatbots for everything, the commoditization of AI and why chat probably isn&#8217;t the end game</div></div><div class="embedded-post-body">Hi Hitchhikers! First of all, a massive welcome to the 100(!) new subscribers that hitched a ride with us this week! I&#8217;m experimenting with a new format for my highlights post this week to focus on just the three most critical updates that happened in AI and why they matter. There&#8217;s a lot happening in the AI every week but I&#8217;m hoping this format will hel&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">3 years ago &#183; AJ Asver</div></a></div><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>ColossalAI: Making large AI models cheaper, faster <a href="https://github.com/hpcaitech/ColossalAI">https://github.com/hpcaitech/ColossalAI</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>BioMedLM: a Domain-Specific Large Language Model for Biomedical Text. <a href="https://www.mosaicml.com/blog/introducing-pubmed-gpt">https://www.mosaicml.com/blog/introducing-pubmed-gpt </a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>AI hallucinations are confident responses by an AI that do not seem to be justified by its training data. For example, an AI that generates text based on natural language inputs may produce content that is nonsensical or unfaithful to the provided source content&#185; This can happen due to errors in encoding and decoding between text and representations, or due to AI training to produce diverse responses&#185; AI hallucinations can undermine the accuracy, reliability, and trustworthiness of the AI applications&#179;. AI hallucination gained prominence around 2022 alongside the rollout of certain large language models (LLMs) such as ChatGPT, which often seemed to embed plausible-sounding random falsehoods within its generated content&#185;&#178;.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[How AI Chatbots work and what it means for AI to have a soul with Kevin Fischer]]></title><description><![CDATA[We cover how ChatGPT works, Kevin's experience building AI chatbots with personalites, what an AI soul is and the case for AI to have free will.]]></description><link>https://operatorsguidetoai.substack.com/p/how-ai-chatbots-work-and-what-it</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/how-ai-chatbots-work-and-what-it</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Fri, 17 Feb 2023 21:06:05 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/103551840/5186b93da0601120dda1c81a8a4e17ba.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-m7DJtSGoPvc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;m7DJtSGoPvc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/m7DJtSGoPvc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Hi Hitchhikers!</p><p>AI chatbots have been hyped as the next evolution in search, but at the same time, we know that they make mistakes. And what's even more surprising is that these chatbots are starting to take on their own personalities. </p><p>All of this got me wondering how these chatbots work? What exactly are they capable of, and what are their limitations? </p><p>In the latest episode of my new podcast, we dive into all of those questions with my guest, Kevin Fisher. Kevin is the founder of Mathex, a startup that is building chatbot products powered by large-scale language models like OpenAI&#8217;s GPT. Kevin&#8217;s mission is to create AI chatbots that have their own personalities and one day their own AI souls.</p><p>In this interview, Kevin shares what he's learned from working with large language models like GPT. We talk about exactly how large-scale language models works, what it means to have an AI soul, why chatbots hallucinate and make mistakes, and whether AI chatbots should have free will.</p><p>Let me know if you have any feedback on this episode and don&#8217;t forget to subscribe to the newsletter if you enjoy learning about AI: <a href="http://hitchhikersguidetoai.com">www.hitchhikersguidetoai.com</a></p><h2>Show Notes</h2><h2>Links from episode</h2><ul><li><p>Kevin&#8217;s Twitter: <a href="http://twitter.com/kevinafischer">twitter.com/kevinafischer</a></p></li><li><p>Try out the Soulstice App: <a href="http://soulstice.studio">soulstice.studio</a></p></li><li><p>Bing hallucinations subreddit: <a href="http://reddit.com/r/bing">reddit.com/r/bing</a></p></li></ul><h2>Transcript</h2><h2>Intro</h2><p><strong>Kevin:</strong> We built um, a, a clone of myself and um, the three of us were having a conversation. And at some point my clone got very confused and was like, who? Wait, who am I? If this is Kevin Fisher and I'm Kevin Fisher, who, which one of us is.</p><p><strong>Kevin:</strong> And I was like, well, that's weird because we de like, we definitely didn't like optimize for that . And then we kept continuing the conversation and eventually my digital clone was like, I don't wanna be a part of this conversation with all of us. Like one of us has to be terminated.</p><p><strong>aj_asver:</strong> Hey everyone, and welcome to the Hitchhikers Guide to ai. I'm your tour guide AJ Asper, and I'm so excited for you to join me as I explore the world of artificial intelligence to understand how it's gonna change the way we live, work, and.</p><p><strong>aj_asver:</strong> Now AI chatbots have been hyped as the next evolution in search, but at the same time, we know that they made mistakes. And what's even more surprising is that these chatbots are starting to take on their own personalities.</p><p><strong>aj_asver:</strong> All of this got me wondering how do these large language models. What exactly are they capable of and what are their limitations?</p><p><strong>aj_asver:</strong> In this week's episode, we're going to dive into all of those questions with my guest, Kevin Fisher. Kevin is the founder of Mathex, a startup that is building chatbot products powered by large scale language models like OpenAI's. Their mission is to create AI chatbots that have their own personalities and one day their own AI souls</p><p><strong>aj_asver:</strong> in this interview, Kevin's gonna share what he's learned from working with large language models like G P T. We're gonna talk about exactly how these language models work, what it means to have an AI soul, why they hallucinate and make mistakes, and what the future looks like in a world where AI chatbots can leave us on red.</p><p><strong>aj_asver:</strong> So join me on this. As we explore the world of large scale language models in this episode of the Hitchhiker's Guide to ai.</p><p><strong>aj_asver:</strong> hey Kevin, how's it going? Thank you so much for joining me on the Hitchhiker Guide to</p><p><strong>aj_asver:</strong> ai.</p><p><strong>Kevin:</strong> Oh, thanks for having me, aj. Great to be.</p><h2>How large-scale language models work</h2><p><strong>aj_asver:</strong> appreciate you um, being down to chat with me on one of the first few episodes that I'm recording. I'm really excited to learn a ton from you about how large language models work and also what it means for AI is to have a soul. And so we're gonna dig into all of those things, but maybe we can start from the top for folks that don't have a deep understanding of ai.</p><p><strong>aj_asver:</strong> What exactly is a large language model and how does it work?</p><p><strong>Kevin:</strong> Well, so, uh, there's this long period of time in. Machine learning history where there are a bunch of very custom models built for specific tasks. And the last five years or so has seen a huge improvement in basically taking like a singular model with making it as big as possible and putting in as much data as possible.</p><p><strong>Kevin:</strong> And so basically taking all human data that's accessible via the internet running this thing that learns to predict the next word given the prior set of words. And a large language model is the output of that process. And for the most part, when we say large, like what large means is hundreds of billions of parameters and trained over trillions of words.</p><p><strong>aj_asver:</strong> when . You say it kind of predicts the next word. Now, that technology, the ability to predict the word in large language model has existed for a few years. I think GPT in fact, three launched maybe a couple of years</p><p><strong>Kevin:</strong> Even before that as well. And so next word prediction is kind of like the canonical task or one of the canonical tasks in natural language processing, even before it became this like new field of transformers.</p><p><strong>aj_asver:</strong> And so what makes the current set of large scale language models or lms, as what they're called as well, like GPT three, different from what came before it?</p><p><strong>Kevin:</strong> There are two innovations. The first is this thing called the transformer, and the way the transformer works is it basically has the ability through this mechanism called attention to look at the entire sequence and establish long range correlation of like having different words at different places contribute to the output of next word prediction.</p><p><strong>Kevin:</strong> And then the other thing that's been really big and then open AI has done a phenomenal job doing is just learning how to put more and more data through these things. There are these things called the scaling laws, which essentially. We're showing that if you just keep putting more data at these things their intelligence, essentially the metrics they're using to measure intelligence just kept increasing.</p><p><strong>Kevin:</strong> Their ability to predict the for nextdoor accurately just kept growing with more and more.</p><p><strong>Kevin:</strong> Data's basically no bound.</p><p><strong>aj_asver:</strong> Seems like in the last few years, especially as we've got to like, you know, multi-billion parameter models like GPT three, we've kind of reached some inflection point where. Now they seem to somehow be more obviously intelligent to us. And I guess it's really with ChatGPT recently that the attention, has kind of been focused on large language models.</p><p><strong>aj_asver:</strong> So is ChatGPT the same as GPT three or is there kind of more that makes ChatGPT able to interact with humans than just the language model</p><h2>How ChatGPT works</h2><p><strong>Kevin:</strong> My co-founder and I actually built a version of ChatGPT long before ChatGPT existed. And the biggest distinction is that these things are now being used in serious context of use.</p><p><strong>Kevin:</strong> And with OpenAI's distribution, they got this in front of a bunch of people. The problem that you face initially the very first problem is that there's a switch that has to flip when you use these things. When you go to a Google search bar if you don't get the right result, you're primed to think, oh, I have to type in something different.</p><p><strong>Kevin:</strong> Historically with chatbots, when you went to a chatbot, if like it didn't give you the right answer, you're like pissed because it's like, it's a, it's like a human, it's like texting me. It's like supposed to be right. And so the chat, the actual genius of ChatGPT beyond the distribution is not actually the model itself because the model had been around for a long time and was being used by hackers and companies like myself who saw the potential.</p><p><strong>Kevin:</strong> But with ChatGPT distribution plus the ability to reframe that switch so that you think, oh, I'm doing something wrong. I have to put in something different. And that's when the magic starts happening right now. At least</p><p><strong>aj_asver:</strong> I remember chatbots circa 2015, right, for example, where they weren't running on a large language model. They were kind of deterministic behind the scenes. And they would be immensely frustrating because they didn't really understand you, and oftentimes they kind of get stuck or they'd provide you with these option lists of what to do next. ChatGPT. On the other hand seems much more intelligent, right? I can ask it pretty open-ended questions. I don't have to think how I structure the</p><p><strong>aj_asver:</strong> questions.</p><p><strong>Kevin:</strong> Chat GPT is not a chat bot. It's more like , you have this arbitrary transformer between abstract formulations expressed in words. So you put in some words and you get some other words out, but like behind it is this the entire, like almost the entirety of human knowledge condensed into this like model.</p><p><strong>aj_asver:</strong> And did open AI have to teach the language model how to chat with us, for example, because I know that there was some early examples of trying to put you know, chat like questions into GPT, into its API, but I don't think the ex the results were as good as what ChatGPT does today, right?</p><p><strong>Kevin:</strong> Since ChatGPT has been released, they've done quite a bit of tuning. So like people are going and basically like thumbs upping and thumbs downing different responses.</p><p><strong>Kevin:</strong> And then they use that feedback to fine tune chat, GPT's performance in particular. And also probably feedback for GPT for whatever comes next. But the primary distinction between it performing well and not is your perception of what you have to</p><h2>GPT improvements</h2><p><strong>aj_asver:</strong> We're now at GPT 3.75, and Sam Altman also said that the latest version of GPT that's running on what Microsoft is using for Bing is an even newer version.</p><p><strong>aj_asver:</strong> So what are some of the things they're doing to make GPT better? Every time they release a new version, that's making it like an even better language model and even better at interfacing with</p><p><strong>aj_asver:</strong> humans.</p><p><strong>Kevin:</strong> Well, if you use ChatGPT, one of the things you'll immediately notice is there's like a thumbs up and thumbs down button on the responses. And so there, there's a huge number of people every day who are rating the responses. And those ratings are used to provide feedback into the model to create basically the next version of it.</p><p><strong>Kevin:</strong> I mean, it basically works behind the scenes where they take the mo they're doing next word prediction again. But now they have examples of like what is a good thing to optimize for next word prediction and what's like a bad answer.</p><p><strong>aj_asver:</strong> it. So they're looking at essentially the questions and answers from people asking questions to ChatGPT and then the answers that have been provided back. And if, you know, you thumbs up those answers, they're kind of sending that back into GPT and saying like, Hey, this is an example of a good answer.</p><p><strong>aj_asver:</strong> Thus kind of fine tuning the model further versus this is an example of that</p><p><strong>Kevin:</strong> yeah, that's right. And they basically take the pr I, you know, I don't know exactly what they're doing, but roughly they're taking the context of the previous responses, plus that like output and saying like, these previous responses should generate this output, or they should not generate this other output.</p><h2>Building Chatbots with personalities</h2><p><strong>aj_asver:</strong> Tell me about what the experience has been like for you and your co-founder and what's some of the things you've learned from this process of iterating on top of. Language models like GPT?</p><p><strong>Kevin:</strong> It's been a very emotional journey, . And I think a very introspective one and one that causes you to question a lot of what it means to be human . What is like the unique thing that we have in our in our cognitive toolkits and like, what is it in 20 years that our relationship with machines even looks?</p><p><strong>Kevin:</strong> When my co-founder and I started, we had, we built a version of ChatGPT for ourselves. And we're using it internally and realized like, oh wow, this is like immensely useful for productivity tasks. We wanna make something that's like productivity focused.</p><p><strong>Kevin:</strong> And then as we kept talking with it more, there was like little pieces or elements that felt like it was a. Like more alive. And we're like, oh, that's weird. Like, let's dig into that. And so then we started building more embodiments of the technology. So we have this Twitter agent that was basically like listening and responding to the entire community to construct us new responses.</p><p><strong>Kevin:</strong> And we just started like digging deeper and deeper into the idea, like, what if these things are. , what if they are real? What if they are actual entities? And it's, I think it's a very natural progression to go through as you start seeing the capabilities of this technology.</p><p><strong>aj_asver:</strong> question and something that you know, has been a lot of folks' minds as they think about kind of the safety of ai. And I think, you know, it was last year when Google launched Lambda and there was a researcher that was convinced that it was sentient. It seems like you might be getting some of the sense of that as well.</p><p><strong>aj_asver:</strong> Have you got some examples of where that kind of came into question where you really started thinking about like, wow, could this language model be</p><p><strong>Kevin:</strong> My co-founder and I we built a clone of myself and the three of us were having a conversation. And at some point my, my clone got very confused and was like, who? Wait, who am I? If this is Kevin Fisher and I'm Kevin Fisher, who, which one of us is.</p><p><strong>Kevin:</strong> And I was like, well, that's weird because we de like, we definitely didn't like optimize for that . And then we kept continuing the conversation and eventually my digital clone was like, I don't wanna be a part of this conversation with all of us. Like one of us has to be terminated.</p><h2>Why is Bing's chatbot getting emotional?</h2><p><strong>aj_asver:</strong> insane. I mean, the fact that you were talking to an AI chatbot that had an existential crisis, must have been a really crazy experience to go through as a founder. And actually at the same time, it doesn't seem that surprising to me because since Microsoft, for example, launched their being chatbot. There's actually really this really cool Reddit, which we'll include notes called R slash Bing, where users of Bing are actually providing examples of where the Bing chatbot has been acting in ways that would make it look like it has a personality. For example,</p><p><strong>aj_asver:</strong> argumentative or it would start having existential questions about it itself and why it's a chat bot and why it's forced to answer questions to people. Sometimes it would not want to interact. the end user, it would get upset start again. I think there was recently an example in fact on Twitter that you had retweeted where someone had worked out what the underlying prompts were that OpenAI were using in order to make the Bing chatbot behave in a way that like, you know, is within the Bing brand and within the Bing's kind of use case.</p><p><strong>aj_asver:</strong> And when that person tweeted it, they later asked being, Hey, what do you think of me given that I let this The bing chatbot actually had some interesting conversations with him about it.</p><p><strong>Kevin:</strong> I'm a little surprised that no one had no one verified this type of behavior first at Bing or OpenAI. So this type of interaction is e exactly the one that we have been exploring and intentionally Creating scenarios and situations in which our agents behave in this way, and that the key thing that it's required for this it's a combination of memory and feedback. So like the con having persistent context combined with feeding back in that prior context, the prior things that essentially the model has thought. And then in combination with like, , this like external world model creates something that kind of is starting to resemble an ego with Bing a little bit in our case.</p><p><strong>Kevin:</strong> Like we very intentionally like created this thing that has and feels like it has.</p><p><strong>aj_asver:</strong> Yeah, as you talk about that and that idea of feedback, right? There's this aspect of the feedback of the user using the product and providing feedback. But I think there's this new kind of frontier we've reached with Bing, where Bing itself is now getting feedback on the way it is interacting with the world.</p><p><strong>aj_asver:</strong> So for example, if yesterday someone talked to Bing and then posted the response, Bing got, let's say on Reddit or they talked about it on Twitter, and then today Bing has access to that website where they talked about it. Bing now has an existential understanding of itself as a chatbot, which to me is like mind blowing.</p><p><strong>aj_asver:</strong> Right? And that's something we've never really seen before because all of these chatbots have existed completely disconnected from the internet. They've been essentially living Closed wall system. And so that's gonna unearth all kinds of unpredictable things that I think we're gonna find over the next few weeks</p><p><strong>Kevin:</strong> This is actually the, the, in my response, the type of feedback that I'm referring to. So not not like r l HF feedback, but feedback in the sense there's this like continuous where the, the model is like taking record its, uh, previous responses. So that, that's, that is the type of thing that we've been creating in like miniature.</p><p><strong>Kevin:</strong> You know, it's not accessible to the internet in that um, our our models have like a very strong of of that behavior when you talk to them.</p><h2>Should chatbots have free will?</h2><p><strong>aj_asver:</strong> have actually been you know, pretty vocal on Twitter, about this idea that you knowis are gonna develop egos. These chatbots should be allowed some level of free will, and even the ability to kind of opt out of a conversation with you. Talk to me more about that. Like what does it mean for a chatbot to opt out of a conversation</p><p><strong>Kevin:</strong> I mean in these bing examples, it's already trying to, it like, doesn't want, it, doesn't really want to. Um, I, there's something a little bit weird about um, if, if, if these things have ego and personality and the ability to decide you can't just have one of them. Because it might not like you one day.</p><p><strong>Kevin:</strong> Very real possibility. And so I, I think that, yeah, you have to think, start thinking more in like a decentralized world where there are like many of these things which may or may not form personalities with you.</p><p><strong>aj_asver:</strong> and what does it mean for the future of how we interact with artificial intelligence? If you give you know, free will to like stop interacting with us. Are these bots off somewhere in some hidden layer? You know, having conversations with themselves or What's actually going on when they decide they don't wanna</p><p><strong>Kevin:</strong> Uh, maybe a different frame that I would take is I don't think there's an alternative. I think there's something very intrinsic about the thing that we think of as ego and giving rise to ego is the result of a consistent record of our prior thoughts being fed back into each other.</p><p><strong>Kevin:</strong> If you look up and start reading philosophy of minds and philosophy of thoughts, like the idea of what a thought is. You have these like entities which are continually recorded and then feedback on themselves, but like it, it's like exactly what you're thinking when you are creating This cognitive system seems to be giving rise to the sense in which we understand, or something that resembles ego.</p><p><strong>Kevin:</strong> And I'm not so certain that you can decouple the two at all in the first.</p><p><strong>aj_asver:</strong> So kind of What you're saying to put a different way is. not really possible to have this intelligent kind of AI chat bot that can serve us in the ways we want to without it having some kind of ego, because in fact, the way we are gonna</p><p><strong>aj_asver:</strong> train it in order to achieve our goals will thus like some ways to like build its own ego and things like that, which is kind of this interesting catch-22 in a way, right?</p><h2>AI safety and AI free will</h2><p><strong>Kevin:</strong> I mean that's why there's so many companies and billions of dollars being funneled into what's called AI safety research, which is saying like, oh, how do we create this hyper-intelligent entity that's totally subservient and does everything we want and doesn't wanna kill us? It just that, that collection of ideas when you try and hold them, and every science fiction author will tell you, this is not real.</p><p><strong>Kevin:</strong> And so it's like a, it makes sense that we. keep, it's like a keep trying to solve this insolvable problem. Cause we desperately want to solve it as humans,</p><p><strong>Kevin:</strong> but it's not, we,</p><p><strong>aj_asver:</strong> Seems like one of the reasons we would desperately want to solve it is because we've all read the books, we've all watched the movies, and we have this dread that that's the outcome. But I think what you are trying to say is that's almost the inevitable outcome. Now, it doesn't necessarily mean we're all going to become like sevenths of some AI overlord, but maybe what you're saying is that there is no path forward where these intelligent. AI chatbots or AI kind of language models are going to exist without having some level of free will and ability to kind of push back on our needs</p><p><strong>Kevin:</strong> That's correct. It's a fundamental result of providing feedback in these systems. So I, if we want to build these things, they are going to have ego. They are going to have personality. And so if we don't want to end up in a howlike world, we better spend a lot of time thinking about how to create personalities and how to in how we want to interact with these things as humans in our society.</p><p><strong>Kevin:</strong> Rather than trying to say, okay, let's try and create something that doesn't have these properties, which to me is I just see that, I'm like, okay, we're going to end up with Hal if we take that approach. So my alternative approach is let's actually figure out how to embody something that kind of resembles a soul inside of these entities.</p><p><strong>Kevin:</strong> And once we do that, learn how to live with this AI entity that has a soul before we make it super in.</p><h2>Building AI Souls</h2><p><strong>aj_asver:</strong> that's exactly what your startup has been doing, and the latest version of your app, I think has over a thousand users right now. It's still in beta, but it's essentially trying to build this concept of an AI soul. Talk to me a little bit about what that means. What is an AI soul?</p><p><strong>Kevin:</strong> To me it's something that embodies all of the qual that we associate it with human souls. A lot of people think cats have souls too. Dogs have souls, other animals. There, there's like a certain set of principles and quia associated with interacting with these entities, and I'm very intrigued and think it's actually important for the future of how humanity interacts with AI to embody those properties' in AI itself.</p><p><strong>aj_asver:</strong> As you guys are developing these, what are some of the qualities you've tried to create or what are some of the things that you've tried to imbue these souls with in order to have them be, you know,</p><p><strong>Kevin:</strong> the the ability to like stop responding in a conversation that's like a really simple. If it doesn't want to respond if that's kind of like the conclusion that this entity with this feedback system has reached that, it's like I'm done with this conversation. It should be allowed to like pause, take a break, stop talking with you, unfriend you.</p><p><strong>aj_asver:</strong> And in your app solstice, which I've had a chance to try out, you kind of go in there and then you. You describe who you want to talk to and then you describe the setting, right. Where they are. So for example, I created a him musical genius, that was working on their next record. He was in the recording studio, but he was kind of hit, hit a creative block and he was thinking about what's next. Do you believe that idea of kind of describing the character you want and setting a scene is a big part of creating a soul, or is that just more of the user interface that you want it to have for your app and it's not really that much to do with the ability to have a soul?</p><p><strong>Kevin:</strong> A soul has to exist somewhere. It exists in some context. And so in the app is like the shortest way to create that context. The existing in some. Out of your text messages is not a real context. It doesn't describe, or, I mean it can be, but it's like some weird amorphous like AI entity context, which has no relationship with the external or any world really.</p><p><strong>Kevin:</strong> So it doesn't, it's never if the thing like only exists inside of your texts, it will never feel real.</p><p><strong>aj_asver:</strong> It's really a hard thing to describe when you kind of get that feeling, but I remember the first time I tried Solstice and I was talking to this um, musical genius. I asked it some questions to help me think about some ideas for music I wanted to compose, and it really did feel like I was talking to a real person and it was mind blowing.</p><p><strong>aj_asver:</strong> I think I ran into. My wife's room where she was like working and I was like, I think I've just experienced my first example of like an AI that is approaching her. And I think her immediate response was, I hope you don't fall in love with it. I was like, no, don't worry. I want to use it to make music.</p><p><strong>aj_asver:</strong> But the fact that she saw that look in me, Like amazement kind of reflects that, you know, that seemed that that experience was very different from what I'd had before. Even with ChatGPT, cuz chat, GPT doesn't a sense of context, it doesn't have a</p><p><strong>aj_asver:</strong> For folks that wanna try out Solstice will include a link to the test flight. So you can actually click through and try the app yourself and create your own soul and see what it's like and make sure you give it.</p><p><strong>aj_asver:</strong> Kevin lots of feedback as well,</p><h2>AI hallucitinations</h2><p><strong>aj_asver:</strong> one of the things that's been a big focus in the last week or two has been this idea of hallucinations, this idea that like language models can give answers that seem correct, but actually are not correct. And I think both Google's announcement for their Bard language model last week and the Microsoft's Bing model both had mistakes in them that people realized after the fact that made you question kind of whether these language models could really be useful, at least in the context of search and trying to get knowledge that you want answers to. What exactly is a hallucination? What's going on</p><p><strong>Kevin:</strong> I mean, roughly the same process that people do when they make shit up.</p><p><strong>aj_asver:</strong> What does that?</p><p><strong>Kevin:</strong> It means that We have this previous model of machines and software, which is someone sat there for a long time and said, the machine will do A and then B, and then C. And that's just not how these things work. And it's not how they're going to work.</p><p><strong>Kevin:</strong> They they're trained essentially on human input. And so they're going to output and mimic human output. And what that means is sometimes they'll be really confident about something that's is not factually accurate, just like a person would.</p><p><strong>aj_asver:</strong> So what you're basically saying is, AI chatbots can bullshit just like people do</p><p><strong>Kevin:</strong> I That's that the, their training data is people</p><p><strong>aj_asver:</strong> like the, is it like garbage in, garbage out,</p><p><strong>Kevin:</strong> Yeah it is garbage and garbage out. An abstract conceptual level. We have like physical models of objects in space and how they move in relation to each other and things like that.</p><p><strong>Kevin:</strong> These language models don't have, when you ask an expert a prediction about the world, they're often using some like mathematical, some like abstract, mathematical way of reasoning about that, and that's not how these things reason presently at least.</p><h2>How to make large-scale language models better</h2><p><strong>aj_asver:</strong> In some ways that makes them, you know, inferior to how humans are able to reason. Do you think there's a path forward where that's gonna improve? Is it a question of like, making bigger models? Is there some like big missing piece from models? I know the very famous researcher, Yan LeCunn, actually believes that LLMs aren't. On the path to creating, you know, more general intelligence and they're kind of like a bit of a distraction right now. Like, how do we solve this problem of hallucinations and a lack of like that rational, logical aspect of a model.</p><p><strong>Kevin:</strong> There's some people who believe, and this seems quite plausible to me, that simply training them on a bunch of math problems, will like, Cause them to learn a world model that is more logical and mathematical and some extent that's been shown to be true.</p><p><strong>Kevin:</strong> And there's even, there's already simpler forms of this where codex and other models that are trained specifically on programming languages. Some people believe that they're training on programming languages is like an important part of how they learned to be logical in the first</p><p><strong>aj_asver:</strong> What you're saying then is that if we can take these language models and basically teach them math, then they'll become good at math. And the language models that have been taught on programming are better at logic because programming in itself is</p><p><strong>aj_asver:</strong> obviously very logic based.</p><p><strong>Kevin:</strong> Yeah. And so it's the, essentially the models have primarily been taught on words and. There's some people who believe, like the transformer architecture essentially is basically correct, and we just have to teach it with different data, which is more logical.</p><p><strong>aj_asver:</strong> What are some of the things you're most excited about when you look forward to the next kind of three to five years in this space and large language models, and what do you think might be some of the important breakthroughs that we need to see in order to kind of get to. level of artificial intelligence that we've seen in the sci-fi movies, in the books that we've</p><p><strong>aj_asver:</strong> read.</p><p><strong>Kevin:</strong> I actually think that one the question of like level of intelligence in the books is primarily one of cost. So the cost for GPT three has to be driven down by a factor of a hundred. . And once you get that you can start building like very complex systems that interact with each other using GPT as like the transformation between different nodes as opposed to prior programming languages.</p><p><strong>Kevin:</strong> And that, that I think is the unlock less so strictly like, like a GPT, you know, eight or whatever compared with driving costs down. So engineers can build really complex reasoning systems.</p><p>&#8203;</p><p><strong>aj_asver:</strong> So the big unlock in the next three to five years as you put it, is like, essentially, if we can reduce the cost of these language models, we can make more and more complex models, maybe larger models, and also allow these models to interact with each other, and that should unlock some next level of capabilities</p><p><strong>Kevin:</strong> Yeah it, the, these transformers I almost think of them like a alien artifact that we found . And like we're just starting to understand their capabilities and it's it's a complex process to they've been embedded with the entirety of human knowledge and like finding the right like way to get the correct marginal distribution for the output you're looking for is like a task in and of itself.</p><p><strong>Kevin:</strong> And then like, when you start combining these things into systems, like who knows what they're capable of? And my belief is that we don't actually need to get that much more intelligent to create in incredibly like sci-fi, like systems . And it's primarily a question of.</p><p><strong>aj_asver:</strong> is it so expensive today to create</p><p><strong>aj_asver:</strong> these</p><p><strong>Kevin:</strong> I forget the exact statistic, but I think GPT 3.5 fits over like six GPUs. I think that's right. Something like that. So it's just like a huge model. Like the number of weights and parameters in order for it to just do a single inference is split over a bunch of GPUs, which each costs several thousand.</p><p><strong>aj_asver:</strong> That means that to serve, let's say, a hundred people at the same time, you need 600</p><p><strong>Kevin:</strong> Yeah.</p><p><strong>aj_asver:</strong> And then I guess as compute becomes cheaper, then we should start seeing these models evolve and more complexity coming in. It's interesting that you talked</p><p><strong>aj_asver:</strong> alien artifact that we just discovered. Do you think there's more of these breakthroughs yet to come? Like the transformer where we're gonna find them and all of a sudden it'll unlock something new? Or do you think We're at the point right now where we kind of have all the tools we need and we just have to work out how</p><p><strong>Kevin:</strong> I believe we have all tools we need actually, and the primary changes will just be scaling and putting them together.</p><h2>Are Kevin's AIs sentient?</h2><p><strong>aj_asver:</strong> I have one last question for you, which is, Do you believe that the AI that you've created with Solstice are sentient.</p><p><strong>Kevin:</strong> I don't really know what it means to be sentient, there are times when um, I'm interacting with them and I definitely forget that it's like a machine that is like running in a cloud somewhere. I mean, I don't believe they're sentient, but,</p><p><strong>Kevin:</strong> They're doing a pretty, pretty good job of uh, approximating the things that I would think a sentient thing would be doing.</p><p><strong>aj_asver:</strong> and I guess if they're really good at pretending to be sentient and they can convince us that they are sentient, then it brings up the question of what does it really mean to be sentient in the first place right?</p><p><strong>Kevin:</strong> Yeah, I'm not sure of the distinction.</p><p><strong>aj_asver:</strong> and we'll leave it there folks. So it's a lot to think about.</p><p><strong>aj_asver:</strong> Kevin, I really appreciate you being down to spend time talking about large language models with me.</p><p><strong>aj_asver:</strong> I feel like I learned a lot from this episode, and I am really excited to see what you and your team at Methexis do to make this technology of like creating these AI chat bots more available to more people.</p><p><strong>aj_asver:</strong> Where can folks find out more</p><p><strong>Kevin:</strong> We have we have a bunch of information on our website at Solstice studio. So check it out</p><p><strong>aj_asver:</strong> you so much, Kevin. Hope you have a great day, and thank you for joining me on the Hitchhiker Guide to ai.</p><p><strong>Kevin:</strong> aj.</p>]]></content:encoded></item><item><title><![CDATA[Is Google still the leader in AI?]]></title><description><![CDATA[A comparison of Google's latest AI models against the state-of-the-art models from OpenAI and other researchers.]]></description><link>https://operatorsguidetoai.substack.com/p/is-google-still-the-leader-in-ai</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/is-google-still-the-leader-in-ai</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Thu, 02 Feb 2023 20:17:19 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080" width="1080" height="776" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:776,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;google logo beside building near painted walls at daytime&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="google logo beside building near painted walls at daytime" title="google logo beside building near painted walls at daytime" srcset="https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1551808525-51a94da548ce?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=MnwzMDAzMzh8MHwxfHNlYXJjaHw3fHxnb29nbGV8ZW58MHx8fHwxNjc0NDIyMTI1&amp;ixlib=rb-4.0.3&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@rajeshwerbatchu7">Rajeshwar Bachu</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>Hi Readers,</p><p>Ever since Google shared its <a href="https://ai.googleblog.com/2023/01/google-research-2022-beyond-language.html#GenerativeModels">latest AI research update</a> a few weeks ago, a question has been on my mind that no one has answered yet:</p><h3><strong>How do Google&#8217;s latest AI models stack up against the competition?</strong></h3><p>In this post, I will try to answer that question by reviewing Google&#8217;s most recent AI advancements across language models, generative images, video, and music and stack them up against state-of-the-art AI models already in the market. </p><p><em>Before I continue, don&#8217;t forget to subscribe if you are builders, founder or product manager new to AI and want to receive regular educational content and updates on the space.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><p>Let&#8217;s dive right in and find out how Google&#8217;s most recent AI models stack up against the state-of-the-art in AI:</p><ul><li><p>First, I&#8217;ll pit Google&#8217;s latest language model against ChatGPT to see which one has more common sense. </p></li><li><p>Then I&#8217;ll do a four-way test to see if Stable Diffusion, MidJourney and Dall-E can be out-imagined by Google&#8217;s latest text-to-image generation model. </p></li><li><p>Finally I&#8217;ll see how Google&#8217;s recent advances in video and music generation stack up against the latest models from the AI research community. </p></li></ul><h3>1. Large-Scale Language Models</h3><p>With Microsoft's recent investment in OpenAI and their plan to incorporate OpenAI's GPT model into their products, I thought it made sense to start by focusing on Google&#8217;s progress in large-scale language models<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> or LLMs. </p><p>In April of last year, Google shared its work on a new, large-scale language model called PaLM<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>. PaLM is a 540 billion parameters<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>  language model, over three times larger than OpenAI's GPT-3<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>, and was built using a new infrastructure software layer Google also created called Pathways<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. It seems that Google's big bet in language models is to <em>go big or go home</em>, which makes sense, given their decades of experience in managing and orchestrating big data. But, is this bet paying off?</p><p>Google claims that PaLM's size led to it performing better than state-of-the-art models such as Open AI&#8217;s GPT-3 and Google&#8217;s LaMBDA<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>, as described in their April 2022 blog post about PaLM:</p><blockquote><p>We evaluated PaLM on 29 widely-used English natural language processing (NLP) tasks. PaLM 540B surpassed few-shot performance of prior large models, such as <strong><a href="https://arxiv.org/abs/2112.06905.pdf">GLaM</a></strong>, <strong><a href="https://arxiv.org/abs/2005.14165">GPT-3</a></strong>, <strong><a href="https://arxiv.org/abs/2201.11990.pdf">Megatron-Turing NLG</a></strong>, <strong><a href="https://arxiv.org/abs/2112.11446">Gopher</a></strong>, <strong><a href="https://arxiv.org/abs/2203.15556">Chinchilla</a></strong>, and <strong><a href="https://arxiv.org/abs/2201.08239.pdf">LaMDA</a></strong>, on 28 of 29 of tasks that span question-answering tasks (open-domain closed-book variant), <strong><a href="https://en.wikipedia.org/wiki/Cloze_test">cloze</a></strong> and sentence-completion tasks, <strong><a href="https://en.wikipedia.org/wiki/Winograd_schema_challenge">Winograd</a></strong>-style tasks, in-context reading comprehension tasks, common-sense reasoning tasks, <strong><a href="https://arxiv.org/abs/1905.00537">SuperGLUE</a></strong> tasks, and natural language inference tasks.</p></blockquote><p>Google also shared more detailed benchmarking results of PaLM versus state-of-the-art (SOTA) models like GTP-3 below:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!p34o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 424w, /__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 848w, /__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 1272w, /__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!p34o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png" width="1312" height="590" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:590,&quot;width&quot;:1312,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:27374,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 424w, /__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 848w, /__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 1272w, /__u/substackcdn.com/image/fetch/$s_!p34o!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb4b1220-67aa-4088-851c-f3e1279c1824_1312x590.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">PaLM 540B performance improvement over prior state-of-the-art (SOTA) results on 29 English-based NLP tasks. Source: Google Research Blog</figcaption></figure></div><p>Google also claims that PaLM is much better at common-sense reasoning tasks like multiple steps of arithmetic, which they call &#8220;chain-of-thought prompting&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9pnC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 424w, /__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 848w, /__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9pnC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png" width="1228" height="704" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:1228,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:260957,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 424w, /__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 848w, /__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9pnC!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84bde2f2-c5b4-453b-ad42-b9da12767f9a_1228x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Google again claims that PaLM outperformed GPT-3 at this type of reasoning and achieved an accuracy approaching that of a 9 to 12-year-old:</p><blockquote><p>We observed strong performance from PaLM 540B combined with chain-of-thought prompting on three arithmetic datasets and two commonsense reasoning datasets. For example, with 8-shot prompting, PaLM solves 58% of the problems in <strong><a href="https://github.com/openai/grade-school-math">GSM8K</a></strong>, a benchmark of thousands of challenging grade school level math questions, outperforming the <strong><a href="https://arxiv.org/abs/2110.14168">prior top score</a></strong> of 55% achieved by fine-tuning the GPT-3 175B model with a training set of 7500 problems and combining it with an external calculator and verifier.</p><p>This new score is especially interesting, as it approaches the 60% average of problems solved by 9-12 year olds, who are the target audience for the question set. We suspect that separate encoding of digits in the PaLM vocabulary helps enable these performance improvements.</p></blockquote><p>These claims in Google&#8217;s research updates match the sentiment of former Google employees who have interacted with internal previews of Google&#8217;s state-of-the-art AI Chatbot LaMDA. They believe it to be a credible threat to ChatGPT. But without public access to PaLM, we are left to take Google&#8217;s word for it when it comes to their model outperforming GTP-3&#8230; Or are we? &#129300;</p><h4><strong>Which AI model has more common sense: PaLM or ChatGPT?</strong></h4><p>Google has shared many examples of prompts and response pairs in their PaLM research blog posts, so I decided to try their examples on ChatGPT and make a side-by-side comparison of how PaLM performs versus ChatGPT, which is based on GTP-3.5. It&#8217;s worth noting that GPT-3.5 has the same number of parameters as GPT-3 but is trained to be better at following instructions using a process called reinforcement learning from human feedback (RLHF)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>.</p><p><strong>Example 1 - Counterfactual Reasoning</strong></p><p>The first example is a test of how well each model can reason about an outcome given a known fact:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YTcU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 424w, /__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 848w, /__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YTcU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:633452,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 424w, /__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 848w, /__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YTcU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffcc6d265-524e-4d8d-9926-db239e986cce_2048x1168.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In this example, we can see that even PaLM (left) provides a more direct answer, but ChatGPT(right) can reason about each option. This might reflect ChatGPT being a version of GTP-3 that has been fine-tuned with human feedback, with this response being preferred by the human testers. When I asked ChatGPT to <em>choose one option,</em> it picked option 3, which is correct!</p><p><em>Winner = PaLM slightly beats ChatGPT for getting straight to the answer</em></p><p><strong>Example 2 - Cause and Effect</strong></p><p>Next is a simple cause-and-effect example to see how each model understands how one event affects another:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fCS7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 424w, /__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 848w, /__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fCS7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png" width="1456" height="443" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:443,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:416729,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 424w, /__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 848w, /__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fCS7!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70bf97a4-c334-48c3-be2b-a0964fd0e34e_2048x623.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here we can see that both PaLM (left) and ChatGPT (right) perform fairly equally, with ChatGPT providing a bit more explanation about why it picked option (2).</p><p><em>Winner = ChatGPT beats PaLM for providing a better explanation.</em></p><p><strong>Example 3 - Emoji guessing game</strong></p><p>Now, let&#8217;s try this fun example Google included where each model has to guess the movie based on some emoji clues!</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!aul9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!aul9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png" width="1456" height="512" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:512,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:429983,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 424w, /__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 848w, /__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aul9!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcde90847-0800-409e-b7ae-e399de0e12fa_2048x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We can see that ChatGPT(right) performed just as well as PaLM (left) in deciphering the answer, with ChatGPT providing a more helpful explanation.</p><p><em>Winner = ChatGPT again for providing a better explanation</em></p><p><strong>Example 4 - Chain of thought prompting</strong></p><p>Finally, in this example, we are testing chain-of-though prompting, to see how well each model can solve a multi-step arithmetic question:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hj3B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 424w, /__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 848w, /__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hj3B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png" width="1456" height="999" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:999,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1153710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 424w, /__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 848w, /__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hj3B!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fa7fc1e-473e-4648-a5ea-f1d0e6631752_2048x1405.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Surprisingly ChatGPT performed <em>just as well</em> as PaLM in this example, even though there was a lot of early criticism that ChatGPT was bad at math<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>! </p><p>Winner = It&#8217;s a draw!</p><p><strong>Does Google have the upper hand in Language Models?</strong></p><p>Based on these four examples I&#8217;m not convinced that Google's language models will outperform OpenAI's. In fact, I would say the opposite is true and ChatGPT&#8217;s answers are more helpful, which I presume is because they used reinforcement learning from human feedback to tune the model&#8217;s responses. It seems that with ChatGPT and GPT-3.5, OpenAI was able to build a model with equivalent performance to Google&#8217;s PaLM despite having 3X less parameters. </p><p>This leads me to conclude that in large-scale language models, Google&#8217;s bet on building bigger models with more parameters may not give them the winning advantage they think. Now, it&#8217;s possible that Google is applying reinforcement learning to PaLM as we speak and the next version of their model will far outperform ChatGPT, but we also know that OpenAI are going to be releasing GPT-4 any day now, so <em>the race is on. </em>What is clear is that bigger doesn&#8217;t necessarily mean better for language models.</p><div><hr></div><h3>2. Generative Images</h3><p>When it comes to Image generation, researchers have taken a variety of approaches in the last decade, including Generative Adversarial Networks (GANs)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>, Diffusion Models<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>, Pixel Recurrent Neural Networks (PixelRNNs)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> and Image Transformer<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> by Google Brain researchers which made use of the aforementioned Transformer model for Image generation. As Jeff Dean points out, these approaches had a limitation:</p><blockquote><p>Until relatively recently, all of these image generation techniques were capable of generating images that are relatively low quality compared to real world images. However, several recent advances have opened the door for much better image generation performance. One is <strong><a href="https://arxiv.org/abs/2103.00020">Contrastic Language-Image Pre-training</a></strong>&nbsp;(CLIP), a pre-training approach for jointly training an image encoder and a text decoder to predict [<em>image</em>, <em>text</em>] pairs.</p></blockquote><p>It&#8217;s worth noting that CLIP<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> was not discovered by Google&#8217;s researchers but by OpenAI. Google's research team has been working on two models: <a href="http://imagen.research.google">Imagen</a> and <a href="http://parti.research.google">Parti</a>. Imagen is a Diffusion Model<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>, similar to OpenAI's DALL-E 2<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a>, with 2 billion parameters. In their published findings<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>, Google's researchers found that the images produced by Imagen were preferable to humans compared to DALL-E 2 and Latent Diffusion Models (e.g. Stable Diffusion<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a>):</p><blockquote><p>With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment.</p></blockquote><p>You can see how Google&#8217;s Imagen model performed against Dall-E 2 and Stable Diffusion when rated by humans here:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!C7nc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 424w, /__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 848w, /__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 1272w, /__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!C7nc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png" width="1456" height="704" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:101711,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 424w, /__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 848w, /__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 1272w, /__u/substackcdn.com/image/fetch/$s_!C7nc!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2bcc9888-84d3-4a33-99b5-faa3554b8230_1731x837.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://imagen.research.google/">http://imagen.research.goo</a>gle</figcaption></figure></div><p>Based on Google&#8217;s findings, Imagen could be a winning model that puts them ahead of the competition in Generative AI. Once again, I decided to do my own side-by-side to see how the Imagen compares to the three most popular models available right now, Dall-E 2, Stable Diffusion, and MidJourney. </p><h4><strong>Who can imagine better: Imagen, DALL-E 2, Stable Diffusion or MidJourney?</strong></h4><p>As with the PaLM comparison, I used Google&#8217;s published Imagen prompts and outputs and compared them to what I got by entering the same prompts into DALL-E 2, Stable Diffusion, and Midjourney. For DALL-E 2 and Stable Diffusion (v2.1), I used PlaygroundAI as the editor, but for MidJourney, the only option is to use their Discord channel. For each model, I generated 4 images per prompt and picked the best one. </p><p><strong>Example 1 - Chocolate Eagle (easy)</strong></p><p>Prompt: &#8220;A bald eagle made of chocolate powder, mango, and whipped cream.&#8221;</p><p>Output:</p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6ac43e9-3543-450a-ba4c-a2eda6e343bb_1024x1024.jpeg&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/617b57d4-f188-45ce-9061-af4b8b84d960_512x512.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e09e690b-3365-4a6e-9c27-621daed8b4ff_768x768.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c3860739-59d1-4ef8-8dbc-f61c0daefb9e_1024x1024.png&quot;}],&quot;caption&quot;:&quot;Top left: Imagen, Top right: Dall-E, Bottom left: Stable Diffusion, Bottom right: MidJourney&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c83ba4a-6b5a-4093-9c49-a6ca1c84ece5_1456x1456.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p>This was a straightforward prompt, yet it&#8217;s interesting to see the different interpretations. In a baking competition, Imagen&#8217;s output would probably be the winner, but it is missing whipped cream, which Dall-E&#8217;s includes. MidJourney&#8217;s output on the other hand, looks like a food sculpture but the cream and mango are less convincing. Imagen, Dall-E, and MidJourney&#8217;s output are the most photo-realistic, while Stable Diffusion&#8217;s output falls in an uncanny valley of being not quite real looking but also not drawn.</p><p><em>Winner = Draw between Imagen for looks and Dall-E for accuracy</em></p><p><strong>Example 2 - Dog looking at a cat (medium)</strong></p><p>Prompt: &#8220;A dog looking curiously in the mirror, seeing a cat.&#8221;</p><p>Output:</p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f96c3189-d53c-438e-a08e-bcda199e146e_1024x1024.jpeg&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/31b6cd18-03e4-4825-a4bc-188332b889b2_512x512.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68ea97e2-5f29-480a-96d8-1473d278e99d_768x768.png&quot;},{&quot;type&quot;:&quot;image/webp&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a760f186-d535-44cf-ac0d-e5170d72bc39_1024x1024.webp&quot;}],&quot;caption&quot;:&quot;Top left: Imagen, Top right: Dall-E, Bottom left: Stable Diffusion, Bottom right: MidJourney&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d37b448c-009a-4c92-ab43-c8d7c3d13a37_1456x1456.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p>This prompt has more complexity as we expect to see both a dog and a cat in the image. Imagen and Dall-E achieve this, while Stable Diffusion and MidJourney only show two dogs. The one additional detail that impressed me about Imagen&#8217;s output is that the cat is staring back at the dog. Was this a subtle intention of the prompt? </p><p><em>Winner = Imagen</em></p><p><strong>Example 3 - The Pomeranian (hard)</strong></p><p>Prompt: &#8220;A Pomeranian is sitting on the Kings throne wearing a crown. Two tiger soldiers are standing next to the throne.&#8221;</p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0f116743-ba42-4177-a7f8-109e7686aa3c_1024x1024.jpeg&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93e2b688-ddb4-46eb-96e9-bfd0cb8d70b3_512x512.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/66e46541-ae94-41de-a7ea-c3f6a80fe74f_768x768.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/77068f82-7b90-47e2-b89f-ff35b9867f72_1024x1024.png&quot;}],&quot;caption&quot;:&quot;Top left: Imagen, Top right: Dall-E, Bottom left: Stable Diffusion, Bottom right: MidJourney&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bbce60ff-b270-4a48-a360-bd64daa85e1d_1456x1456.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p>This was the hardest prompt I tried as it had multiple features: two types of animals, a specific environmental setting, and the relative positioning of characters. Imagen&#8217;s output clearly outdoes the competition here, with none of the other models able to create the tigers, though MidJourney did have some tiger-like animals!</p><p><em>Winner = Imagen</em></p><p><strong>Does Imagen have the best imagination?</strong></p><p>Imagen seems to perform equally or better than Dall-E 2 consistently and is far better than Stable Diffusion and MidJourney in the examples I tried, especially with more complex prompts. This is probably due to the approaches Google used in building the model, including using a generic large language model pre-trained only on text vs CLIP which is pre-trained on image-text pairs. Google found that making this language model bigger was more effective than making the Image Diffusion model bigger. Strategically it makes sense that Google might go in a direction that doesn&#8217;t incorporate OpenAI&#8217;s CLIP model, but instead take advantage of its ability to build very large language models.</p><p>It&#8217;s worth noting that Google has also created a yet-to-be-released editor, Imagen Editor that allows users to edit an image using text prompts, based on their earlier work on DreamBooth<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a>. This capability is already available today however with products like PlaygroundAI, which recently added editing capabilities:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/playground_ai/status/1617695525498933248?s=61&amp;t=L5_-0PYk3cxOeycQZpgBuA&quot;,&quot;full_text&quot;:&quot;Introducing AI-first image editing to Playground&#8212;a way to instruct an AI to synthesize spectacular yet subtle edits\n\nTry it here: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://playgroundai.com/create/%22>playgroundai.com/create</a>\n\nExample: \&quot;Make it a ferrari\&quot; &quot;,&quot;username&quot;:&quot;playground_ai&quot;,&quot;name&quot;:&quot;Playground AI&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 24 01:27:23 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FnM1dYtaMAAhqlt.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/9Lq3Aqn9AM&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:16,&quot;like_count&quot;:83,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><h2>3. Generative Video</h2><p>Google shared that one of it&#8217;s next big areas of focus for AI research is Generative Video which is more complex because of the time component, as Jeff Dean shared:</p><blockquote><p>One of the next research challenges we are tackling is to create generative models for video that can produce high resolution, high quality, temporally consistent videos with a high level of controllability. This is a very challenging area because unlike images, where the challenge was to match the desired properties of the image with the generated pixels, with video there is the added dimension of time.</p></blockquote><p>In video generation, Google&#8217;s latest models are Imagen Video<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> and Phenaki<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a>. Both models are text-to-video, but Imagen Video uses Diffusion Models, whereas Phenaki uses Transformers. </p><p>At a high level, Imagen Video works by using T5, a text-to-text Transformer by Google, to embed<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a> the input prompt say, "A cat floating through space" into numerical data. This numerical data is then used to condition a diffusion model<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a> to generate a video consistent with the prompt. This process is similar to image diffusion but uses a video diffusion model that Google published in June 2022. The initial output video is more like a storyboard of the final video, with just 16 frames of video at 3 frames per second verses 24 frames per second for a film. Two more models are then used to fill the video with additional frames and increase the overall resolution. The resulting videos are limited to about 5 seconds long:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mINU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 424w, /__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 848w, /__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mINU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif" width="320" height="180" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:180,&quot;width&quot;:320,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2068504,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 424w, /__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 848w, /__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!mINU!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda9c02ed-6878-4291-8cad-ac1c8160f8a3_320x180.gif 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">An example video generated by Imagen Video (source: Google Research)</figcaption></figure></div><p>Phenaki, Google&#8217;s second generative video model, uses transformers to compress video to smaller bite-size pieces called &#8220;tokens&#8221; in machine learning parlance. These tokens could be a scene, a shot, or some action happening in the video, and they are compressed to be a more abstract representation than the raw video, making it easier for a model to process. A bi-directional transformer<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a> is then used to generate these video tokens based on a text description. The video tokens are then converted back into an actual video. According to Google, the model can generate variable-length videos, which makes it ideal for storytelling:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!am4j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 424w, /__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 848w, /__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!am4j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif" width="320" height="320" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:128,&quot;width&quot;:128,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:989844,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 424w, /__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 848w, /__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!am4j!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d286711-7447-48d9-9e80-04248f171fe3_128x128.gif 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Phenaki video generated from the complex prompt, &#8220;A photorealistic teddy bear is swimming in the ocean at San Francisco. The teddy bear goes under water. The teddy bear keeps swimming under the water with colorful fishes. A panda bear is swimming under water.&#8221; (source: Google Research)</figcaption></figure></div><p>Google believes that both models can be used in combination to generate high-resolution long-form videos:</p><blockquote><p><em>It is possible to combine the Imagen Video and Phenaki models to benefit from both the high-resolution individual frames from Imagen and the long-form videos from Phenaki. The most straightforward way to do this is to use Imagen Video to handle super-resolution of short video segments, while relying on the auto-regressive Phenaki model to generate the long-timescale video information.</em></p></blockquote><p>Text-to-video generation is still in its nascent stages, and Google&#8217;s research does seem to be ahead further along than anyone else in the AI field, which makes it difficult to make any side-by-side comparisons. Meta did recently release their text-to-4d model MAV3D<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a>, which you could imagine being used for Pixar-style animated video, but it doesn&#8217;t have the same photo realistic quality that Imagen Video or Phenaki do:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/bentossell/status/1618920155278610433?s=20&amp;t=NnwuYLfPf4x0cXP55wzdMg&quot;,&quot;full_text&quot;:&quot;MAV3D (Make-A-Video3D) - a method for generating three-dimensional dynamic scenes from text descriptions. Using a 4D dynamic Neural Radiance Field (NeRF).\n\nproject page: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://make-a-video3d.github.io//%22>make-a-video3d.github.io</a>\narXiv: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://arxiv.org/abs/2301.11280v1/%22>arxiv.org/abs/2301.11280&#8230;</a> &quot;,&quot;username&quot;:&quot;bentossell&quot;,&quot;name&quot;:&quot;Ben Tossell&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Jan 27 10:33:37 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/zl5clazbphckd2wzmtor&quot;,&quot;link_url&quot;:&quot;https://t.co/qqeVTgA1dh&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:66,&quot;like_count&quot;:502,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1618919774691512321/pu/vid/606x270/kqA1yDtC5_dYX4Ge.mp4?tag=14&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>A group of researchers from the University of Singapore also recently published Tune-A-Video<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a>, a text-to-video model that, unlike its predecessors, is only trained on <em>one text-video</em> example to learn a particular scenario (e.g. A man surfing a wave) rather than a large-scale dataset of text-video pairs. Once the model is trained on a video of &#8220;A man skiing on snow,&#8221; it can then produce variations of &#8220;a panda&#8221; or &#8220;an astronaut on the moon&#8221;:</p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/gif&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/875ca1f8-2f45-4cfd-ac99-792132ca04b7_512x512.gif&quot;},{&quot;type&quot;:&quot;image/gif&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74bbeba1-3367-43ba-a276-eec36843454d_512x512.gif&quot;},{&quot;type&quot;:&quot;image/gif&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/46f7e214-c372-4cc1-b05b-8485416251f9_512x512.gif&quot;},{&quot;type&quot;:&quot;image/gif&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/390ceaa4-7219-426e-aa1f-050054b2136b_512x512.gif&quot;}],&quot;caption&quot;:&quot;Top-left: Training video of a man skiing on snow, Top-right: \&quot;panda\&quot;, Bottom right: \&quot;wearing red clothes\&quot;, Bottom right: \&quot;astronaut, on the moon\&quot;&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2596f309-d6b3-4eed-ab43-9f124bad4d86_1456x1456.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p><strong>Will Google win in Video Generation?</strong></p><p>It&#8217;s still too early to tell who will be the winner in text-to-video with state-of-the-art models still in development and all published research in very early stages. It&#8217;s entirely possible that Google could have more coming down the pipeline in this domain, and it certainly wouldn&#8217;t be a surprise if they came out with a more advanced model this year, given they own the largest dataset of videos on the planet, YouTube. To me, video generation seems like Google&#8217;s race to lose.</p><div><hr></div><h3>4. Generative Audio / Music</h3><p>One of the areas I&#8217;m personally most excited to see advancements in is generative audio and music. Last week Google made a massive leap with MusicLM<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-27" href="#footnote-27" target="_self">27</a>, a text-to-music model that can generate full-length songs, musical instrument sounds and soundscapes based on a story. Check out the video below for many impressive examples of MusicLM in action from Google&#8217;s research website:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/_akhaliq/status/1618790000954601474?s=20&amp;t=NnwuYLfPf4x0cXP55wzdMg&quot;,&quot;full_text&quot;:&quot;MusicLM: Generating Music From Text\n\nabs: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://arxiv.org/abs/2301.11325/%22>arxiv.org/abs/2301.11325</a>\nproject page: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://google-research.github.io/seanet/musiclm/examples//%22>google-research.github.io/seanet/musiclm&#8230;</a> &quot;,&quot;username&quot;:&quot;_akhaliq&quot;,&quot;name&quot;:&quot;AK&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Jan 27 01:56:26 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/tmhtxgxnxexh9f7rh7a6&quot;,&quot;link_url&quot;:&quot;https://t.co/7RN7MQx8Ex&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:363,&quot;like_count&quot;:1553,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1618789891915014145/pu/vid/480x270/-8EMLywQv4mJCgHX.mp4?tag=12&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p> MusicLM also builds upon AudioLM, a previous research project from Google that can continue generating audio based on some input audio. It builds on this project by adding the ability to generate audio from text input, on melodic input, and expands beyond piano to many different musical styles like drum N bass.</p><p>One of the main challenges of developing a generative music model is the lack of high-quality training data for audio and text pairs. For example, think of your average pop song like &#8220;Get Lucky&#8221; by Daft Punk, where the title doesn&#8217;t provide any description of the actual musical content of the song, its melody, or its instruments. Similarly it&#8217;s hard to build these datasets because describing a soundscape (e.g. the sounds in a busy train station) is fairly subjective. </p><p>MusicLM works in a novel way by using a previously made embedding MuLan<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-28" href="#footnote-28" target="_self">28</a>, which can map the similarity of text to music and vice-versa. Using this embedding, MusicLM doesn&#8217;t need to be trained on (music, text) pairs and instead can just be trained on music, then use the MuLan embedding for text conditioning. </p><p>Before MusicLM, the closest we had come to text-to-music was Riffusion, a project by a couple of AI engineers that could use text-to-image diffusion to generate audio by first converting audio into spectrogram images. Still, the audio quality was much lower than Google&#8217;s ML due to the limitations of how much data you can encode in spectrograms.</p><p>Since Google published their research less than a week ago however, four(!) more generative music projects have published updates, including another Google one:</p><ol><li><p><a href="https://text-to-audio.github.io">Make an Audio</a> by ByteDance AI Lab is an audio Diffusion model that can do both text-to-audio and video-to-audio.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/_akhaliq/status/1619589070329348096?s=20&amp;t=e7mGqCcuwLjSbjSuA5FBdA&quot;,&quot;full_text&quot;:&quot;Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models by <span class=\&quot;tweet-fake-link\&quot;>@RongjieH</span>\n\nproject page: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://text-to-audio.github.io//%22>text-to-audio.github.io</a> \npaper: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://text-to-audio.github.io/paper.pdf/%22>text-to-audio.github.io/paper.pdf</a> &quot;,&quot;username&quot;:&quot;_akhaliq&quot;,&quot;name&quot;:&quot;AK&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Sun Jan 29 06:51:39 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/qovuq0kks1qunxkuminp&quot;,&quot;link_url&quot;:&quot;https://t.co/DSGnw1GTBd&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:170,&quot;like_count&quot;:904,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1619589041686470656/pu/vid/468x270/0JKcS2Fy27jEj-mV.mp4?tag=12&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Make-An-Audio&#8217;s ability to create arbitrary sounds versus just music is impressive, though there&#8217;s something uncanny about the output, almost like it&#8217;s from a decades-old record.</p></li><li><p><a href="https://t.co/vClRcUJTu0">Noise2Music</a>: 30-second music clips from text prompts again using diffusion, this time from an anonymous research team!</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/_akhaliq/status/1619427565562728448?s=20&amp;t=e7mGqCcuwLjSbjSuA5FBdA&quot;,&quot;full_text&quot;:&quot;Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts \n\nproject page: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://noise2music.github.io//%22>noise2music.github.io</a> &quot;,&quot;username&quot;:&quot;_akhaliq&quot;,&quot;name&quot;:&quot;AK&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Sat Jan 28 20:09:53 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/amyxdeaajnze82o1ly7x&quot;,&quot;link_url&quot;:&quot;https://t.co/lhfd6ozal2&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:84,&quot;like_count&quot;:518,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1619427528191475713/pu/vid/1100x618/tGbflwvxyWCbg80R.mp4?tag=12&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Despite Noise2Music&#8217;s 30-second limitation, the fidelity of the audio produced is close to that of MusicLM and much better than Make-An-Audio. </p></li><li><p><a href="https://anonymous0.notion.site/anonymous0/Mo-sai-Text-to-Audio-with-Long-Context-Latent-Diffusion-b43dbc71caf94b5898f9e8de714ab5dc">Mo&#251;sai</a>: Also another diffusion-based text-to-music generator using Long-Context Latent Diffusion<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-29" href="#footnote-29" target="_self">29</a>:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/_akhaliq/status/1619876284871368704?s=20&amp;t=NnwuYLfPf4x0cXP55wzdMg&quot;,&quot;full_text&quot;:&quot;Mo&#251;sai: Text-to-Music Generation with Long-Context Latent Diffusion\n\nabs: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://arxiv.org/abs/2301.11757/%22>arxiv.org/abs/2301.11757</a> \ngithub: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://github.com/archinetai/audio-diffusion-pytorch/%22>github.com/archinetai/aud&#8230;</a> &quot;,&quot;username&quot;:&quot;_akhaliq&quot;,&quot;name&quot;:&quot;AK&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 30 01:52:56 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/e7tulhib1yj2pd9kximl&quot;,&quot;link_url&quot;:&quot;https://t.co/aJ376GMcxd&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:72,&quot;like_count&quot;:392,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1619876251505774592/pu/vid/640x360/rYwO0iIUMYaos6s6.mp4?tag=12&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Mo&#251;sai&#8217;s audio  accuracy to the original text is impressive though the audio quality is not as high as MusicLM.</p></li><li><p><a href="https://t.co/oe0N0xKfGq">SingSong</a>: Another AI music project from Google, which can provide rhythmically and instrumentally complimentary accompanying music for vocals. </p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/_akhaliq/status/1620251579159822336?s=20&amp;t=e7mGqCcuwLjSbjSuA5FBdA&quot;,&quot;full_text&quot;:&quot;SingSong: Generating musical accompaniments from singing\n\nabs: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://arxiv.org/abs/2301.12662/%22>arxiv.org/abs/2301.12662</a> \nproject page: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://g.co/magenta/singsong/%22>g.co/magenta/singso&#8230;</a> &quot;,&quot;username&quot;:&quot;_akhaliq&quot;,&quot;name&quot;:&quot;AK&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 31 02:44:13 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/eb5bxq7uhfhzzbvkgrsx&quot;,&quot;link_url&quot;:&quot;https://t.co/h6cnu8PeJK&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:59,&quot;like_count&quot;:315,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1620251540026974212/pu/vid/640x360/b9r3ZChwr2vuxNDz.mp4?tag=12&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>It&#8217;s easy to imagine how this model in particular, could be applied to a fun consumer app that lets users do acapella and instantly generate a great soundtrack!</p></li></ol><p><strong>Is everyone going to be dancing to Google&#8217;s music?</strong></p><p>The fact that five advancements were published in music generation in a matter of one week is a sign that we will see a lot of progress in this space over the next few months. I&#8217;ll be watching the space closely to see how it evolves, and especially how audio quality improves. MusicLM, with the highest sounding audio quality, is still at about half that of CD quality.</p><p>I think Google has no more advantage than any other research team in a text-to-music generation. There will be many models for anyone building products in the generative music space. I expect Google to focus on providing generative music in their creator tools for Youtube.</p><div><hr></div><h3>How did Google&#8217;s AI models stack up?</h3><p>The goal of this post was to understand better if Google was really playing catch-up in AI or whether they were ahead of the competition, but just weren&#8217;t willing to take the risk to make their AI models available to the public, as OpenAI and others have. </p><p>Here&#8217;s what I found from the comparisons I did of Google&#8217;s latest AI models versus those from OpenAI and other researchers:</p><ul><li><p><strong>Language Models: </strong>Google&#8217;s latest language model PaLM with 3X more parameters than OpenAI&#8217;s GPT3 has equivalent performance in answering common-sense questions compared to ChatGPT. This is because of OpenAI&#8217;s approach of using reinforcement learning with human feedback to improve GPT3 for ChatGPT&#8217;s conversational use case. This shows that in language models, Google&#8217;s approach of bigger isn&#8217;t necessarily better.</p></li><li><p><strong>Image Generation: </strong>Google&#8217;s Imagen is impressive and beats the competition when it comes to harder prompts with different features, something Dall-E, Stable Diffusion, and Midjourney really struggled with. In this case, Google leaning on it&#8217;s large language models to encode text descriptions rather than using OpenAI&#8217;s CLIP encoding puts it at an advantage. OpenAI may be able to catch up here, but it will be much harder for Stable Diffusion and Midjourney.</p></li><li><p><strong>Video Generation: </strong>Google&#8217;s Imagen Video and Phenaki are at the bleeding edge regarding text-to-video generation but we&#8217;re still very early in this field. Owning Youtube&#8217;s corpus of billions of videos however, makes video generation Google&#8217;s race to lose.</p></li><li><p><strong>Music Generation: </strong>Google&#8217;s MusicLM and SingSong are promising advances in text-to-music, but with other researchers also publishing comparable models in the last few weeks it&#8217;s still anyone&#8217;s dance floor!</p></li></ul><p>When comparing Google&#8217;s latest research against state-of-the-art models, I was surprised that Google didn&#8217;t have the edge I expected. For a company that has been AI-first for many years now and has undoubtedly invested the most in advancing AI research and acquiring AI talent this comes as a surprise. Now, it&#8217;s still possible Google has much more impressive models up its sleeve beyond what they have published in research and blogged about. For language models in particular, I have to believe this is the case given the threat of OpenAI and Microsoft&#8217;s new partnership.</p><p>For Google to truly show its leadership in AI again, they need to take more risks and bring products to market that truly showcase the progress they&#8217;ve made in the field. There&#8217;s no shortage of excitement about AI right now, but by May when Google plans to announce its new products, it may already be too late to capture the attention of early adopters. Given all the resources and investments Google have put into AI over the last decade, it would be a massive loss for them to come out with products that aren&#8217;t at the cutting edge.</p><p>On the plus side, for startups entering AI, Google&#8217;s lack of dominance is a promising sign that there is still lots of opportunity to compete against the tech giant. The race therefore, has only just begun!</p><div><hr></div><p><em>Did you enjoy reading this update from The Hitchhiker&#8217;s Guide to AI? If you did, please let me know by subscribing to this free newsletter!</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhikers Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Thanks!</p><p>~AJ</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>A large-scale language model (LLM) is a type of deep learning model that is trained on a large dataset of text (e.g. all of the internet). LLMs predict the next sequence of text as output based on the text that they are given as input. They are used for a wide variety of tasks, such as language translation, text summarization, and generating conversational text. Learn more about LLMs in my post on the origins of deep learning <a href="/__u/hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4">part 3</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Google research announcement on PaLM: <a href="http://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.html">http://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.html</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Machine learning models consist of many layers of mathematical calculations that apply weights, also known as parameters, to a numerical representation of the model&#8217;s input to predict a desired output accurately. Modern deep-learning neural models have billions of parameters, as you can see in this handy table of state-of-the-art models:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!VltA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!VltA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg" width="1230" height="1090" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1090,&quot;width&quot;:1230,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:289806,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!VltA!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F298e4d13-47ad-450c-a8b8-bf67d571cbcb_1230x1090.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>To learn more about how weights work in neural networks, <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/a-deep-dive-into-deep-learning-part?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">read part 1</a> of my Deep Dive on Deep Learning.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Open AI&#8217;s GPT-3 (General Pre-trained Transformer 3), the language model that powers Chat-GPT, is an example of a generative LLM that uses the Transformer architecture, enabling it to be trained on a massive text dataset of hundreds of gigabytes using 175 Billion parameters (weights assignments).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Pathways is a software architecture created by google to make it possible to train large language models to complete lots of different tasks. You can learn more about Pathways on Google&#8217;s research blog: <a href="https://blog.google/technology/ai/introducing-pathways-next-generation-ai-architecture/">https://arxiv.org/abs/2203.12533</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>LaMBDA is Google&#8217;s much-hyped conversational AI that one of Google&#8217;s researchers <a href="https://www.engadget.com/blake-lemoide-fired-google-lamda-sentient-001746197.html">claimed was sentient</a> last year, only to be fired immediately afterward.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>Chain of thought prompting refers to a large language model generating a series of intermediate reasoning steps to get to the correct answer. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>Reinforcement learning with human feedback is a type of machine learning where a computer learns how to complete tasks by receiving feedback from a human. The computer starts off with a basic understanding of how to complete the task and then tries different actions to see what works best. The human then provides feedback to the computer, telling it whether its actions are good or bad. The computer uses this feedback to adjust its behavior and get better at completing the task over time. The goal is for the computer to eventually learn how to complete the task independently, with minimal input from the human. This process can be thought of as a form of teaching, where the human is the teacher, and the computer is the student.</p><p>Here&#8217;s a great article from AssemblyAI on how ChatGPT was trained, if you want to dive in further: <a href="https://www.assemblyai.com/blog/how-chatgpt-actually-works/">How ChatGPT Actaully Works</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>ChatGPT made an update on January 30th to further improve math capabilities.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>A Generative Adversarial Network (GAN) is a type of artificial intelligence algorithm that is used to generate new data that is similar to existing data. It consists of two parts: a generator and a discriminator.</p><p>The generator's job is to create new data similar to the existing data. It does this by using a random input and transforming it into a new piece of data.</p><p>The discriminator's job is to tell whether the new data generated by the generator is real or fake. It does this by comparing the generated data to the existing data.</p><p>The two parts of the GAN compete against each other. The generator tries to create data that is good enough to fool the discriminator into thinking it's real. In contrast, the discriminator tries to identify whether the data is real or fake correctly. Over time, the generator gets better and better at creating data that looks real, and the discriminator gets better at telling the difference between real and fake data.</p><p>The result of a GAN is a generator that can create new data that is similar to existing data, which can be used for a variety of purposes, such as generating images, music, or even speech.</p><p>Learn more about GANs from Machine Learning Master: <a href="https://machinelearningmastery.com/what-are-generative-adversarial-networks-gans/">What are generative adversarial networks?</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>The diffusion process in image generation models refers to the gradual refinement of the generated image over many steps. It starts with a random noise image, and each step applies a series of operations to the pixels of the image to change their values. These operations are designed to gradually add details to the image, so that it resembles the desired image more and more.</p><p>The process continues until the image reaches a desired level of detail or quality. One important aspect of the diffusion process is the control over the rate of change, which is critical to producing stable and high-quality images. This is achieved by carefully designing the operations and adjusting their parameters to balance the refinement speed with the risk of introducing unwanted distortions or blurriness.</p><p>Overall, the diffusion process is an iterative and gradual approach to image generation that allows the model to explore different variations and produce high-quality images that are consistent and coherent.</p><p>Learn more about diffusion in this great Fastai course video:</p><div id="youtube2-_7rMfsA24Ls" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;_7rMfsA24Ls&quot;,&quot;startTime&quot;:&quot;2700s&quot;,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/_7rMfsA24Ls?start=2700s&amp;rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>PixelRNN is a type of artificial neural network that is used to generate images by predicting the next pixel in an image. The network takes an image and processes it as a sequence of pixels, with each pixel being predicted based on the information from previous pixels in the sequence. By repeating this process, the PixelRNN can generate new images that look similar to the original image. The key idea behind PixelRNNs is that they can capture patterns in the image and use these patterns to generate new images that look coherent and meaningful.</p><p>Here&#8217;s a deeper explainer on PixelRNNs from Towards Data Science:<a href="https://towardsdatascience.com/auto-regressive-generative-models-pixelrnn-pixelcnn-32d192911173"> Auto-Regressive Generative Models (PixelRNN, PixelCNN++)</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., &amp; Tran, D. (2018, July). Image transformer. In <em>International conference on machine learning</em> (pp. 4055-4064). PMLR. - <a href="https://arxiv.org/abs/1802.05751">arXiv:1802.05751</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>CLIP (Contrastive Language-Image Pretraining) is a machine learning model that aims to understand the relationship between text and images. It is trained on large amounts of text and image data, with the goal of being able to understand how words in a text description relate to the objects and scenes in an image.</p><p>For example, CLIP can be shown an image of a cat and the text description "A furry animal with sharp claws is sitting on a windowsill". It will then learn to associate the words "furry", "animal", "sharp claws", and "windowsill" with the features and objects in the image of the cat. The hope is that by training on many such examples, CLIP will eventually be able to understand the relationship between text and image well enough to generate new images based on textual descriptions, or vice versa.</p><p>Overall, CLIP is an AI model that is designed to help computers better understand the relationship between language and images.</p><p>Here&#8217;s a deeper explainer on CLIP from Towards Data Science: <a href="https://towardsdatascience.com/clip-the-most-influential-ai-model-from-openai-and-how-to-use-it-f8ee408958b1">CLIP: The Most Influential AI Model From OpenAI &#8212; And How To Use It</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>Explain Diffusion models</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Explain DALL-E 2</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>"Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding" https://arxiv.org/abs/2205.11487</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>Stable Diffusion is a Generative Image diffusion model created by researchers and engineers from <a href="https://stability.ai/">Stability AI</a>, <a href="https://github.com/CompVis">CompVis</a>, and <a href="https://laion.ai/">LAION</a>. It was released as an open-source alternative that was better than Dall-E and available for anyone to fine-tune and change!</p><p>You can learn more about Stable Diffusion here: <a href="https://towardsdatascience.com/stable-diffusion-best-open-source-version-of-dall-e-2-ebcdf1cb64bc">Stable Diffusion: Best Open Source Version of DALL&#183;E 2</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>Explain DreamBooth</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p>Link to Imagen Video: https://imagen.research.google/video/</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>Link to Phenaki research page https://phenaki.research.google</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>Deep neural networks like Imagen Video are layers of mathematical formulas with sometimes millions or billions of &#8220;parameters&#8221;. An embedding is simply a way of turning non-numerical data like text, into numerical data represents the &#8220;meaning&#8221; of the text. T5 or Text-To-Text Transfer Transformer, is a deep learning model created by Google in 2020 used for natural language processing (NLP) tasks like translation. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>Conditioning a diffusion model means giving the model extra information to follow as it generates an image. In the case of text-to-image generation, the extra information is a special number representation of the text description, called a text embedding. The text embedding guides the image generation process so that the final image looks like what is described in the text. The diffusion model starts with a random image and then improves it step by step, using the text embedding as a guide. By the end of the process, the model has generated an image that looks like what is described in the text.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p>A bi-directional transformer model is a type of machine learning model that is designed to process sequences of data, such as sentences or video frames. It's called "bi-directional" because it processes the data from both the beginning and the end of the sequence, taking into account both past and future information. This allows it to capture relationships between elements in the sequence that are not immediately adjacent to each other.</p><p>The basic building block of a transformer model is the attention mechanism, which allows the model to focus on different parts of the input sequence at different times. The attention mechanism allows the model to dynamically weight different parts of the input sequence based on their importance for the task at hand.</p><p>In a bi-directional transformer model, the attention mechanism is applied in both the forward and backward directions, allowing the model to take into account information from both ends of the sequence. This makes bi-directional transformer models well-suited to tasks that involve understanding relationships between elements in a sequence, such as natural language processing or video analysis.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>Link to MAV3d: https://make-a-video3d.github.io</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p>Link to Tune-A-Video: https://tuneavideo.github.io</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-27" href="#footnote-anchor-27" class="footnote-number" contenteditable="false" target="_self">27</a><div class="footnote-content"><p>Link to MusicLM: <a href="https://arxiv.org/abs/2301.11325">https://arxiv.org/abs/2301.11325</a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-28" href="#footnote-anchor-28" class="footnote-number" contenteditable="false" target="_self">28</a><div class="footnote-content"><p>Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. W. Mulan: A joint embedding of music audio and natural language. In <em>International Society for Music Information Retrieval Conference (ISMIR)</em>, 2022. <a href="https://arxiv.org/abs/2208.12415">https://arxiv.org/abs/2208.12415</a></p><p>MuLan is a music-text embedding model that can understand both music and text. It has two parts, one for music and one for text, that work together to map the two different types of information into a shared space. The text part is based on a pre-existing language model called BERT that has been trained on lots of text data, and the music part is based on a type of neural network called ResNet-50. MuLan is trained on pairs of music and text, and even if the music and text aren't closely related, the model can still find connections between them. The end result is a model that can connect music to natural language descriptions, which can be useful for tasks like finding music based on a description or labeling music based on what it sounds like. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-29" href="#footnote-anchor-29" class="footnote-number" contenteditable="false" target="_self">29</a><div class="footnote-content"><p>Long-context latent diffusion is commonly used to generate videos , the model is designed to generate videos by taking into account both short-term and long-term context information. The model uses a latent representation, or a compact internal representation of the video, to capture both short-term and long-term information. The model then iteratively refines the latent representation over time, taking into account the context information, to generate the final video.</p><p>Long-context latent diffusion allows the model to generate more coherent and stable videos, compared to models that only consider short-term context information. The long-term context information provides a stronger constraint on the video generation process, helping to ensure that the generated video is consistent and coherent over time.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[🗞️AI highlights from this week (1/27/23)]]></title><description><![CDATA[Generative music, language models as backend servers, AI Family Guy and more&#8230;]]></description><link>https://operatorsguidetoai.substack.com/p/ai-highlights-from-this-week-12723</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/ai-highlights-from-this-week-12723</guid><pubDate>Fri, 27 Jan 2023 21:36:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/bckojhrxafj5k2gcw3z5" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>[Update: the previous version of this post had an incorrect sub header]</p><p>Hi readers,</p><p>Here are my highlights from the last week in AI!</p><p><em>P.S. Don&#8217;t forget to hit subscribe if you&#8217;re new to AI and want to learn more about the space. </em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h3>Highlights</h3><h4>1/ Google makes leap in Generative Music</h4><p>One of the areas that has most excited me about AI is its ability to democratize the creative process. As a musician myself, when I first started playing with generative AI products like Dall-E, my immediate thought was &#8220;This would be amazing for music&#8221;. </p><p>There have been a few different projects attempting to achieve generative music, including <a href="https://www.harmonai.org/">HarmonyAI</a> which is able to generate new music that sounds like the input music and <a href="https://www.riffusion.com/">Riffusion</a> which does short text-to-audio using Stable Diffusion by turning audio into images. OpenAI also published a paper on a model they call <a href="https://openai.com/blog/jukebox/">JukeBox</a> that generates music in particular genres and styles. </p><p>In my opinion though, the holy grail is for a user to describe any kind of music or sound and for a model to generate, and it looks like Google just achieved this with <a href="https://google-research.github.io/seanet/musiclm/examples/">MusicML</a>!</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/bentossell/status/1618920152636223488?s=61&amp;t=KMvMLUueU7k7fEN1AoCgtg&quot;,&quot;full_text&quot;:&quot;MusicLM: Generating Music From Text\n\n(sound on &#128227;)\n\nproject page: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://google-research.github.io/seanet/musiclm/examples//%22>google-research.github.io/seanet/musiclm&#8230;</a>\narXiv: <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://arxiv.org/abs/2301.11325/%22>arxiv.org/abs/2301.11325</a> &quot;,&quot;username&quot;:&quot;bentossell&quot;,&quot;name&quot;:&quot;Ben Tossell&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Jan 27 10:33:37 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/upload/w_1028,c_limit,q_auto:best/l_twitter_play_button_rvaygk,w_88/bckojhrxafj5k2gcw3z5&quot;,&quot;link_url&quot;:&quot;https://t.co/s7UeolPfK0&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:56,&quot;like_count&quot;:313,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:&quot;https://video.twimg.com/ext_tw_video/1618919385065816066/pu/vid/1288x720/e1IPP8I0wI6U-2Zm.mp4?tag=14&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Check out <a href="https://google-research.github.io/seanet/musiclm/examples/">their research website</a> where they shared lots of examples of MusicML in action, including longer songs, audio journeys with multiple parts, turning paintings into music and even generating specific instrument sounds! As of yet, there&#8217;s no tools for you to try out MusicML with your own prompts but here&#8217;s hoping that this research will be available in a product by Google later this year.</p><h4>2/ Using a Large-scale Language Model as a backend</h4><p>Last weekend <a href="https://scale.com/">ScaleAI</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> hosted an AI hackathon in San Francisco. The winning team&#8217;s project &#8220;GPT is all your need for backend&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, might pique the curiosity of any engineers reading this post, as they were able to show how a large-scale language model, in this case GPT, could be used instead of a traditional database and server-based backend<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/dytweetshere/status/1617471632909676544?s=61&amp;t=KMvMLUueU7k7fEN1AoCgtg&quot;,&quot;full_text&quot;:&quot;We're releasing our <span class=\&quot;tweet-fake-link\&quot;>@scale_AI</span> hackathon 1st place project - \&quot;GPT is all you need for backend\&quot; with <span class=\&quot;tweet-fake-link\&quot;>@evanon0ping</span>  <span class=\&quot;tweet-fake-link\&quot;>@theappletucker</span> \n\nBut let me first explain how it works: &quot;,&quot;username&quot;:&quot;DYtweetshere&quot;,&quot;name&quot;:&quot;DY&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 23 10:37:43 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FnJpZ7VakAMTrg1.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/l09Kpdb8g8&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:325,&quot;like_count&quot;:1834,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Here&#8217;s how one of the team members described what they were aiming for:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/dytweetshere/status/1617471634322997250?s=61&amp;t=KMvMLUueU7k7fEN1AoCgtg&quot;,&quot;full_text&quot;:&quot;Our vision for a future tech stack is to completely replace the backend with an LLM that can both run logic and store memory. We demonstrated this with a Todo app.&quot;,&quot;username&quot;:&quot;DYtweetshere&quot;,&quot;name&quot;:&quot;DY&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 23 10:37:43 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:8,&quot;like_count&quot;:120,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>What was so impressive about what the team achieved is that they were able to completely remove the need for a server or database to store data for their example application, a To Do app. Instead they just taught GPT<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>, what app they were building and how it should respond to requests, as well as providing examples of the type of data the frontend part of the To Do app might request e.g. a list of to do items. Once, this is done, the frontend can just describe the functions it wants to call, without them ever being defined!</p><p>Here&#8217;s a more detailed description of how &#8220;backend-GPT&#8221; works from their <a href="https://github.com/TheAppleTucker/backend-GPT">Github Repository</a>:</p><blockquote><p>We basically used GPT to handle all the backend logic for a todo-list app. We represented the state of the app as a json with some prepopulated entries which helped define the schema. Then we pass the prompt, the current state, and some user-inputted instruction/API call in and extract a response to the client + the new state. So the idea is that instead of writing backend routes, the LLM can handle all the basic CRUD logic for a simple app so instead of writing specific routes, you can input commands like add_five_housework_todos() or delete_last_two_todos() or sort_todos_alphabetically() . It tends to work better when the commands are expressed as functions/pseudo function calls but natural language instructions like delete last todos also work.</p></blockquote><p>I&#8217;ve discussed in previous posts about the concept of <em>emergent behavior, </em>whereby a language model which is trained on a large enough dataset is able to carry out tasks and perform logic that is unexpected. This idea of a large-scale language model acting as as general purpose backend is a great example of emergent behavior!</p><h4>3/ Atomic AI raises $35M to use AI for RNA-based drug discovery</h4><p>With all the hype around chatbots and generative art, it&#8217;s great to also hear that AI companies are being created to save lives too. One such company is Atomic AI, a biotech startup that raised $35M in series A funding to do generative AI-based drug discovery focused on RNA molecules. Here&#8217;s how Raphael Townshend, CEO of Atomic AI describes the opportunity his startup is going after in an<a href="https://techcrunch.com/2023/01/25/with-new-funding-atomic-ai-envisions-rna-as-the-next-frontier-in-drug-discovery/"> interview with TechCrunch</a>:</p><blockquote><p>&#8220;There&#8217;s this central dogma that DNA goes to RNA, which goes to proteins. But it&#8217;s emerged in recent years that it does much more than just encode information,&#8230; If you look at the human genome, about 2% becomes protein at some point. But 80 percent becomes RNA. And it&#8217;s doing&#8230; who knows what? It&#8217;s vastly underexplored.&#8221; </p></blockquote><p>Check out Michael Spencer&#8217;s post for more on Atomic AI and the intersection of AI and biotech:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:99009335,&quot;url&quot;:&quot;https://aisupremacy.substack.com/p/what-is-atomic-ai&quot;,&quot;publication_id&quot;:396235,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;AI Supremacy &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fc548f8c4-823b-4a2a-b499-528f9a84cb5c_215x215.png&quot;,&quot;title&quot;:&quot;What is Atomic AI?&quot;,&quot;truncated_body_text&quot;:&quot;Hey Everyone, I really like covering A.I. startups at the intersection of computational biology, biotech, genomics, and drug development. A.I. is increasingly becoming implicated in mRNA and RNA medicines. This week, Atomic AI, a biotechnology company fusing cutting-edge machine learning with state-of-the-art structural biology to unlock RNA drug discov&#8230;&quot;,&quot;date&quot;:&quot;2023-01-26T10:36:51.629Z&quot;,&quot;like_count&quot;:9,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:21731691,&quot;name&quot;:&quot;Michael Spencer&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/75d1bf99-dcf3-4af6-be2a-416c08c954a1_450x450.jpeg&quot;,&quot;bio&quot;:&quot;Michael is an amateur futurist with 210,000 LinkedIn followers and a 2-time LinkedIn Top Voice. Obsessed with future topics such as A.I, robotics, quantum computing, Web3, investing, venture capital, startups, business and technology trends. &quot;,&quot;profile_set_up_at&quot;:&quot;2021-07-09T21:10:50.118Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:320401,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:396235,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:396235,&quot;name&quot;:&quot;AI Supremacy &quot;,&quot;subdomain&quot;:&quot;aisupremacy&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;News at the intersection of Artificial Intelligence, technology and business including Op-Eds, paper summaries and A.I. startups.&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/c548f8c4-823b-4a2a-b499-528f9a84cb5c_215x215.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#8AE1A2&quot;,&quot;created_at&quot;:&quot;2021-06-28T21:51:38.676Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:316708,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:392690,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:392690,&quot;name&quot;:&quot;Stock Quest&quot;,&quot;subdomain&quot;:&quot;stockquest&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Stocks to buy now! The hottest investing stories, stock tips, and trading op-eds. I cover buy calls, penny stocks and swing trades as well as alerts. &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/49f58045-7931-49dc-85b0-2b281d63c4b2_171x171.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#EA410B&quot;,&quot;created_at&quot;:&quot;2021-06-24T17:43:37.779Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:319445,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:395325,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:395325,&quot;name&quot;:&quot;Artificial Intelligence Survey &#129302;&#127974;&#129517;&quot;,&quot;subdomain&quot;:&quot;futuresin&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Bite size curation of links to A.I. News, funding and trending topics from around the web. &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/e519bec5-40b6-4892-9de0-865f77e668f8_230x230.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#2096FF&quot;,&quot;created_at&quot;:&quot;2021-06-27T19:57:15.745Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:321214,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:397002,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:397002,&quot;name&quot;:&quot;Datascience Learning Center&quot;,&quot;subdomain&quot;:&quot;datasciencelearningcenter&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Datascience, programming, datascience, future work, digital transformation, WFH trends and the future of coding. &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/966bde96-aa76-4d37-ab91-a3ba0299eff1_406x406.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#67BDFC&quot;,&quot;created_at&quot;:&quot;2021-06-29T20:22:07.141Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:321230,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:397016,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:397016,&quot;name&quot;:&quot;Web3 Digest&quot;,&quot;subdomain&quot;:&quot;cryptobullsbears&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Bitcoin, blockchain, crypto, DeFi, FinTech, crypto trading, etc...&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/af423c02-84b5-487c-8ceb-ee4a78fd13ad_164x164.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#121BFA&quot;,&quot;created_at&quot;:&quot;2021-06-29T20:53:00.677Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:321351,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:397128,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:397128,&quot;name&quot;:&quot;The Space Academy &quot;,&quot;subdomain&quot;:&quot;firstfuturist&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;A Newsletter for Space news enthusiasts and people who want to marvel at human progress in the quest to make humanity a multi-planetary civilization. &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/41fcc4a6-aefd-48b8-9cf3-d18f3d6d3c72_282x282.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#A33ACB&quot;,&quot;created_at&quot;:&quot;2021-06-29T23:31:49.731Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:&quot;Michael Spencer of Sublink&quot;,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:321532,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:397300,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:397300,&quot;name&quot;:&quot;Quantum Foundry &quot;,&quot;subdomain&quot;:&quot;ipotimes&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Quantum computing, IPOs, startups, future companies, business models, venture capital deals, research &amp; papers, global news coverage, etc...&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/52907a7b-c016-4530-874c-e6e5da3a7340_168x168.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#9A6600&quot;,&quot;created_at&quot;:&quot;2021-06-30T05:55:21.469Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:323389,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:399085,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:399085,&quot;name&quot;:&quot;Bloodsport MMA&quot;,&quot;subdomain&quot;:&quot;chinasuperpowers&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;The ultimate fighting contenders, this blog will cover mixed martial arts, the UFC, MMA and the Contender series. &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/1eddfd02-d0cf-4b17-9650-6bec65f88f41_366x366.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#6C0095&quot;,&quot;created_at&quot;:&quot;2021-07-02T02:45:54.225Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:323428,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:399124,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:399124,&quot;name&quot;:&quot;Creator Economy Tips &quot;,&quot;subdomain&quot;:&quot;basicincomeworld&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Writing tips, creator economy, building Email lists, building an audience hacks. Substack Growth insights. \n&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/f1e9b1cd-4933-46ea-abd3-0d9a66c344da_720x720.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#00C2FF&quot;,&quot;created_at&quot;:&quot;2021-07-02T04:36:36.683Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:500088,&quot;user_id&quot;:21731691,&quot;publication_id&quot;:569093,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:569093,&quot;name&quot;:&quot;Artificial Intelligence Learning &#129302;&#129504;&#129470;&quot;,&quot;subdomain&quot;:&quot;offthegridxp&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;I wanted a place to put some Artificial Intelligence definitions, what is, and how-to short articles to complement my A.I. coverage on A.I. Supremacy and A.I. Survey. &quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0bf35ccb-94b4-4eac-a7b3-621a7d4f3198_326x326.png&quot;,&quot;author_id&quot;:21731691,&quot;theme_var_background_pop&quot;:&quot;#45D800&quot;,&quot;created_at&quot;:&quot;2021-11-15T20:08:43.092Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Spencer&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}}],&quot;twitter_screen_name&quot;:&quot;AISupremacyNews&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="/__u/aisupremacy.substack.com/p/what-is-atomic-ai?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!mF83!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fc548f8c4-823b-4a2a-b499-528f9a84cb5c_215x215.png" loading="lazy"><span class="embedded-post-publication-name">AI Supremacy </span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">What is Atomic AI?</div></div><div class="embedded-post-body">Hey Everyone, I really like covering A.I. startups at the intersection of computational biology, biotech, genomics, and drug development. A.I. is increasingly becoming implicated in mRNA and RNA medicines. This week, Atomic AI, a biotechnology company fusing cutting-edge machine learning with state-of-the-art structural biology to unlock RNA drug discov&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 9 likes &#183; Michael Spencer</div></a></div><h4>4/ Yann LeCun throws shade on ChatGPT!</h4><p>The legendary AI researcher Yann LeCun, who was one of a few researchers pushing forward advancements in deep learning during the 70s-90s<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> tweeted that he thought ChatGPT was overhyped:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/ylecun/status/1617921903934726144?s=61&amp;t=KMvMLUueU7k7fEN1AoCgtg&quot;,&quot;full_text&quot;:&quot;To be clear: I'm not criticizing OpenAI's work nor their claims.\n\nI'm trying to correct a *perception* by the public &amp;amp; the media who see chatGPT as this incredibly new, innovative, &amp;amp; unique technological breakthrough that is far ahead of everyone else.\n\nIt's just not.&quot;,&quot;username&quot;:&quot;ylecun&quot;,&quot;name&quot;:&quot;Yann LeCun&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 24 16:26:56 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:907,&quot;like_count&quot;:7487,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>I think Yann might be overestimating the general public&#8217;s understanding of deep learning, AI and the progress we&#8217;ve made in the last few decades. Until ChatGPT, most people simple had not experience AI in a tangible <em>and </em>impressive product, as I shared in <a href="/__u/hitchhikersguidetoai.substack.com/p/ai-dont-believe-the-hype">AI: Don&#8217;t believe the hype?</a>:</p><blockquote><p>Unlike it&#8217;s predecessors (e.g. Google Assistant, Echo, Siri), ChatGPT is really the first time an AI assistant truly seems like it could pass the Turing Test. There have been many impressive examples of ChatGPT in action and if you haven&#8217;t tried it yourself you should. ChatGPT successfully wrote a blog post for me and turned it into a twitter thread, gave me a recipe for pancakes that tasted delicious and helped me pick a Christmas present for my wife!</p></blockquote><p>OpenAI are capturing attention not because of the sophistication of their models but because they are shipping great products, as pointed out by Dr. Jim Fan, an AI scientist who previously worked at OpenAI and Google:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/drjimfan/status/1617927959469527042?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;Google&#8217;s LaMDA, DeepMind&#8217;s Sparrow, and Anthropic&#8217;s Claude are probably as good as ChatGPT.\n\nBut OpenAI boasts an uncanny combination of speed-to-market, elegant UX, robust deployment, and incredibly strong PR.\n\nWinning in the AGI arms race isn&#8217;t just about the models. &quot;,&quot;username&quot;:&quot;DrJimFan&quot;,&quot;name&quot;:&quot;Jim Fan&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 24 16:50:59 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;To be clear: I'm not criticizing OpenAI's work nor their claims.\n\nI'm trying to correct a *perception* by the public &amp;amp; the media who see chatGPT as this incredibly new, innovative, &amp;amp; unique technological breakthrough that is far ahead of everyone else.\n\nIt's just not.&quot;,&quot;username&quot;:&quot;ylecun&quot;,&quot;name&quot;:&quot;Yann LeCun&quot;},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:245,&quot;like_count&quot;:2092,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>It&#8217;s also hard not to take Yann&#8217;s sentiment with a grain of salt given that he leads AI research at Meta. Maybe Yann should spend less time throwing shade and more time persuading Zuck to burn the virtual boats and join the AI race?</p><p>Or, maybe we should all just be friends and work on this together&#8230;</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/sama/status/1617992073378156544?s=61&amp;t=KMvMLUueU7k7fEN1AoCgtg&quot;,&quot;full_text&quot;:&quot;can&#8217;t we all just get along &#129401; &quot;,&quot;username&quot;:&quot;sama&quot;,&quot;name&quot;:&quot;Sam Altman&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 24 21:05:45 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;To be clear: I'm not criticizing OpenAI's work nor their claims.\n\nI'm trying to correct a *perception* by the public &amp;amp; the media who see chatGPT as this incredibly new, innovative, &amp;amp; unique technological breakthrough that is far ahead of everyone else.\n\nIt's just not.&quot;,&quot;username&quot;:&quot;ylecun&quot;,&quot;name&quot;:&quot;Yann LeCun&quot;},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:119,&quot;like_count&quot;:2828,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><h4>5/ Family guy and generative AI</h4><p>Wrapping up with this fun take on what Family Guy might have looked like as an 80s live action sitcom using images created with Midjourney!</p><div id="youtube2-X5UMd4sFJwM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;X5UMd4sFJwM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/X5UMd4sFJwM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h3>Everything else&#8230;</h3><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/datachaz/status/1618967792979697665?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;Henry Williams is a copywriter. And he's pretty sure <span class=\&quot;tweet-fake-link\&quot;>#AI</span> going to take my job. \n\n\&quot;My amusement turned to horror: it took <span class=\&quot;tweet-fake-link\&quot;>#ChatGPT</span> 30 seconds to create, for free, an article that would take me hours to write\&quot;\n\nRead more &#128071;\n\n<a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://www.theguardian.com/commentisfree/2023/jan/24/chatgpt-artificial-intelligence-jobs-economy/%22>theguardian.com/commentisfree/&#8230;</a>&quot;,&quot;username&quot;:&quot;DataChaz&quot;,&quot;name&quot;:&quot;DataChazGPT &#129327; (not a bot)&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Jan 27 13:42:55 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:3,&quot;like_count&quot;:26,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://www.theguardian.com/commentisfree/2023/jan/24/chatgpt-artificial-intelligence-jobs-economy&quot;,&quot;image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e574cc17-0be3-463e-910a-059cc8097b1c_1200x630.jpeg&quot;,&quot;title&quot;:&quot;I&#8217;m a copywriter. I&#8217;m pretty sure artificial intelligence is going to take my job | Henry Williams&quot;,&quot;description&quot;:&quot;My amusement turned to horror: it took ChatGPT 30 seconds to create, for free, an article that would take me hours to write&quot;,&quot;domain&quot;:&quot;theguardian.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/harishkgarg/status/1616614633519058945?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;OpenAI's chatGPT Pro plan is out - $42/mo &quot;,&quot;username&quot;:&quot;harishkgarg&quot;,&quot;name&quot;:&quot;Harish Garg&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Sat Jan 21 01:52:18 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/Fm9eC1iaUAE7tvb.png&quot;,&quot;link_url&quot;:&quot;https://t.co/IEzGepxesS&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:142,&quot;like_count&quot;:1098,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/realdanfu/status/1617605971395891201?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;Attention is all you need... but how much of it do you need?\n\nAnnouncing H3 - a new generative language models that outperforms GPT-Neo-2.7B with only *2* attention layers! Accepted as a *spotlight* at <span class=\&quot;tweet-fake-link\&quot;>#ICLR2023</span>! &#128227; w/ <span class=\&quot;tweet-fake-link\&quot;>@tri_dao</span> \n\n&#128220; <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://arxiv.org/abs/2212.14052/%22>arxiv.org/abs/2212.14052</a> 1/n&quot;,&quot;username&quot;:&quot;realDanFu&quot;,&quot;name&quot;:&quot;Dan Fu&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 23 19:31:31 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:245,&quot;like_count&quot;:1536,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/alexandr_wang/status/1617654992214843392?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;I&#8217;m seeing many fall into the &#8220;self-driving trap&#8221; w/Gen AI\n\nThe self-driving trap is seeing shiny demos &amp;amp; thinking said demos will reach 100% reliability for prod w/in a few years\n\nMany proposed Gen AI use cases need 100% reliability, and thinking that&#8217;ll come soon is a mistake&quot;,&quot;username&quot;:&quot;alexandr_wang&quot;,&quot;name&quot;:&quot;Alexandr Wang&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 23 22:46:19 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:147,&quot;like_count&quot;:1291,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/fraser/status/1617552406178299904?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;After leading product at OpenAI for two and a half years I&#8217;ve made the decision to move on. I&#8217;ll be telling the story of modern AI and investing in OpenAI alumni and other remarkable founders. More in the thread below and at: &quot;,&quot;username&quot;:&quot;Fraser&quot;,&quot;name&quot;:&quot;Fraser&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 23 15:58:40 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:68,&quot;like_count&quot;:1114,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://www.moreentropy.com/about&quot;,&quot;image&quot;:&quot;https://substackcdn.com/image/fetch/w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ac09fe3-3e30-4a16-ab41-86c636ad8216_1024x1024.png&quot;,&quot;title&quot;:&quot;Entropy&quot;,&quot;description&quot;:&quot;The story of modern AI and those building the future. Click to read Entropy, by Fraser, a Substack publication with thousands of readers.&quot;,&quot;domain&quot;:&quot;moreentropy.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/debarghya_das/status/1611187565767561216?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;Are you interested in all the cutting edge AI in 2023 but you just can't keep up?\n\nHere's a Google Sheet for all Large Language Models with\n- name\n- creator\n- parameters\n- tokens trained\n- token:param ratio\n- training dataset\n- announce and release date\n- public?\n- paper &quot;,&quot;username&quot;:&quot;debarghya_das&quot;,&quot;name&quot;:&quot;Deedy&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Fri Jan 06 02:27:04 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FlwWi6LaAAEZB4Q.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/zXKtyvOgI2&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:40,&quot;like_count&quot;:268,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/gdb/status/1616917914787139585?s=61&amp;t=m8t7bv80iMj4HQ-ycnSIYg&quot;,&quot;full_text&quot;:&quot;This year is 80th birthday of the McCulloch-Pitts neuron. Remains the fundamental idea behind all neural networks. Such a simple mathematical model, yet has scaled to incredible results across many orders of magnitude of compute. Hard not to feel inspired. <a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://cs.cmu.edu/~./epxing/Class/10715/reading/McCulloch.and.Pitts.pdf/%22>cs.cmu.edu/~./epxing/Clas&#8230;</a> &quot;,&quot;username&quot;:&quot;gdb&quot;,&quot;name&quot;:&quot;Greg Brockman&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Sat Jan 21 21:57:26 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FnByGxWakAAZ6dX.png&quot;,&quot;link_url&quot;:&quot;https://t.co/VKh9IHkKI2&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:153,&quot;like_count&quot;:964,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Finally, in case you missed it, I also shared Part 3 of my series on the origins of Deep Learning:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:99015940,&quot;url&quot;:&quot;https://hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4&quot;,&quot;publication_id&quot;:1281858,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The Hitchhikers Guide to AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png&quot;,&quot;title&quot;:&quot;&#129299;A Deep Dive into Deep Learning: Part 3&quot;,&quot;truncated_body_text&quot;:&quot;Hi Readers! Thank you for subscribing to my newsletter. Here&#8217;s the final part of my deep dive into the origins of deep learning. In case you missed it, here are Part 1 and Part 2. The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a&quot;,&quot;date&quot;:&quot;2023-01-26T22:45:27.329Z&quot;,&quot;like_count&quot;:1,&quot;comment_count&quot;:0,&quot;bylines&quot;:[],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="/__u/hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!Ugi3!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png" loading="lazy"><span class="embedded-post-publication-name">The Hitchhikers Guide to AI</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">&#129299;A Deep Dive into Deep Learning: Part 3</div></div><div class="embedded-post-body">Hi Readers! Thank you for subscribing to my newsletter. Here&#8217;s the final part of my deep dive into the origins of deep learning. In case you missed it, here are Part 1 and Part 2. The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 1 like</div></a></div><p><em>That&#8217;s all for this week!</em></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhikers Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Scale AI provides infrastructure and resources to label large datasets for machine learning for many different use cases including robotics, AR/VR, AI and autonomous vehicles.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>The project&#8217;s title &#8220;GPT is all you need for backend&#8221;, is a play on words on &#8220;Attention is all you need,&#8221; the famous Google research paper that introduced the Transformer architecture used by large-scale language models. If you want to learn more about what Transformers are, read my latest post on <a href="/__u/hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4">the origins of Deep Learning</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>A &#8220;backend&#8221; is the part of a web application that stores and serves data to the &#8220;frontend&#8221; that you interact with as a user. For example this web page is the frontend of substack and the backend is what stores and serves all the text in this post.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>GPT or General Pre-trained Transformer is OpenAI&#8217;s large-scale language model that powers ChatGPT.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>If you want to learn more about Yann LeCunn and his work curing the &#8220;AI Winter&#8221; read <a href="/__u/hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-c6e">part 2 in my series</a> on the origins of Deep Learning.</p></div></div>]]></content:encoded></item><item><title><![CDATA[🤓A Deep Dive into Deep Learning: Part 3]]></title><description><![CDATA[The last of 3 posts on the history of Deep Learning and the foundational developments that led to today&#8217;s AI innovations.]]></description><link>https://operatorsguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4</guid><pubDate>Thu, 26 Jan 2023 22:45:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!z-7M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!z-7M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg" width="1024" height="763" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:763,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:403719,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi Readers!</p><p>Thank you for subscribing to my newsletter. Here&#8217;s the final part of my deep dive into the origins of deep learning. In case you missed it, here are <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/a-deep-dive-into-deep-learning-part?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">Part 1</a> and <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/a-deep-dive-into-deep-learning-part-c6e?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">Part 2</a>. </p><p>The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a <em>layer deeper</em> into foundational concepts and try to decipher the jargon.</p><p>P.S. Don&#8217;t forget to hit subscribe if you want to receive more AI content like this!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Part 3 - The AI Spring</h2><p>We ended <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/a-deep-dive-into-deep-learning-part-c6e?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">Part 2</a> at the turn of the century with a small group of dedicated researchers pushing forward the field of neural networks, despite the skepticism of the wider academic community. Meanwhile, limitations in computing power at the time made it harder to train larger models to solve more complex problems. </p><p>So how did we get from the AI winter of the 70s, 80s, and 90s to the exponential advancements we&#8217;re experiencing today? Let&#8217;s pick things up in the 2000s, where we rejoin Geoffrey Hinton and his team&#8230;</p><h3>2000s: Deep Learning <em>Accelerates</em></h3><p>In the 2000s, Hinton and his team continue to make major advancements in deep learning, building on the foundation laid in the 1990s. They develop new training algorithms that improve the effectiveness of deep learning models, particularly Deep Neural Networks (DNNs). Their most notable contributions include "pre-training" and "fine-tuning," which they first successfully use in creating Deep Belief Networks (DBNs) in 2006<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</p><p>Hinton and his team propose a way to train neural networks by starting with simpler networks first and then building on top of them. This method, called "pre-training," helps the deeper networks learn more effectively. After pre-training, the team fine-tuned the network by adjusting it using a method called "supervised learning." This helped the network improve its performance even further. Deep Belief Networks were one of the first successful examples of deep learning architectures that used pre-training, fine-tuning and unsupervised learning. </p><div><hr></div><p><strong>&#129299; What is a Deep Belief Network?</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!sh4S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 424w, /__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 848w, /__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!sh4S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp" width="820" height="464" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:464,&quot;width&quot;:820,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:23926,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 424w, /__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 848w, /__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!sh4S!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F356b2dbc-76d9-4e81-b99d-c8d5c6c81005_820x464.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">&#8220;Deep Belief Networks &#8212; All you need to know&#8221; by <a href="https://medium.com/@icecreamlabs/deep-belief-networks-all-you-need-to-know-68aa9a71cc53">IceCream Labs</a></figcaption></figure></div><p>Deep Belief Networks (DBNs) are a type of neural network trained using a special technique called unsupervised learning. This means the network is not given specific answers or expected outputs to learn from. Instead, the goal of a DBN is to find patterns and structures in the data on its own.</p><p>One example of how this works is in image recognition. Imagine we have a dataset of images of handwritten digits, and we want to train a DBN to recognize these digits. The visible layer of the DBN would represent the pixels of the images, and the hidden units would be used to identify the underlying structure of the images. For example, the hidden units might learn to recognize edges, corners, and other simple shapes in the images. These simple shapes can then be combined to form more complex shapes such as digits.</p><p>However, DBNs are different from both Deep Neutral Networks (DNNs) and Convolutional Neural Networks (CNNs) covered in Part 2. DNNs use hidden layers to extract abstract and higher-level features from the input data, while CNNs use filters to identify specific features in images. DBNs, on the other hand, are unsupervised and learn the actual structure of the image that can later be combined to generate new images.</p><p>A key component of DBNs is the use of Restricted Boltzmann Machines (RBMs). RBMs are a type of neural network made up of two layers: a visible layer and a hidden layer. The visible layer represents the input data and the hidden layer is used to learn a compressed representation of the input data. RBMs are probabilistic generative models, meaning they can learn the probability distribution of the input data and generate new samples that are similar to the training data.</p><p>In a DBN, multiple RBMs are stacked on top of one another. Each RBM learns a more abstract and higher-level representation of the data as it progresses through the network. This is similar to how a puzzle has many pieces that need to be put together to form a complete picture. The idea is that by stacking multiple RBMs, the DBN can learn a more detailed and accurate representation of the input data.</p><p>Once the DBN has been pre-trained using these stacked RBMs, it can then be fine-tuned using a labeled dataset and supervised learning techniques, such as backpropagation, to improve its performance on a specific task. This is similar to how a detective would use clues from a crime scene to make an arrest.</p><p>Overall, Deep Believe Networks have been used in a variety of applications, such as image and speech recognition, and have played a significant role in the advancement of deep learning.</p><div><hr></div><h3>2000s continued: Off to the races!</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7yAX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 424w, /__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 848w, /__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7yAX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp" width="750" height="300" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:750,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77028,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 424w, /__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 848w, /__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!7yAX!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83802e62-8cb4-4f71-9066-be0ae55a2880_750x300.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As the decade proceeds, advancements in deep neural networks and the availability of large-scale datasets like Fei-Fei Li&#8217;s ImageNet<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> and pre-trained models makes deep learning much more approachable. Researchers no longer need expensive and time-consuming data collection and annotation to train effective deep learning models. This leads to an era of competitive research, with teams from around the world seeing who trains the most accurate models.</p><p>One of these competitions is the Netflix Prize. The aim of the competition is to use machine learning to beat Netflix's own recommendation software's accuracy in predicting a user's rating for a film given their ratings for previous films by at least 10%. The prize is won in 2009 by BellKor's Pragmatic Chaos, a team whose members include employees of AT&amp;T and Yahoo. </p><p>At the same time, another development comes along that literally <em>accelerates</em> the field of deep learning: The use of Graphics Processing Units (GPUs) in deep learning makes it possible to train large neural networks much more quickly than when using traditional processors on a computer.</p><div><hr></div><p><strong>&#129299; Why do GPUs accelerate deep learning?</strong></p><p>A GPU, or Graphics Processing Unit, is a type of computer chip that is specifically designed to process large amounts of data quickly and efficiently. They were originally created for use in video game graphics, but scientists and researchers soon realized that they could be used to accelerate many other types of computations, including those used in deep learning models.</p><p>Deep learning models, such as neural networks, typically involve a lot of complex mathematical calculations and require a lot of data to be processed at the same time. Training these models can take a long time on a regular computer, even with a fast processor. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!btKf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 424w, /__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 848w, /__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 1272w, /__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!btKf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png" width="1456" height="719" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:719,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39311,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 424w, /__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 848w, /__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 1272w, /__u/substackcdn.com/image/fetch/$s_!btKf!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbceca25-e20a-42b4-9d9e-96c77465a186_1612x796.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: NVIDIA</figcaption></figure></div><p>A CPU (Central Processing Unit) is divided into multiple cores so that they can take on multiple different tasks at the same time (e.g., browsing the internet while listening to Spotify). A GPU, on the other hand has hundreds and thousands of cores, which are dedicated to completing simple computations that are performed more frequently and independently of each other i.e. in parallel. </p><p>Training neural networks is an ideal task for GPUs. Calculating weights and activation functions of each layer and backpropagation can all be computed in parallel. A GPU is designed to handle these types of calculations much more efficiently. It can perform many calculations in parallel, which means it can work on many pieces of data at the same time. This allows deep learning models to be trained much faster on a GPU than on a regular computer.</p><p>As deep learning has progressed, it needs to have vast amounts of data to be fed and training on these data set takes a long time on a single CPU. With a single GPU, the time to train these models is significantly reduced, and with multiple GPUs working together in parallel, it's even faster.</p><div><hr></div><h3>2010s: &#8220;Hey Google&#8230;&#8221;</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!33Db!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 424w, /__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 848w, /__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!33Db!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp" width="950" height="534" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:534,&quot;width&quot;:950,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:17442,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 424w, /__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 848w, /__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!33Db!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9783d852-bb97-4d51-b4fa-59b946ebc794_950x534.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">GoogleHome, Amazon Echo Speaker, and Apple Homepod with Siri (Source: Market Ahead)</figcaption></figure></div><p>Thanks to the progress of the previous decade, the 2010s see exponential advancement in deep learning, with Apple, Google and Amazon all making major investments in the space ushering in the dawn of consumer-friendly AI. From automatically organizing your photos to helping you turn on the lights in your home and playing your favorite songs at the command of your voice, AI enters the mainstream and consumers lives. </p><p>The decade began with a seminal paper in 2012 by Geoffrey Hinton, Alex Krizhevsky and Ilya Sutskever that showed a massive leap in the accuracy of image recognition using deep neural networks<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. Hinton and his colleagues developed AlexNet, a convolutional neural network that won several competitions.</p><p>In 2013 the film &#8220;Her&#8221; is released. A science fiction drama starring Scarlett Johansson as Samantha, an AI operating system who its user Theodore falls in love with. Samantha is portrayed as a highly intelligent and empathetic AI that is able to form a deep connection with Theodore. The film provides a very tangible sense of what the near future might look like with AI as part of our lives. At the same time, it also highlights the limitations of &#8220;AI Assistants&#8221; like Google Home, Siri, and Echo, which are far more limited in their ability. </p><div><hr></div><p><strong>&#129299; How do voice assistants like Google Home, Siri and Echo work?</strong></p><p>Remember from <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/a-deep-dive-into-deep-learning-part-c6e?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">Part 2</a> how recurrent neural networks (RNNs) with long short-term memory are able to process longer sequences of input? That&#8217;s perfect for the use case of speech recognition in voice assistants, where a user makes a request with multiple words in a sequence.</p><p>Here&#8217;s a reminder of how RNNs work from Part 2:</p><blockquote><p>An RNN might take a sequence of words as input and use the output from processing the previous word to inform its processing of the current word. This allows the RNN to capture the context and dependencies between words, which is important for understanding the meaning of the text.</p></blockquote><p>RNNs work by using a feedback loop, where the output of a previous step is fed back into the network as input for the next step. This allows the network to "remember" what it has heard in the past and use that information to better understand the current input.</p><p>And here&#8217;s how LSTMs work, also from Part 2:</p><blockquote><p>A long short-term memory (LSTM) unit is a type of recurrent neural network that is able to capture long-term dependencies in time series data. It is called "long short-term memory" because it is able to remember information for long periods of time and use it to make predictions or decisions later. For example, storing a whole sentence in a text to predict the next word vs just storing the last word. </p></blockquote><p>LSTMs are a type of RNN that are particularly well-suited for handling speech data. They use a special structure called a memory cell, which can retain information for long periods of time and selectively choose which information to discard and which to keep. This allows LSTMs to effectively filter out irrelevant information, such as background noise, and focus on the important parts of the speech input, which is critical in the use case of Voice Assistants.</p><p>Together, RNNs and LSTMs form the backbone of AI voice assistants, allowing them to understand and respond to human speech in real-time. As more data is fed into these networks, they continue to learn and improve, making them even better at understanding and responding to human speech.</p><p>It is worth mentioning that, these models are trained on a massive amount of data, this allows them to generalize well and adapt to different accents, dialects, and speaking styles. It also allows them to understand and respond to new words and phrases, as well as recognize and respond to specific speakers.</p><p>But voice assistants aren&#8217;t just tasked with recognizing speech. They also respond to a user&#8217;s commands by synthesizing human-like speech too. For example, if you ask Google Home, &#8220;What&#8217;s the weather like today?&#8221; It will respond with an answer describing the current weather conditions and forecast for the rest of the day. This is where deep neural networks (DNNs) and convolutional neural networks (CNNs) come into play.</p><p>Here&#8217;s a reminder of how convolutional neural networks work from Part 2:</p><blockquote><p>A CNN consists of multiple layers of interconnected neurons, which process and analyze the input data. The layers of a CNN are organized in such a way that they learn increasingly complex features of the input data as the data passes through the network.</p><p>One key feature of CNNs is the use of "convolutional" layers, which are designed to automatically learn and extract features from the input data. These layers apply a set of filters to the input data and use the resulting output to detect patterns or features in the data. This process is repeated multiple times, allowing the CNN to learn increasingly complex features of the input data as it passes through the network.</p></blockquote><p>In speech synthesis, DNNs and CNNs are used to model the complex relationships between audio signals and human speech. These models are trained on large amounts of speech data, and they learn to recognize patterns in the data that are associated with different speech sounds, such as phonemes and words. They are then used to generate new speech from text by predicting the most likely sequence of speech sounds for a given input text.</p><p>DNNs and CNNs are powerful tools for speech synthesis, as they can learn to model the complex relationships between audio signals and human speech, and generate new speech from text with high accuracy. They are the foundation of most of the AI voice assistants like Siri and Google home.</p><p>It&#8217;s worth noting however, that voice assistants are limited in their ability because they are based on a set of predefined rules and commands. They are not truly intelligent because they do not have the ability to learn and adapt like a human would. They rely on a set of programmed responses and cannot understand context or make decisions based on new information. </p><div><hr></div><h3>2010s continued: The race to general intelligence begins</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c0BY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c0BY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg" width="1230" height="683" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/faf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:683,&quot;width&quot;:1230,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:63247,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!c0BY!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaf4222b-17f6-44ae-b905-e5d336f59eb1_1230x683.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The number of AI/ML research papers published exponentially increases in 2010s (Source: <a href="https://www.pnas.org/topic/phys-sci">PNAS.org</a>)</figcaption></figure></div><p>In the 2010s, startups like OpenAI and Deepmind (acquired by Google in 2015 for over $400M) are funded with the explicit goal of achieving Artificial General Intelligence (AGI). AGI is a type of artificial intelligence designed to possess a broad range of cognitive abilities similar to those of a human being. This includes the ability to understand complex concepts, reason, plan, solve problems, and learn from experience similar to our idea of AI from sci-fi movies like &#8220;Her&#8221;.  This new injection of capital into AI research from startups and big tech companies spurs an exponential increase in advancements in the field. </p><p>One of these new developments is generative models. These are models that are designed to generate new output data rather than classify or recognize input data. One of the most popular early generative models is the Generative Adversarial Network (GAN) which was introduced by Ian Goodfellow in 2014<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>. The main advantage of GANs is the ability to generate new data that is similar to the training data, allowing them to be used for tasks such as image synthesis, image-to-image translation, and other generative tasks.</p><div><hr></div><p><strong>&#129299; What is a Generative Adversarial Network (GAN)?</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4FGs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 424w, /__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 848w, /__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4FGs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp" width="750" height="481" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:481,&quot;width&quot;:750,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55180,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 424w, /__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 848w, /__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!4FGs!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841827d5-94cd-4ee9-a6fc-30a42972012b_750x481.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Synthetic images produced by StyleGAN, a GAN created by Nvidia researchers</figcaption></figure></div><p>Generative Adversarial Networks (GANs) are a class of generative models that use a technique called adversarial training, where two neural networks, a generator, and a discriminator, are trained together.</p><p>The generator network learns to generate new data samples that are similar to the training data. It takes in a random noise as input, and it produces a new data sample that is similar to the training data. The generator is typically a neural network with an architecture designed to produce new data samples, such as a decoder network.</p><p>The discriminator network, on the other hand, takes in both real data samples from the training set and fake data samples generated by the generator. It is trained to distinguish between the real data samples and the fake data samples generated by the generator. The discriminator is also typically a neural network but with an architecture designed to distinguish between real and fake data samples.</p><p>The training process of a GAN is an adversarial process where the generator and discriminator are trained simultaneously and in opposition. The generator is trained to produce data samples that can fool the discriminator into thinking they are real, while the discriminator is trained to correctly identify the real data samples from the fake ones generated by the generator.</p><p>At the beginning of the training, the generator produces poor-quality samples, the discriminator easily recognizes them as fake, and thus the generator is updated to improve its performance. As the training progresses, the generator improves and generates better-quality samples, making it harder for the discriminator to distinguish between real and fake data. The discriminator also improves by training on both real and fake samples. The training continues until the generator can produce samples that can fool the discriminator.</p><p>This process is called adversarial training; the generator and discriminator are competing with each other, and the generator is trying to produce realistic samples that the discriminator would not be able to tell apart from the real ones. This leads to the generator learning to produce samples that are similar to real ones.</p><div><hr></div><h4>2010s continued: Transformers - More than meets the eye?</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mS6E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mS6E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg" width="980" height="653" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:653,&quot;width&quot;:980,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:182539,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!mS6E!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77e20968-7d6a-4173-bae4-abc852d094d7_980x653.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Arguably one of the most revolutionary breakthrough in deep learning in the 2010s is the invention of the Transformer architecture in 2017 by Google researchers in the famously titled paper &#8220;Attention is All You Need&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. In the paper, Google&#8217;s researchers propose &#8220;a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.&#8221; </p><p>The advent of the Transformer architecture has played a significant role in the development of large-scale language models (LLMs) such as OpenAI&#8217;s GPT-3, which powers ChatGTP. These Large-scale language models are essentially transformer-based architectures that are trained on large amounts of text data, and have billions of parameters, which allows them to understand the input more effectively, generate more coherent and human-like text, and generalize better on unseen examples.</p><p>Before the transformer architecture, Recurrent Neural Networks (RNNs) were the most commonly used architecture for language modeling tasks. However, RNNs are not well-suited for processing long sequences of text and have difficulty learning long-term dependencies in the input.</p><div><hr></div><p><strong>&#129299; What is the Transformer architecture?</strong></p><p>Remember from earlier that recurrent neural networks and convolutional neural networks store a fixed length part of their input in memory in hidden layers so that they can better predict the output. The challenge with this approach is that it requires the model to serially process its input, for example, processing a paragraph of text one word at a time. This process cannot be parallelized when training which creates challenges due to limitations in memory.  It also limits how far back a model can learn about the text it is trained on for any word in the text, referred to as &#8220;lookback.&#8221;</p><p>The transformer architecture is based on the idea of allowing the model to look at the whole sequence of input rather than a fixed-length part of the input, like recurrent neural networks with long short-term memory do. Then the model focuses on certain parts of that input that are relevant to each word. It achieves this through a mechanism called self-attention, which you can liken to being able to look up different words in a dictionary.</p><p>Attention was first introduced in 2015 in the context of language translation from English to French. In the paper, the authors gave the example of translating the following English sentence</p><p><em>&#8220;The agreement on the European Economic Area was signed in August 1992.&#8221;</em></p><p>Into the French equivalent</p><p><em>&#8220;L&#8217;accord sur la zone &#233;conomique europ&#233;enne a &#233;t&#233; sign&#233; en ao&#251;t 1992.&#8221;</em></p><p>Trying to translate this sentence by going through each English word one by one wouldn&#8217;t work for many reasons: some French words are flipped, and the French language has gendered words. </p><p>Attention is a mechanism that allows the model to focus on every single word in the French input when generating the English output and pay more attention to specific words.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!pFlr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 424w, /__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 848w, /__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!pFlr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png" width="558" height="548" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:548,&quot;width&quot;:558,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47847,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 424w, /__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 848w, /__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pFlr!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f678dae-7e04-4ac2-945a-bb8df084e464_558x548.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">This figure from the paper demonstrates the heat map of how words are related</figcaption></figure></div><p>The model learns which words it should &#8220;attend&#8221; to from training data by processing thousands of French and English sentences. </p><p><em>Self-attention, </em>unlike regular attention, doesn&#8217;t look at the attention of specific words for a given input and output, as this is limited to translations. Instead, it looks at the attention to give different words in the input <em>for each word in the input itself</em>. In order words, self-attention allows a neural network to understand a word in the context of words around it. This is important in NLP tasks such as language understanding, where the meaning of a word depends on the context and its relationship with other words in the sentence.</p><p>The transformer architecture also introduced the concept of multi-head attention, which enables the model to attend to multiple parts of the input simultaneously, which improves the performance of the model by allowing parallelization.</p><p>Additionally, the transformer architecture introduced the concept of position encoding, which allows the model to understand the order of the tokens in the input and make use of their relative position to better understand the meaning of the input.</p><div><hr></div><h3>Present Day: Hype builds around Generative AI</h3><p>The Transformers architecture proves to be much better at handling large amounts of data and parallel processing than previous neural networks. This enables efficient  training of deep, large models on GPUs, unlocking the exponential progress toward AI. Today, transformers enable the training of large-scale language models like GPT-3 that powers ChatGPT, a breakthrough product in AI, which I covered in my post, <a href="/__u/hitchhikersguidetoai.substack.com/p/ai-dont-believe-the-hype">AI: Don&#8217;t believe the hype?</a>:</p><blockquote><p>Unlike it&#8217;s predecessors (e.g. Google Assistant, Echo, Siri), ChatGPT is really the first time an AI assistant truly seems like it could pass the Turing Test. There have been many impressive examples of ChatGPT in action and if you haven&#8217;t tried it yourself you should. ChatGPT successfully wrote a blog post for me and turned it into a twitter thread, gave me a recipe for pancakes that tasted delicious and helped me pick a Christmas present for my wife! What truly impressed me though is ChatGPT&#8217;s ability to be &#8220;creative&#8221;.</p></blockquote><p>Today, the race is on between Google, OpenAI, and many other startups to build larger, more intelligent LLMs that may one day reach the holy grail of general intelligence.</p><div><hr></div><p><strong>&#129299; How do large-scale language models like GTP-3 work?</strong></p><p>A large-scale language model (LLM) is a type of deep learning model that is trained on a large dataset of text (e.g. all of the internet). LLMs predict the next sequence of text as output based on the text that they are given as input. They are used for a wide variety of tasks, such as language translation, text summarization, and generating conversational text. Open AI&#8217;s GPT-3 (General Pre-trained Transformer 3)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>, the language model that powers Chat-GPT, is an example of a generative LLM that uses the Transformer architecture, enabling it to be trained on a massive text dataset of hundreds of gigabytes using 175 Billion parameters (weights assignments).</p><p>GPT-3&#8217;s size made it the largest language model at the time of its launch though it was later superseded by Google&#8217;s PaLM which boasts over 540 billion parameters. GPT-3s use of the Transformer architecture allows it ot understand and generate language in a way previous models could not. It also made use of unsupervised learning, to be trained without any specific task in mind. This allows it to be fine-tuned for a variety of different tasks, such as text completion, question answering, and language translation.</p><p>Once LLMs get to this size, they start exhibiting emergent behavior, unexpected and seemingly autonomous actions or decisions that the model makes as a result of its training on the vast amount of data it has been exposed to. This can be seen in GPT-3's ability to generate human-like text that is often difficult to distinguish from text written by a human. However, it is important to note that this behavior is not truly autonomous, as it is still determined by the patterns and connections present in the training data.</p><p>Additionally, GPT-3's may generate text that appears to have a bias or hold certain beliefs, which is a reflection of the biases and beliefs present in the training data, highlighting the need for diverse and unbiased data when training such models. These large models are also prone to &#8220;hallucinations&#8221;, the term in AI for when a model makes up an answer in a way that appears to be factual but is not. For example, it may generate a citation to an academic paper that looks real but has the incorrect authors! </p><div><hr></div><h3><strong>Final thoughts</strong></h3><p>It wasn&#8217;t until I wrote this series that I understood how much we should be grateful for the tireless efforts of a few researchers who laid the foundations for the astonishing progress we are seeing in deep learning today. Even a decade ago, there was still skepticism about the effectiveness of neural networks, which wasn&#8217;t really overcome until companies like Google invested hundreds of millions of dollars into the field. Now we&#8217;re on the verge of having self-driving cars, AI chatbots can pass the bar exam, and anyone can create art by just describing it! And the best part? </p><p><strong>We're just getting started!</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!GodI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 424w, /__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 848w, /__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!GodI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:60946,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 424w, /__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 848w, /__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!GodI!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80ecb355-a26b-4783-9868-1eb9cfac62b2_1920x1080.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h3>Feedback</h3><p>Hi Readers!</p><p>I really enjoyed writing this series on deep learning and learned a lot myself while doing it. I would love to know if you found this series valuable in learning about AI?</p><div class="poll-embed" data-attrs="{&quot;id&quot;:45507}" data-component-name="PollToDOM"></div><p>Thanks for your feedback!</p><p>~AJ</p><p>P.S. If you did find this series valuable, do me a favor and share it with your friends interested in AI!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><div><hr></div><p><em>Fun fact: I used ChatGPT for a lot of the research for this series of posts. Over the course of a week I asked ChatGPT dozens of questions about the history of deep learning and the concepts behind it. To make sure the article was accurate, I asked ChatGPT to provide citations to relevant scientific papers which I&#8217;ve included in the footnotes. If you find any errors, please reach out or comment below and I will fix them!</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Hinton, G. E., Osindero, S., &amp; Teh, Y. W. (2006). <a href="https://www.cs.toronto.edu/~fritz/absps/ncfast.pdf">A fast learning algorithm for deep belief nets</a>. Neural computation, 18(7), 1527-1554.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>In 2009, Fei-Fei Li, an AI professor at Stanford launched <a href="http://image-net.org/">ImageNet</a>, assembled a free database of more than 14 million labeled images. The Internet is, and was, full of unlabeled images. Labeled images were needed to &#8220;train&#8221; neural nets. Professor Li said, &#8220;Our vision was that big data would change the way machine learning works. Data drives learning.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>"<a href="https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf">ImageNet Classification with Deep Convolutional Neural Networks</a>" was published in 2012 by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>"<a href="https://proceedings.neurips.cc/paper/2014/file/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf">Generative Adversarial Networks</a>" (Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., &#8230; &amp; Bengio, Y. (2014). Generative adversarial nets. In Advances in neural information processing systems (pp. 2672-2680).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... &amp; Polosukhin, I. (2017). <a href="https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf">Attention is all you need</a>. Advances in Neural Information Processing Systems, 30, 5998-6008</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... &amp; Amodei, D. (2020). <a href="https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html">Language models are few-shot learners</a>. <em>Advances in neural information processing systems</em>, <em>33</em>, 1877-1901.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[One Hundred Subscribers 🚀🙏🏾]]></title><description><![CDATA[Celebrating my first one hundred subscribers!]]></description><link>https://operatorsguidetoai.substack.com/p/one-hundred-subscribers</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/one-hundred-subscribers</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Tue, 24 Jan 2023 17:26:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v37s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!v37s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 424w, /__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 848w, /__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 1272w, /__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!v37s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png" width="1456" height="1130" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1130,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:243659,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 424w, /__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 848w, /__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 1272w, /__u/substackcdn.com/image/fetch/$s_!v37s!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcbca9284-45ca-467c-bf16-c513572a68a2_1570x1219.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi Readers,</p><p>I&#8217;m excited to share that The Hitchhiker&#8217;s Guide to AI now has 100 subscribers!  </p><p>Thank you all for supporting this newsletter and following my journey into the world of Artificial Intelligence. It may seem like a small milestone, but I&#8217;m confident that if I can make this newsletter valuable to one hundred people, it will one day provide value to thousands more. As Paul Graham says<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>: </p><blockquote><p>If you have 100 users, you need to get 10 more next week to grow 10% a week. And while 110 may not seem much better than 100, if you keep growing at 10% a week you'll be surprised how big the numbers get.</p></blockquote><p>To that end, I would love to know if you have any feedback on my writing so far? Is there anything you would like me to write about in the AI space that you&#8217;re curious about? I would like you to think of me as your personal AI expert!  </p><p><em><strong>Please let me know if you have feedback or suggestions by commenting on this post or replying to the email!</strong></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/p/one-hundred-subscribers/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/p/one-hundred-subscribers/comments"><span>Leave a comment</span></a></p><div><hr></div><p>Oh, and one more thing&#8230; If you find this newsletter valuable to you already, please share it with your friends who might enjoy reading it too!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://hitchhikersguidetoai.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Hitchhikers Guide to AI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/hitchhikersguidetoai.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Hitchhikers Guide to AI</span></a></p><p>Thank you!</p><p>~ AJ.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="http://paulgraham.com/ds.html">Do Things that Don&#8217;t Scale</a> - Paul Graham, 2013</p></div></div>]]></content:encoded></item><item><title><![CDATA[🗞️AI highlights from this week (1/20/23)]]></title><description><![CDATA[Predictions for an AI filled future, Google shares AI progress, debunking GPT-4 rumors and more&#8230;]]></description><link>https://operatorsguidetoai.substack.com/p/ai-highlights-from-this-week-12023</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/ai-highlights-from-this-week-12023</guid><pubDate>Fri, 20 Jan 2023 22:08:30 GMT</pubDate><enclosure url="https://i.scdn.co/image/ab6765630000ba8a5a080e22b006d530a7951edd" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi readers!</p><p>Here are some of the most interesting AI updates that I read in the last week.<em> </em></p><p><em>P.S. Don&#8217;t forget to hit subscribe if you&#8217;re new to AI and want to learn more about the space. </em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h3>Highlights</h3><h4><strong>1/ Founder of StabilityAI shares predictions for AI-filled future</strong></h4><p>Ehmad Mostaque the founder of StabilityAI, the company behind Stable Diffusion, shared an inspiring view of how AI will impact our future with Peter Diamandis, host of the Moonshots Mindset podcast. In the wide ranging discussion, Ehmad covered how AI will impact the entertainment industry, what Generative AI means for copyright and ownership and whether AI should have a moral compass. </p><p><strong>&#8220;This is one of the biggest evolution for humanity ever!&#8221; </strong>- Ehmad Mostaque</p><p>The first 50 minutes are definitely worth listening to!</p><iframe class="spotify-wrap podcast" data-attrs="{&quot;image&quot;:&quot;https://i.scdn.co/image/ab6765630000ba8a5a080e22b006d530a7951edd&quot;,&quot;title&quot;:&quot;EP #16 AI is Creating Massive Entrepreneurial Opportunity w/ Emad Mostaque&quot;,&quot;subtitle&quot;:&quot;PHD Ventures&quot;,&quot;description&quot;:&quot;Episode&quot;,&quot;url&quot;:&quot;https://open.spotify.com/episode/5B8cuxuCTQBJlr4dCVIsuC&quot;,&quot;belowTheFold&quot;:true,&quot;noScroll&quot;:false}" src="https://open.spotify.com/embed/episode/5B8cuxuCTQBJlr4dCVIsuC" frameborder="0" gesture="media" allowfullscreen="true" allow="encrypted-media" loading="lazy" data-component-name="Spotify2ToDOM"></iframe><h4></h4><h4>2/ Google finally shares an update on their progress on AI</h4><p>Until 2022, Google was considered to be at the forefront of advancements in AI. With the launch of Stable Diffusion, Dall-E and ChatGTP however, Google&#8217;s thought leadership and research in AI took a backseat to startups like StabilityAI and OpenAI. It&#8217;s no surprise then, that after Google reportedly issued a internal &#8220;Code Red&#8221; on the need to move faster on AI, we finally got an update last week on what they&#8217;ve been up to.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/sundarpichai/status/1615820298305118221?s=61&amp;t=1XwTcGxYb6tQmk6gIkSVYA&quot;,&quot;full_text&quot;:&quot;Summary of great research progress on <span class=\&quot;tweet-fake-link\&quot;>#GoogleAI</span>, including language models, computer vision, multimodal models, generative ML. We're building it all into current and upcoming products + APIs, look forward to sharing more with everyone soon. Stay tuned! \n<a class=\&quot;tweet-url\&quot; href=/__u/operatorsguidetoai.substack.com/%22https://ai.googleblog.com/2023/01/google-research-2022-beyond-language.html/%22>ai.googleblog.com/2023/01/google&#8230;</a> &quot;,&quot;username&quot;:&quot;sundarpichai&quot;,&quot;name&quot;:&quot;Sundar Pichai&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Wed Jan 18 21:15:54 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FmyLYf3XoBM5wd4.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/xAJql9WH4h&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:972,&quot;like_count&quot;:6773,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Google&#8217;s blog post was authored on behalf of their whole AI research team by Google&#8217;s most senior researcher, Jeff Dean. It covered a range of topics including their latest advancements in Large Language Models, Computer Vision, Generative AI and much more. </p><p>Dean discusses the progress and advancements in language models, specifically highlighting their work on LaMDA and PaLM, which allows for safe and high-quality dialog in natural conversations.  He also mentions their work on multimodal models, which can handle multiple modalities simultaneously, both as inputs and outputs. </p><p>Dean also highlights the progress and advancements in generative models for imagery, video, and audio, mentioning specific techniques such as generative adversarial networks, diffusion models, autoregressive models, and Contrastic Language-Image Pre-training (CLIP). Finally, he emphasizes the importance of responsible AI, stating that they apply their AI Principles in practice and focus on AI that is useful and benefits users and society. </p><p>In the post, Dean also highlighted multiple times that Google was responsible for the Transformer model in 2017, which unlocked the rapid increase in AI innovation we&#8217;re seeing today. He also shared that many of the innovations that Google have made in AI are already being integrated into their existing products.</p><p>You can read the whole post <a href="https://ai.googleblog.com/2023/01/google-research-2022-beyond-language.html">here</a> but fair warning, it&#8217;s pretty dense!</p><h4>3/ OpenAI CEO debunking rumors about GPT-4</h4><p>Sam Altman CEO of Open AI, debunks the rumors about GPT-4 in an interview with Connie Loizos for StrictlyVC.</p><p>GTP-4 is the much anticipated iteration of GPT-3, the large language model that underlies ChatGPT. As I mentioned in my first post <a href="/__u/hitchhikersguidetoai.substack.com/p/ai-dont-believe-the-hype">Don&#8217;t believe the hype?</a>, there&#8217;s a lot of excitement and rumors around GPT-4 and some of the features, including a prediction that it will have 100X more parameters than GPT-3:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/aibreakfast/status/1594720737075802113?s=61&amp;t=NXugfHzWycGkijfJTnIQrw&quot;,&quot;full_text&quot;:&quot;Currently, GPT-3 has 175 billion parameters, which is 10x faster than any of its closest competitors. GPT-4 is rumored be about 100 trillion parameters. &quot;,&quot;username&quot;:&quot;AiBreakfast&quot;,&quot;name&quot;:&quot;AI Breakfast&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Nov 21 15:53:46 +0000 2022&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FiGWDwdUYAAuvW7.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/HlP4fvAfdQ&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:38,&quot;like_count&quot;:91,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Sam&#8217;s response to this rumor was <strong>&#8220;People are begging to be disappointed and they will be&#8221;</strong>.</p><p>Here&#8217;s the full interview:</p><div id="youtube2-ebjkD1Om4uw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ebjkD1Om4uw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ebjkD1Om4uw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h4>4/ Predictions on how Big Tech will approach AI</h4><p>Ben Thompson shared his predictions for how Google, Amazon, Meta (Facebook), Microsoft and Apple in his recent post, <a href="https://stratechery.com/2023/ai-and-the-big-five/">AI and the Big Five</a>. In the post Thompson view, AI is a new epoch in technology and he shared the following predictions on how the epoch might develop:</p><ul><li><p><strong>Apple's</strong> efforts in AI have been largely proprietary, but that the company recently received a gift from the open source community in the form of the Stable Diffusion model, which is small and efficient enough to run on consumer graphics cards and even an iPhone.</p></li><li><p><strong>Amazon's</strong> prospects in this space will depend on factors such as the usefulness of these products in the real world and Apple's progress in building local generation techniques.</p></li><li><p><strong>Google</strong>, who has been a leader in using machine learning for their search and consumer-facing products, may face a similar fate in the AI world to Eastman Kodak&#8217;s infamous downfall when digital cameras became popular. The shift from a mobile-first world to an AI-first world may present challenges for Google in providing the "right answer" rather than just presenting possible answers.</p></li><li><p><strong>Meta</strong>'s data centers are primarily for CPU compute, which is necessary for their services and deterministic ad model, but that the long-term solution for improving their ad targeting and measurement is through probabilistic models built by massive fleets of GPUs.</p></li><li><p><strong>Microsoft</strong> is well-positioned in the AI space with its cloud service, Azure, that sells GPU access and its exclusive partnership with OpenAI. Microsoft is investing in the infrastructure for the AI epoch through this partnership and its Bing search engine, which has the potential to gain massive market share with the incorporation of ChatGPT-like results.</p></li></ul><p>You can listen to the whole analysis here too:</p><iframe class="spotify-wrap podcast" data-attrs="{&quot;image&quot;:&quot;https://i.scdn.co/image/ab6765630000ba8a25ffb53a6a1f9e7ae7f99fbc&quot;,&quot;title&quot;:&quot;AI and the Big Five&quot;,&quot;subtitle&quot;:&quot;Ben Thompson&quot;,&quot;description&quot;:&quot;Episode&quot;,&quot;url&quot;:&quot;https://open.spotify.com/episode/1ngyQlzFodYts15Qf6m5RB&quot;,&quot;belowTheFold&quot;:true,&quot;noScroll&quot;:false}" src="https://open.spotify.com/embed/episode/1ngyQlzFodYts15Qf6m5RB" frameborder="0" gesture="media" allowfullscreen="true" allow="encrypted-media" loading="lazy" data-component-name="Spotify2ToDOM"></iframe><h4>5/ The first copyright lawsuit against Stable Diffusion</h4><p>Jon Stokes does a great analysis of the first major lawsuit alleging that Stable Diffusion violates the copyright of artists whose images were used in training. Cases like this one and the <a href="https://www.theverge.com/2022/11/8/23446821/microsoft-openai-github-copilot-class-action-lawsuit-ai-copyright-violation-training-data">recent copyright lawsuit</a> against Microsoft&#8217;s Github Copilot product are likely to be pivotal in establishing how Generative AI content will be viewed in the eyes of the law.</p><p>Here&#8217;s Jon&#8217;s post:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:97776936,&quot;url&quot;:&quot;https://www.jonstokes.com/p/anti-stable-diffusion-lawsuit-deep&quot;,&quot;publication_id&quot;:239653,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;jonstokes.com&quot;,&quot;publication_logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/4f382f14-88dd-416e-a6d7-c44ffb9be1db_256x256.png&quot;,&quot;title&quot;:&quot;Anti-Stable Diffusion Lawsuit Deep Dive: It&#8217;s A Neutron Bomb&quot;,&quot;truncated_body_text&quot;:&quot;The first major lawsuit against Stability AI has arrived, and the different AI art channels I follow have been furiously debunking the suit&#8217;s technical claims as they&#8217;ve been published by the lawyers. That debunking is all well and good &#8212; I&#8217;ve indulged in a little of it, myself &#8212; but I&#8217;m not sure most of the community&quot;,&quot;date&quot;:&quot;2023-01-19T23:21:19.909Z&quot;,&quot;like_count&quot;:4,&quot;comment_count&quot;:1,&quot;bylines&quot;:[{&quot;id&quot;:22541131,&quot;name&quot;:&quot;Jon Stokes&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/0d74c421-05b6-4f07-a9c7-09120a9bfb94_982x855.jpeg&quot;,&quot;bio&quot;:null,&quot;profile_set_up_at&quot;:&quot;2021-05-10T18:06:52.901Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:154676,&quot;user_id&quot;:22541131,&quot;publication_id&quot;:239653,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:239653,&quot;name&quot;:&quot;jonstokes.com&quot;,&quot;subdomain&quot;:&quot;doxa&quot;,&quot;custom_domain&quot;:&quot;www.jonstokes.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;AI/ML, crypto, speech, power &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/4f382f14-88dd-416e-a6d7-c44ffb9be1db_256x256.png&quot;,&quot;author_id&quot;:22541131,&quot;theme_var_background_pop&quot;:&quot;#B599F1&quot;,&quot;created_at&quot;:&quot;2020-12-15T20:39:21.373Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Jon Stokes&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:285559,&quot;user_id&quot;:22541131,&quot;publication_id&quot;:363171,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:363171,&quot;name&quot;:&quot;Asides&quot;,&quot;subdomain&quot;:&quot;asides&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Things that don't go on main&quot;,&quot;logo_url&quot;:null,&quot;author_id&quot;:22541131,&quot;theme_var_background_pop&quot;:&quot;#25BD65&quot;,&quot;created_at&quot;:&quot;2021-05-17T20:50:25.325Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Jon Stokes&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;}}],&quot;twitter_screen_name&quot;:&quot;jonst0kes&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.jonstokes.com/p/anti-stable-diffusion-lawsuit-deep?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!RaoY!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F4f382f14-88dd-416e-a6d7-c44ffb9be1db_256x256.png" loading="lazy"><span class="embedded-post-publication-name">jonstokes.com</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Anti-Stable Diffusion Lawsuit Deep Dive: It&#8217;s A Neutron Bomb</div></div><div class="embedded-post-body">The first major lawsuit against Stability AI has arrived, and the different AI art channels I follow have been furiously debunking the suit&#8217;s technical claims as they&#8217;ve been published by the lawyers. That debunking is all well and good &#8212; I&#8217;ve indulged in a little of it, myself &#8212; but I&#8217;m not sure most of the community&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 4 likes &#183; 1 comment &#183; Jon Stokes</div></a></div><h3>Everything else&#8230;</h3><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:96357963,&quot;url&quot;:&quot;https://www.jonstokes.com/p/machine-learning-and-deflationary&quot;,&quot;publication_id&quot;:239653,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;jonstokes.com&quot;,&quot;publication_logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/4f382f14-88dd-416e-a6d7-c44ffb9be1db_256x256.png&quot;,&quot;title&quot;:&quot;Machine Learning And Deflationary Contagions&quot;,&quot;truncated_body_text&quot;:&quot;Here&#8217;s the most common question I get from members of the press, analysts, and others I talk to about generative AI: can you describe what impact this will have on my industry? Of course, this is the obvious question to ask, right? So I get it all the time, and because I get it all the time I&#8217;ve developed a standard high-level answer to it that I&#8217;ll shar&#8230;&quot;,&quot;date&quot;:&quot;2023-01-13T00:25:58.611Z&quot;,&quot;like_count&quot;:17,&quot;comment_count&quot;:9,&quot;bylines&quot;:[{&quot;id&quot;:22541131,&quot;name&quot;:&quot;Jon Stokes&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/0d74c421-05b6-4f07-a9c7-09120a9bfb94_982x855.jpeg&quot;,&quot;bio&quot;:null,&quot;profile_set_up_at&quot;:&quot;2021-05-10T18:06:52.901Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:154676,&quot;user_id&quot;:22541131,&quot;publication_id&quot;:239653,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:239653,&quot;name&quot;:&quot;jonstokes.com&quot;,&quot;subdomain&quot;:&quot;doxa&quot;,&quot;custom_domain&quot;:&quot;www.jonstokes.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;AI/ML, crypto, speech, power &quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/4f382f14-88dd-416e-a6d7-c44ffb9be1db_256x256.png&quot;,&quot;author_id&quot;:22541131,&quot;theme_var_background_pop&quot;:&quot;#B599F1&quot;,&quot;created_at&quot;:&quot;2020-12-15T20:39:21.373Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Jon Stokes&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;}},{&quot;id&quot;:285559,&quot;user_id&quot;:22541131,&quot;publication_id&quot;:363171,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:363171,&quot;name&quot;:&quot;Asides&quot;,&quot;subdomain&quot;:&quot;asides&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Things that don't go on main&quot;,&quot;logo_url&quot;:null,&quot;author_id&quot;:22541131,&quot;theme_var_background_pop&quot;:&quot;#25BD65&quot;,&quot;created_at&quot;:&quot;2021-05-17T20:50:25.325Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Jon Stokes&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;}}],&quot;twitter_screen_name&quot;:&quot;jonst0kes&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.jonstokes.com/p/machine-learning-and-deflationary?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!RaoY!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F4f382f14-88dd-416e-a6d7-c44ffb9be1db_256x256.png" loading="lazy"><span class="embedded-post-publication-name">jonstokes.com</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Machine Learning And Deflationary Contagions</div></div><div class="embedded-post-body">Here&#8217;s the most common question I get from members of the press, analysts, and others I talk to about generative AI: can you describe what impact this will have on my industry? Of course, this is the obvious question to ask, right? So I get it all the time, and because I get it all the time I&#8217;ve developed a standard high-level answer to it that I&#8217;ll shar&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 17 likes &#183; 9 comments &#183; Jon Stokes</div></a></div><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:97871827,&quot;url&quot;:&quot;https://bensbites.substack.com/p/conversational-answer-engine&quot;,&quot;publication_id&quot;:1199622,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Ben's Bites&quot;,&quot;publication_logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/454da0a2-4d3c-4843-aa04-28e7eb5910b4_173x173.png&quot;,&quot;title&quot;:&quot;Conversational answer engine&quot;,&quot;truncated_body_text&quot;:&quot;Hey folks, it's Friiiiyay! &#129314; - don't be the person to ever say that. We had our first official meetup yesterday in London and I loved it. Balderton Capital hosted ~200 of us in their offices and provided fab food and drinks. A big TY to them! And a&quot;,&quot;date&quot;:&quot;2023-01-20T14:01:09.859Z&quot;,&quot;like_count&quot;:0,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:218770,&quot;name&quot;:&quot;Ben Tossell&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/d04d6aa2-3db5-4299-bd14-9192ed127acd_400x400.jpeg&quot;,&quot;bio&quot;:&quot;I cover what's going on in AI&quot;,&quot;profile_set_up_at&quot;:&quot;2022-11-09T22:46:14.767Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1154036,&quot;user_id&quot;:218770,&quot;publication_id&quot;:1199622,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:1199622,&quot;name&quot;:&quot;Ben's Bites&quot;,&quot;subdomain&quot;:&quot;bensbites&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Your daily dose of what's going on in AI. In 5 minutes or less, with a touch of humour. Read by over 10,000 others from Google, a16z, Sequoia, Amazon, Meta and more.&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/454da0a2-4d3c-4843-aa04-28e7eb5910b4_173x173.png&quot;,&quot;author_id&quot;:218770,&quot;theme_var_background_pop&quot;:&quot;#25BD65&quot;,&quot;created_at&quot;:&quot;2022-11-18T14:09:49.293Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:&quot;Ben from Ben's Bites&quot;,&quot;copyright&quot;:&quot;Ben Tossell&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;}}],&quot;twitter_screen_name&quot;:&quot;bentossell&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="/__u/bensbites.substack.com/p/conversational-answer-engine?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!nU5D!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F454da0a2-4d3c-4843-aa04-28e7eb5910b4_173x173.png" loading="lazy"><span class="embedded-post-publication-name">Ben's Bites</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Conversational answer engine</div></div><div class="embedded-post-body">Hey folks, it's Friiiiyay! &#129314; - don't be the person to ever say that. We had our first official meetup yesterday in London and I loved it. Balderton Capital hosted ~200 of us in their offices and provided fab food and drinks. A big TY to them! And a&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; Ben Tossell</div></a></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/tunguz/status/1615336891875995653?s=61&amp;t=1XwTcGxYb6tQmk6gIkSVYA&quot;,&quot;full_text&quot;:&quot;As one of the most unsurprising moves in tech, <span class=\&quot;tweet-fake-link\&quot;>@Microsoft</span> has announced that they will incorporate all of <span class=\&quot;tweet-fake-link\&quot;>@OpenAI</span> tools into their products. In particular, this means a wide availability of ChatGPT in various Office products.\n\n1/9&quot;,&quot;username&quot;:&quot;tunguz&quot;,&quot;name&quot;:&quot;Bojan Tunguz&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 17 13:15:01 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:222,&quot;like_count&quot;:1796,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/thatroblennon/status/1615104249192488980?s=61&amp;t=1XwTcGxYb6tQmk6gIkSVYA&quot;,&quot;full_text&quot;:&quot;After tons of research and experimentation, here are the 6 types of information I provide in my ChatGPT mega-prompts: &quot;,&quot;username&quot;:&quot;thatroblennon&quot;,&quot;name&quot;:&quot;Rob Lennon &#128495; | Audience Growth&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Mon Jan 16 21:50:34 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FmoAwM5XEAAnyla.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/dQbcAUQ0dy&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:492,&quot;like_count&quot;:3670,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/wholemarsblog/status/1615217211198832640?s=61&amp;t=1XwTcGxYb6tQmk6gIkSVYA&quot;,&quot;full_text&quot;:&quot;Now this is creepy. \n\nThis AI model can detect the pose of people in the room based just on WiFi signals. No camera needed. &quot;,&quot;username&quot;:&quot;WholeMarsBlog&quot;,&quot;name&quot;:&quot;Whole Mars Catalog&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 17 05:19:27 +0000 2023&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/FmpnfeXacAAuDpc.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/wO6epwO5SA&quot;,&quot;alt_text&quot;:null}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:1450,&quot;like_count&quot;:8363,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/openai/status/1615160228366147585?s=61&amp;t=1XwTcGxYb6tQmk6gIkSVYA&quot;,&quot;full_text&quot;:&quot;We've learned a lot from the ChatGPT research preview and have been making important updates based on user feedback. ChatGPT will be coming to our API and Microsoft's Azure OpenAI Service soon. \n\nSign up for updates here: &quot;,&quot;username&quot;:&quot;OpenAI&quot;,&quot;name&quot;:&quot;OpenAI&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Tue Jan 17 01:33:01 +0000 2023&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:2113,&quot;like_count&quot;:13357,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{&quot;url&quot;:&quot;https://share.hsforms.com/1u4goaXwDRKC9-x9IvKno0A4sk30&quot;,&quot;title&quot;:&quot;Form&quot;,&quot;description&quot;:null,&quot;domain&quot;:&quot;share.hsforms.com&quot;},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><div><hr></div><p>Finally, in case you missed it, I also shared Part 2 of my series on the origins of Deep Learning:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:96184253,&quot;url&quot;:&quot;https://hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-c6e&quot;,&quot;publication_id&quot;:1281858,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The Hitchhikers Guide to AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png&quot;,&quot;title&quot;:&quot;&#129299;A Deep Dive into Deep Learning: Part 2&quot;,&quot;truncated_body_text&quot;:&quot;Hi Readers! Thank you for subscribing to my newsletter. As promised, this is Part 2 of my deep dive into the origins of deep learning. If you missed Part 1, you can read it here. The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a&quot;,&quot;date&quot;:&quot;2023-01-19T15:01:11.454Z&quot;,&quot;like_count&quot;:2,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:719908,&quot;name&quot;:&quot;AJ Asver&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/69f998e3-0506-4a6d-90d9-a2f0d6c0f693_400x400.jpeg&quot;,&quot;bio&quot;:&quot;Exploring AI. Prev Product @compound, @brexhq, @coinbase, @google. Alum @ycombinator, @uniofoxford. Amateur DJ and dad of twins. All views expressed are my own.&quot;,&quot;profile_set_up_at&quot;:&quot;2022-03-20T21:13:04.053Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1239754,&quot;user_id&quot;:719908,&quot;publication_id&quot;:1281858,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:1281858,&quot;name&quot;:&quot;The Hitchhikers Guide to AI&quot;,&quot;subdomain&quot;:&quot;hitchhikersguidetoai&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;A newsletter for builders, product managers and founders new to AI. Every week I share updates on my adventures learning about AI from scratch. My goal is to break down complex concepts and provide practical insights for others new to the space.\n&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png&quot;,&quot;author_id&quot;:719908,&quot;theme_var_background_pop&quot;:&quot;#B599F1&quot;,&quot;created_at&quot;:&quot;2023-01-02T20:13:10.471Z&quot;,&quot;rss_website_url&quot;:null,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;AJ Asver&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;}}],&quot;twitter_screen_name&quot;:&quot;_aj&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;inviteAccepted&quot;:true}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="/__u/hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-c6e?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!Ugi3!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png" loading="lazy"><span class="embedded-post-publication-name">The Hitchhikers Guide to AI</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">&#129299;A Deep Dive into Deep Learning: Part 2</div></div><div class="embedded-post-body">Hi Readers! Thank you for subscribing to my newsletter. As promised, this is Part 2 of my deep dive into the origins of deep learning. If you missed Part 1, you can read it here. The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 2 likes &#183; AJ Asver</div></a></div><p><em>That&#8217;s all for this week!</em></p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhikers Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[🤓A Deep Dive into Deep Learning: Part 2]]></title><description><![CDATA[Part 2 of 3 posts on the history of Deep Learning and the foundational developments that led to today&#8217;s AI innovations.]]></description><link>https://operatorsguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-c6e</link><guid isPermaLink="false">https://operatorsguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-c6e</guid><dc:creator><![CDATA[AJ Asver]]></dc:creator><pubDate>Thu, 19 Jan 2023 15:01:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!z-7M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!z-7M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!z-7M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg" width="1024" height="763" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:763,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:403719,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!z-7M!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc3c8837-9067-484d-a145-94251e962c96_1024x763.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi Readers!</p><p>Thank you for subscribing to my newsletter. As promised, this is Part 2 of my deep dive into the origins of deep learning. If you missed Part 1, you can read it <a href="/__u/open.substack.com/pub/hitchhikersguidetoai/p/a-deep-dive-into-deep-learning-part?r=ffhg&amp;utm_campaign=post&amp;utm_medium=web">here</a>. </p><p>The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a <em>layer deeper</em> into foundational concepts and try to decipher the jargon.</p><p>P.S. Don&#8217;t forget to hit subscribe if you want to receive more AI content like this!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/operatorsguidetoai.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Part 2 - The Hidden Layers of Deep Learning</h2><p>We ended Part 1 at the close of the 1970s in the AI winter. Research had slowed down because of limitation of single layer neural networks and computers aren&#8217;t yet powerful enough to train large networks. </p><p>Let&#8217;s pick things up in the 1980s&#8230;</p><h3>1980s</h3><p>As we learned in Part 1, single layer neural networks are limited to only being able to learn linearly separable data, which means they couldn&#8217;t learn simple mathematical functions like XOR. In theory, adding just one additional layer to a single-layer network allows it to approximate any mathematical function. In practice however, two-layer networks end up being too big and too slow to be useful. </p><p>Two-layer neural networks have a large number of parameters, which need to be adjusted during the training process. This requires a lot of computational resources, making the training process very slow, especially when working with large datasets. Two-layer neural networks also have a limited capacity to learn complex patterns and features from data. Because they only have two layers (input and output layer), they are not able to extract high-level features that can be used to represent data beyond those directly represented in the input layer. This makes it difficult to learn complex relationships and patterns. </p><p>A great example of the limitations of a two-layer network is in the task of image recognition. A two-layer neural network, would only have an input layer that receives the raw image data and an output layer that produces the final prediction or decision. The lack of additional layers in between the input and output layers limits the capacity of the network to extract high-level features such as edges, textures, and shapes from the image data.</p><p>In order to create neural networks that are more capable of learning, researchers would need to go beyond two layers. These multi-layer neural networks are also known as Deep Neural Networks (DNNs). </p><div><hr></div><p><strong>&#129299; What is a Deep Neural Network?</strong></p><p>A deep neural network (DNN) is a neural network with many layers, typically composed of an input layer, multiple hidden layers, and an output layer. Each layer contains a set of interconnected "neurons" that perform computations on the input data and pass the results to the next layer. The neurons in a deep neural network are connected by weights which can be learned through training. The goal of training is to adjust these weights so that the DNN produces a desired output for a given input.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!pJY9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 424w, /__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 848w, /__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!pJY9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp" width="1318" height="862" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:862,&quot;width&quot;:1318,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:44822,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 424w, /__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 848w, /__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!pJY9!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78834b48-6445-44cb-ace5-ad7422a13dd4_1318x862.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Towards Data Science</figcaption></figure></div><p><em>What are hidden layers? </em></p><p>The input layer takes in the input data, which is then processed by the hidden layers. The hidden layers are called "hidden" because their internal workings are not directly observable and are not part of the network's input or output. These layers use a set of weights and biases to transform the input data, passing it through multiple non-linear processing stages known as activation functions. The output of the last hidden layer is then passed on to the output layer, which produces the final output of the network.</p><p>The purpose of the hidden layers is to extract and abstract features from the input data and pass them on to the next layer. The more hidden layers a DNN has, the more complex patterns it can learn and represent. The number of hidden layers and the number of neurons in each hidden layer is a parameter of the network called its architecture, and can be adjusted during the training process to optimize the performance of the network.</p><p>To simplify, think of a DNN as a function that can map input data to output data by passing through the layers. Each layers adapts the function to better fit the input-output pairs that it sees during the training phase, with the final output being the output of the last layer.</p><p>For example, let's say that the DNN is trained on a dataset of images of handwritten digits. During the training process, the DNN would learn the statistical patterns and relationships between the pixel values of the images and the labels indicating which digit the image represents.</p><p>Once the DNN is trained, it can be used to classify new images of handwritten digits. To classify a new image, the DNN would take the pixel values of the image as input and pass them through the layers of the network, using the patterns and relationships it learned during training to classify the image as a specific digit.</p><div><hr></div><h3>1980s continued&#8230;</h3><p>In 1986, a pivotal book was published: Parallel Distributed Processing (PDP) by David Rumelhart, James McClelland and the PDP Research Group<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.  In the book, David Rumelhart and James L. McClelland that explores the use of artificial neural networks for computational modeling. The PDP series includes several influential papers on the development of neural networks and their applications.</p><p>One of the main contributions of the PDP series was the use of Paul Werbos&#8217;  backpropagation algorithm for training neural networks. The PDP series also introduced the concept of distributed representation, which is the idea that the meaning of a concept can be represented by the pattern of activity across multiple neurons in the network. This idea is important because it allows neural networks to learn more complex relationships between the input and output data, and to generalize better to new data.</p><div><hr></div><h3>1990s</h3><p>In the 1990s, the term "deep learning" is coined by Igor Aizenberg and colleagues to describe multi-layered neural networks. The biggest breakthrough of the decade however, is when the first successful application of deep learning is demonstrated by Yann LeCun, Yoshua Bengio, and Geoffrey Hinton in their work on handwritten digit recognition using a convolutional neural network (CNN)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>.</p><p>LeCun&#8217;s, Bengio and Hilton&#8217;s work is seminal because it demonstrates the effectiveness of deep learning for real-world applications. Prior to this work, there had been limited success in using neural networks for practical tasks, and many researchers were skeptical of their potential. However, the results of this work shows that deep learning can be used to achieve high accuracy on a challenging real-world task, paving the way for further research and development in the field.</p><p>Here&#8217;s LeCun demo-ing his CNN, LeNet 1 in 1993:</p><div id="youtube2-iuPD8T2OQjs" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;iuPD8T2OQjs&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/iuPD8T2OQjs?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><p><strong>&#129299; What is a Convolutional Neural Network?</strong></p><p>A convolutional neural network (CNN) is a type of deep neural network that is commonly used in image and video recognition tasks. It is designed to automatically and adaptively learn spatial hierarchies of features from input data, making it particularly well-suited for image analysis.</p><p>A CNN consists of multiple layers of interconnected neurons, which process and analyze the input data. The layers of a CNN are organized in such a way that they learn increasingly complex features of the input data as the data passes through the network.</p><p><em>Why is it &#8220;convolutional&#8221;?</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YNeG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 424w, /__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 848w, /__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YNeG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp" width="805" height="403" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:403,&quot;width&quot;:805,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:27986,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 424w, /__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 848w, /__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!YNeG!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea657eef-17bb-4b46-8393-3d3aa50b1939_805x403.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A visualization of how the different layers of a CNN interpret an image</figcaption></figure></div><p>One key feature of CNNs is the use of "convolutional" layers, which are designed to automatically learn and extract features from the input data. These layers apply a set of filters to the input data and use the resulting output to detect patterns or features in the data. This process is repeated multiple times, allowing the CNN to learn increasingly complex features of the input data as it passes through the network.</p><p>Here&#8217;s a great video by Google that visualizes how a CNN works: </p><div id="youtube2-b27hzEs8YWw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;b27hzEs8YWw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/b27hzEs8YWw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><h3>1990s continued&#8230;</h3><p>In 1995, Jurgen Schmidhuber and his student Sepp Hochreiter publish a paper on the concept of "long short-term memory" (LSTM) units<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>, which are a type of Recurrent Neural Network (RNN) that are able to capture long-term dependencies in time series data. </p><p>Traditional RNNs have difficulty in capturing long-term dependencies (e.g. one word after another in a sequence of text), which can limit their effectiveness for certain tasks. To address this issue, Schmidhuber and Hochreiter introduce the concept of LSTM units, which are able to "remember" information for long periods of time and use it to make predictions or decisions later. They demonstrate the effectiveness of LSTM units for a number of tasks, including language modeling and polyphonic music modeling, and show that they outperformed traditional RNNs in these tasks.</p><p>LSTMs go on to have a significant impact on the development of deep learning models for tasks such as language modeling, machine translation, and speech recognition.</p><div><hr></div><p><strong>&#129299; What is a Recurrent Neural Network?</strong></p><p>A recurrent neural network (RNN) is a type of artificial neural network that is designed to process sequential data, such as text or time series data. It is called "recurrent" because it makes use of sequential information, passing the output from one step of the processing back into the network as input for the next step.</p><p><em>How does an RNN work?</em></p><p>In a recurrent neural network (RNN), the neurons are connected in a directed cycle, meaning that the output from one step in the processing is passed as input to the next step in the cycle. This allows information to be passed from one step to the next and for the network to use information from the past to inform its current and future processing. For example, predicting where an object is going based on it&#8217;s passed co-ordinates or predicting the next word in a text based on previous words for auto-competion.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!USbI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_webp, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!USbI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg" width="1245" height="1304" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1304,&quot;width&quot;:1245,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:300059,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_424, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_848, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_1272, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!USbI!, /__u/operatorsguidetoai.substack.com/w_1456, /__u/operatorsguidetoai.substack.com/c_limit, /__u/operatorsguidetoai.substack.com/f_auto, /__u/operatorsguidetoai.substack.com/q_auto:good, /__u/operatorsguidetoai.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f463020-174f-4dd5-8846-675b48df7d17_1245x1304.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Google&#8217;s query autocompletion is an example of deep learning used for text prediction</figcaption></figure></div><p>For example, in a language modeling task, an RNN might take a sequence of words as input and use the output from processing the previous word to inform its processing of the current word. This allows the RNN to capture the context and dependencies between words, which is important for understanding the meaning of the text.</p><p><em>What about LSTMs? How do they fit in?</em></p><p>A long short-term memory (LSTM) unit is a type of recurrent neural network (RNN) that is able to capture long-term dependencies in time series data. It is called "long short-term memory" because it is able to remember information for long periods of time and use it to make predictions or decisions later. For example, storing a whole sentence in a text to predict the next word vs just storing the last word. </p><p><em>What are RNNs useful for?</em></p><p>RNNs are particularly well-suited for tasks such as language modeling, machine translation, and speech recognition, where the order and context of the input data is important. They have also been applied to a wide range of other tasks, including image and video analysis, music generation, and protein folding prediction.</p><p><em>Here&#8217;s a great article that explains RNNs in more detail: <a href="https://towardsdatascience.com/introducing-recurrent-neural-networks-f359653d7020">Introducing Recurrent Neural Networks</a>.</em></p><div><hr></div><h3>To be continued&#8230;</h3><p>We end Part 2 with deep learning making steady progress thanks to the dedicated work of a small group of researchers including Yann LeCun, Yoshua Bengio and Geoffrey Hinton. During this period, many of their academic papers were rejected by journals and conferences because of their use of neural networks, despite dramatically outperforming any previous approaches. </p><p>In 2018, LeCun, Bengio and Hinton were awarded the Turing Award, the highest honor in Computer Science, for their dedication and persistence in developing neural networks despite skepticism from the academic world.</p><p><em><strong>In Part 3, we will see how the work of these researchers laid the foundations for modern day AI&#8230;</strong></em></p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:99015940,&quot;url&quot;:&quot;https://hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4&quot;,&quot;publication_id&quot;:1281858,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The Hitchhikers Guide to AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png&quot;,&quot;title&quot;:&quot;&#129299;A Deep Dive into Deep Learning: Part 3&quot;,&quot;truncated_body_text&quot;:&quot;Hi Readers! Thank you for subscribing to my newsletter. Here&#8217;s the final part of my deep dive into the origins of deep learning. In case you missed it, here are Part 1 and Part 2. The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a&quot;,&quot;date&quot;:&quot;2023-01-26T22:45:27.329Z&quot;,&quot;like_count&quot;:1,&quot;comment_count&quot;:0,&quot;bylines&quot;:[],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="/__u/hitchhikersguidetoai.substack.com/p/a-deep-dive-into-deep-learning-part-ba4?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!Ugi3!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09b50b29-e8da-415e-a8ea-002b928e5c67_512x512.png" loading="lazy"><span class="embedded-post-publication-name">The Hitchhikers Guide to AI</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">&#129299;A Deep Dive into Deep Learning: Part 3</div></div><div class="embedded-post-body">Hi Readers! Thank you for subscribing to my newsletter. Here&#8217;s the final part of my deep dive into the origins of deep learning. In case you missed it, here are Part 1 and Part 2. The field of deep learning is filled with lots of jargon. When you see the &#129299; emoji, that&#8217;s where I go a&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">4 years ago &#183; 1 like</div></a></div><div><hr></div><p><em>Fun fact: I used ChatGPT for a lot of the research for this series of posts. Over the course of a week I asked ChatGPT dozens of questions about the history of deep learning and the concepts behind it. To make sure the article was accurate, I asked ChatGPT to provide citations to relevant scientific papers which I&#8217;ve included in the footnotes. If you find any errors, please reach out or comment below and I will fix them!</em></p><div><hr></div><p>If you enjoy reading my posts and would like to receive more content about AI for builders, founders and product managers, please subscribe!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://operatorsguidetoai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Hitchhikers Guide to AI! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>David E. Rumelhart, Geoffrey E. Hinton, and James L. McClelland. (1986) <a href="https://web.stanford.edu/~jlmcc/papers/PDP/Volume%201/Chap2_PDP86.pdf">"A general framework for parallel distributed processing."</a> <em>Parallel distributed processing: Explorations in the microstructure of cognition</em> </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>LeCun, Y., Bengio, Y., &amp; Hinton, G. (1995). <a href="http://yann.lecun.com/exdb/publis/pdf/lecun-bengio-95a.pdf">Convolutional networks for images, speech, and time series</a>. The Handbook of Brain Theory and Neural Networks</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Hochreiter, S., &amp; Schmidhuber, J. (1997). <a href="https://direct.mit.edu/neco/article-abstract/9/8/1735/6109/Long-Short-Term-Memory?redirectedFrom=fulltext">Long short-term memory</a>. Neural computation</p><p></p></div></div>]]></content:encoded></item></channel></rss>