<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Practical Data Scientist]]></title><description><![CDATA[A newsletter focused on applying data science to real-world problems, with practical insights on building impactful solutions and growing a successful career in the field.]]></description><link>https://thepracticaldatascientist.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!-hnq!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb48b4ac0-533f-4cf8-8f0b-e7321abf793a_1024x1024.png</url><title>The Practical Data Scientist</title><link>https://thepracticaldatascientist.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 16:47:30 GMT</lastBuildDate><atom:link href="/__u/thepracticaldatascientist.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Gowthami Peri]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[thepracticaldatascientist@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[thepracticaldatascientist@substack.com]]></itunes:email><itunes:name><![CDATA[Gowthami Peri]]></itunes:name></itunes:owner><itunes:author><![CDATA[Gowthami Peri]]></itunes:author><googleplay:owner><![CDATA[thepracticaldatascientist@substack.com]]></googleplay:owner><googleplay:email><![CDATA[thepracticaldatascientist@substack.com]]></googleplay:email><googleplay:author><![CDATA[Gowthami Peri]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Collaborative Filtering: Learning From What Others Like]]></title><description><![CDATA[How User Behavior Powers Personalized Recommendations]]></description><link>https://thepracticaldatascientist.substack.com/p/collaborative-filtering-learning</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/collaborative-filtering-learning</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 01 Sep 2026 14:02:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Wimx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Wimx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 424w, /__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 848w, /__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Wimx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png" width="1456" height="801" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:801,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1706238,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/213435989?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 424w, /__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 848w, /__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Wimx!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc95b4a67-2eb2-4060-a482-e01b61f8683b_1738x956.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>In the previous article, we looked at <strong>content-based filtering</strong>. The idea was simple: understand what a user already likes, look at the characteristics of those items, and recommend other items with similar features.</p><p>It works well, but it also has an important limitation.</p><p>If you&#8217;ve watched several science-fiction movies, a content-based system may keep recommending more science fiction. If you frequently buy running gear, you may continue seeing products that look very similar to things you&#8217;ve already purchased.</p><p>The recommendations can be highly relevant, but they can also become predictable.</p><p>This is known as <strong>overspecialization</strong>. By relying heavily on a user&#8217;s existing preferences, the system may keep reinforcing what it already knows instead of helping the user discover something new.</p><p>That&#8217;s a problem for the user, but it can also become a problem for the business.</p><p>A retailer may have thousands of products that a customer would potentially love but never see because they don&#8217;t resemble previous purchases. A streaming platform may have an entire catalog of content that remains undiscovered because recommendations stay within a narrow set of genres.</p><p>Good recommendations shouldn&#8217;t just help users find <strong>more of the same</strong>. They should also create opportunities for discovery.</p><p>So how can we recommend something outside a user&#8217;s usual preferences without simply guessing?</p><p>One answer is to learn from <strong>other users</strong>.</p><p>Imagine you and another user have enjoyed many of the same movies. You both liked <em>Interstellar</em>, <em>Inception</em>, and <em>The Martian</em>. That overlap suggests that your tastes may be similar.</p><p>Now suppose that user also loved a movie you&#8217;ve never watched.</p><p>Even if the movie doesn&#8217;t look particularly similar to your previous favorites, their behavior gives us a new signal:</p><blockquote><p><strong>People who like many of the same things you do also liked this. Maybe you will too.</strong></p></blockquote><p>This is the intuition behind <strong>collaborative filtering</strong>.</p><p>Instead of relying primarily on the characteristics of the items, collaborative filtering learns from patterns in how users interact with them. Those interactions might include ratings, purchases, clicks, views, listens, or likes.</p><p>This allows the system to uncover relationships that aren&#8217;t obvious from item features alone. Two movies don&#8217;t need to share a genre to be related. Two products don&#8217;t need to belong to the same category. If similar users consistently interact with them, the system can learn that connection.</p><p>In this article, we&#8217;ll look at how collaborative filtering works, explore <strong>user-based and item-based approaches</strong>, and see how collective user behavior can turn into personalized recommendations.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Is Collaborative Filtering?</h2><p>Collaborative filtering makes recommendations by learning from patterns in <strong>user behavior</strong>.</p><p>Unlike content-based filtering, it doesn&#8217;t necessarily need to know anything about the items themselves. We don&#8217;t need a movie&#8217;s genre, actors, director, or description. Instead, we look at how users have interacted with different movies.</p><p>The underlying assumption is simple:</p><blockquote><p><strong>Users who behaved similarly in the past may have similar preferences in the future.</strong></p></blockquote><h3>The User-Item Interaction Matrix</h3><p>A common way to represent this information is with a <strong>user-item interaction matrix</strong>.</p><p>Imagine four users and five movies:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_1qV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 424w, /__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 848w, /__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_1qV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png" width="1286" height="418" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:418,&quot;width&quot;:1286,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:35655,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/213435989?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 424w, /__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 848w, /__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_1qV!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719c43cb-54f3-45d6-a82d-a240ebc99a1c_1286x418.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Each row represents a <strong>user</strong>, while each column represents an <strong>item</strong>. The values capture interactions&#8212;in this case, movie ratings from 1 to 5.</p><p>The question marks are where things become interesting. User 1 hasn&#8217;t rated Movies C or E, so we don&#8217;t know whether they would enjoy them.</p><p>But look at User 1 and User 2.</p><p>Both gave Movie A a rating of 5, Movie B a 4, and Movie D a 1. Their preferences appear to be quite similar.</p><p>User 2 also gave <strong>Movie C a rating of 5</strong>.</p><p>That gives us evidence that Movie C could be a good recommendation for User 1&#8212;even without knowing anything about what Movie C is actually about.</p><h3>Learning From the Crowd</h3><p>This is where the word <strong>collaborative</strong> comes from. The system uses the collective behavior of many users to learn which recommendations might be relevant to an individual.</p><p>The same idea works with more than ratings. An e-commerce platform could use purchases, a music platform could use listens, and a news platform could use clicks or reading history.</p><p>These are all different ways of building the interaction matrix.</p><p>In practice, most of these matrices are extremely sparse. A user may interact with only a tiny fraction of the millions of items available, leaving most user-item combinations empty.</p><p>The challenge is to use the interactions we <strong>do</strong> have to make useful predictions about the ones we don&#8217;t.</p><p>There are two classic ways to approach this problem:</p><p><strong>User-based collaborative filtering</strong> asks:</p><blockquote><p><em>Which users are similar to this user?</em></p></blockquote><p><strong>Item-based collaborative filtering</strong> asks:</p><blockquote><p><em>Which items tend to receive similar interactions from users?</em></p></blockquote><p>We&#8217;ll start with the first approach.</p><div><hr></div><h2>User-Based Collaborative Filtering</h2><p>User-based collaborative filtering starts with a simple idea:</p><blockquote><p><strong>Find users with similar preferences and use their behavior to recommend something new.</strong></p></blockquote><p>Let&#8217;s return to our movie example.</p><p>Suppose User 1 has rated:</p><p><strong>Movie A:</strong> 5<br><strong>Movie B:</strong> 4<br><strong>Movie D:</strong> 1</p><p>User 2 has rated those same movies:</p><p><strong>Movie A:</strong> 5<br><strong>Movie B:</strong> 4<br><strong>Movie D:</strong> 1</p><p>Their preferences look very similar. They both liked Movies A and B and disliked Movie D.</p><p>But User 2 has also watched <strong>Movie C</strong> and rated it 5.</p><p>Since User 1 and User 2 have shown similar preferences in the past, Movie C becomes a strong candidate to recommend to User 1.</p><h3>How Do We Find Similar Users?</h3><p>With millions of users, we need a mathematical way to measure similarity.</p><p>One option is <strong>cosine similarity</strong>, the same concept we used in content-based filtering. Instead of comparing item-feature vectors, however, we&#8217;re now comparing <strong>user interaction vectors</strong>.</p><p>For example:</p><p><strong>User 1:</strong> [5, 4, ?, 1]<br><strong>User 2:</strong> [5, 4, 5, 1]</p><p>We compare their ratings on the items they have both interacted with. The more similar their interaction patterns are, the more similar we consider the users.</p><p>Other measures, such as <strong>Pearson correlation</strong>, can also be used. Pearson correlation can be particularly useful with ratings because different users may use rating scales differently.</p><h3>From Similar Users to Recommendations</h3><p>In practice, we don&#8217;t usually rely on just one similar user. We find a group of the target user&#8217;s closest neighbors and look at the items they liked.</p><p>Users who are more similar can be given more influence when estimating how much the target user might like an unseen item.</p><p>The process is essentially:</p><p><strong>Find similar users &#8594; Look at what they liked &#8594; Score unseen items &#8594; Rank &#8594; Recommend</strong></p><p>This is often called <strong>user-user collaborative filtering</strong> or <strong>user-based nearest-neighbor collaborative filtering</strong>.</p><p>The important idea is that we still don&#8217;t need to know anything about Movie C itself. We recommend it because <strong>people whose preferences resemble User 1&#8217;s preferences liked it</strong>.</p><div><hr></div><h2>Item-Based Collaborative Filtering</h2><p>User-based collaborative filtering finds <strong>people with similar preferences</strong>. Item-based collaborative filtering flips the idea around and looks for <strong>items that users tend to interact with in similar ways</strong>.</p><p>Suppose many users who gave Movie A a high rating also gave Movie C a high rating. Over time, the system learns that these two movies have similar interaction patterns.</p><p>Now imagine User 1 has already enjoyed Movie A but hasn&#8217;t watched Movie C. Because other users tend to like both movies, Movie C becomes a good recommendation.</p><p>The important point is that we don&#8217;t need to know anything about the movies themselves.</p><h3>How Are Similar Items Identified?</h3><p>Just as users can be represented by their interactions with items, items can be represented by the users who interacted with them.</p><p>For example:</p><p><strong>Movie A:</strong> [5, 5, 1, 2]<br><strong>Movie C:</strong> [?, 5, 1, 2]</p><p>Each position represents a different user&#8217;s rating. Since Movies A and C receive similar ratings from the users who have seen both, the system may consider them similar.</p><p>Once again, measures such as <strong>cosine similarity</strong> can be used to compare these interaction vectors.</p><p>If a user likes Movie A, we can find movies with similar interaction patterns and use them as candidates for recommendation.</p><h3>Isn&#8217;t This the Same as Content-Based Filtering?</h3><p>They may sound similar, but there&#8217;s an important difference.</p><p><strong>Content-based filtering</strong> might decide that two movies are similar because they share the same genre, actors, director, or themes.</p><p><strong>Item-based collaborative filtering</strong> might decide that those same movies are similar because <strong>the same types of users tend to like them</strong>.</p><p>This means two movies could have completely different characteristics and still be considered similar by collaborative filtering.</p><p>A science-fiction movie and a historical drama may not look similar based on their content. But if users who enjoy one consistently enjoy the other, the system can discover that relationship.</p><p>That&#8217;s one of the strengths of collaborative filtering: <strong>it can uncover connections that aren&#8217;t obvious from the item features alone.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>From Similarity to Recommendations</h2><p>Finding similar users or items is useful, but similarity alone doesn&#8217;t give us a recommendation. We still need to decide which unseen items are most likely to interest the user.</p><p>This is where the <strong>neighbors</strong> come in.</p><h3>Using Similar Users</h3><p>Suppose we want to estimate whether User 1 would like Movie C.</p><p>We first find the users most similar to User 1 who have already rated Movie C. Instead of treating all of their ratings equally, we can give more weight to users whose preferences are more similar to User 1.</p><p>For example:</p><p><strong>User 2:</strong> Similarity = 0.90 &#8594; Movie C rating = 5<br><strong>User 3:</strong> Similarity = 0.70 &#8594; Movie C rating = 4<br><strong>User 4:</strong> Similarity = 0.20 &#8594; Movie C rating = 2</p><p>User 2&#8217;s rating should influence our prediction more than User 4&#8217;s because User 2 has a much more similar interaction history.</p><p>By combining these ratings and weighting them by similarity, we can estimate a preference score for User 1.</p><h3>Using Similar Items</h3><p>The same idea works with item-based collaborative filtering.</p><p>Suppose User 1 hasn&#8217;t watched Movie C but has given high ratings to Movies A and B. If Movie C has interaction patterns similar to Movies A and B, that&#8217;s evidence that User 1 may also enjoy it.</p><p>Again, the most similar items contribute more to the estimated score than weakly related ones.</p><h3>Turning Scores Into Recommendations</h3><p>We repeat this process for items the user hasn&#8217;t interacted with and estimate a preference score for each one.</p><p>Those scores can then be ranked:</p><p><strong>Movie C &#8594; 4.7</strong><br><strong>Movie F &#8594; 4.3</strong><br><strong>Movie E &#8594; 3.8</strong><br><strong>Movie G &#8594; 2.1</strong></p><p>The highest-scoring unseen items become the top recommendations.</p><p>The overall process is:</p><p><strong>Find neighbors &#8594; Estimate preference scores &#8594; Rank unseen items &#8594; Recommend</strong></p><p>This is the core idea behind neighborhood-based collaborative filtering. We use the behavior of similar users&#8212;or the interaction patterns of similar items&#8212;to fill in some of the missing information in the user-item matrix.</p><div><hr></div><h2>Advantages and Limitations of Collaborative Filtering</h2><p>Collaborative filtering is powerful because it doesn&#8217;t need to understand what an item actually is. It learns from the patterns created by users interacting with items.</p><p>That gives it some important advantages over content-based filtering, but it also introduces a different set of challenges.</p><h3>It Can Help Users Discover Something New</h3><p>One of the biggest advantages of collaborative filtering is <strong>discovery</strong>.</p><p>With content-based filtering, someone who watches science-fiction movies may continue receiving more science-fiction recommendations because the system looks for items with similar features.</p><p>Collaborative filtering can move beyond those obvious similarities. If people with similar tastes also enjoy a historical drama, the system can recommend that movie even though it looks very different from the user&#8217;s previous choices.</p><p>This can benefit both the user and the business. Users discover more of the catalog, while businesses have an opportunity to surface relevant items that might otherwise receive very little exposure.</p><h3>It Doesn&#8217;t Require Detailed Item Features</h3><p>Collaborative filtering can also work without detailed information about the items.</p><p>We don&#8217;t necessarily need to know a movie&#8217;s genre or actors, or a product&#8217;s category and description. If we have enough user-item interactions, the system can learn useful relationships directly from behavior.</p><p>This is particularly valuable when item features are difficult to define or don&#8217;t fully capture why people like something.</p><h3>The Cold-Start Problem</h3><p>Cold start is one of the biggest challenges for collaborative filtering.</p><p>A <strong>new user</strong> has no interaction history, so the system doesn&#8217;t yet know which existing users have similar preferences. Until the user starts clicking, watching, purchasing, or rating items, personalized recommendations are difficult.</p><p>A <strong>new item</strong> creates a similar problem. If nobody has interacted with it yet, collaborative filtering has no behavioral pattern from which to learn who might like it.</p><p>This is different from content-based filtering, which can often recommend a new item immediately using its features.</p><p>In practice, platforms may use popular items, onboarding preferences, content-based recommendations, or other signals until enough interaction data becomes available.</p><h3>Data Sparsity</h3><p>Even established users interact with only a tiny fraction of the available catalog.</p><p>Imagine a platform with one million users and 100,000 items. The user-item matrix contains an enormous number of possible interactions, but most of those cells will be empty.</p><p>This <strong>sparsity</strong> makes it harder to find reliable similarities, especially for users or items with relatively few interactions.</p><h3>Scalability</h3><p>Neighborhood-based collaborative filtering can also become expensive as the platform grows.</p><p>Finding similar users among millions of users&#8212;or comparing large numbers of items&#8212;requires significant computation. Techniques such as approximate nearest-neighbor search can help, but at very large scale, other approaches may become more practical.</p><p>This is one reason recommendation systems often move beyond traditional neighborhood methods to techniques such as <strong>matrix factorization and embeddings</strong>.</p><h3>The Trade-Off</h3><p>Collaborative filtering can uncover relationships that item features alone would never reveal. It is especially useful for discovery because recommendations are informed by the collective behavior of many users.</p><p>But that strength depends on having enough interaction data. New users, new items, sparse matrices, and large-scale similarity calculations can all make the approach more difficult.</p><p>There is no universally better choice between content-based and collaborative filtering. In practice, many recommendation systems combine both because their strengths and weaknesses complement each other.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Key Takeaways</h2><p>Collaborative filtering makes recommendations by learning from patterns in <strong>user-item interactions</strong>. Instead of relying on detailed information about the items themselves, it uses the collective behavior of users to discover relationships.</p><p>Here are the main ideas to remember:</p><ul><li><p><strong>The user-item interaction matrix is the foundation.</strong> It can contain explicit feedback such as ratings or implicit signals such as clicks, purchases, views, or likes.</p></li><li><p><strong>User-based collaborative filtering finds similar users.</strong> If people with preferences similar to yours liked something you haven&#8217;t seen, that item may be a good recommendation for you.</p></li><li><p><strong>Item-based collaborative filtering finds similar items based on interaction patterns.</strong> Two items can be considered similar even if their actual characteristics are completely different.</p></li><li><p><strong>Similarity can be measured in different ways.</strong> Cosine similarity is common, but the right measure depends on the type of interaction data.</p></li><li><p><strong>Neighbor preferences can be used to estimate scores for unseen items.</strong> Those scores are then ranked to generate the final recommendations.</p></li><li><p><strong>Collaborative filtering can improve discovery.</strong> Because it isn&#8217;t limited to item features, it can surface relevant items outside a user&#8217;s usual preferences.</p></li><li><p><strong>Cold start and sparsity are major challenges.</strong> New users and new items have little interaction history, while most users interact with only a small fraction of the available catalog.</p></li></ul><p>The central idea is simple:</p><blockquote><p><strong>Don&#8217;t just look at what an item is. Learn from how people interact with it.</strong></p></blockquote><p>That shift&#8212;from understanding item characteristics to learning from collective behavior&#8212;is what makes collaborative filtering so powerful.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>Collaborative filtering gives us a powerful way to generate recommendations without needing to understand the content of every item. By learning from user-item interactions, we can discover relationships that may not be obvious from item features alone.</p><p>But as the number of users and items grows, the interaction matrix becomes enormous&#8212;and mostly empty. Comparing users or items directly can become difficult and expensive.</p><p>What if we could take that large, sparse matrix and represent every user and item using a much smaller set of hidden characteristics?</p><p>That&#8217;s the idea behind <strong>matrix factorization</strong>.</p><p>Instead of working directly with millions of interactions, matrix factorization learns compact representations&#8212;or <strong>latent factors</strong>&#8212;for users and items. Those representations can then be used to estimate preferences and generate recommendations.</p><p>In the next article, we&#8217;ll explore what those latent factors actually mean, how matrix factorization learns them, and why this approach became such an important part of modern recommendation systems.</p><p><em><strong>Until next time, keep learning, keep building!</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h3></h3>]]></content:encoded></item><item><title><![CDATA[Where’s the Money in Data?]]></title><description><![CDATA[The Highest-Paying Roles and Companies in 2026]]></description><link>https://thepracticaldatascientist.substack.com/p/wheres-the-money-in-data</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/wheres-the-money-in-data</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Mon, 31 Aug 2026 14:34:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5MLY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><a href="https://dataford.io/">Dataford</a> shared a report from over 85,000 job listings and more than 600,000 interviews with me. Here&#8217;s what it shows about which data jobs pay the most, and how hard they are to get.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5MLY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 424w, /__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 848w, /__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5MLY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png" width="1418" height="942" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:942,&quot;width&quot;:1418,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1389857,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/213479679?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 424w, /__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 848w, /__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5MLY!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59341f8-1aa0-4c4a-907e-4a23f757d8d1_1418x942.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>A few weeks ago, <a href="https://dataford.io/">Dataford</a> sent me some data that surprised me. For years, &#8220;Data Scientist&#8221; was the top job in data. The best title, the best pay, the one everyone wanted. That&#8217;s no longer true. Two newer job titles now pay more: Research Engineer and Machine Learning Engineer.</p><p>Here&#8217;s what the data shows, and what it means if you&#8217;re building a career in data.</p><h2>The pay gap is bigger than you&#8217;d think</h2><p>Dataford looked at median pay across the US, using over 85,000 real job listings. They also scored how hard each job is to interview for, using more than 600,000 real interview reports, on a scale of 1 to 5.</p><p>The best-paid data job, Research Engineer, has a median pay of $200,000 a year. The lowest-paid, Research Analyst, pays under $93,000. That&#8217;s more than double, even though both jobs have &#8220;research&#8221; in the name.</p><p>The lesson here: a job title alone doesn&#8217;t tell you much about the pay. Two people can both call themselves an &#8220;Analyst&#8221; and earn very different salaries.</p><h2>Data jobs are splitting into two groups</h2><p>A clear pattern shows up in the data. A small group of jobs pays very well: Research Engineer, Machine Learning Engineer, Quantitative Analyst, and Data Scientist. These all pay over $170,000.</p><p>Below that is a much bigger group of jobs, all packed into a narrower pay range, roughly $93,000 to $141,000. This group includes Data Analyst, Business Intelligence Analyst, Analytics Engineer, and several others. Even though these jobs use different skills, like writing SQL, building dashboards, or running statistics, they all land in a similar, lower pay range.</p><p>In plain terms: <strong>jobs that build and ship machine learning systems are pulling far ahead in pay. Jobs based on analysis and reporting are staying behind.</strong></p><p>If you&#8217;re early in your data career, this is the biggest thing to take away. The fastest way to a higher salary is through machine learning and engineering skills, not analytics tools alone.</p><h2>The hardest job to get isn&#8217;t the best paid</h2><p>Here&#8217;s the most surprising finding in the data. Quantitative Analyst has the hardest interview of any job Dataford measured, out of 53 different roles. But it doesn&#8217;t pay the most. It actually pays less than both Research Engineer and Machine Learning Engineer.</p><p>So Quantitative Analysts go through the toughest interviews in tech and still end up earning less than people in some easier-to-get roles. If you&#8217;re deciding where to focus your job search energy, Machine Learning Engineer looks like a better trade: strong pay, for a somewhat easier interview.</p><h2>The company matters more than the job title</h2><p>One more pattern stood out. How hard a company&#8217;s interview is barely changes based on the job you&#8217;re applying for. But it changes a lot from company to company.</p><p>A few examples:</p><p>Anthropic and OpenAI both pay very well for data and machine learning roles, over $338,000 a year on average. Their interviews are also some of the hardest in the dataset.</p><p>Netflix pays the most of any company Dataford measured, at $473,000 a year. But its interview difficulty is only average, not the hardest.</p><p>Google, on the other hand, has one of the toughest interview processes of any company, but its pay for the same kind of role is less than half of what Netflix offers.</p><p>The lesson: two companies can pay very differently for the same job, at a similar interview difficulty. It pays to research the company, not just the job title.</p><h2>What this means for you</h2><p>If you&#8217;re planning your next move in data, a few things are worth remembering:</p><p>The top of the pay scale has shifted from Data Scientist to Machine Learning Engineer and Research Engineer. If higher pay is your goal, that&#8217;s the direction to build your skills in.</p><p>The middle of the market, roughly $93,000 to $141,000, is crowded. Moving from one analyst role to another is usually a small pay bump. The bigger jump in pay comes from moving into machine learning or engineering work.</p><p>A hard interview doesn&#8217;t mean a higher salary. It often says more about the company than about the job. It&#8217;s worth researching companies as carefully as you prepare for their interviews.</p><h2>About Dataford</h2><p>That last point, about interviews telling you more about the company than the job, is exactly what <a href="https://dataford.io/">Dataford</a> is built to help with. Founded in 2023 by Amney Mounir, a former Lead Analyst at Meta, Dataford is an interview prep platform built specifically for tech, AI, and data roles. Rather than generic study guides, it uses real candidate interview reports to build prep that&#8217;s tailored to the specific company and role you&#8217;re interviewing for &#8212; a question bank of 19,000+ practice questions, over 40,000 company-specific guides across 50+ roles, AI-run mock interviews with automated feedback, and built-in SQL and Python practice spaces. The 613,000+ interview reports behind this article&#8217;s difficulty scores are the same kind of data that powers Dataford&#8217;s prep tools. You can check it out at <a href="https://dataford.io/">dataford.io</a>.</p><p><em>Data shared by Dataford. Pay figures are the midpoint of each job listing&#8217;s stated salary range, based on US job postings. Interview difficulty is scored from real interview reports, from 1 (very easy) to 5 (very hard).</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Building an End-to-End Content-Based Recommendation System]]></title><description><![CDATA[From product metadata to recommendations using TF-IDF and cosine similarity]]></description><link>https://thepracticaldatascientist.substack.com/p/building-an-end-to-end-content-based</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/building-an-end-to-end-content-based</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 25 Aug 2026 12:45:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!I4Go!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!I4Go!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 424w, /__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 848w, /__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 1272w, /__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!I4Go!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png" width="1440" height="956" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:956,&quot;width&quot;:1440,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2121963,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/212648125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 424w, /__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 848w, /__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 1272w, /__u/substackcdn.com/image/fetch/$s_!I4Go!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde05d8c9-0aa1-4b6d-8141-a39a9a17a8b3_1440x956.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>In my previous article, I explored how <strong>content-based recommendation systems work</strong>&#8212;how they use the characteristics of items to identify and recommend similar products.</p><p>But understanding the idea is only the first step.</p><p>What does it actually take to build one end-to-end?</p><p>In this article, we&#8217;ll move from theory to implementation by building a content-based recommendation system using the <strong>H&amp;M Personalized Fashion Recommendations dataset</strong>. We&#8217;ll start with raw product metadata and work our way through creating product profiles, converting them into numerical representations, measuring similarity, and generating recommendations.</p><p>But we won&#8217;t stop once the algorithm produces a list of similar products.</p><p>We&#8217;ll inspect the recommendations, identify where the first version falls short, improve the recommendation logic, and finally evaluate the system using actual customer transaction data.</p><p>By the end, we&#8217;ll have worked through the complete pipeline:</p><p><strong>Product Data &#8594; Product Profiles &#8594; TF-IDF Vectors &#8594; Cosine Similarity &#8594; Recommendations &#8594; Refinement &#8594; Evaluation</strong></p><p>The goal isn&#8217;t just to build a recommender that works technically. It&#8217;s to understand the decisions that turn a similarity algorithm into a more useful recommendation experience.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>1. Understanding the Dataset</h2><p>For this project, we&#8217;ll use the <strong>H&amp;M Personalized Fashion Recommendations</strong> dataset from Kaggle. It contains information about H&amp;M&#8217;s product catalog, customers, and historical transactions.</p><p>The dataset includes three main files:</p><ul><li><p><code>articles.csv</code> &#8212; product information and attributes</p></li><li><p><code>customers.csv</code> &#8212; customer-level information</p></li><li><p><code>transactions_train.csv</code> &#8212; historical customer purchases</p></li></ul><p>We&#8217;ll start with <code>articles.csv</code>, since a content-based recommender relies primarily on information about the <strong>items themselves</strong>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a5f561d2-7e94-4524-bcdd-bce5e66f7e3b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">articles = pd.read_csv("data/articles.csv")

print(articles.shape)
articles.head()</code></pre></div><p>The dataset contains more than <strong>100,000 products</strong>, with attributes describing different aspects of each article.</p><p>Some of the most useful fields for our recommender are:</p><ul><li><p><code>product_type_name</code> &#8212; dress, trousers, shirt, etc.</p></li><li><p><code>product_group_name</code> &#8212; broader product category</p></li><li><p><code>graphical_appearance_name</code> &#8212; solid, striped, patterned, etc.</p></li><li><p><code>colour_group_name</code> &#8212; product color</p></li><li><p><code>department_name</code> &#8212; department the product belongs to</p></li><li><p><code>section_name</code> &#8212; product section</p></li><li><p><code>garment_group_name</code> &#8212; garment category</p></li><li><p><code>detail_desc</code> &#8212; natural-language product description</p></li></ul><p>For example, a product in the catalog might look something like:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0gZv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 424w, /__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 848w, /__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0gZv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png" width="1246" height="160" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:160,&quot;width&quot;:1246,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:27993,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/212648125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 424w, /__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 848w, /__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0gZv!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43a87179-47f6-4171-9c31-5bc7a8e6f95d_1246x160.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>These attributes give us the raw material for describing what each product <strong>is</strong>.</p><p>That&#8217;s important because our recommender won&#8217;t initially know anything about customer preferences or what other shoppers purchased. It will make recommendations entirely by comparing the characteristics of one product with another.</p><p>Later in the article, we&#8217;ll bring in <code>transactions_train.csv</code>&#8212;not to build the recommendations, but to evaluate them against actual customer behavior.</p><p><strong>Dataset:</strong> <a href="https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data">H&amp;M Personalized Fashion Recommendations &#8212; Kaggle</a></p><p>With the product catalog loaded, the next step is deciding <strong>which attributes should define similarity between two products</strong>.</p><div><hr></div><h2>2. Preparing the Product Data</h2><p>Not every column in the product catalog needs to be part of the recommendation system. The goal is to select attributes that meaningfully describe a product and can help us determine whether two items are similar.</p><p>For this project, we&#8217;ll use a combination of product category, appearance, color, department, and description.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9da8c95e-b174-4241-948f-db2caa80ec3d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">content_features = [
    "product_type_name",
    "product_group_name",
    "graphical_appearance_name",
    "colour_group_name",
    "department_name",
    "section_name",
    "garment_group_name",
    "detail_desc"
]</code></pre></div><p>We also retain <code>article_id</code>, <code>product_code</code>, and <code>prod_name</code>. The <code>article_id</code> uniquely identifies an article, while <code>product_code</code> will become useful later when we deal with multiple variants of the same underlying product.</p><h3>Handling Missing Values</h3><p>Before combining these attributes, we need to handle missing values. Since we&#8217;re going to represent the product information as text, we can replace missing values with an empty string.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;89653f79-32c2-486f-9a39-ea68364c5fac&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">products[content_features] = (
    products[content_features].fillna("")
)
</code></pre></div><h3>Preparing Categorical Features</h3><p>There is one more small but important preprocessing step.</p><p>Consider the color <strong>Light Blue</strong>. If we treat all our attributes as ordinary text, the vectorizer will see <code>light</code> and <code>blue</code> as two separate terms. Similarly, <code>Garment Upper Body</code> would become three independent words.</p><p>For categorical attributes, we want to preserve the complete category instead:</p><p><code>Light Blue</code> &#8594; <code>light_blue</code></p><p><code>Garment Upper Body</code> &#8594; <code>garment_upper_body</code></p><p>We can do that with a simple transformation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7194da4c-047b-4f2f-8f21-5925452bcacf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for col in categorical_features:
    products[col] = (
        products[col]
        .str.lower()
        .str.replace(" ", "_")
    )</code></pre></div><p>Product descriptions are different. They contain natural language, so individual words can carry useful information. We therefore keep the description as regular text and apply only light cleaning.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4e8762bc-fc67-4a1a-a266-797fec8a57e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">products["detail_desc"] = (
    products["detail_desc"]
    .str.lower()
    .str.strip()
)</code></pre></div><p>At this point, we have a clean set of attributes describing each product.</p><p>Next, we&#8217;ll combine them to create a <strong>content profile for every item in the catalog</strong>&#8212;the representation our recommendation system will use to compare products.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>3. Building Product Content Profiles</h2><p>Now that the product attributes are clean, we need to bring them together into a single representation for each item.</p><p>Think of this as creating a <strong>profile that describes each product</strong> using everything we know about it.</p><p>First, we combine the categorical attributes:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;33a00ec6-c205-46c1-8f2c-86da474d7b80&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">products["categorical_content"] = (
    products[categorical_features]
    .agg(" ".join, axis=1)
)</code></pre></div><p>Then we add the product description:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4ae5363b-f3de-432f-a0f9-df41c9c38dfb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">products["content_profile"] = (
    products["categorical_content"]
    + " "
    + products["detail_desc"]
)</code></pre></div><p>A simplified content profile might now look something like:</p><pre><code><code>dress garment_full_body solid black dark
divided_basics jersey_basic
short fitted dress with narrow shoulder straps</code></code></pre><p>Instead of knowing the product only as an <code>article_id</code>, our system now has information about <strong>what the product is, what it looks like, where it belongs in the catalog, and how it is described</strong>.</p><h3>Converting Product Profiles into Vectors</h3><p>The content profiles are still text, so we need to convert them into a numerical form before we can compare products.</p><p>We&#8217;ll use <strong>TF-IDF (Term Frequency&#8211;Inverse Document Frequency)</strong>.</p><p>TF-IDF gives more importance to terms that help distinguish one product from another while reducing the influence of terms that occur across many products.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a5ef6819-3435-48f7-ab5c-50416edcf1a6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">tfidf = TfidfVectorizer(
    stop_words="english",
    max_features=10000
)

tfidf_matrix = tfidf.fit_transform(
    products["content_profile"]
)</code></pre></div><p>The result is a sparse matrix where:</p><p><strong>Rows &#8594; Products</strong></p><p><strong>Columns &#8594; Product attributes and descriptive terms</strong></p><p>Each product is now represented as a numerical vector containing information about its content.</p><h3>Why Use TF-IDF?</h3><p>We could simply one-hot encode all the categorical attributes. But our product profiles contain both <strong>structured attributes and natural-language descriptions</strong>.</p><p>TF-IDF gives us a simple way to represent both in the same feature space without introducing a much more complex modeling approach.</p><p>With every product represented as a vector, we&#8217;re ready for the key step: <strong>measuring the similarity between products and generating recommendations.</strong></p><div><hr></div><h2>4. Generating Recommendations with Cosine Similarity</h2><p>Now that every product has a numerical representation, we can start comparing them.</p><p>We&#8217;ll use <strong>cosine similarity</strong>, which measures how similar two vectors are based on the angle between them. A score closer to <strong>1</strong> indicates greater similarity, while a score closer to <strong>0</strong> indicates that the products share fewer characteristics.</p><p>Instead of calculating a similarity matrix between every pair of products&#8212;which would be unnecessarily large for a catalog of more than 100,000 items&#8212;we calculate similarity <strong>only when a recommendation is requested</strong>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d0d15846-f0c6-44f1-9bbf-a6acc49522c7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">similarity_scores = cosine_similarity(
    tfidf_matrix[idx],
    tfidf_matrix
).flatten()
</code></pre></div><p>We then rank the products from most to least similar and return the top results.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;658289a6-9117-44ac-b61b-477a810bbc16&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">similar_indices = similarity_scores.argsort()[::-1]
top_indices = similar_indices[:5]
</code></pre></div><p>The basic recommendation process is therefore quite simple:</p><p><strong>Select a Product &#8594; Find its TF-IDF Vector &#8594; Calculate Cosine Similarity &#8594; Rank Products &#8594; Return Top Recommendations</strong></p><h3>Our First Recommendations</h3><p>Let&#8217;s test the recommender using the <strong>Alcazar Strap Dress</strong>, a black dress from the <code>Divided Basics</code> section.</p><p>Our initial Top-5 recommendations looked like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!r77k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 424w, /__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 848w, /__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 1272w, /__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!r77k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png" width="1268" height="508" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:508,&quot;width&quot;:1268,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66785,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/212648125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 424w, /__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 848w, /__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 1272w, /__u/substackcdn.com/image/fetch/$s_!r77k!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7357a4-522a-4d39-91db-4b7dc674ac40_1268x508.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>At first glance, these are excellent similarity scores. The recommender clearly understands that these products are closely related to the selected dress.</p><p>But there is also an obvious problem.</p><p><strong>Every recommendation is essentially the same dress in a different color.</strong></p><p>We observed the same behavior when testing a pair of jogger bottoms&#8212;the recommendations were dominated by the same or nearly identical Jerry Jogger products.</p><p>Technically, the recommender is doing exactly what we asked it to do. Different variants of the same product have almost identical metadata and descriptions, so cosine similarity naturally ranks them at the top.</p><p>From a recommendation-experience perspective, however, we need to distinguish between two different use cases:</p><p><strong>Other Colors / Variants</strong></p><blockquote><p>Show me another version of the product I&#8217;m already viewing.</p></blockquote><p>versus</p><p><strong>You May Also Like</strong></p><blockquote><p>Show me different products that are similar to the one I&#8217;m viewing.</p></blockquote><p>For this project, we&#8217;re trying to build the second experience.</p><p>This is an important reminder that a mathematically correct recommendation isn&#8217;t necessarily a <strong>useful recommendation</strong>.</p><p>So before evaluating the system, we need to make one more improvement to our recommendation logic.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>5. Improving Recommendation Diversity</h2><p>Our first recommender was good at finding similar products&#8212;but <strong>too good at finding near-duplicates</strong>.</p><p>To make the results more useful for product discovery, we need to prevent multiple variants of essentially the same product from occupying our recommendation slots.</p><p>Fortunately, the H&amp;M dataset gives us a useful field: <code>product_code</code>.</p><p>While <code>article_id</code> identifies a specific article or variant, <code>product_code</code> helps us identify products belonging to the same underlying product family.</p><p>We can use this information as a <strong>post-ranking rule</strong>. After ranking products by cosine similarity, we:</p><ol><li><p>Exclude the product family of the selected item.</p></li><li><p>Keep only one recommendation from each product family.</p></li><li><p>Avoid repeating the same product name.</p></li></ol><p>The core filtering logic looks like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8b991f2c-0a44-4f51-98e2-0d3431477655&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">seen_codes = {selected_product_code}
seen_names = {selected_name.lower()}

for i in similar_indices:
    code = products.loc[i, "product_code"]
    name = products.loc[i, "prod_name"].lower()

    if code not in seen_codes and name not in seen_names:
        filtered_indices.append(i)
        seen_codes.add(code)
        seen_names.add(name)
</code></pre></div><p>This doesn&#8217;t change the underlying similarity model. Instead, it changes <strong>how we use its rankings</strong> to create the final recommendation experience.</p><h3>Did It Improve the Recommendations?</h3><p>Let&#8217;s return to the black <strong>Alcazar Strap Dress</strong>.</p><p>After applying the diversity rules, the recommendations changed from different colors of the same dress to products such as:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!wEnK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 424w, /__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 848w, /__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!wEnK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png" width="1252" height="434" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:434,&quot;width&quot;:1252,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53244,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/212648125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 424w, /__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 848w, /__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wEnK!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3afd8a-9284-4d59-8a53-10045def6710_1252x434.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The similarity scores are lower than before&#8212;and that&#8217;s perfectly reasonable.</p><p>Previously, scores above 0.90 came from products that were nearly identical. Now the system is finding <strong>different products that still share important characteristics</strong> with the selected item.</p><p>We observed the same improvement with the jogger example. Instead of filling the results with Jerry Jogger variants, the recommender began surfacing related products such as shorts and sweatpants.</p><p>This highlights an important aspect of building recommendation systems:</p><blockquote><p><strong>The highest similarity score does not always produce the best recommendation experience.</strong></p></blockquote><p>The model provides a ranking, but business rules and post-ranking logic can determine whether those results are actually useful to the customer.</p><p>Now that we&#8217;re satisfied with the recommendation behavior, the next question is harder:</p><p><strong>How do we know whether the recommender is actually good?</strong></p><div><hr></div><h2>6. Evaluating the Recommender</h2><p>Generating recommendations is relatively easy. Evaluating whether those recommendations are actually good is much harder.</p><p>So far, we have evaluated our system qualitatively by looking at the recommendations and asking whether they make sense. To add a quantitative perspective, we&#8217;ll use the historical purchases available in <code>transactions_train.csv</code>.</p><p>The basic idea is:</p><p><strong>Past Purchase &#8594; Generate Top-10 Recommendations &#8594; Compare Against Future Purchases</strong></p><p>To avoid using future information, we create a <strong>time-based holdout</strong> rather than randomly splitting transactions.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;39ea8a10-fdfa-49dc-8f06-adf3f22c17e8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cutoff_date = pd.Timestamp("2020-09-01")

history = transactions[transactions["t_dat"] &lt; cutoff_date]
future = transactions[transactions["t_dat"] &gt;= cutoff_date]
</code></pre></div><p>For each customer, we use their most recent purchase before the cutoff as the <strong>seed product</strong>. Products purchased after the cutoff become our relevant items.</p><p>Because the transaction dataset is large, we evaluate the recommender on a sample of <strong>500 customers</strong> who made purchases both before and after the cutoff.</p><h3>Choosing the Metrics</h3><p>We&#8217;ll look at four common recommendation metrics:</p><ul><li><p><strong>Precision@10</strong> &#8212; How many of the Top-10 recommendations were later purchased?</p></li><li><p><strong>Recall@10</strong> &#8212; How many of the customer&#8217;s future purchases appeared in the Top 10?</p></li><li><p><strong>NDCG@10</strong> &#8212; Did relevant products appear near the top of the ranking?</p></li><li><p><strong>Hit Rate@10</strong> &#8212; For how many customers did at least one recommendation match a future purchase?</p></li></ul><p>Our first evaluation produced:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vmQ8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 424w, /__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 848w, /__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vmQ8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png" width="1262" height="420" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:420,&quot;width&quot;:1262,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:42289,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/212648125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 424w, /__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 848w, /__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vmQ8!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf48b704-acb6-4b13-ac41-877deb1350ec_1262x420.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These numbers are extremely low.</p><p>Does that mean our recommender is bad?</p><p><strong>Not necessarily.</strong></p><p>Consider a customer whose most recent purchase was a black dress. Our content-based recommender might return ten highly relevant dresses. But if the customer&#8217;s next purchase is a handbag or a pair of trousers, every recommendation receives zero credit.</p><p>The problem is that our recommender answers:</p><blockquote><p><strong>&#8220;What products are similar to this product?&#8221;</strong></p></blockquote><p>Our evaluation, however, is asking:</p><blockquote><p><strong>&#8220;Can this one product predict anything the customer will purchase next?&#8221;</strong></p></blockquote><p>Those are fundamentally different objectives.</p><p>This brings us to an important lesson in evaluating recommendation systems:</p><blockquote><p><strong>Before interpreting a metric, make sure the evaluation task matches the recommendation experience you&#8217;re trying to build.</strong></p></blockquote><p>For our <strong>&#8220;You May Also Like&#8221;</strong> recommender, we therefore need an evaluation that is better aligned with its purpose.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>7. Aligning Evaluation with the Recommendation Use Case</h2><p>Since our system is designed to recommend <strong>similar products</strong>, a more appropriate evaluation is to compare recommendations against future purchases within the <strong>same product category</strong>.</p><p>For example, if the seed product is a dress, we evaluate whether the Top-10 recommendations include another dress that the customer later purchased.</p><p>Conceptually, the evaluation becomes:</p><p><strong>Past Purchase (Dress) &#8594; Recommend Similar Products &#8594; Compare Against Future Dress Purchases</strong></p><p>We keep the same time-based holdout and the same ranking metrics. The only difference is how we define a <strong>relevant item</strong>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fa5c4838-ef11-4c64-9b06-a4e87abd5547&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">relevant = customer_future[
    customer_future["product_type_name"] == seed_type
]["article_id"].unique()
</code></pre></div><p>Of the original 500 customers, <strong>103 had at least one future purchase in the same product category</strong>, giving us a smaller but more relevant evaluation set.</p><p>The results changed considerably:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5nBi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 424w, /__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 848w, /__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5nBi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png" width="1256" height="414" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:414,&quot;width&quot;:1256,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:54282,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/212648125?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 424w, /__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 848w, /__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5nBi!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6175d6e6-abd6-416e-9363-19c61d300640_1256x414.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every metric improves when the evaluation is better aligned with the intended recommendation experience. Recall@10, for example, increases from approximately <strong>0.03% to 1.46%</strong>.</p><p>The absolute numbers are still low, and that&#8217;s worth acknowledging.</p><p>Exact-purchase evaluation is a very strict test. If we recommend a highly relevant black dress but the customer purchases a different black dress outside our Top 10, the recommendation receives no credit. We&#8217;re also selecting just 10 items from a catalog containing more than 100,000 products.</p><p>More importantly, content similarity alone cannot capture everything that influences a purchase. Price, availability, individual taste, previous behavior, trends, and many other factors can affect what a customer ultimately chooses.</p><p>The goal here isn&#8217;t to make the metrics look impressive. It&#8217;s to evaluate the recommender in a way that reflects what it was actually designed to do.</p><blockquote><p><strong>A recommendation metric is only meaningful when the definition of relevance matches the experience you&#8217;re trying to build.</strong></p></blockquote><p>This is why offline recommendation metrics should always be interpreted alongside the business objective and the type of recommendation system being evaluated.</p><div><hr></div><h2>8. Limitations and Key Takeaways</h2><p>We now have an end-to-end content-based recommendation system: starting with raw product metadata, creating product representations, generating and refining recommendations, and finally evaluating them against actual customer behavior.</p><p>But like any recommendation approach, content-based filtering has limitations.</p><h3>Where Does Content-Based Filtering Fall Short?</h3><p><strong>It is only as good as the product metadata.</strong><br>Our model understands products through attributes such as type, color, department, garment group, and description. If an important characteristic isn&#8217;t captured in the data, the recommender can&#8217;t use it.</p><p><strong>Similarity doesn&#8217;t necessarily mean preference.</strong><br>Knowing that two dresses are similar doesn&#8217;t tell us whether a particular customer will like both. Our model understands products, but it doesn&#8217;t really understand the customer.</p><p><strong>It can become repetitive.</strong><br>We saw this firsthand when our initial recommendations were dominated by different variants of the same product. Post-ranking rules helped, but this highlights the tendency of content-based systems to stay close to what they already know.</p><p><strong>Offline metrics don&#8217;t tell the whole story.</strong><br>A recommendation can be useful even if the customer doesn&#8217;t eventually purchase that exact product. Metrics such as Precision, Recall, and NDCG are valuable, but they need to be interpreted in the context of the recommendation experience.</p><h3>Key Takeaways</h3><p>There are a few lessons from this project that extend beyond content-based filtering:</p><ul><li><p><strong>Item representation matters.</strong> Better product features lead to better similarity comparisons.</p></li><li><p><strong>The highest similarity score isn&#8217;t always the best recommendation.</strong> Ranking often needs additional business logic.</p></li><li><p><strong>Inspect your recommendations.</strong> Our first model worked mathematically, but qualitative inspection exposed a clear product-experience problem.</p></li><li><p><strong>Evaluation should match the objective.</strong> Changing how we defined relevance substantially changed how we interpreted the same recommender.</p></li><li><p><strong>A model is only one part of a recommendation system.</strong> Feature engineering, catalog structure, ranking rules, evaluation, and business context all contribute to the final experience.</p></li></ul><p>Perhaps the biggest takeaway is that building a recommender doesn&#8217;t end when <code>cosine_similarity()</code> returns a score.</p><p><strong>The real work begins when we ask whether those recommendations are actually useful.</strong></p><div><hr></div><h2>Final Thoughts</h2><p>In my previous article, we looked at how content-based recommendation systems work. In this one, we took the next step and built one end-to-end using a real-world product catalog.</p><p>What I enjoyed most about this project was that the interesting part wasn&#8217;t getting cosine similarity to work&#8212;it was everything that came after. We had to inspect the recommendations, deal with repetitive product variants, think about what a useful recommendation actually means, and make sure our evaluation matched the experience we were trying to build.</p><p>If you&#8217;d like to explore the complete implementation or experiment with it yourself, I&#8217;ve included the full notebook and code in the <strong><a href="https://github.com/gowthami-peri/content-based-recommendation-system/tree/main">GitHub repository</a></strong>.</p><p>I&#8217;ll continue sharing practical applications of data science and breaking down how these systems work beyond just the algorithms.</p><p><strong>Until next time, keep learning, keep building, and keep solving real problems with data.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Content-Based Filtering for Data Scientists]]></title><description><![CDATA[Similarity, User Profiles, and Personalized Recommendations]]></description><link>https://thepracticaldatascientist.substack.com/p/content-based-filtering-for-data</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/content-based-filtering-for-data</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 18 Aug 2026 11:55:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SGSt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!SGSt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 424w, /__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 848w, /__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!SGSt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png" width="1456" height="975" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:975,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2426629,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/211490646?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 424w, /__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 848w, /__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SGSt!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5106a765-ff1b-4d77-8054-a2123d7ebe48_1526x1022.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>You watch a movie on a streaming platform and enjoy it. When you come back the next day, you see several recommendations that feel similar&#8212;maybe they belong to the same genre, have similar themes, or feature some of the same actors.</p><p>How did the system decide what to show you?</p><p>One way is to look at the movies you&#8217;ve already enjoyed, understand what they have in common, and find other movies with similar characteristics.</p><p>This is the basic idea behind <strong>content-based filtering</strong>.</p><p>Unlike approaches that rely on what other users are watching or buying, content-based filtering focuses primarily on <strong>you and the items you&#8217;ve interacted with</strong>. If you frequently read articles about machine learning, for example, the system might recommend other articles covering similar topics.</p><p>The idea sounds simple, but it raises a few interesting questions. How does a machine understand that two movies or products are similar? Which characteristics should it compare? And how do we turn those similarities into a ranked list of recommendations?</p><p>That&#8217;s what we&#8217;ll explore in this article.</p><p>We&#8217;ll look at how items are represented using features, how similarity is measured, and how those pieces come together to generate personalized recommendations. We&#8217;ll also discuss where content-based filtering works well and where it starts to struggle.</p><p>By the end, you&#8217;ll understand the intuition behind one of the simplest and most useful approaches to building recommendation systems.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Is Content-Based Filtering?</h2><p>Content-based filtering recommends items based on the characteristics of things a user has liked in the past. The idea is fairly intuitive. If we know what you liked before, we can look for other items that share similar characteristics.</p><p>Suppose you watched and enjoyed several science-fiction movies. Those movies might share features such as <strong>science fiction, space, adventure, futuristic themes, or certain actors and directors</strong>.</p><p>A content-based recommendation system uses this information to build an understanding of your preferences. It can then search the catalog for other movies with similar features and rank them based on how closely they match your interests.</p><p>The same idea works outside of movies. An e-commerce platform might use product category, brand, price range, or product description. A news platform could use topics, keywords, or article content.</p><p>The important point is that the system is learning from the <strong>content of the items themselves</strong>.</p><h4>A Simple Example</h4><p>Imagine you&#8217;ve liked these three movies:</p><ul><li><p>Movie A: <strong>Sci-Fi, Space, Adventure</strong></p></li><li><p>Movie B: <strong>Sci-Fi, Space, Thriller</strong></p></li><li><p>Movie C: <strong>Sci-Fi, Adventure, Futuristic</strong></p></li></ul><p>From this history, the system might learn that you have a strong preference for <strong>science fiction and space-related content</strong>.</p><p>Now suppose there are two movies you haven&#8217;t watched:</p><ul><li><p>Movie D: <strong>Sci-Fi, Space, Futuristic</strong></p></li><li><p>Movie E: <strong>Romance, Comedy, Drama</strong></p></li></ul><p>Movie D is much more similar to the movies you&#8217;ve enjoyed before, so it would likely receive a higher recommendation score.</p><p>At a high level, the process looks like this:</p><p><strong>Items you liked &#8594; Identify their features &#8594; Find similar items &#8594; Recommend</strong></p><p>This also highlights what makes content-based filtering different from collaborative filtering. We don&#8217;t need to know whether thousands of other users liked Movie D. The recommendation can be made based on the relationship between <strong>your preferences and the characteristics of the movie</strong>.</p><p>Of course, a machine needs a mathematical way to represent those characteristics before it can decide that two items are similar.</p><p>That&#8217;s where <strong>item features and vectors</strong> come in.</p><div><hr></div><h2>How Do We Represent Items?</h2><p>For content-based filtering to work, the system needs a way to describe each item.</p><p>As humans, we can look at two movies and recognize that they&#8217;re similar. A machine needs that information represented in a form it can compare.</p><p>This is where <strong>item features</strong> come in.</p><p>For a movie, useful features might include genre, actors, director, language, or keywords from the description. For a product, they could include category, brand, price range, color, or information from the product description.</p><p>The choice of features matters. If the features don&#8217;t capture what makes two items meaningfully similar, the recommendations won&#8217;t be very useful.</p><h4>Turning Features Into Vectors</h4><p>Once we&#8217;ve identified the features, we need to convert them into numbers.</p><p>Suppose we describe movies using four genres:</p><p><strong>Sci-Fi | Adventure | Romance | Comedy</strong></p><p>A science-fiction adventure movie could be represented as:</p><p><strong>[1, 1, 0, 0]</strong></p><p>A romantic comedy might look like:</p><p><strong>[0, 0, 1, 1]</strong></p><p>These numerical representations are called <strong>item vectors</strong>. Once every movie is represented as a vector, the system has a consistent way to compare them.</p><p>Real recommendation systems can have hundreds or thousands of features, but the basic idea remains the same.</p><h4>What About Text?</h4><p>Sometimes the most useful information isn&#8217;t stored in neat categories. It might be hidden inside a movie synopsis, product description, or article.</p><p>One common approach is <strong>TF-IDF</strong>, which converts text into numerical features based on the words that are important within each document. Items that contain similar important words will end up with similar representations.</p><p>More modern systems often use <strong>embeddings</strong> instead. Embeddings can capture meaning beyond exact word matches, allowing two items to be considered similar even when they use different words to describe the same idea.</p><p>For now, the important idea isn&#8217;t which technique we use. It&#8217;s that we need to turn each item into a numerical representation that captures something meaningful about it.</p><p>Once we have those vectors, the next question becomes:</p><blockquote><p><strong>How do we measure how similar two vectors are?</strong></p></blockquote><p>That&#8217;s where <strong>cosine similarity</strong> comes in.</p><div><hr></div><h2>How Do We Measure Similarity?</h2><p>Once every item has been represented as a vector, we need a way to compare those vectors.</p><p>In other words, we need to answer a simple question: <strong>How similar are these two items?</strong></p><p>One of the most common ways to do this is <strong>cosine similarity</strong>.</p><h4>The Intuition Behind Cosine Similarity</h4><p>Cosine similarity looks at the <strong>angle between two vectors</strong> rather than simply measuring the distance between them.</p><p>If two vectors point in roughly the same direction, the items they represent are considered similar. If they point in very different directions, the items are less similar.</p><p>The score usually ranges from <strong>0 to 1</strong> when the features are non-negative. A score close to 1 means the items are very similar, while a score closer to 0 means they have little in common.</p><h4>A Simple Example</h4><p>Let&#8217;s go back to our movie example.</p><p>Suppose a user enjoyed a movie represented by:</p><p><strong>Movie A:</strong> [1, 1, 0, 0]<br><em>Sci-Fi, Adventure</em></p><p>Now we have two possible recommendations:</p><p><strong>Movie B:</strong> [1, 1, 1, 0]<br><em>Sci-Fi, Adventure, Romance</em></p><p><strong>Movie C:</strong> [0, 0, 1, 1]<br><em>Romance, Comedy</em></p><p>Movie B points in a similar direction to Movie A because they share two important features. Movie C shares none of those features, so its cosine similarity with Movie A would be much lower.</p><p>The system can calculate this similarity across many items and use the scores to identify the closest matches.</p><h4>Why Cosine Similarity?</h4><p>Another option would be to calculate the straight-line distance between vectors. That can work, but cosine similarity is often useful when we care more about the <strong>pattern of features</strong> than their absolute magnitude.</p><p>This is especially helpful with text representations such as TF-IDF, where one document may contain many more words than another. Cosine similarity focuses on whether the two documents emphasize similar features rather than simply comparing their size.</p><p>Cosine similarity isn&#8217;t the only way to measure similarity. Depending on the problem, we could also use Euclidean distance, Jaccard similarity, or other measures. But for content-based recommendation systems, cosine similarity is a common and intuitive place to start.</p><p>Now we have the two pieces we need: <strong>a way to represent items and a way to measure how similar they are.</strong></p><p>The next step is putting them together to generate recommendations.</p><div><hr></div><h2>How Content-Based Recommendations Are Generated</h2><p>So far, we&#8217;ve seen how to represent items as vectors and measure the similarity between them. Now we can put those pieces together to generate recommendations for a user.</p><p>Let&#8217;s continue with our movie example.</p><p>Suppose we&#8217;re representing movies using these features:</p><p><strong>[Sci-Fi, Adventure, Romance, Comedy, Thriller, Space, Drama]</strong></p><h4>Step 1: Look at the User&#8217;s History</h4><p>We start with movies the user has already interacted with. These could be movies they watched, liked, or rated highly.</p><p>Suppose the user liked these two movies:</p><p><strong>Movie A:</strong> [1, 1, 0, 0, 0, 1, 0]<br><em>Sci-Fi, Adventure, Space</em></p><p><strong>Movie B:</strong> [1, 0, 0, 0, 1, 1, 0]<br><em>Sci-Fi, Thriller, Space</em></p><p>Looking at these movies, we can already see some patterns. Both contain <strong>Sci-Fi and Space</strong>, while Adventure and Thriller appear once each.</p><h4>Step 2: Build a User Profile</h4><p>Next, we combine the vectors of the movies the user liked to create a representation of their preferences.</p><p>One simple approach is to average the item vectors.</p><p>From Movie A and Movie B, we get:</p><p><strong>User Profile:</strong> [1, 0.5, 0, 0, 0.5, 1, 0]</p><p>This tells us that <strong>Sci-Fi and Space</strong> are the strongest signals in the user&#8217;s history, while Adventure and Thriller also contribute to their profile.</p><p>In a real system, the calculation could be more sophisticated. A movie the user rated five stars, for example, could receive more weight than one they simply watched.</p><h4>Step 3: Compare the Profile With New Items</h4><p>Now suppose we have three unseen movies:</p><p><strong>Movie C:</strong> [1, 1, 0, 0, 0, 1, 0]<br><em>Sci-Fi, Adventure, Space</em></p><p><strong>Movie D:</strong> [1, 0, 0, 0, 1, 0, 0]<br><em>Sci-Fi, Thriller</em></p><p><strong>Movie E:</strong> [0, 0, 1, 1, 0, 0, 1]<br><em>Romance, Comedy, Drama</em></p><p>We can compare each movie vector with the user&#8217;s profile using <strong>cosine similarity</strong>.</p><p>Movie C matches several of the user&#8217;s strongest preferences, so we would expect a high similarity score. Movie D also has some overlap, while Movie E looks very different from the user&#8217;s history.</p><h4>Step 4: Rank the Items</h4><p>Suppose the similarity scores are:</p><p><strong>Movie C &#8594; 0.94</strong><br><strong>Movie D &#8594; 0.67</strong><br><strong>Movie E &#8594; 0.00</strong></p><p>The system can now rank the movies:</p><p><strong>1. Movie C</strong><br><strong>2. Movie D</strong><br><strong>3. Movie E</strong></p><p>Movie C would be the strongest recommendation because its feature vector most closely matches the user&#8217;s preference vector.</p><p>The entire process can be summarized as:</p><p><strong>Item vectors &#8594; User profile vector &#8594; Calculate similarity &#8594; Rank items &#8594; Recommend</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!niar!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 424w, /__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 848w, /__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 1272w, /__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!niar!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png" width="1456" height="501" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:501,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1676321,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/211490646?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 424w, /__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 848w, /__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 1272w, /__u/substackcdn.com/image/fetch/$s_!niar!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f1d9a4-dcd2-411f-af64-fdfd53e4ed65_1818x626.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is the core idea behind content-based filtering. Instead of relying on what other users liked, the system builds a representation of <strong>your preferences</strong> and looks for items that are mathematically similar to that profile.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Advantages and Limitations of Content-Based Filtering</h2><p>Content-based filtering is relatively simple to understand and can produce highly personalized recommendations. But like any recommendation approach, it comes with trade-offs.</p><p>Understanding those trade-offs helps us decide when content-based filtering is a good fit and when another approach might work better.</p><h4>Why Content-Based Filtering Works Well</h4><p>One major advantage is that the system doesn&#8217;t need information about other users. Recommendations are based on an individual&#8217;s preferences and the characteristics of the available items.</p><p>This also makes the recommendations relatively easy to explain. If a user likes science-fiction movies and receives another movie with similar themes, genres, or actors, there is a clear reason behind the recommendation.</p><p>Content-based filtering can also work well for <strong>new items</strong>. A newly added movie may have no views or ratings, but if we know its genre, actors, description, and other features, the system can still determine which users might find it relevant.</p><h4>The Cold-Start Problem</h4><p>A <strong>cold-start problem</strong> occurs when the system doesn&#8217;t have enough information to make useful recommendations. This can happen with either a new user or a new item.</p><p><strong>New-user cold start</strong> is a challenge for content-based filtering. If someone has just joined the platform, they haven&#8217;t watched, purchased, liked, or rated anything yet. Without that history, there isn&#8217;t enough information to build a meaningful user profile.</p><p>Platforms can address this in several ways. A streaming service might ask new users to choose a few movies or genres they like during onboarding. Until enough interactions are collected, the system might also rely on popular or trending items.</p><p><strong>New-item cold start is different.</strong> Content-based filtering can often handle new items better because it doesn&#8217;t need other users to interact with them first.</p><p>Suppose a new movie has this feature vector:</p><p><strong>New Movie:</strong> [1, 1, 0, 0, 0, 1, 0]<br><em>Sci-Fi, Adventure, Space</em></p><p>Even if nobody has watched it yet, we can compare it with an existing user profile:</p><p><strong>User Profile:</strong> [1, 0.5, 0, 0, 0.5, 1, 0]</p><p>Because the vectors share several features, the movie could still become a strong recommendation for that user.</p><p>This is an important distinction: <strong>content-based filtering struggles with new users, but can often handle new items as long as useful item features are available.</strong></p><h4>Overspecialization</h4><p>Another limitation is that content-based filtering tends to recommend <strong>more of what the user already likes</strong>.</p><p>If someone frequently watches science-fiction movies, the system may continue recommending similar science-fiction titles. Those recommendations may be relevant, but they don&#8217;t necessarily help the user discover something different.</p><p>This is known as <strong>overspecialization</strong>. The system can become so focused on existing preferences that it limits novelty and discovery.</p><p>A user who loves science fiction might also enjoy a historical drama, but a purely content-based system may never discover that connection if the items don&#8217;t share enough features.</p><h4>Dependence on Item Features</h4><p>The quality of the recommendations also depends heavily on how well the items are represented.</p><p>If important characteristics are missing, the system may conclude that two genuinely similar items are unrelated. Adding more features doesn&#8217;t automatically solve the problem either; those features need to capture what actually matters to users.</p><p>This can be particularly challenging for products where preferences are difficult to describe using structured attributes alone. Two books, songs, or movies may feel similar to a person even when their obvious metadata looks quite different.</p><h4>The Trade-Off</h4><p>Content-based filtering works particularly well when items have rich, meaningful features and users have enough history to reveal their preferences. It can personalize recommendations without needing behavior from a large community of users and can surface new items before they accumulate interactions.</p><p>Its biggest weakness is also closely connected to its strength: <strong>it learns from what the user already likes.</strong> That makes recommendations personalized, but it can also make them predictable.</p><p>This is where <strong>collaborative filtering</strong> takes a very different approach. Instead of asking <em>&#8220;What items are similar to the things you already like?&#8221;</em>, it asks what we can learn from the preferences of other users.</p><div><hr></div><h2>Key Takeaways</h2><p>Content-based filtering is one of the most intuitive ways to build a recommendation system. It learns what a user prefers from the items they&#8217;ve interacted with and then searches for other items with similar characteristics.</p><p>Here are the main ideas to remember:</p><ul><li><p><strong>Items need to be represented using meaningful features.</strong> These could be categories, genres, product attributes, TF-IDF features, or embeddings.</p></li><li><p><strong>Item features are converted into vectors</strong>, giving the system a numerical representation that it can compare.</p></li><li><p><strong>A user profile can be built from previously liked items.</strong> A simple approach is to average their item vectors, although interactions can also be weighted based on ratings or other signals.</p></li><li><p><strong>Cosine similarity is commonly used to compare vectors.</strong> Items that are more similar to the user&#8217;s preference vector receive higher recommendation scores.</p></li><li><p><strong>Content-based filtering doesn&#8217;t depend on other users&#8217; behavior.</strong> This makes it possible to recommend new items as long as useful item features are available.</p></li><li><p><strong>Cold start affects users and items differently.</strong> New users are difficult because there isn&#8217;t enough history to understand their preferences. New items are easier to handle because their features can still be compared with existing user profiles.</p></li><li><p><strong>Overspecialization is an important limitation.</strong> Recommending only items similar to what a user already likes can reduce diversity and discovery.</p></li></ul><p>The central idea is simple:</p><blockquote><p><strong>Understand what the user likes, represent those preferences as features, and find other items that look similar.</strong></p></blockquote><p>That simplicity is what makes content-based filtering a useful starting point for understanding recommendation systems.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>So far, we&#8217;ve focused on understanding how content-based filtering works&#8212;from representing items as vectors to building user profiles and using similarity to find relevant recommendations.</p><p>But understanding the idea is only the first step. The best way to make it stick is to actually build one.</p><p>In the next article, we&#8217;ll create a <strong>content-based recommendation system end to end using Python</strong>.</p><p>We&#8217;ll start with a real dataset, prepare the item features, convert them into numerical representations, calculate similarity, and generate top recommendations for a user. Along the way, we&#8217;ll connect each piece of code back to the concepts we covered in this article.</p><p>The goal isn&#8217;t just to get a recommender running.</p><p>It&#8217;s to understand exactly <strong>how raw data turns into a personalized list of recommendations.</strong></p><p>After that, we&#8217;ll move on to <strong>collaborative filtering</strong> and see how recommendations change when we start learning from the behavior of other users.</p><p><em><strong>Until then, keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h3></h3><p></p>]]></content:encoded></item><item><title><![CDATA[Introduction to Recommendation Systems]]></title><description><![CDATA[How Netflix, Spotify, and Amazon Personalize What You See]]></description><link>https://thepracticaldatascientist.substack.com/p/introduction-to-recommendation-systems</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/introduction-to-recommendation-systems</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 11 Aug 2026 13:02:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yR2K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!yR2K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 424w, /__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 848w, /__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!yR2K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png" width="1456" height="966" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:966,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2426039,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/210539151?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 424w, /__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 848w, /__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yR2K!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a79ea3d-af42-4f85-9a81-ad3bb76a255a_1522x1010.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>Open Netflix, Spotify, YouTube, or your favorite shopping app, and you&#8217;ll notice something interesting: <strong>your experience is uniquely yours.</strong></p><p>The movies on your homepage aren&#8217;t the same ones your friend sees. Spotify introduces you to artists you may never have searched for. An online retailer suggests a product you didn&#8217;t know you needed&#8212;yet somehow, it feels relevant.</p><p>None of this happens by accident.</p><p>Behind these experiences are <strong>recommendation systems</strong>, designed to solve a simple but challenging problem:</p><blockquote><p><strong>Out of thousands&#8212;or even millions&#8212;of options, what should we show this user next?</strong></p></blockquote><p>Think about an online retailer with millions of products. Showing every customer the same bestsellers would be easy, but not particularly useful. Someone shopping for running gear has very different interests from someone furnishing a new apartment.</p><p>The challenge is figuring out those preferences from the signals users leave behind.</p><p>What did they view? What did they buy? What did they skip? Which products did they return to? What do people with similar interests seem to like?</p><p>Those signals help recommendation systems narrow an enormous catalog down to a small set of options that might actually matter to an individual user.</p><p>But getting recommendations right involves more than predicting what someone is likely to click.</p><p>What do you recommend to a brand-new user with no history? How do new products get discovered when nobody has interacted with them yet? Should you keep showing users more of what they already like, or occasionally introduce something different?</p><p>And perhaps most importantly:</p><blockquote><p><strong>How do you know whether your recommendations are actually working?</strong></p></blockquote><p>These questions are what make recommendation systems such an interesting machine learning problem. They sit at the intersection of algorithms, user behavior, experimentation, and business outcomes.</p><p>In this article, we&#8217;ll build the foundation: what recommendation systems are, the major approaches behind them, how recommendations are generated and ranked, and how we evaluate whether they&#8217;re delivering value.</p><p>Because ultimately, a good recommendation system isn&#8217;t about showing users more options.</p><blockquote><p><strong>It&#8217;s about finding the right item for the right user at the right time.</strong></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Is a Recommendation System?</h2><p>At its simplest, a recommendation system helps users find relevant items from a large set of choices.</p><p>Those &#8220;items&#8221; could be almost anything: movies, songs, products, articles, restaurants, jobs, or even people to follow. Most recommendation problems involve three basic pieces:</p><p><strong>Users</strong> &#8212; the people receiving recommendations<br><strong>Items</strong> &#8212; the things that could be recommended<br><strong>Interactions</strong> &#8212; the signals that connect users and items</p><p>Interactions can be <strong>explicit</strong>, such as giving a movie five stars or liking a song. But often, they&#8217;re implicit. A user clicks on a product, watches a video until the end, skips a song after ten seconds, adds something to a cart, or makes a purchase.</p><p>Each action tells us something about preference&#8212;although not always as clearly as we might think. For example, purchasing a product usually indicates interest. But what about viewing the same product five times without buying it? Or watching an entire movie and never rating it?</p><p>Recommendation systems have to make sense of these imperfect signals.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oEG-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 424w, /__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 848w, /__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oEG-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png" width="1456" height="954" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:954,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1448414,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/210539151?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 424w, /__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 848w, /__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oEG-!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec36ffc-2edc-4706-9466-3fc1b837c7b4_1498x982.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>From Prediction to Ranking</h3><p>There&#8217;s another important distinction between recommendation systems and many traditional machine learning problems.</p><p>Suppose a streaming platform has 50,000 movies available. The goal isn&#8217;t simply to predict whether you would like one particular movie.</p><p>The more useful question is:</p><blockquote><p><strong>Which movies should appear at the top of your recommendations?</strong></p></blockquote><p>That makes recommendation systems fundamentally a <strong>ranking problem</strong> in many real-world applications. The system might estimate your preference for thousands of items, narrow those possibilities down, and then rank the strongest candidates so that the most relevant ones appear first.</p><p>And relevance isn&#8217;t necessarily the same as popularity.</p><p>The most popular movie on a platform may not interest you at all. Meanwhile, a lesser-known movie that matches your viewing history could be a much better recommendation.</p><p>That&#8217;s what personalization is trying to achieve: moving from</p><blockquote><p><strong>&#8220;What is popular?&#8221;</strong></p></blockquote><p>to</p><blockquote><p><strong>&#8220;What is relevant to this user?&#8221;</strong></p></blockquote><p>Of course, we still need a way to determine that relevance. Some recommendation systems look at the characteristics of the items you&#8217;ve liked before. Others learn from the behavior of users who seem similar to you. Many real-world systems combine several approaches.</p><p>We&#8217;ll look at those approaches next.</p><div><hr></div><h2>The Three Main Approaches to Recommendation Systems</h2><p>There isn&#8217;t one universal way to build a recommendation system.</p><p>The right approach depends on the data you have, the type of recommendations you&#8217;re making, and the problem you&#8217;re trying to solve.</p><p>Most recommendation systems are built around three broad approaches: <strong>content-based filtering, collaborative filtering, and hybrid systems.</strong></p><h3>1. Content-Based Filtering</h3><p>Content-based filtering starts with the items a user has already shown interest in.</p><p>It looks at the characteristics of those items and searches for others that are similar.</p><p>Imagine you&#8217;ve watched several science-fiction movies. A content-based system might look at features such as genre, director, actors, or themes and recommend other movies with similar characteristics.</p><p>The basic idea is:</p><blockquote><p><strong>You liked this, so you might like something similar.</strong></p></blockquote><p>This approach can work well even when users have very different tastes because recommendations are based on each person&#8217;s individual history.</p><p>However, it can also become repetitive. If the system only recommends things similar to what you&#8217;ve already consumed, discovering something genuinely different becomes harder.</p><h3>2. Collaborative Filtering</h3><p>Collaborative filtering takes a different approach.</p><p>Instead of asking <em>&#8220;What items are similar?&#8221;</em>, it looks for patterns across users and their interactions.</p><p>Suppose you and another user have watched and enjoyed many of the same movies. If that person loved a movie you haven&#8217;t seen yet, it could be a good recommendation for you.</p><p>The intuition is:</p><blockquote><p><strong>People with similar preferences may enjoy similar things.</strong></p></blockquote><p>One advantage is that the system doesn&#8217;t necessarily need detailed information about the items themselves. Patterns in user behavior can be enough to uncover useful recommendations.</p><p>The challenge is that collaborative filtering needs interaction data. New users and new items have very little history, creating what is commonly known as the <strong>cold-start problem</strong>.</p><h3>3. Hybrid Recommendation Systems</h3><p>In practice, recommendation systems don&#8217;t always fit neatly into one category.</p><p>A hybrid system combines multiple approaches to take advantage of their strengths while reducing some of their weaknesses.</p><p>For example, a system might use collaborative filtering to learn from behavior across users while also incorporating information about the products themselves. It may even bring in additional signals such as popularity, recency, location, context, or business rules.</p><p>This is why real-world recommendation systems are often much more than a single machine learning algorithm.</p><p>The important thing at this stage isn&#8217;t to memorize every technique. It&#8217;s to understand the different sources of information each approach relies on:</p><p><strong>Content-based filtering</strong> learns from the items you like.</p><p><strong>Collaborative filtering</strong> learns from patterns across users.</p><p><strong>Hybrid systems</strong> combine multiple signals to produce better recommendations.</p><p>In the next section, we&#8217;ll zoom out from individual algorithms and look at how these pieces come together in an end-to-end recommendation system.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>How Does a Recommendation System Work?</h2><p>A recommendation system doesn&#8217;t usually compare every user with every possible item and simply pick the best matches. For a platform with millions of users and products, that would quickly become too slow and computationally expensive.</p><p>Instead, most real-world recommendation systems break the process into stages. Each stage narrows the possibilities until we&#8217;re left with a small set of relevant recommendations for the user.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nCTp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 424w, /__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 848w, /__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nCTp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png" width="1456" height="709" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:709,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2106529,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/210539151?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 424w, /__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 848w, /__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nCTp!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F064dd9d3-a82f-4903-80d9-1f4ef6a7a627_2016x982.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Step 1: Learn From User Interactions</h3><p>It starts with user behavior. Every time someone views a product, watches a video, makes a purchase, rates an item, or skips a song, they leave behind a signal about their preferences.</p><p>Not all signals mean the same thing. Purchasing a product usually indicates stronger interest than viewing it once, while repeatedly skipping certain content may suggest the opposite. Over time, these interactions help the system build a better understanding of each user&#8217;s interests.</p><h3>Step 2: Generate Candidates</h3><p>Imagine an online retailer with millions of products. Evaluating every product for every customer would be inefficient, so the system first creates a much smaller set of potentially relevant items. This is known as <strong>candidate generation</strong>.</p><p>Candidates can come from several sources. The system might retrieve products similar to previous purchases, items popular among users with similar behavior, or products that are currently trending.</p><p>At this stage, the goal isn&#8217;t to identify the perfect recommendation. It&#8217;s simply to reduce millions of possibilities to a manageable set of promising candidates.</p><h3>Step 3: Rank the Candidates</h3><p>Once the candidate set is smaller, the system can evaluate each item more carefully. A ranking model estimates how relevant each candidate is likely to be for that particular user.</p><p>The model might consider the user&#8217;s past behavior, item characteristics, popularity, recency, price, and other contextual information. Items receiving stronger scores are then placed higher in the recommendation list.</p><p>This is why ranking matters so much. A highly relevant product appearing near the top of the page is much more likely to be discovered than the same product buried much further down.</p><h3>Step 4: Apply Business Rules</h3><p>The highest-ranked items aren&#8217;t always the ones that ultimately reach the user. Real-world systems also need to account for practical constraints and business objectives.</p><p>For example, a retailer may remove products that are out of stock. A streaming platform may avoid showing several nearly identical recommendations, while another system may intentionally introduce newer items to encourage discovery.</p><p>These rules help connect the machine learning model with the actual experience the business wants to create.</p><h3>Step 5: Learn From What Happens Next</h3><p>Once recommendations are shown, users respond to them. They might click, purchase, watch, save, skip, or simply ignore what they see.</p><p>Those responses become new information that the system can learn from. Over time, this creates a continuous feedback loop: the system makes recommendations, observes what users do, and uses those interactions to improve future recommendations.</p><p>This is an important distinction. A machine learning model may produce a relevance score, but the <strong>recommendation system</strong> determines what the user ultimately sees.</p><div><hr></div><h2>How Do We Know if Recommendations Are Good?</h2><p>Building a recommendation system is only half the problem. Once we have one, we still need to answer a more important question: <strong>Are the recommendations actually useful?</strong></p><p>At first, this might seem easy to measure. If users click on more recommendations, the system must be getting better. In reality, recommendation quality is more nuanced than any single metric can capture.</p><h3>Offline Evaluation</h3><p>Before exposing a new recommendation model to users, data scientists typically evaluate it using historical data. One common approach is to hide some of the items a user actually interacted with and see whether the model can successfully recommend them.</p><p>This is where metrics such as <strong>Precision@K, Recall@K, and NDCG</strong> come in. Each looks at recommendation quality from a slightly different perspective.</p><p><strong>Precision@K</strong> measures how many of the top K recommendations are relevant. If we recommend 10 products and 4 are relevant to the user, Precision@10 helps capture that quality.</p><p><strong>Recall@K</strong> looks at the problem from the other direction. Instead of asking how many recommendations were relevant, it asks how many of all the relevant items we successfully retrieved.</p><p><strong>NDCG</strong> also considers ranking. A relevant item appearing near the top of the list is more valuable than the same item appearing much further down.</p><p>These metrics are useful for comparing models quickly before deployment. But they only tell us how well the model performs against historical behavior, which may not always reflect how users will respond to recommendations in the real world.</p><h3>Online and Business Metrics</h3><p>Once a recommendation system reaches real users, evaluation becomes much more practical. We want to know whether the recommendations are actually changing user behavior in a useful way.</p><p>The metrics we care about depend on the product and business objective. A retailer might track click-through rate, conversion, or revenue, while a streaming platform may care more about watch time, engagement, or retention.</p><p>This is why A/B testing plays such an important role in recommendation systems. Instead of assuming that a model with better offline performance will create a better experience, we can test it with real users and measure the actual impact.</p><h3>When Better Offline Doesn&#8217;t Mean Better Online</h3><p>Imagine that a new recommendation model improves NDCG by 5%. Based on offline evaluation, it looks like a clear improvement.</p><p>Then you run an A/B test and discover that click-through rate has dropped.</p><p>There could be several reasons. Perhaps the model became very good at recommending items users already knew about. Maybe the recommendations became too similar to one another, reducing discovery. Or the model may have learned historical patterns that don&#8217;t translate well to current user behavior.</p><p>This is an important lesson in recommendation systems: <strong>a better model metric doesn&#8217;t automatically mean a better recommendation experience.</strong></p><p>A strong evaluation strategy therefore looks at the system from multiple angles. Offline metrics help us evaluate ranking quality, online experiments tell us how users respond, and business metrics tell us whether those changes ultimately create value.</p><p><span>&#127908; </span><strong>Interview Q: </strong>How would you evaluate a recommendation system?</p><p><em><strong>A:</strong> </em>I would start with offline ranking metrics such as Precision@K, Recall@K, and NDCG to compare candidate models. Once a promising model is identified, I would validate it through an A/B test using relevant product and business metrics such as click-through rate, conversion, engagement, or retention. I wouldn&#8217;t rely on offline metrics alone because better offline performance doesn&#8217;t always translate into better user outcomes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Real-World Challenges in Recommendation Systems</h2><p>Recommendation systems can look straightforward on paper: understand what a user likes, find relevant items, and rank them.</p><p>Real-world systems are much messier. Users change their preferences, new products constantly enter the catalog, interaction data is often sparse, and the recommendations themselves can influence what users do next.</p><p>Let&#8217;s look at some of the most common challenges.</p><h3>The Cold-Start Problem</h3><p>Imagine a new user joining a streaming platform. They haven&#8217;t watched, rated, or searched for anything yet, so the system has very little information about their preferences.</p><p>The same problem occurs with new items. A newly added movie or product has no interaction history, making it difficult for collaborative filtering approaches to know who might like it.</p><p>Platforms often address this by using information available at the start, such as item characteristics, popular content, onboarding preferences, or contextual signals. As more interactions are collected, recommendations can gradually become more personalized.</p><h3>Data Sparsity</h3><p>Most users interact with only a tiny fraction of the items available on a platform.</p><p>Someone may have watched 100 movies, for example, while the catalog contains tens of thousands. In an e-commerce setting with millions of products, the gap can be even larger.</p><p>This creates a very sparse user-item interaction matrix. Finding meaningful patterns becomes harder when most user-item combinations contain no interaction at all.</p><h3>Popularity Bias</h3><p>Recommendation systems learn from historical behavior, and popular items naturally generate more interactions. That gives the system more evidence about those items, which can lead to them being recommended even more frequently.</p><p>Over time, this can create a feedback loop where already-popular items receive most of the exposure while less popular or newer items struggle to be discovered.</p><p>A good recommendation experience therefore needs to balance <strong>relevance with discovery</strong> rather than simply reinforcing what is already popular.</p><h3>Filter Bubbles and Lack of Diversity</h3><p>If someone watches several crime documentaries, recommending more crime documentaries might initially seem like a great strategy.</p><p>But if every recommendation becomes increasingly similar, the experience can quickly feel repetitive. The system may become very good at predicting existing preferences while becoming worse at helping users discover something new.</p><p>This is why recommendation systems often consider <strong>diversity, novelty, and serendipity</strong> alongside relevance. Sometimes a slightly unexpected recommendation can create more value than another perfectly predictable one.</p><h3>Scalability</h3><p>Finally, recommendation systems need to work quickly.</p><p>A platform may have millions of users and millions of items, creating an enormous number of possible user-item combinations. Calculating detailed scores for every combination whenever someone opens an app isn&#8217;t practical.</p><p>This is one reason real-world systems often separate <strong>candidate generation from ranking</strong>. Candidate generation quickly narrows the catalog, allowing more sophisticated models to focus on a much smaller set of promising items.</p><p>These challenges highlight an important point: <strong>a recommendation system isn&#8217;t successful simply because its model predicts preferences accurately.</strong></p><p>It also needs to handle new users and items, surface enough variety, operate at scale, and ultimately create an experience that people find useful.</p><p>&#127908; <strong>Interview Q:  </strong>What are some of the biggest challenges when building recommendation systems?</p><p><em><strong>A:</strong></em> A strong answer might discuss <strong>cold start, data sparsity, popularity bias, diversity, and scalability</strong>. Rather than simply listing them, explain why each challenge matters and how it could affect the recommendations users actually see.</p><div><hr></div><h2>Key Takeaways</h2><p>Recommendation systems are ultimately about helping users navigate an overwhelming number of choices. The algorithms may vary, but the goal remains the same: surface items that are relevant to each user.</p><p>Here are the main ideas to remember:</p><ul><li><p><strong>Recommendation systems personalize what users see</strong> by learning from users, items, and the interactions between them.</p></li><li><p><strong>Content-based filtering</strong> recommends items similar to those a user has liked before, while <strong>collaborative filtering</strong> learns from patterns across many users. Real-world systems often combine multiple approaches.</p></li><li><p>At scale, recommendations are commonly generated in stages. <strong>Candidate generation</strong> narrows millions of possible items to a smaller set, and <strong>ranking</strong> determines which of those candidates should appear first.</p></li><li><p><strong>Offline metrics are only part of the story.</strong> Precision@K, Recall@K, and NDCG help compare models, but online experiments are needed to understand how recommendations affect real user behavior and business outcomes.</p></li><li><p><strong>Real-world recommenders face challenges beyond model accuracy</strong>, including cold start, sparse data, popularity bias, lack of diversity, and scalability.</p></li></ul><p>Perhaps the most important idea is that a recommendation system shouldn&#8217;t be judged only by how accurately it predicts what a user already likes.</p><p>A useful system also helps users <strong>discover what they might like next</strong>.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>We&#8217;ve covered the foundation of recommendation systems: what they are, how they work, the major approaches, how they&#8217;re evaluated, and some of the challenges that appear in the real world.</p><p>But we&#8217;ve only scratched the surface.</p><p>In the next article, we&#8217;ll take a closer look at <strong>Content-Based Filtering</strong>, one of the most intuitive approaches to building recommendations.</p><p>We&#8217;ll explore how a system can use information about items and a user&#8217;s past preferences to answer a simple question:</p><blockquote><p><strong>&#8220;If you liked this, what else might you like?&#8221;</strong></p></blockquote><p>We&#8217;ll also look at how similarity is calculated, walk through a practical example, and discuss where content-based recommendations work well&#8212;and where they begin to fall short.</p><p>From there, we&#8217;ll continue building our understanding of recommendation systems one concept at a time.</p><p>Because before building sophisticated recommenders, it helps to understand the ideas that make personalization possible in the first place.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[K-Means Clustering for Data Scientists]]></title><description><![CDATA[Introduction]]></description><link>https://thepracticaldatascientist.substack.com/p/k-means-clustering-for-data-scientists</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/k-means-clustering-for-data-scientists</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 28 Jul 2026 12:16:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ED4-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ED4-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 424w, /__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 848w, /__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ED4-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png" width="1456" height="966" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:966,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1821063,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/208729977?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 424w, /__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 848w, /__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ED4-!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b31f91c-4369-4f6a-9161-241c6aa7b8d5_1510x1002.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>Imagine you&#8217;re a data scientist at an online retailer. You have data on millions of customers, including:</p><ul><li><p>Their purchase history</p></li><li><p>Average order value</p></li><li><p>Shopping frequency</p></li><li><p>Product categories they buy</p></li><li><p>Time since their last purchase</p></li></ul><p>Your marketing team asks a simple question:</p><blockquote><p><strong>&#8220;How can personalize our marketing campaigns? We want to send different marketing to different groups of customers.&#8221;</strong></p></blockquote><p>There&#8217;s just one problem.</p><p>No customer has been labelled as a &#8220;high-value customer,&#8221; &#8220;bargain shopper,&#8221; or &#8220;occasional buyer.&#8221; The data simply contains customer behaviour&#8212;nothing more.</p><p>So how do we discover groups of similar customers?</p><p>This is where <strong>clustering</strong> comes in.</p><p>Unlike supervised learning, where models learn from labelled examples to make predictions, clustering looks for hidden patterns within the data itself. Instead of predicting an outcome, it groups together observations that are similar to one another.</p><p>One of the most popular algorithms for this task is <strong>K-Means Clustering</strong>.</p><p>K-Means has applications far beyond customer segmentation. It can be used to group stores with similar sales patterns, identify products with similar purchasing behaviour, organize documents by topic, detect unusual observations, and even compress images.</p><p>Despite its simplicity, K-Means remains one of the most widely used unsupervised learning algorithms because it is intuitive, scalable, and often provides valuable insights before any predictive modeling begins.</p><p>In this article, we&#8217;ll explore how K-Means works, how to choose the right number of clusters, where it performs well, where it struggles, and how it&#8217;s used to solve real-world business problems.</p><p>By the end of this issue, you&#8217;ll understand why one of the most powerful ways to learn from data isn&#8217;t by making predictions&#8212;it&#8217;s by discovering patterns you didn&#8217;t know were there.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Is Clustering?</h2><p>Most machine learning algorithms I&#8217;ve written about so far have likely been <strong>supervised learning</strong> algorithms. Models such as Linear Regression, Logistic Regression, Decision Trees, and Random Forests all have one thing in common&#8212;they learn from labelled data. </p><p>For example, if you are building a model that identifies which customer will make a purchase, your predictions would be based on historical data that actually tells which customer made a purchase and which ones didn&#8217;t.</p><p>But, what do you do when those labels don&#8217;t exist.</p><p>Imagine you&#8217;re analysing customer purchasing behaviour for the first time. You know how often customers shop, how much they spend, and the types of products they buy, but you don&#8217;t know which customers belong to similar groups. There are no labels such as &#8220;loyal customers,&#8221; &#8220;budget shoppers,&#8221; or &#8220;seasonal buyers.&#8221;</p><p>This is where <strong>clustering</strong> becomes useful.</p><p>Clustering is an <strong>unsupervised learning</strong> technique that automatically groups similar observations based on their characteristics. Instead of predicting a known outcome, it searches for natural patterns within the data.</p><p>Think of it as organising a box of mixed LEGO bricks. Without knowing anything about them beforehand, you naturally begin sorting them by colour, size, or shape because similar pieces belong together. Clustering algorithms do something similar&#8212;they group observations that are more alike than they are different.</p><p>The key difference between supervised and unsupervised learning is the goal:</p><ul><li><p><strong>Supervised learning</strong> answers, <em>&#8220;Can I predict what will happen?&#8221;</em></p></li><li><p><strong>Unsupervised learning</strong> answers, <em>&#8220;What patterns already exist in my data?&#8221;</em></p></li></ul><p>These patterns often reveal insights that aren&#8217;t immediately obvious. A retailer might discover distinct customer segments, a bank might identify groups of customers with different spending behaviours, or a manufacturer might find stores with similar inventory patterns. These insights can then guide marketing strategies, operational decisions, or even serve as inputs to future predictive models.</p><p>One of the simplest and most widely used algorithms for discovering these hidden groups is <strong>K-Means Clustering</strong>.</p><div><hr></div><h2>What Is K-Means Clustering?</h2><p>Now that we understand what clustering is, let&#8217;s look at one of the most popular clustering algorithms: <strong>K-Means</strong>.</p><p>At its core, K-Means tries to answer a simple question:</p><blockquote><p><strong>How can we divide our data into groups so that observations within the same group are as similar as possible?</strong></p></blockquote><p>The algorithm does this by creating <strong>K</strong> distinct clusters, where <strong>K</strong> is a number chosen by the user.</p><p>For example, if you&#8217;re segmenting customers, you might decide to create:</p><ul><li><p>3 customer segments</p></li><li><p>5 customer segments</p></li><li><p>8 customer segments</p></li></ul><p>Each cluster is represented by a point called a <strong>centroid</strong>. You can think of a centroid as the &#8220;centre&#8221; of a cluster. Customers that are most similar to one another are grouped around the same centroid.</p><p>The goal of K-Means is straightforward:</p><ul><li><p>Make observations within the same cluster as similar as possible.</p></li><li><p>Make different clusters as distinct as possible.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dL-f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 424w, /__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 848w, /__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dL-f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png" width="1434" height="822" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:822,&quot;width&quot;:1434,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1184571,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/208729977?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 424w, /__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 848w, /__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dL-f!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21692067-9900-419b-aaf1-ab16df1b9539_1434x822.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>To achieve this, the algorithm follows a simple, iterative process. It begins with an initial set of centroids, assigns every observation to its nearest centroid, recalculates the centroid based on the observations assigned to it, and then repeats the process until the clusters stop changing significantly.</p><p>Although this may sound computationally intensive, modern implementations are highly efficient and can cluster millions of observations in a relatively short amount of time.</p><p>Let&#8217;s return to our customer segmentation example.</p><p>Suppose you have three hundred thousand customers and choose <strong>K = 4</strong>. After running K-Means, you might discover four distinct groups:</p><ul><li><p>Customers who shop frequently and spend a lot.</p></li><li><p>Customers who make occasional large purchases.</p></li><li><p>Price-sensitive customers who mainly buy during promotions.</p></li><li><p>Infrequent customers with low overall spending.</p></li></ul><p>Notice that these labels weren&#8217;t provided to the algorithm. K-Means simply grouped customers based on similarities in their behaviour. It&#8217;s up to the data scientist and business stakeholders to interpret each cluster and decide how to use those insights.</p><p>This highlights an important point: <strong>K-Means discovers patterns, but humans give those patterns meaning.</strong></p><p>In the next section, we&#8217;ll walk through the algorithm step by step to see exactly how these clusters are formed.</p><div><hr></div><h2>How K-Means Works</h2><p>Although K-Means may sound sophisticated, the algorithm follows a surprisingly simple process. It performs two tasks iteratively: assigning observations to clusters and updating the centre of each cluster.</p><p>Let&#8217;s walk through the process using our customer segmentation example.</p><h4>Step 1: Choose the Number of Clusters (K)</h4><p>The first step is deciding how many clusters you want the algorithm to create.</p><p>Suppose you believe your customers can be grouped into four distinct segments. You would set <strong>K = 4</strong>.</p><p>At this point, the algorithm doesn&#8217;t know what those four groups look like&#8212;it only knows that it needs to create four of them. The first K can be chosen at random based on intuition.</p><h4>Step 2: Initialize the Centroids</h4><p>Next, K-Means randomly selects <strong>K</strong> observations to serve as the initial centroids.</p><p>Think of these centroids as temporary &#8220;cluster centres.&#8221; Since they are chosen randomly, they usually aren&#8217;t the final centres, but they provide a starting point for the algorithm.</p><h4>Step 3: Assign Every Observation to the Nearest Centroid</h4><p>The algorithm now looks at every observation in the dataset and calculates which centroid it is closest to.</p><p>Each observation is assigned to the nearest centroid, forming the first version of the clusters.</p><p>At this stage, the clusters are almost always imperfect because the centroids were selected randomly.</p><h4>Step 4: Update the Centroids</h4><p>Once all observations have been assigned, the algorithm recalculates the centre of each cluster. The new centroid is simply the average location of all the observations assigned to that cluster.</p><p>Because the centre has changed, some observations may now be closer to a different centroid than before.</p><h4>Step 5: Repeat Until the Clusters Stabilize</h4><p>The algorithm repeats the previous two steps:</p><ul><li><p>Assign observations to the nearest centroid.</p></li><li><p>Recalculate each centroid.</p></li></ul><p>With every iteration, the clusters become more stable and the centroids move less and less. Eventually, the centroids stop changing significantly, indicating that the algorithm has converged.</p><p>At this point, K-Means has found a set of clusters where observations within the same cluster are as similar as possible while remaining as different as possible from observations in other clusters.</p><p>Although the algorithm doesn&#8217;t guarantee the perfect clustering solution, it often produces highly useful groupings that can reveal valuable business insights.</p><p><span>&#127908; </span><strong>Interview Q: </strong>Why does K-Means repeat the assignment and update steps?</p><p><em><strong>A:</strong></em> Assigning observations changes the composition of each cluster, which changes the centroid. Once the centroids move, some observations may become closer to a different centroid. The algorithm repeats these steps until the clusters stop changing significantly.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Choosing the Right Value of K</h2><p>One of the first questions you&#8217;ll encounter when using K-Means is surprisingly simple:</p><blockquote><p><strong>How many clusters should I create?</strong></p></blockquote><p>Unfortunately, there isn&#8217;t a single correct answer.</p><p>If you choose too few clusters, you may group together observations that are actually quite different. If you choose too many, you may split meaningful groups into several smaller clusters, making the results difficult to interpret.</p><p>So, how do data scientists decide?</p><h4>1. Start with the Business Problem</h4><p>Before looking at any metrics, ask yourself why you&#8217;re clustering the data in the first place.</p><p>For example:</p><ul><li><p>A retailer may want <strong>three</strong> customer segments for a simple marketing strategy.</p></li><li><p>A bank may need <strong>five</strong> risk groups to support different lending policies.</p></li><li><p>A logistics company may cluster stores into <strong>four</strong> operational profiles for inventory planning.</p></li></ul><p>The &#8220;best&#8221; value of <strong>K</strong> should support the business decision you&#8217;re trying to make.</p><h4>2. Use Evaluation Metrics as a Guide</h4><p>Once you have a reasonable range of values for <strong>K</strong>, evaluation metrics can help narrow down your options.</p><p>The most common approaches include:</p><p><strong>Elbow Method</strong></p><p>The Elbow Method measures how much variation exists within each cluster. As you increase <strong>K</strong>, the clusters become tighter and this variation decreases.</p><p>At some point, however, adding another cluster provides only a small improvement. This point, often called the <strong>elbow</strong>, is a good candidate for the optimal value of <strong>K</strong>.</p><p><strong>Silhouette Score</strong></p><p>The Silhouette Score measures two things at the same time:</p><ul><li><p>How similar an observation is to others within its own cluster.</p></li><li><p>How different it is from observations in neighbouring clusters.</p></li></ul><p>Higher scores generally indicate well-separated, well-defined clusters.</p><p><strong>Gap Statistic</strong></p><p>The Gap Statistic compares your clustering results to what would be expected if the data had no natural grouping at all.</p><p>If your clusters are significantly better than random, the chosen value of <strong>K</strong> is likely meaningful.</p><h4>3. Validate the Results</h4><p>Even if a metric suggests an &#8220;optimal&#8221; value of <strong>K</strong>, the final decision shouldn&#8217;t stop there. Ask yourself:</p><ul><li><p>Do these clusters make business sense?</p></li><li><p>Can you explain what makes each cluster unique?</p></li><li><p>Would different actions be taken for each cluster?</p></li></ul><p>If the answer is no, then the clustering may not be useful&#8212;even if the metrics look excellent.</p><p>The ultimate goal of clustering isn&#8217;t to maximise a mathematical score. It&#8217;s to discover groups that lead to better business decisions.</p><p>&#127908; <strong>Interview Q: </strong>How do you choose the value of K?</p><p><em><strong>A:</strong></em> I start with the business objective to determine a reasonable range of values for K. Then I use techniques such as the Elbow Method, Silhouette Score, or Gap Statistic to compare different options. Finally, I validate whether the resulting clusters are meaningful and actionable from a business perspective.</p><div><hr></div><h2>Applications, Advantages, and Limitations</h2><p>K-Means is one of the most widely used clustering algorithms because it strikes a balance between simplicity, speed, and effectiveness. While it isn&#8217;t the right solution for every problem, it often serves as an excellent starting point for exploring patterns in data.</p><h4>Where Is K-Means Used?</h4><p>One of the most common applications of K-Means is <strong>customer segmentation</strong>. Instead of treating every customer the same, businesses can group customers based on purchasing behaviour, spending habits, or engagement levels. Marketing teams can then tailor promotions and campaigns for each segment rather than adopting a one-size-fits-all approach.</p><p>Retailers also use K-Means to <strong>cluster stores</strong> with similar sales patterns, inventory needs, or customer demographics. This allows operational strategies to be customised for different groups of stores instead of applying the same approach across the entire network.</p><p>Beyond retail, K-Means has applications in <strong>fraud detection</strong>, where unusual transactions may form their own cluster, <strong>document clustering</strong>, where similar articles or documents are grouped together, and <strong>image compression</strong>, where similar colours are combined to reduce file size.</p><h4>Why Is K-Means So Popular?</h4><p>K-Means remains popular because it is:</p><ul><li><p><strong>Simple</strong> to understand and implement.</p></li><li><p><strong>Fast and scalable</strong>, making it suitable for large datasets.</p></li><li><p><strong>Easy to interpret</strong>, with each cluster representing a meaningful group of similar observations.</p></li><li><p>Often a strong <strong>baseline algorithm</strong> before exploring more advanced clustering techniques.</p></li></ul><p>For many business problems, these advantages make K-Means the first algorithm data scientists try.</p><h4>When Should You Avoid K-Means?</h4><p>Like every machine learning algorithm, K-Means makes certain assumptions about the data.</p><p>It works best when clusters are relatively compact and well separated. If the clusters have irregular shapes or overlap significantly, K-Means may struggle to identify meaningful groups.</p><p>The algorithm is also sensitive to <strong>outliers</strong>. A single extreme observation can pull a centroid away from the true centre of a cluster, affecting the final results.</p><p>Another important consideration is <strong>feature scaling</strong>. Since K-Means relies on distance to measure similarity, variables measured on larger scales can dominate those measured on smaller scales. Standardising or normalising the data before clustering is therefore an important preprocessing step.</p><p>Finally, K-Means requires you to choose the number of clusters in advance. While methods such as the Elbow Method and Silhouette Score can provide useful guidance, they don&#8217;t replace business understanding and domain expertise.</p><p>&#127908; <strong>Interview Q: </strong>What are the main limitations of K-Means?</p><p><em><strong>A:</strong></em> K-Means requires the number of clusters to be specified beforehand, is sensitive to outliers and feature scaling, and performs best when clusters are compact and well separated. It may not perform well on datasets with irregularly shaped or overlapping clusters.</p><div><hr></div><h2>Key Takeaways</h2><p>Let&#8217;s recap the key concepts from this article:</p><ul><li><p><strong>K-Means is an unsupervised learning algorithm</strong> that groups similar observations into clusters without requiring labelled data.</p></li><li><p><strong>The objective of K-Means is to maximise similarity within a cluster while keeping different clusters as distinct as possible.</strong></p></li><li><p><strong>The algorithm works iteratively</strong> by assigning observations to the nearest centroid and updating the centroids until the clusters stabilise.</p></li><li><p><strong>Choosing the right value of K requires both data and domain knowledge.</strong> Techniques such as the Elbow Method, Silhouette Score, and Gap Statistic provide useful guidance, but the final decision should align with the business problem.</p></li><li><p><strong>K-Means is simple, fast, and scalable</strong>, making it one of the most widely used clustering algorithms in industry.</p></li><li><p><strong>Like any algorithm, K-Means has limitations.</strong> It performs best on compact, well-separated clusters and is sensitive to feature scaling, outliers, and the choice of K.</p></li></ul><p>The most important lesson to remember is this:</p><blockquote><p><strong>K-Means doesn&#8217;t predict outcomes&#8212;it discovers patterns. Those patterns become valuable only when they&#8217;re interpreted in the context of the business problem you&#8217;re trying to solve.</strong></p></blockquote><div><hr></div><h2>What&#8217;s Next?</h2><p>K-Means is an excellent introduction to clustering and remains one of the most widely used unsupervised learning algorithms. However, it isn&#8217;t always the best choice.</p><p>What happens if your data contains outliers? Or if the clusters aren&#8217;t neat, circular groups? What if you don&#8217;t know how many clusters exist in the first place?</p><p>These are situations where K-Means can struggle.</p><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll explore <strong>DBSCAN</strong>, a density-based clustering algorithm that takes a very different approach. Unlike K-Means, DBSCAN doesn&#8217;t require you to specify the number of clusters in advance, can identify clusters of arbitrary shapes, and naturally detects outliers.</p><p>By the end of the next article, you&#8217;ll understand not only how DBSCAN works, but also when to choose it over K-Means in real-world data science projects.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Mastering Feature Engineering]]></title><description><![CDATA[The Skill Behind Great Machine Learning Models]]></description><link>https://thepracticaldatascientist.substack.com/p/mastering-feature-engineering</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/mastering-feature-engineering</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 21 Jul 2026 13:01:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!TWXw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!TWXw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!TWXw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1524084,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/207857762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TWXw!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf30b95a-5fdb-45ac-aa8c-4557b705d84c_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Imagine two data scientists are given the same dataset and asked to predict customer churn.</p><p>The first data scientist spends hours tuning hyperparameters and experimenting with increasingly sophisticated machine learning models.</p><p>The second spends their time understanding the data.</p><p>Instead of using raw variables, they create features such as:</p><ul><li><p>Days since the customer&#8217;s last purchase</p></li><li><p>Average monthly spending</p></li><li><p>Number of purchases in the last 30 days</p></li><li><p>Change in spending over time</p></li><li><p>Customer tenure</p></li></ul><p>Surprisingly, the second data scientist achieves better results using a much simpler model.</p><p>How is that possible?</p><p>Because great machine learning models aren&#8217;t built with algorithms alone. They&#8217;re built with meaningful representations of the underlying problem.</p><p>This is the essence of <strong>feature engineering</strong>.</p><p>Feature engineering is often described as the process of transforming raw data into features that machine learning models can learn from. While that&#8217;s true, it doesn&#8217;t fully capture its importance.</p><p>At its core, feature engineering is about helping models understand what matters.</p><p>Consider the difference between these two inputs:</p><p><strong>Purchase Date:</strong> January 15, 2026</p><p>and</p><p><strong>Days Since Last Purchase:</strong> 7</p><p>The second feature immediately provides meaningful information about customer behavior. Similarly, a customer&#8217;s date of birth may be less informative than their age, and a list of transaction timestamps may be less useful than their average monthly spending or purchase frequency.</p><p>The goal isn&#8217;t simply to create more features&#8212;it&#8217;s to create better ones.</p><p>In many real-world machine learning projects, improvements in feature engineering can produce larger performance gains than switching from one algorithm to another. In fact, it&#8217;s not uncommon for a well-engineered Logistic Regression model to outperform a poorly designed gradient boosting model.</p><p>This is why feature engineering remains one of the most valuable skills for data scientists. It sits at the intersection of domain knowledge, business understanding, and machine learning.</p><p>In this article, we&#8217;ll explore how data scientists transform raw data into meaningful signals, discuss common feature engineering techniques, examine real-world examples, and learn how better features often lead to better predictions.</p><p>By the end of this issue, you&#8217;ll understand why great models are frequently built long before the training process even begins.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Is Feature Engineering?</h2><p>Before we explore specific techniques, let&#8217;s answer an important question:</p><blockquote><p><strong>What exactly is feature engineering?</strong></p></blockquote><p>Feature engineering is the process of transforming raw data into meaningful inputs that help machine learning models identify patterns and make better predictions.</p><p>Simply put, it&#8217;s the art of representing a problem in a way that a model can understand.</p><p>Consider the following examples:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XK6Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 424w, /__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 848w, /__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XK6Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png" width="1036" height="538" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:538,&quot;width&quot;:1036,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:61258,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/207857762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 424w, /__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 848w, /__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XK6Q!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd52d2a87-ca7a-47b1-98d1-e60d0063fd8b_1036x538.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Notice that none of the engineered features introduced any new information. Instead, they transformed the existing data into representations that are often more meaningful and informative.</p><p>This is what makes feature engineering so powerful.</p><p>Machine learning models don&#8217;t understand business context. They don&#8217;t know that customers who haven&#8217;t made a purchase in six months may be more likely to churn or that spending patterns often matter more than individual transactions.</p><p>Feature engineering helps us incorporate that understanding into the data.</p><h3>Why Raw Data Isn&#8217;t Always Enough</h3><p>Imagine we&#8217;re building a model to predict customer churn.</p><p>Suppose our dataset contains the following variables:</p><ul><li><p>Customer ID</p></li><li><p>Date of Birth</p></li><li><p>Purchase Dates</p></li><li><p>Transaction Amounts</p></li><li><p>Account Creation Date</p></li></ul><p>These variables may be useful, but they don&#8217;t necessarily capture customer behavior.</p><p>By engineering additional features, we can provide the model with much more meaningful information:</p><ul><li><p>Customer Age</p></li><li><p>Customer Tenure</p></li><li><p>Average Monthly Spending</p></li><li><p>Days Since Last Purchase</p></li><li><p>Purchase Frequency</p></li><li><p>Change in Spending Over Time</p></li></ul><p>These features tell a richer story about the customer and often make it easier for the model to learn meaningful patterns.</p><h3>Feature Engineering Is More Than Data Transformation</h3><p>A common misconception is that feature engineering is simply applying techniques such as one-hot encoding or scaling numerical variables.</p><p>While those techniques are important, feature engineering begins much earlier.</p><p>It starts by asking questions such as:</p><ul><li><p>What behavior are we trying to capture?</p></li><li><p>Which variables are likely to influence the outcome?</p></li><li><p>How can we represent this business problem more effectively?</p></li></ul><p>For example, when predicting whether a customer will churn, a customer&#8217;s age may matter less than how frequently they use the product. Similarly, when predicting fraudulent transactions, the transaction amount may be less informative than whether it is unusually large for that particular customer.</p><p>The best features are often those that reflect an understanding of both the business problem and the data itself.</p><h3>Important Insight</h3><p>Feature engineering isn&#8217;t about creating more variables.</p><p>It&#8217;s about creating variables that better represent the patterns we want our models to learn.</p><p>In many cases, improving the representation of the problem can have a greater impact on model performance than changing the algorithm itself.</p><h4>Interview Pro Tip</h4><p>&#127908; <strong>Interview Q:</strong> What is feature engineering, and why is it important?</p><p><em><strong>A:</strong></em> Feature engineering is the process of transforming raw data into meaningful features that help machine learning models learn more effectively. It is important because better feature representations often improve model performance, interpretability, and generalization more than simply choosing a more sophisticated algorithm.</p><div><hr></div><h2>Why Feature Engineering Matters</h2><p>It&#8217;s easy to assume that better predictions come from better algorithms.</p><p>When a model performs poorly, our first instinct is often to try something more sophisticated.</p><blockquote><p>&#8220;Let&#8217;s use XGBoost instead of Logistic Regression.&#8221;</p><p>&#8220;Maybe we should tune the hyperparameters.&#8221;</p><p>&#8220;Perhaps we need a larger neural network.&#8221;</p></blockquote><p>Sometimes, those changes help. But surprisingly often, the biggest improvements come from better features rather than better models.</p><p>This is because machine learning models can only learn from the information we provide. If the features fail to capture meaningful patterns in the data, even the most sophisticated algorithm will struggle to produce accurate predictions.</p><p>Consider the following example.</p><p>Suppose we&#8217;re building a model to predict customer churn.</p><p>One approach might be to provide the model with raw transactional data such as:</p><ul><li><p>Account creation date</p></li><li><p>Purchase timestamps</p></li><li><p>Individual transaction amounts</p></li><li><p>Website visit logs</p></li></ul><p>A better approach might be to engineer features that summarize customer behavior, such as:</p><ul><li><p>Customer tenure</p></li><li><p>Days since the last purchase</p></li><li><p>Average monthly spending</p></li><li><p>Number of website visits in the last 30 days</p></li><li><p>Change in spending over time</p></li></ul><p>Although both models are built using the same underlying data, the second set of features provides much richer information about customer behavior.</p><p>The algorithm hasn&#8217;t changed. The representation of the problem has.</p><h3>Feature Engineering Helps Models Learn Patterns</h3><p>Machine learning models are excellent at identifying patterns, but they don&#8217;t understand business context.</p><p>For example, consider the following feature:</p><blockquote><p><strong>Transaction Timestamp:</strong> <code>2026-07-15 14:35:26</code></p></blockquote><p>On its own, this value isn&#8217;t particularly meaningful.</p><p>However, we can transform it into features such as:</p><ul><li><p>Weekend vs. Weekday</p></li><li><p>Month of the Year</p></li><li><p>Days Since Last Purchase</p></li><li><p>Number of Transactions This Week</p></li><li><p>Time Since Last Login</p></li></ul><p>These engineered features provide signals that are often much more informative than the original variable.</p><p>Similarly, suppose we&#8217;re predicting whether a loan application should be approved.</p><p>Instead of using only:</p><ul><li><p>Annual Income</p></li><li><p>Monthly Expenses</p></li></ul><p>we might create features such as:</p><ul><li><p>Debt-to-Income Ratio</p></li><li><p>Savings Rate</p></li><li><p>Percentage of Income Spent on Housing</p></li></ul><p>These engineered features may better capture a customer&#8217;s financial behavior than the raw variables themselves. </p><p>One of the most valuable lessons in machine learning is that sophisticated models cannot compensate for poorly designed features. It is not uncommon for a well-engineered Logistic Regression model to outperform a poorly designed gradient boosting model.</p><p>This is why experienced data scientists spend a significant amount of their time understanding the data and designing meaningful features before training a model.</p><p>Feature engineering isn&#8217;t an afterthought&#8212;it&#8217;s often one of the most important parts of the entire machine learning pipeline.</p><p>Machine learning models learn from the features we give them. If those features fail to represent the underlying problem effectively, no amount of hyperparameter tuning or model complexity can fully compensate for that limitation.</p><p>Great models aren&#8217;t built by asking:</p><blockquote><p><strong>&#8220;Which algorithm should I use?&#8221;</strong></p></blockquote><p>They&#8217;re built by first asking:</p><blockquote><p><strong>&#8220;Am I teaching the model what actually matters?&#8221;</strong></p></blockquote><p>Better features don&#8217;t simply improve predictions&#8212;they help models learn the right patterns in the data.</p><h4>Interview Pro Tip</h4><p>&#127908; <strong>Interview Q:</strong>  Can feature engineering improve model performance more than changing the algorithm?</p><p>A: Yes. Feature engineering often has a significant impact on model performance because it determines how well the underlying problem is represented. In many cases, meaningful features can provide larger improvements than switching to a more sophisticated algorithm or tuning additional hyperparameters.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Common Feature Engineering Techniques</h2><p>Feature engineering isn&#8217;t a single technique&#8212;it&#8217;s a collection of approaches for representing data more effectively.</p><p>The best data scientists don&#8217;t simply ask, <em>&#8220;What variables do I have?&#8221;</em> They ask:</p><blockquote><p><strong>&#8220;What information am I trying to help the model learn?&#8221;</strong></p></blockquote><p>Sometimes we&#8217;re interested in understanding how often something happens. Other times, we&#8217;re trying to capture changes over time, identify relationships between variables, or provide additional business context.</p><p>Let&#8217;s look at some of the most commonly used feature engineering techniques.</p><h4>Numerical Features</h4><p>Numerical variables often benefit from transformations that make patterns easier for models to learn.</p><p>For example, we might transform:</p><ul><li><p>Annual salary into salary brackets</p></li><li><p>Transaction amounts using logarithmic transformations</p></li><li><p>Customer age into age groups</p></li><li><p>Numerical variables through standardization or normalization</p></li></ul><p>These transformations can make the data more interpretable and improve model performance.</p><h4>Categorical Features</h4><p>Machine learning models generally cannot work directly with categorical values such as:</p><ul><li><p>Product category</p></li><li><p>Payment method</p></li><li><p>Country</p></li><li><p>Customer segment</p></li></ul><p>These variables are typically transformed using techniques such as:</p><ul><li><p>One-hot encoding</p></li><li><p>Label encoding</p></li><li><p>Frequency encoding</p></li><li><p>Target encoding</p></li></ul><p>Choosing the appropriate encoding strategy often depends on both the model being used and the characteristics of the data.</p><h4>Date and Time Features</h4><p>Date and time variables frequently contain valuable information that isn&#8217;t immediately apparent.</p><p>For example, a purchase timestamp can be transformed into:</p><ul><li><p>Day of the week</p></li><li><p>Month of the year</p></li><li><p>Weekend vs. weekday</p></li><li><p>Days since the last purchase</p></li><li><p>Customer tenure</p></li><li><p>Time since the last login</p></li></ul><p>These engineered features often capture behavioral patterns more effectively than the original timestamp itself.</p><h4>Aggregated Features</h4><p>Some of the most valuable features summarize behavior across multiple observations. For example, rather than using individual transactions, we might create features such as:</p><ul><li><p>Average monthly spending</p></li><li><p>Number of purchases in the last 30 days</p></li><li><p>Average basket size</p></li><li><p>Customer lifetime value</p></li><li><p>Number of support tickets submitted</p></li></ul><p>Aggregation allows us to transform large amounts of raw data into meaningful summaries that models can learn from more effectively.</p><h4>Interaction Features</h4><p>Sometimes, individual variables aren&#8217;t particularly informative on their own, but become highly predictive when combined.</p><p>For example:</p><ul><li><p>Debt-to-income ratio combines income and expenses.</p></li><li><p>Revenue per customer combines total revenue and customer count.</p></li><li><p>Conversion rate combines purchases and website visits.</p></li></ul><p>These interaction features often capture relationships that would otherwise remain hidden in the raw data.</p><p>Feature engineering is rarely about applying a single technique. Instead, it involves combining domain knowledge, business understanding, and statistical intuition to create features that better represent the underlying problem.</p><p>When experienced data scientists look at a dataset, they don&#8217;t simply see columns and rows&#8212;they think about the behaviors, relationships, and patterns those variables might represent.</p><h4>Interview Pro Tip</h4><p>&#127908; <strong>Interview Q:</strong> What are some common feature engineering techniques?</p><p><em><strong>A:</strong></em> Common feature engineering techniques include scaling and transforming numerical variables, encoding categorical variables, extracting information from dates and timestamps, creating aggregated features, and engineering interaction features that capture relationships between variables. The appropriate technique depends on both the problem being solved and the characteristics of the data.</p><div><hr></div><h2>Feature Engineering in the Real World</h2><p>Feature engineering rarely begins with the question:</p><blockquote><p><strong>&#8220;What transformations should I apply?&#8221;</strong></p></blockquote><p>Instead, it begins with a much more important question:</p><blockquote><p><strong>&#8220;What behavior am I trying to capture?&#8221;</strong></p></blockquote><p>Experienced data scientists don&#8217;t simply look at datasets&#8212;they look for signals that represent the underlying business problem.</p><p>Let&#8217;s consider a few examples.</p><h3>Predicting Customer Churn</h3><p>Suppose we&#8217;re building a model to predict whether customers are likely to stop using a product. Our dataset contains the following variables:</p><ul><li><p>Customer ID</p></li><li><p>Purchase history</p></li><li><p>Transaction amounts</p></li><li><p>Website activity</p></li><li><p>Account creation date</p></li></ul><p>A beginner might use these variables exactly as they appear in the dataset.</p><p>An experienced data scientist might instead ask:</p><blockquote><p><strong>&#8220;What behaviors are typically associated with churn?&#8221;</strong></p></blockquote><p>Possible answers include:</p><ul><li><p>Customers are purchasing less frequently.</p></li><li><p>Customers haven&#8217;t logged in recently.</p></li><li><p>Customers are spending less than they used to.</p></li><li><p>Customers are becoming less engaged over time.</p></li></ul><p>That way of thinking naturally leads to features such as:</p><ul><li><p>Days since the last purchase</p></li><li><p>Customer tenure</p></li><li><p>Average monthly spending</p></li><li><p>Change in spending over time</p></li><li><p>Number of website visits in the last 30 days</p></li></ul><p>Notice that none of these features existed in the original dataset. They were created by combining domain knowledge with an understanding of customer behavior.</p><h3>Detecting Fraudulent Transactions</h3><p>Feature engineering is equally important in fraud detection.</p><p>Suppose we have the following information:</p><ul><li><p>Transaction amount</p></li><li><p>Transaction timestamp</p></li><li><p>Customer location</p></li><li><p>Payment method</p></li></ul><p>These variables may be useful, but they don&#8217;t necessarily tell us whether a transaction is suspicious.</p><p>Instead, we might create features such as:</p><ul><li><p>Is the transaction amount unusually large for this customer?</p></li><li><p>How many transactions occurred in the last hour?</p></li><li><p>Is this transaction occurring in a new location?</p></li><li><p>How much time has passed since the previous transaction?</p></li></ul><p>These features provide much stronger signals about potentially fraudulent behavior than the raw variables themselves.</p><h3>Building Recommendation Systems</h3><p>Consider a product recommendation model.</p><p>Rather than using individual purchases alone, we might engineer features such as:</p><ul><li><p>Purchase frequency</p></li><li><p>Average order value</p></li><li><p>Categories purchased most frequently</p></li><li><p>Products viewed recently</p></li><li><p>Time since the customer&#8217;s last interaction</p></li></ul><p>These features help us better understand customer preferences and purchasing patterns.</p><h3>Thinking Like a Data Scientist</h3><p>Notice that feature engineering looks very different across these examples.</p><p>That&#8217;s because there is no universal recipe for creating good features.</p><p>The best features are almost always driven by the business problem we&#8217;re trying to solve.</p><p>Experienced data scientists don&#8217;t begin by asking:</p><blockquote><p><strong>&#8220;Which algorithm should I use?&#8221;</strong></p></blockquote><p>They begin by asking:</p><blockquote><p><strong>&#8220;What information would help me make this prediction?&#8221;</strong></p></blockquote><p>That question often leads to better features&#8212;and ultimately, better models.</p><p>&#127908; <strong>Interview Q: </strong>How do you approach feature engineering for a new machine learning problem?</p><p><em><strong>A:</strong></em> I begin by understanding the business problem and identifying the behaviors or patterns that are likely to influence the outcome I&#8217;m trying to predict. I then use domain knowledge to engineer features that better represent those patterns before selecting and training a model.</p><div><hr></div><h2>Common Mistakes to Avoid</h2><p>Feature engineering can significantly improve model performance, but it can also introduce problems when applied without careful thought.</p><p>Ironically, some of the most common mistakes aren&#8217;t technical&#8212;they stem from focusing on the data instead of the problem we&#8217;re trying to solve.</p><h3>Creating Features Without Understanding the Business Problem</h3><p>One of the biggest mistakes data scientists make is engineering features simply because they seem useful.</p><p>Feature engineering should always begin with the question:</p><blockquote><p><strong>&#8220;What information would help explain or predict the outcome?&#8221;</strong></p></blockquote><p>Without understanding the business problem, it&#8217;s easy to create features that add complexity without adding meaningful information.</p><p>For example, when predicting customer churn, customer engagement may be far more informative than demographic information. Similarly, in fraud detection, unusual behavior may matter more than transaction amounts alone.</p><p>The most useful features are rarely the most complicated&#8212;they&#8217;re often the most meaningful.</p><h3>Assuming More Features Are Always Better</h3><p>More features do not necessarily lead to better models.</p><p>Adding hundreds of variables can introduce noise, increase computational complexity, and sometimes even reduce model performance.</p><p>A smaller set of meaningful features will often outperform a large collection of poorly designed ones.</p><p>The goal of feature engineering isn&#8217;t to maximize the number of features&#8212;it&#8217;s to maximize the amount of useful information they provide.</p><h3>Ignoring Data Leakage</h3><p>Data leakage occurs when information that would not be available at prediction time unintentionally makes its way into the training data.</p><p>For example, suppose we&#8217;re predicting whether a customer will churn next month.</p><p>Features such as:</p><ul><li><p>Purchases made after the prediction date</p></li><li><p>Future account activity</p></li><li><p>Information collected after the customer has already churned</p></li></ul><p>would all introduce information that the model would never have access to in production.</p><p>This often leads to unrealistically high model performance during development and disappointing results once the model is deployed.</p><h3>Forgetting That Features Must Generalize</h3><p>A feature may work exceptionally well on one dataset but fail when applied to new data.</p><p>Good features should capture meaningful and consistent patterns rather than accidental relationships that exist only in the training data.</p><p>This is one reason why domain knowledge remains so valuable in machine learning. Understanding why a feature should be predictive is often just as important as observing that it is predictive.</p><p>Feature engineering isn&#8217;t simply about improving performance&#8212;it&#8217;s about creating features that continue to be useful when the model encounters new and unseen data.</p><p>Feature engineering is often described as both an art and a science, and for good reason. The best features are rarely created by applying a checklist of transformations. They emerge from understanding the business problem, asking the right questions, and carefully considering what information is truly useful for making predictions.</p><h4>Interview Pro Tip</h4><p>&#127908; <strong>Interview Q: What are some common mistakes in feature engineering?&#8221;</strong></p><p><em><strong>A:</strong></em> Common mistakes include creating features without understanding the business problem, assuming that more features always improve performance, introducing data leakage, and engineering features that do not generalize well to new data. Effective feature engineering requires both domain knowledge and careful validation.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Final Takeaway</h2><p>Let&#8217;s return to the question that started this article.</p><blockquote><p><strong>What makes a great machine learning model?</strong></p></blockquote><p>It&#8217;s tempting to believe that the answer lies in more sophisticated algorithms, larger datasets, or extensive hyperparameter tuning.</p><p>While all of those things matter, they aren&#8217;t where great models usually begin.</p><p>Great models begin with a deep understanding of the problem we&#8217;re trying to solve.</p><p>Feature engineering is much more than transforming variables or applying preprocessing techniques. It&#8217;s the process of helping machine learning models understand what matters by representing the underlying problem more effectively.</p><p>Throughout this article, we&#8217;ve seen that:</p><ul><li><p>Better features often matter more than better algorithms.</p></li><li><p>Meaningful features are driven by business understanding and domain knowledge.</p></li><li><p>Feature engineering is about representing behaviors and patterns&#8212;not simply creating more variables.</p></li><li><p>The best features are often those that generalize well to new and unseen data.</p></li></ul><p>Perhaps the most important lesson is that experienced data scientists don&#8217;t approach feature engineering by asking:</p><blockquote><p><strong>&#8220;What transformations should I apply?&#8221;</strong></p></blockquote><p>They begin by asking:</p><blockquote><p><strong>&#8220;What information would help me make this prediction?&#8221;</strong></p></blockquote><p>That shift in thinking is what separates feature engineering from data preprocessing.</p><p>Machine learning models don&#8217;t understand customers, products, or businesses. They understand patterns in data. Feature engineering is what allows us to translate real-world problems into representations that models can learn from effectively.</p><p>Great models aren&#8217;t simply trained.</p><blockquote><p><strong>They&#8217;re engineered.</strong></p></blockquote><p>And that process often begins long before the first model is ever trained.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>So far, we&#8217;ve focused primarily on supervised machine learning&#8212;from building models and tuning hyperparameters to regularization and feature engineering.</p><p>But not every data science problem comes with labeled data or a clearly defined outcome variable.</p><p>Sometimes, the goal isn&#8217;t to predict what will happen next. Instead, we&#8217;re trying to answer questions such as:</p><ul><li><p>Are there natural groups hidden within our data?</p></li><li><p>Which customers behave similarly?</p></li><li><p>How can we reduce complexity without losing important information?</p></li><li><p>What patterns exist that we haven&#8217;t discovered yet?</p></li></ul><p>This is where <strong>unsupervised learning</strong> comes in.</p><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll begin a new series exploring some of the most widely used unsupervised learning techniques and the business problems they help solve.</p><p>We&#8217;ll start with one of the most popular clustering algorithms in machine learning&#8212;<strong>K-Means Clustering</strong>&#8212;and learn how data scientists use it to uncover meaningful patterns hidden within their data.</p><p>Because sometimes, the most valuable insights aren&#8217;t the ones we predict&#8212;they&#8217;re the ones we discover.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Difference-in-Differences Explained]]></title><description><![CDATA[A Business Case Study - Did the Price Increase Really Hurt Sales?]]></description><link>https://thepracticaldatascientist.substack.com/p/difference-in-differences-explained</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/difference-in-differences-explained</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 14 Jul 2026 14:03:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FkPx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FkPx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FkPx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1457159,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/206962894?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FkPx!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae240c43-f208-4a63-b3a1-c127c7df01d0_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Imagine you&#8217;re a data scientist at a global retail company.</p><p>To improve profitability, leadership decides to increase prices by 5% <strong>in Canada</strong> while keeping prices unchanged <strong>in the United States</strong>.</p><p>A month later, the results are in.</p><p>Sales in Canada have declined.</p><p>Leadership immediately asks:</p><blockquote><p><strong>&#8220;Did the price increase reduce sales?&#8221;</strong></p></blockquote><p>At first glance, answering this question seems straightforward.</p><p>Simply compare sales before and after the price increase.</p><p>If sales decreased, the pricing strategy must have been responsible.</p><p>Or was it?</p><p>What if consumer spending declined in both countries because of broader economic conditions?</p><p>What if inflation reduced discretionary spending during the same period?</p><p>Or what if demand naturally softened after the holiday season?</p><p>These factors can influence sales regardless of whether prices changed.</p><p>Simply comparing sales in Canada before and after the intervention doesn&#8217;t tell us whether the pricing strategy caused the decline.</p><p>To answer that question, we need to separate the impact of the price increase from all the other factors affecting sales over time.</p><p>This is where <strong>Difference-in-Differences (DiD)</strong> comes in.</p><p>Difference-in-Differences is one of the most widely used causal inference techniques for estimating the impact of a policy, business decision, or intervention when randomized experiments aren&#8217;t possible. Rather than looking only at how outcomes change over time, it compares those changes with a similar group that was <strong>not</strong> exposed to the intervention.</p><p>In the previous article, we explored how <strong>Propensity Score Matching</strong> creates fair comparisons between similar individuals. In this issue, we&#8217;ll tackle a different kind of problem&#8212;one where we have observations <strong>before and after</strong> an intervention.</p><p>Using the pricing strategy example throughout this article, we&#8217;ll see how Difference-in-Differences helps answer one of the most common questions in business analytics:</p><blockquote><p><strong>Did the price increase actually reduce sales, or would sales have declined even without it?</strong></p></blockquote><p>By the end of this issue, you&#8217;ll understand the intuition behind Difference-in-Differences, how it works, the assumptions it relies on, and when it should be used.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Why Before-and-After Comparisons Can Be Misleading</h2><p>Let&#8217;s return to our pricing example.</p><p>Suppose the company increased prices in Canada on January 1.</p><p>At the end of January, the analytics team reports that monthly sales have declined by <strong>8%</strong> compared to December.</p><p>Leadership immediately concludes:</p><blockquote><p><strong>&#8220;The price increase reduced sales by 8%.&#8221;</strong></p></blockquote><p>But can we really make that claim?</p><p>Not necessarily.</p><p>The problem is that many factors can influence sales over time.</p><p>For example:</p><ul><li><p>Changes in the economy</p></li><li><p>Seasonal shopping patterns</p></li><li><p>Competitor promotions</p></li><li><p>Consumer confidence</p></li><li><p>Inflation</p></li></ul><p>Any one of these factors&#8212;or several acting together&#8212;could have affected sales during the same period. As a result, simply comparing sales before and after the price increase mixes together two effects:</p><ul><li><p>The impact of the pricing strategy</p></li><li><p>All the other factors that changed over time</p></li></ul><p>Without separating these effects, we cannot confidently determine how much of the decline was actually caused by the price increase.</p><h3>Why This Matters</h3><p>Imagine that consumer spending declined across North America because of worsening economic conditions.</p><p>Even if the company had <strong>not</strong> increased prices in Canada, sales might still have fallen.</p><p>In that case, attributing the entire 8% decline to the pricing strategy would overestimate its impact.</p><p>A before-and-after comparison tells us <strong>that something changed</strong>. It does <strong>not</strong> tell us <strong>why</strong> it changed.</p><p>To estimate the true impact of an intervention, we need a way to account for changes that would have occurred even if the intervention had never happened.</p><p>This naturally raises the next question:</p><blockquote><p><strong>How can we separate the effect of the price increase from all the other factors influencing sales?</strong></p></blockquote><p>Difference-in-Differences provides one answer.</p><p><strong>&#127908; Interview Q: </strong>Why isn&#8217;t a before-and-after comparison sufficient for measuring causal effects?</p><p><em><strong>A:</strong></em> Because many external factors can change over time alongside the intervention. A before-and-after comparison cannot distinguish the effect of the intervention from other events that occurred during the same period, making it difficult to estimate the true causal effect.</p><div><hr></div><h2>The Core Idea Behind Difference-in-Differences</h2><p>We&#8217;ve seen that a simple before-and-after comparison isn&#8217;t enough to estimate the impact of the price increase.</p><p>So how can we do better?</p><p>The key idea behind <strong>Difference-in-Differences (DiD)</strong> is surprisingly simple.</p><p>Instead of looking only at how sales changed in Canada, we compare that change with what happened in a similar country where prices did <strong>not</strong> change.</p><p>In our example:</p><ul><li><p><strong>Canada</strong> is the <strong>treatment group</strong> because prices were increased.</p></li><li><p><strong>The United States</strong> is the <strong>control group</strong> because prices remained unchanged.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tvBd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tvBd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1359753,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/206962894?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tvBd!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529d367b-9b05-42c3-bd54-6bd39fc1b9b0_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Now, instead of asking:</p><blockquote><p><strong>&#8220;Did sales change after the price increase?&#8221;</strong></p></blockquote><p>we ask:</p><blockquote><p><strong>&#8220;Did sales change more in Canada than they did in the United States?&#8221;</strong></p></blockquote><p>This distinction is important.</p><p>If sales declined by a similar amount in both countries, the decline was likely driven by broader economic conditions rather than the pricing strategy.</p><p>However, if sales fell significantly more in Canada than in the United States, the additional decline provides evidence that the price increase contributed to the change.</p><p>In other words, the United States helps us estimate what might have happened in Canada <strong>if prices had never changed</strong>.</p><p>Difference-in-Differences doesn&#8217;t compare sales levels. It compares <strong>changes</strong> in sales.</p><p>By comparing how outcomes evolve over time in both the treatment and control groups, it removes many of the external factors that affect both groups equally.</p><p>This gives us a much clearer picture of the intervention&#8217;s true impact.</p><p>You can think of Difference-in-Differences as answering the question:</p><blockquote><p><strong>&#8220;How much more (or less) did the treatment group change compared with what we would have expected based on the control group?&#8221;</strong></p></blockquote><p>That simple idea is what makes Difference-in-Differences one of the most widely used causal inference methods in business, economics, and public policy.</p><p><strong>&#127908; Interview Q: </strong>What is the intuition behind Difference-in-Differences?</p><p><em><strong>A:</strong></em> Difference-in-Differences estimates the causal effect of an intervention by comparing how outcomes change over time in a treatment group relative to a control group. By focusing on the difference in changes rather than the difference in outcomes, it helps account for external factors that affect both groups.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>How Difference-in-Differences Works</h2><p>Now that we understand the intuition behind Difference-in-Differences, let&#8217;s calculate the treatment effect using our pricing example.</p><p>The figure below summarizes the sales before and after the price increase in both countries.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HmNV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HmNV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1418640,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/206962894?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HmNV!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bff769b-74c8-44f5-bdb1-e9af437a9881_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The first step is to calculate how much sales changed in each country.</p><ul><li><p><strong>Canada:</strong> Sales decreased by <strong>$0.8 million</strong>.</p></li><li><p><strong>United States:</strong> Sales decreased by <strong>$0.4 million</strong>.</p></li></ul><p>Notice that sales declined in both countries.</p><p>This suggests that part of the decline was likely caused by broader economic conditions rather than the pricing strategy itself.</p><p>To estimate the impact of the price increase, we subtract the change observed in the control group from the change observed in the treatment group:</p><p><strong>Difference-in-Differences Estimate = Treatment Change &#8722; Control Change</strong></p><p>For our example:</p><ul><li><p>Treatment change = <strong>&#8722;$0.8M</strong></p></li><li><p>Control change = <strong>&#8722;$0.4M</strong></p></li></ul><p>Estimated treatment effect:</p><p><strong>&#8722;$0.8M &#8722; (&#8722;$0.4M) = &#8722;$0.4M</strong></p><p>In other words, after accounting for the broader decline observed in the United States, the price increase is estimated to have reduced monthly sales in Canada by an additional <strong>$0.4 million</strong>.</p><p>Difference-in-Differences doesn&#8217;t compare sales levels. It compares <strong>how much the treatment group changed relative to the control group</strong>.</p><p>By removing changes that affected both groups, it provides a more credible estimate of the intervention&#8217;s impact.</p><p><strong>&#127908; Interview Q: </strong>Why do we subtract the change in the control group?</p><p><em><strong>A:</strong></em> The control group captures changes that would likely have occurred even without the intervention. Subtracting those changes helps isolate the portion of the outcome that is more likely attributable to the treatment itself.</p><div><hr></div><h2>The Parallel Trends Assumption</h2><p>Difference-in-Differences is a powerful technique, but it relies on one important assumption.</p><p>This assumption is known as the <strong>Parallel Trends Assumption</strong>.</p><p>Simply put, it states that <strong>if the price increase had never happened, sales in Canada and the United States would have followed similar trends over time.</strong></p><p>Notice that this assumption does <strong>not</strong> require the two countries to have identical sales.</p><p>Canada may consistently generate lower sales than the United States because of differences in population, market size, or consumer behavior.</p><p>That&#8217;s perfectly acceptable.</p><p>What matters is that <strong>the pattern of change over time</strong> would have been similar in both countries if the intervention had not occurred.</p><h4>Why This Matters</h4><p>Imagine that sales in Canada had already been declining much faster than sales in the United States <strong>before</strong> the price increase.</p><p>If we later observe a larger decline in Canada, we cannot confidently attribute it to the pricing strategy.</p><p>The decline may simply reflect an existing trend that began long before the intervention.</p><p>In that case, the Difference-in-Differences estimate would be biased.</p><h4>How Do We Check This?</h4><p>Although we can never prove the Parallel Trends Assumption with certainty, we can look for evidence that supports it.</p><p>One common approach is to examine historical data <strong>before</strong> the intervention.</p><p>If sales in the treatment and control groups move in a similar pattern over time before the price change, the assumption becomes more plausible.</p><p>This is why Difference-in-Differences studies often include a graph showing the trends for both groups before the intervention.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ls6Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ls6Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1454367,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/206962894?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ls6Z!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f91761-6461-4730-9fa7-a3baca65c773_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Difference-in-Differences doesn&#8217;t assume that the treatment and control groups are identical. It assumes that <strong>without the intervention, they would have changed in similar ways over time.</strong></p><p>That distinction is what makes the Parallel Trends Assumption the foundation of the entire method.</p><p><strong>&#127908; Interview Q: </strong>What is the Parallel Trends Assumption in Difference-in-Differences?</p><p><em><strong>A:</strong></em> It assumes that, in the absence of the treatment, the treatment and control groups would have followed similar trends over time. This allows the control group to serve as a credible estimate of what would have happened to the treatment group had the intervention never occurred.</p><div><hr></div><h2>Limitations of Difference-in-Differences</h2><p>Difference-in-Differences is a powerful technique, but like every causal inference method, it has limitations.</p><p>Understanding these limitations is just as important as knowing how to apply the method.</p><h4>The Control Group Matters</h4><p>Difference-in-Differences relies on having a control group that provides a reasonable estimate of what would have happened to the treatment group in the absence of the intervention.</p><p>If the control group differs substantially from the treatment group, the estimated treatment effect may be biased.</p><h4>External Events Can Affect the Results</h4><p>The method assumes that no major event affects only one of the groups during the study period.</p><p>For example, suppose a new competitor entered the Canadian market shortly after the price increase while the U.S. market remained unchanged.</p><p>Any decline in Canadian sales could then be caused by the competitor, the price increase, or both.</p><p>Difference-in-Differences alone cannot distinguish between these effects.</p><h4>Timing Is Important</h4><p>Choosing the right time window is critical.</p><p>If the &#8220;after&#8221; period is too short, the full impact of the intervention may not yet be visible.</p><p>If it is too long, additional events may occur that make it harder to isolate the treatment effect.</p><h4>Randomized Experiments Are Still Preferred</h4><p>Whenever randomized experiments are practical and ethical, they remain the gold standard for estimating causal effects.</p><p>Difference-in-Differences is most valuable when randomization is impossible or impractical and suitable observational data are available.</p><h4>Important Insight</h4><p>Difference-in-Differences doesn&#8217;t eliminate every source of bias.</p><p>Instead, it helps account for changes that affect both the treatment and control groups over time.</p><p>The quality of the estimate depends on selecting an appropriate control group, validating the Parallel Trends Assumption, and carefully considering other events that may have influenced the results.</p><p><strong>&#127908; Interview Q: </strong>What are the limitations of Difference-in-Differences?</p><p><em><strong>A:</strong></em> Difference-in-Differences relies on the Parallel Trends Assumption and requires an appropriate control group. It can produce biased estimates if external events affect only one group, if the groups follow different trends before the intervention, or if the timing of the analysis is not chosen carefully.</p><div><hr></div><h2>Final Takeaway</h2><p>Let&#8217;s return to the question that started this article.</p><blockquote><p><strong>Did the price increase reduce sales in Canada?</strong></p></blockquote><p>At first glance, the answer seemed simple.</p><p>Sales declined after the prices increased.</p><p>But as we&#8217;ve seen, that observation alone doesn&#8217;t establish causality.</p><p>Sales can change for many reasons&#8212;economic conditions, seasonality, competitor actions, and other external factors. A simple before-and-after comparison cannot separate the effect of the pricing strategy from everything else happening at the same time.</p><p>That&#8217;s where <strong>Difference-in-Differences</strong> comes in.</p><p>Rather than asking whether sales changed, it asks a more meaningful question:</p><blockquote><p><strong>Did sales change more in the treatment group than they did in a comparable control group?</strong></p></blockquote><p>By comparing changes over time instead of raw outcomes, Difference-in-Differences helps account for common trends affecting both groups and provides a more credible estimate of the intervention&#8217;s impact.</p><p>Like any causal inference technique, it has assumptions and limitations. But when those assumptions are reasonable, Difference-in-Differences becomes one of the most practical and widely used tools for measuring the real-world impact of business decisions.</p><h4>Key Takeaways</h4><ul><li><p>Before-and-after comparisons alone cannot establish causality.</p></li><li><p>Difference-in-Differences compares <strong>changes over time</strong>, not just outcomes.</p></li><li><p>A control group helps account for external factors that affect both groups.</p></li><li><p>The Parallel Trends Assumption is essential for producing credible estimates.</p></li></ul><p>Difference-in-Differences doesn&#8217;t ask:</p><blockquote><p><strong>&#8220;Did the outcome change?&#8221;</strong></p></blockquote><p>It asks:</p><blockquote><p><strong>&#8220;How much more (or less) did the outcome change because of the intervention?&#8221;</strong></p></blockquote><p>That shift in thinking is what makes Difference-in-Differences such a valuable tool for business analytics, economics, public policy, and causal inference.</p><p><strong>&#127908; Interview Q:  </strong>When would you use Difference-in-Differences?</p><p>A: I would use Difference-in-Differences when I have observations before and after an intervention for both a treatment group and a comparable control group. By comparing how outcomes change over time between the two groups, it helps estimate the causal effect of the intervention while accounting for common trends.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>So far in this causal inference series, we&#8217;ve explored two powerful approaches for estimating causal effects using observational data.</p><p>In the previous article, we learned how <strong>Propensity Score Matching</strong> creates fair comparisons between similar individuals. In this issue, we saw how <strong>Difference-in-Differences</strong> compares changes over time to isolate the impact of an intervention.</p><p>But what happens when we can&#8217;t find comparable groups, or when the assumptions behind these methods don&#8217;t hold?</p><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll explore another approach to answering difficult causal questions&#8212;one that helps estimate causal effects even when traditional comparisons become challenging.</p><p>As always, we&#8217;ll work through a real-world business case study, focusing not just on the mathematics behind the method, but on the intuition that helps you decide <strong>when</strong> and <strong>why</strong> to use it.</p><p>By the end of this series, you&#8217;ll have a practical toolkit for choosing the right causal inference technique based on the business problem you&#8217;re trying to solve&#8212;not just the data you have.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Did the Marketing Campaign Actually Increase Sales?]]></title><description><![CDATA[A Case Study in Propensity Score Matching for Data Science Interviews]]></description><link>https://thepracticaldatascientist.substack.com/p/did-the-marketing-campaign-actually</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/did-the-marketing-campaign-actually</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 07 Jul 2026 13:30:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ft0z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ft0z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ft0z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1532510,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/205705764?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ft0z!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F829d2680-38c8-4105-91ce-005c3d000d88_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>Imagine you&#8217;re a data scientist at an e-commerce company.</p><p>The marketing team launches an email campaign promoting a new loyalty program. Over 100,000 customers receive the campaign, while the remaining customers do not.</p><p>A month later, the results look promising.</p><p>Customers who received the email spent <strong>18% more</strong>, on average, than those who didn&#8217;t.</p><p>Leadership is excited and asks a simple question:</p><blockquote><p><strong>&#8220;Great! So how much did the campaign increase sales?&#8221;</strong></p></blockquote><p>At first glance, the answer seems obvious.</p><p>The campaign group spent more money, so the campaign must have worked.</p><p>Or did it?</p><p>What if the marketing team intentionally targeted customers who were already more likely to make a purchase?</p><p>Perhaps they selected customers who had:</p><ul><li><p>Purchased recently</p></li><li><p>Spent more in the past</p></li><li><p>Opened previous marketing emails</p></li><li><p>Shown higher engagement with the brand</p></li></ul><p>If that&#8217;s the case, these customers may have spent more even if they had never received the campaign.</p><p>Suddenly, the question becomes much harder.</p><p>The challenge isn&#8217;t measuring what happened.</p><p>It&#8217;s estimating <strong>what would have happened if those same customers had not received the campaign.</strong></p><p>That is the central challenge of causal inference.</p><p>In the previous issue, we introduced the concepts of correlation and causation and discussed why observational data alone is often insufficient to establish cause-and-effect relationships.</p><p>In this article, we&#8217;ll build on that foundation by exploring one of the most widely used causal inference techniques: <strong>Propensity Score Matching (PSM).</strong></p><p>Using the marketing campaign example throughout the article, we&#8217;ll see how PSM helps us create fair comparisons between treated and untreated customers, allowing us to estimate the true impact of an intervention when randomized experiments aren&#8217;t possible.</p><p>By the end of this issue, you&#8217;ll understand not only how Propensity Score Matching works, but also why it has become one of the most important tools in the causal inference toolbox for data scientists.</p><div><hr></div><h2>The Fundamental Problem of Causal Inference</h2><p>Let&#8217;s return to our marketing campaign.</p><p>Leadership wants to know:</p><blockquote><p><strong>&#8220;How much did the campaign increase sales?&#8221;</strong></p></blockquote><p>To answer that question, we would ideally compare two scenarios for the same customer:</p><p><strong>Scenario 1:</strong> The customer receives the marketing campaign.</p><p><strong>Scenario 2:</strong> The same customer does not receive the marketing campaign.</p><p>The difference in spending between these two scenarios would tell us the true effect of the campaign. Unfortunately, there is one problem.</p><p>We can only observe <strong>one</strong> of these outcomes. For every customer, only one reality exists.</p><p>If a customer received the campaign, we observe how much they spent after receiving it. We can never observe how much that same customer would have spent during the same period had they not received the campaign.</p><p>This unobserved outcome is called the <strong>counterfactual</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KV38!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KV38!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1441250,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/205705764?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KV38!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff25f3d95-291f-4865-bce0-b7c408e1b6ba_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The counterfactual represents the outcome that would have occurred under the alternative scenario. For example, suppose a customer who received the campaign spent <strong>$150</strong>.</p><p>The question we really want to answer is:</p><blockquote><p><strong>How much would this same customer have spent if they had never received the campaign?</strong></p></blockquote><p>If the answer were <strong>$130</strong>, we could conclude that the campaign increased spending by <strong>$20</strong>.</p><p>But because we can never observe both realities for the same customer, the true treatment effect is fundamentally unknowable at the individual level.</p><p>This challenge is known as the <strong>Fundamental Problem of Causal Inference</strong>.</p><h4>Why This Matters</h4><p>If we can&#8217;t compare a customer with themselves, the next best option is to compare them with someone who is as similar as possible.</p><p>That simple idea forms the foundation of many causal inference techniques, including <strong>Propensity Score Matching</strong>.</p><p>Rather than asking:</p><blockquote><p><strong>&#8220;What happened to this customer?&#8221;</strong></p></blockquote><p>we ask:</p><blockquote><p><strong>&#8220;What would likely have happened to a very similar customer who did not receive the treatment?&#8221;</strong></p></blockquote><p>Causal inference is fundamentally about estimating the <strong>missing outcome</strong>&#8212;the counterfactual.</p><p>Since we can never observe both realities for the same individual, we rely on carefully designed methods to approximate what would have happened in the absence of the treatment.</p><p><strong>&#127908; Interview Q: </strong>What is the fundamental problem of causal inference?</p><p><em><strong>A:</strong></em> For any individual, we can observe either the treated outcome or the untreated outcome, but never both. Because the counterfactual is unobservable, causal inference methods aim to estimate it using comparable individuals or carefully designed experiments.</p><div><hr></div><h2>Why Comparing Everyone Doesn&#8217;t Work</h2><p>At this point, you might be wondering:</p><blockquote><p><strong>Why can&#8217;t we simply compare customers who received the campaign with those who didn&#8217;t?</strong></p></blockquote><p>Suppose we calculate the average spending for the two groups.</p><ul><li><p>Average Spending of People who Received Campaign: $126</p></li><li><p>Average Spending of People who Did Not Receive Campaign: $102</p></li></ul><p>At first glance, the conclusion seems straightforward. The campaign group spent <strong>$24 more</strong>.</p><p>Did the campaign increase spending by $24?</p><p>Not necessarily.</p><p>The problem is that the two groups may not have been comparable to begin with.</p><p>Customers who received the campaign may have already been more likely to make a purchase than those who did not. If that&#8217;s true, part&#8212;or even all&#8212;of the difference in spending could simply reflect the characteristics of the customers rather than the impact of the campaign itself.</p><p>In other words, the treatment and control groups started from different baselines.</p><p>As a result, simply comparing the two groups mixes together two effects:</p><ul><li><p>The effect of the campaign itself</p></li><li><p>The pre-existing differences between the customers</p></li></ul><p>This is known as <strong>selection bias</strong>.</p><p>The treatment group was <strong>selected</strong>, rather than randomly assigned, making it difficult to isolate the true effect of the intervention.</p><h4>Why This Matters</h4><p>Imagine comparing the performance of professional athletes with that of casual weekend players after giving only the professionals access to a new training program.</p><p>If the professionals perform better, was it because of the training program?</p><p>Or because they were already better athletes?</p><p>The comparison isn&#8217;t fair because the two groups were fundamentally different from the start.</p><p>The same principle applies to marketing campaigns, loyalty programs, pricing experiments, and countless other business decisions.</p><p>To estimate a causal effect, we need treatment and control groups that are as similar as possible <strong>before</strong> the treatment occurs. Only then can differences in outcomes be more confidently attributed to the intervention itself.</p><p>This naturally leads to the question:</p><blockquote><p><strong>How can we create fair comparisons when treatments weren&#8217;t assigned randomly?</strong></p></blockquote><p>One of the most widely used answers is <strong>Propensity Score Matching</strong>.</p><p><strong>&#127908; Interview Q: </strong>Why can&#8217;t we directly compare treatment and control groups in observational data?</p><p><em><strong>A:</strong></em> Because the treatment and control groups may differ in important ways before the treatment occurs. These pre-existing differences introduce selection bias, making it difficult to determine whether the observed outcome was caused by the treatment or by the characteristics of the individuals themselves.</p><div><hr></div><h2>Enter Propensity Score Matching</h2><p>We&#8217;ve established two important ideas:</p><ul><li><p>We cannot compare a customer with themselves because the counterfactual is unobservable.</p></li><li><p>We cannot simply compare all treated customers with all untreated customers because the groups may be fundamentally different.</p></li></ul><p>So what can we do instead?</p><p>The answer is surprisingly intuitive.</p><p>Rather than comparing every treated customer with every untreated customer, we compare each treated customer with an untreated customer who looked similar <strong>before</strong> the treatment occurred.</p><p>This is the central idea behind <strong>Propensity Score Matching (PSM).</strong></p><p>Instead of asking:</p><blockquote><p><strong>&#8220;How did customers who received the campaign perform compared with everyone else?&#8221;</strong></p></blockquote><p>PSM asks:</p><blockquote><p><strong>&#8220;How did customers who received the campaign perform compared with similar customers who did not receive it?&#8221;</strong></p></blockquote><p>Returning to our marketing campaign example, imagine we identify a customer who:</p><ul><li><p>Is 35 years old</p></li><li><p>Has been a customer for three years</p></li><li><p>Spends about $120 per month</p></li><li><p>Shops once or twice each month</p></li></ul><p>If this customer received the campaign, we would look for another customer with nearly identical characteristics who <strong>did not</strong> receive the campaign.</p><p>By comparing these two customers, we create a much fairer comparison than comparing the entire treatment group with the entire control group.</p><p>Of course, finding a perfect match for every customer is rarely possible.</p><p>Instead, Propensity Score Matching uses a statistical approach to identify customers who have a similar likelihood of receiving the treatment based on their observed characteristics.</p><p>The goal is not to make the customers identical.</p><p>The goal is to make the treatment and control groups <strong>as comparable as possible before the intervention</strong>, so that differences observed afterward are more likely to reflect the effect of the treatment rather than pre-existing differences.</p><h4>Why This Matters</h4><p>Propensity Score Matching doesn&#8217;t eliminate bias completely.</p><p>However, it substantially reduces selection bias by ensuring that treated customers are compared with untreated customers who had similar characteristics before the treatment.</p><p>This produces a much more credible estimate of the treatment effect than a simple comparison of group averages.</p><p>You can think of Propensity Score Matching as creating <strong>apples-to-apples comparisons</strong>.</p><p>Instead of comparing everyone with everyone else, it compares people who were similar before the treatment, allowing us to estimate what might have happened had the treated individuals never received the intervention.</p><p><strong>&#127908; Interview Q:  </strong>What is the intuition behind Propensity Score Matching?</p><p><em><strong>A:</strong></em> Propensity Score Matching reduces selection bias by matching treated individuals with untreated individuals who had similar characteristics before the treatment. This creates fairer comparisons and helps estimate the causal effect of the intervention using observational data.</p><div><hr></div><h2>How Propensity Score Matching Works</h2><p>We&#8217;ve seen that comparing all treated customers with all untreated customers can lead to biased conclusions.</p><p>Instead, Propensity Score Matching compares customers who were similar before the treatment.</p><p>But how do we determine who is &#8220;similar&#8221;?</p><p>Matching customers on every characteristic individually quickly becomes impractical. A customer may differ in age, purchase history, income, geography, browsing behavior, and dozens of other attributes.</p><p>Rather than comparing customers across every variable separately, Propensity Score Matching summarizes these characteristics into a single value called the <strong>propensity score</strong>.</p><p>The propensity score represents the probability that a customer would receive the treatment based on their observed characteristics.</p><p>For our marketing campaign, we might estimate each customer&#8217;s propensity score using variables such as:</p><ul><li><p>Age</p></li><li><p>Purchase history</p></li><li><p>Average monthly spending</p></li><li><p>Website activity</p></li><li><p>Loyalty status</p></li></ul><p>Customers with similar propensity scores had a similar likelihood of receiving the campaign, even if one actually received it and the other did not.</p><h3>Step 1: Estimate the Propensity Score</h3><p>The first step is to build a model that predicts whether a customer would receive the treatment.</p><p>In practice, this is often done using <strong>Logistic Regression</strong>, although other models can also be used.</p><p>The output is a probability between 0 and 1 for every customer.</p><p>For example:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!SoOV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 424w, /__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 848w, /__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!SoOV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png" width="1162" height="482" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:482,&quot;width&quot;:1162,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40813,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/205705764?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 424w, /__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 848w, /__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SoOV!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7eb6f9b5-cb49-4d0e-afd1-973b3db946a7_1162x482.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Step 2: Match Similar Customers</h3><p>Next, each treated customer is matched with one or more untreated customers who have a similar propensity score.</p><p>For example:</p><ul><li><p>Customer A (0.82) &#8594; Customer B (0.80)</p></li><li><p>Customer D (0.25) &#8594; Customer C (0.27)</p></li></ul><p>Because these customers had similar probabilities of receiving the campaign, they form much fairer comparison pairs than randomly selected customers.</p><h3>Step 3: Compare Outcomes</h3><p>Once the matches are created, we compare the outcomes within each matched pair.</p><p>If the treated customer consistently spends more than their matched counterpart, we have stronger evidence that the campaign itself contributed to the increase in spending.</p><p>Repeating this process across all matched pairs allows us to estimate the <strong>Average Treatment Effect</strong> for the campaign.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4zTK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4zTK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7141935e-313b-4617-8803-f89dccea4353_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1548606,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/205705764?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4zTK!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7141935e-313b-4617-8803-f89dccea4353_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The propensity score itself is <strong>not</strong> the quantity we ultimately care about. Its purpose is to help us create treatment and control groups that are comparable before the intervention.</p><p>The causal effect is estimated <strong>after</strong> the matching is complete by comparing the outcomes of these matched groups.</p><p><strong>&#127908; Interview Q: </strong>Why do we estimate a propensity score instead of matching directly on every variable?</p><p><em><strong>A:</strong></em> Matching on many variables simultaneously becomes difficult as the number of features grows. The propensity score summarizes all observed characteristics into a single probability, making it much easier to create comparable treatment and control groups.</p><div><hr></div><h2>Limitations of Propensity Score Matching</h2><p>Propensity Score Matching is a powerful technique, but it is not a magic solution.</p><p>Like every causal inference method, it relies on assumptions and has important limitations that data scientists should understand before applying it.</p><h3>It Only Accounts for Observed Variables</h3><p>Propensity Score Matching can only balance characteristics that are included in the model.</p><p>For our marketing campaign, we might match customers based on variables such as purchase history, average spending, website activity, and loyalty status.</p><p>But what if customer motivation influenced both the likelihood of receiving the campaign and their future spending?</p><p>If motivation wasn&#8217;t measured, it can&#8217;t be included in the propensity score.</p><p>As a result, <strong>unobserved confounding</strong> may still bias the estimated treatment effect.</p><h3>Good Matches Are Essential</h3><p>Propensity Score Matching works best when treated and untreated customers have similar characteristics.</p><p>Suppose a campaign targeted only the company&#8217;s highest-value customers.</p><p>If there are no comparable customers in the control group, meaningful matches cannot be created. Without comparable matches, estimating the treatment effect becomes unreliable.</p><p>This concept is often referred to as <strong>common support</strong> or <strong>overlap</strong>.</p><h3>Matching Reduces the Sample Size</h3><p>Not every customer will find a suitable match. Customers without a good match are typically excluded from the analysis.</p><p>While this improves the quality of the comparisons, it also reduces the amount of data available for estimating the treatment effect.</p><h3>Randomized Experiments Are Still Preferred</h3><p>When randomized experiments are feasible, they remain the preferred approach for estimating causal effects.</p><p>Random assignment naturally balances both observed and unobserved characteristics between the treatment and control groups.</p><p>Propensity Score Matching attempts to mimic this process using observational data, but it cannot fully replace a well-designed randomized experiment.</p><h4>Important Insight</h4><p>Propensity Score Matching does not eliminate bias. Instead, it reduces bias by creating more comparable treatment and control groups based on the information that is available.</p><p>The quality of the results depends not only on the matching algorithm but also on the quality and completeness of the data.</p><p><strong>&#127908; Interview Q: </strong>What are the limitations of Propensity Score Matching?</p><p><em><strong>A:</strong></em> Propensity Score Matching only controls for observed variables. It cannot account for unmeasured confounders, requires sufficient overlap between treatment and control groups, may reduce the sample size after matching, and does not replace randomized experiments when those are feasible.</p><div><hr></div><h2>Final Takeaway</h2><p>Let&#8217;s return to the question that started this article.</p><blockquote><p><strong>Did the marketing campaign actually increase sales?</strong></p></blockquote><p>At first glance, the answer seemed simple.</p><p>Customers who received the campaign spent more than those who did not.</p><p>But as we&#8217;ve seen, that difference alone doesn&#8217;t prove the campaign caused the increase.</p><p>The treatment and control groups may have been different long before the campaign was launched.</p><p>To estimate the true impact of the campaign, we need a fair comparison.</p><p>That&#8217;s exactly what <strong>Propensity Score Matching</strong> helps us achieve.</p><p>By matching treated customers with untreated customers who had similar characteristics before the intervention, PSM reduces selection bias and provides a more credible estimate of the treatment effect.</p><p>While it cannot account for unobserved confounders or replace a randomized experiment, it remains one of the most widely used techniques for estimating causal effects from observational data.</p><h3>Key Takeaways</h3><ul><li><p>We can never observe the counterfactual for the same individual.</p></li><li><p>Simply comparing treatment and control groups can lead to biased conclusions.</p></li><li><p>Propensity Score Matching creates fairer comparisons by matching similar individuals.</p></li><li><p>The quality of the estimate depends on the quality of the matches and the available data.</p></li></ul><p>Machine learning helps us answer:</p><blockquote><p><strong>What is likely to happen?</strong></p></blockquote><p>Propensity Score Matching helps us answer:</p><blockquote><p><strong>What difference did the intervention actually make?</strong></p></blockquote><p>That distinction is what makes causal inference such a valuable skill for modern data scientists.</p><p><strong>&#127908; Interview Q: </strong>When would you use Propensity Score Matching?</p><p><em><strong>A:</strong></em> I would use Propensity Score Matching when I need to estimate the causal effect of a treatment using observational data and a randomized experiment isn&#8217;t possible. By matching treated and untreated individuals with similar characteristics, PSM helps reduce selection bias and create more credible comparisons.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>In this article, we explored how Propensity Score Matching helps us estimate causal effects by creating fair comparisons between treated and untreated groups.</p><p>But what if we have data collected <strong>before and after</strong> an intervention?</p><p>In many real-world situations, businesses, governments, and researchers evaluate policies, product launches, pricing changes, or marketing campaigns over time. Rather than matching similar individuals, they compare how outcomes change before and after an intervention.</p><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll explore this approach and learn how it helps answer one of the most common questions in analytics:</p><blockquote><p><strong>Did the intervention truly change the outcome, or would it have changed anyway?</strong></p></blockquote><p>We&#8217;ll continue building our intuition for causal inference&#8212;one business problem at a time.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Correlation Isn't Causation]]></title><description><![CDATA[A Practical Introduction to Causal Inference for Data Scientists]]></description><link>https://thepracticaldatascientist.substack.com/p/correlation-isnt-causation</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/correlation-isnt-causation</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 30 Jun 2026 13:45:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d0yC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!d0yC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!d0yC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2104248,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/204207630?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!d0yC!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef57be28-b97f-43d1-8505-05a44b4f8668_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>Over the past several issues of <em>The Practical Data Scientist</em>, we&#8217;ve focused on predictive machine learning. We&#8217;ve explored topics ranging from Logistic Regression and Decision Trees to Random Forests, XGBoost, Cross Validation, and Regularization. Together, these techniques help us answer an important question:</p><blockquote><p><strong>What is likely to happen?</strong></p></blockquote><p>Prediction is one of the most powerful applications of data science, but many business decisions require us to answer a different question:</p><blockquote><p><strong>What will happen if we change something?</strong></p></blockquote><p>Suppose a company launches a new marketing campaign and sales increase by 15%.</p><p>Did the campaign actually cause the increase? Or would sales have increased anyway because of seasonal demand?</p><p>Similarly, imagine a retailer introduces free shipping and observes a higher conversion rate.</p><p>Did free shipping drive more purchases, or were customers already more likely to buy because of a holiday shopping season?</p><p>These are not prediction problems. They are <strong>causal inference</strong> problems.</p><p>Causal inference is the field of data science that helps us estimate the effect of an intervention&#8212;whether it&#8217;s a marketing campaign, a pricing change, a product feature, or a public policy. Instead of asking <em>what happened</em>, it asks <em>what would have happened if we had acted differently?</em></p><p>This article marks the beginning of a new series on causal inference. Over the next few issues, we&#8217;ll build an intuitive understanding of how data scientists move beyond correlation to estimate cause and effect. We&#8217;ll cover topics such as propensity score matching, difference-in-differences, instrumental variables, and other techniques commonly used when randomized experiments aren&#8217;t feasible.</p><p>In this first article, we&#8217;ll establish the foundation by exploring one of the most misunderstood concepts in data science:</p><blockquote><p><strong>Correlation does not imply causation.</strong></p></blockquote><p>By the end of this issue, you&#8217;ll understand why identifying relationships in data is only the first step&#8212;and why answering causal questions requires a completely different way of thinking.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Correlation vs. Causation</h2><p>One of the most important lessons in data science is that <strong>correlation does not imply causation</strong>.</p><p>Although the phrase is widely quoted, it is also one of the most misunderstood concepts in analytics. Let&#8217;s start with the definitions.</p><p><strong>Correlation</strong> means that two variables are associated with one another. As one variable changes, the other tends to change as well.</p><p><strong>Causation</strong>, on the other hand, means that a change in one variable directly produces a change in another.</p><p>The difference may seem subtle, but it fundamentally changes the kinds of conclusions we can draw from data. Consider the following example.</p><p>Every summer:</p><ul><li><p>Ice cream sales increase.</p></li><li><p>Drowning incidents also increase.</p></li></ul><p>These two variables are positively correlated. Does this mean eating more ice cream causes drowning?</p><p>Clearly not.</p><p>The fact that two variables move together does not necessarily mean that one causes the other.</p><p>Business examples are just as common. Suppose a retailer launches a new recommendation system, and customer engagement increases.</p><p>Did the recommendation system improve engagement?</p><p>Or suppose a company offers discount coupons, and customers who receive them spend more money.</p><p>Did the coupons increase spending?</p><p>From the observed data alone, we cannot confidently answer these questions.</p><p>We know that two events occurred together, but we do not yet know whether one caused the other.</p><h4>Why This Matters</h4><p>Many machine learning models are designed to identify relationships and make predictions.</p><p>For example, a model might accurately predict which customers are likely to make a purchase or which products are likely to sell well.</p><p>But businesses often need to answer a different type of question:</p><ul><li><p>Should we launch this marketing campaign?</p></li><li><p>Should we introduce free shipping?</p></li><li><p>Should we redesign our website?</p></li><li><p>Should we recommend this product?</p></li></ul><p>These questions require us to estimate the effect of an intervention&#8212;not simply identify patterns in historical data.</p><p>Correlation tells us <strong>what tends to happen together</strong>. Causation tells us <strong>what happens because of an action</strong>.</p><p>Understanding this distinction is the first step toward making better data-driven decisions.</p><p>In the next section, we&#8217;ll explore why determining causation is much more difficult than identifying correlation, and introduce the challenges that make causal inference such an important field in data science.</p><h4>Interview Pro Tip</h4><p>&#127908; Interview Q: What is the difference between correlation and causation?</p><p><em><strong>A:</strong></em> Correlation means two variables are associated with each other, while causation means that a change in one variable directly produces a change in another. Although causal relationships often exhibit correlation, correlated variables do not necessarily have a cause-and-effect relationship.</p><div><hr></div><h2>Why Is Causation So Hard?</h2><p>If correlation doesn&#8217;t prove causation, you might wonder:</p><blockquote><p><strong>Why can&#8217;t we simply compare the outcomes before and after an intervention?</strong></p></blockquote><p>The answer is that many other factors may influence the outcome at the same time.</p><p>These factors can make it difficult&#8212;or even impossible&#8212;to determine whether the intervention itself caused the observed change.</p><p>Let&#8217;s look at some of the most common challenges.</p><h4>Confounding Variables</h4><p>A <strong>confounding variable</strong> is a factor that influences both the treatment and the outcome.</p><p>As a result, it creates the illusion of a causal relationship when one may not actually exist. Consider the ice cream example from the previous section.</p><p>During the summer:</p><ul><li><p>Ice cream sales increase.</p></li><li><p>Drowning incidents increase.</p></li></ul><p>It would be incorrect to conclude that eating ice cream causes drowning.</p><p>The real driver is <strong>temperature</strong>.</p><p>Hot weather encourages more people to buy ice cream and also leads more people to swim, increasing the likelihood of drowning incidents.</p><p>Because temperature affects both variables, it is called a <strong>confounder</strong>.</p><p>Business problems are full of confounders.</p><p>Suppose a retailer offers discount coupons to its most loyal customers. If those customers spend more money, can we conclude that the coupons caused the increase?</p><p>Not necessarily.</p><p>Loyal customers may have spent more even without receiving the coupon. In this case, <strong>customer loyalty</strong> is the confounding variable.</p><h4>Selection Bias</h4><p>Selection bias occurs when the individuals receiving a treatment are systematically different from those who do not. For example, imagine evaluating a premium loyalty program.</p><p>Customers who voluntarily join the program may already be your most engaged customers. If they spend more than non-members, we cannot immediately attribute the difference to the loyalty program.</p><p>Part of the difference may simply be due to who chose to enroll.</p><h4>Reverse Causality</h4><p>Sometimes, we correctly identify a relationship between two variables but misunderstand the direction of the relationship.</p><p>For example, suppose stores with higher sales also employ more staff.</p><p>Did hiring more employees increase sales?</p><p>Or did stores hire more employees because sales were already higher?</p><p>Without additional evidence, both explanations are plausible.</p><h4>Important Insight</h4><p>Observational data can reveal relationships, but it rarely tells us why those relationships exist.</p><p>Before concluding that one variable causes another, we must account for confounders, selection bias, and other factors that can influence the results. This is precisely why causal inference methods are needed.</p><p>&#127908; Interview Q: <strong>What is a confounding variable?&#8221;</strong></p><p><em><strong>A:</strong></em> A confounding variable is a factor that influences both the treatment and the outcome, making it difficult to determine whether the observed relationship is truly causal.</p><div><hr></div><h2>How Do Data Scientists Estimate Causal Effects?</h2><p>Once we&#8217;ve identified the challenges of observational data, the next question becomes:</p><blockquote><p><strong>How can we estimate causal effects despite these challenges?</strong></p></blockquote><p>There isn&#8217;t a single technique that works in every situation. Instead, data scientists choose from a toolbox of causal inference methods depending on the problem, the available data, and whether randomized experiments are feasible.</p><p>Some of the most commonly used approaches include:</p><ul><li><p><strong>Randomized Controlled Trials (A/B Testing)</strong> &#8212; the gold standard when randomization is possible.</p></li><li><p><strong>Propensity Score Matching (PSM)</strong> &#8212; creates comparable treatment and control groups using observational data.</p></li><li><p><strong>Difference-in-Differences (DiD)</strong> &#8212; measures the impact of an intervention by comparing changes over time.</p></li><li><p><strong>Instrumental Variables (IV)</strong> &#8212; helps estimate causal effects when unobserved confounding exists.</p></li><li><p><strong>Regression Discontinuity Design (RDD)</strong> &#8212; leverages treatment assignment based on a threshold or cutoff.</p></li></ul><p>If you&#8217;d like a deeper dive into these techniques, I&#8217;ve previously written about several of them, and you can find those articles below.</p><p>This new series will revisit each method individually, building a more intuitive understanding of when to use it, how it works, and the assumptions behind it.</p><h4>Further Reading</h4><p><a href="/__u/thepracticaldatascientist.substack.com/p/the-math-behind-ab-testing-what-you">The Math Behind A/B Testing: What You Need to Know to Interpret Results</a></p><p><a href="/__u/thepracticaldatascientist.substack.com/p/what-to-do-when-you-cant-run-an-ab">What to Do When You Can&#8217;t Run an A/B Test: Practical Techniques with Examples</a></p><div><hr></div><h2>Final Takeaway</h2><p>Throughout this article, we&#8217;ve explored a simple but powerful idea:</p><blockquote><p><strong>Correlation tells us what happens together. Causation tells us what happens because of an action.</strong></p></blockquote><p>This distinction lies at the heart of many business decisions.</p><p>Organizations rarely want to know whether two variables are related. They want to know whether an intervention&#8212;such as a marketing campaign, a pricing change, a new product feature, or a recommendation algorithm&#8212;actually caused an improvement in outcomes.</p><p>Answering these questions is far more challenging than building a predictive model. It requires careful reasoning, an understanding of bias, and the right analytical methods.</p><p>In this article, we&#8217;ve laid the foundation by understanding:</p><ul><li><p>The difference between correlation and causation</p></li><li><p>Why establishing causality is difficult</p></li><li><p>The challenges posed by confounders, selection bias, and reverse causality</p></li><li><p>The different approaches data scientists use to estimate causal effects</p></li></ul><p>In the coming issues, we&#8217;ll dive deeper into each of these techniques, exploring when to use them, how they work, and the assumptions they rely on.</p><h4>Important Insight</h4><p>Prediction helps us anticipate the future. Causal inference helps us understand how to change it.</p><p>For data scientists, that difference often determines whether an analysis simply describes the world&#8212;or helps improve it.</p><p>&#127908; Interview Q:  Why is causal inference important in data science?</p><p><em><strong>A:</strong></em> Many business decisions require understanding the effect of an intervention rather than simply predicting an outcome. Causal inference provides the framework for estimating whether an action actually caused a change, enabling organizations to make more informed decisions.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>This article marks the beginning of our new series on causal inference.</p><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll take a deeper dive into one of the most widely used approaches for estimating causal effects using observational data.</p><p>Rather than asking whether two variables are related, we&#8217;ll explore how data scientists determine what would have happened if an intervention had never occurred&#8212;a question that lies at the heart of business decision-making.</p><p>As the series progresses, we&#8217;ll continue building on these ideas, developing the intuition behind the methods that organizations use to move from correlation to causation.</p><p>If you&#8217;ve ever wondered how companies measure the true impact of a marketing campaign, a product launch, or a pricing change, this series is for you.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[10 Ways PM–Analyst Collaboration Breaks Down (and How to Fix Them)]]></title><description><![CDATA[This article was a true collaboration.]]></description><link>https://thepracticaldatascientist.substack.com/p/10-ways-pmanalyst-collaboration-breaks</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/10-ways-pmanalyst-collaboration-breaks</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Thu, 25 Jun 2026 13:31:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_iBh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>This article was a true collaboration. I partnered with Anisha Arora, Product Manager at the LCF Group, because she is as excited about product as I am about analytics and we wanted to tell both sides of the story. While I focused on the analyst perspective, she brought the product perspective, helping us explore the habits and communication patterns that quietly break down PM&#8211;Analyst collaboration&#8212;and what both sides can do to build stronger partnerships.</strong></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_iBh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_iBh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1934356,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/203007101?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_iBh!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F756e8e3a-928d-4c3f-b611-a1e44b83738c_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>We like to think collaboration between Product Managers and Analysts is natural.</span></p><p><strong><span>It isn&#8217;t.</span></strong></p><p><span>Even in high-performing teams, this relationship quietly breaks down&#8212;not because either side lacks skill, but because of how they work together day to day.</span></p><p><span>At its best, PM&#8211;Analyst collaboration is investigative. You&#8217;re both trying to understand the same system from different angles. But when collaboration becomes transactional, everyone loses.</span></p><p><span>Here are ten common breakdowns&#8212;and how both sides can work better together.</span></p><h2><strong><span>5 Collaboration Mistakes PMs Make with Analysts</span></strong></h2><p>Most PMs don&#8217;t intentionally create friction with analysts.</p><p>In fact, many of these behaviors come from good intentions&#8212;trying to move quickly, provide clarity, or keep the team aligned. But without realizing it, PMs can sometimes turn analysts into dashboard builders and query executors rather than thought partners.</p><p>The strongest PM&#8211;Analyst relationships aren&#8217;t built on requests and deliverables. They&#8217;re built on shared curiosity and a common goal: making better decisions.</p><p>Here are five ways PMs unintentionally make collaboration harder&#8212;and what better collaboration looks like.</p><h3><strong><span>1. Defining the Goal Without the Analyst</span></strong></h3><p><span>You walk into a meeting with a fully formed question: &#8220;Can we analyze churn for users in segment X over the last 90 days?&#8221;</span></p><p><span>It sounds clear. It&#8217;s actually limiting.</span></p><p><span>When PMs define the goal alone, they often lock into a specific lens before the analyst even gets context. The analyst becomes a query executor, not a thought partner.</span></p><p><em><strong>What to do instead:</strong></em></p><p><span>Bring the analyst in earlier when the problem is still messy. Frame it as:<br>&#8220;We&#8217;re seeing a drop in retention in segment X. I&#8217;d love your help figuring out what&#8217;s really happening.&#8221;</span></p><p><span>You&#8217;ll often get a better question than the one you started with.</span></p><h3><strong><span>2. Focusing on Tasks Instead of the Problem</span></strong></h3><p><span>PMs often default to task-based requests:<br>&#8220;Can you pull this dashboard?&#8221;<br>&#8220;Can you run this cohort analysis?&#8221;</span></p><p><span>But tasks without context force analysts to reverse-engineer intent, which slows everything down and reduces insight quality.</span></p><p><em><strong>What to do instead:</strong></em></p><p><span>Anchor every request in the underlying problem: &#8220;We&#8217;re trying to understand why activation dropped after onboarding changes.&#8221;</span></p><p><span>Now the analyst can challenge assumptions, suggest better approaches, and even tell you if you&#8217;re solving the wrong problem.</span></p><h3><strong><span>3. Treating Data as a Validation Tool, Not a Discovery Tool</span></strong></h3><p><span>A subtle one.</span></p><p><span>Sometimes PMs already have a hypothesis and they look to data to confirm it.<br>&#8220;Can you check if feature X improved engagement?&#8221;</span></p><p><span>This narrows the analysis to validation instead of exploration.</span></p><p><em><strong>What to do instead:</strong></em></p><p><span>Ask open-ended questions alongside your hypothesis: &#8220;My hypothesis is that feature X improved engagement, but I&#8217;d also like to understand what actually changed in user behavior.&#8221;</span></p><p><span>This creates space for unexpected insights the kind that actually move product decisions.</span></p><h3><strong><span>4. Not Closing the Loop After Analysis</span></strong></h3><p><span>The analysis gets delivered. You say thanks. Then&#8230; silence.</span></p><p><span>From the analyst&#8217;s perspective, this is frustrating. They don&#8217;t know:</span></p><ul><li><p><span>Did the work influence a decision?</span></p></li><li><p><span>Was it useful?</span></p></li><li><p><span>Did anything ship because of it?</span></p></li></ul><p><em><strong>What to do instead:</strong></em></p><p><span>Close the loop deliberately:<br>&#8220;Your analysis helped us deprioritize feature Y and focus on onboarding fixes. We&#8217;re rolling out changes next sprint.&#8221;</span></p><p><span>This builds trust and helps analysts refine future work based on real impact.</span></p><h3><strong><span>5. Skipping Retrospectives on Data Work</span></strong></h3><p><span>PMs are used to sprint retros but rarely apply the same thinking to analytics collaboration. So, the same issues repeat:</span></p><ul><li><p><span>Misaligned expectations</span></p></li><li><p><span>Last-minute requests</span></p></li><li><p><span>Over-scoped analyses</span></p></li></ul><p><em><strong>What to do instead:</strong></em></p><p><span>Run lightweight retros specifically for collaboration:</span></p><ul><li><p><span>What worked in how we framed the problem?</span></p></li><li><p><span>Where did we lose time?</span></p></li><li><p><span>What should we do differently next time?</span></p></li></ul><p>None of these mistakes come from a lack of product sense or leadership. They&#8217;re simply habits that emerge when teams become busy and communication becomes transactional.</p><p>But collaboration is a two-way street.</p><p>Just as PMs can unintentionally limit the impact of analytics, analysts can unintentionally make it harder for PMs to move quickly, make decisions, and communicate with stakeholders.</p><p>Let&#8217;s flip the perspective.</p><div><hr></div><h2><strong><span>5 Collaboration Mistakes Analysts Make with PMs</span></strong></h2><p>Analysts often talk about wanting a seat at the table.</p><p>But with that seat comes responsibility.</p><p>The best analysts don&#8217;t just answer questions&#8212;they help shape decisions. And that requires understanding the pressures PMs operate under: ambiguity, competing priorities, stakeholder expectations, and the constant need to balance speed with rigor.</p><p>Many collaboration breakdowns aren&#8217;t caused by technical gaps. They&#8217;re caused by misaligned expectations and communication styles.</p><p>Here are five common mistakes analysts make&#8212;and how to become a better partner to your PM.</p><h3><strong><span>1. Jumping Into Analysis Before Understanding the Decision</span></strong></h3><p><span>Analysts are trained to answer questions. So, when a request comes in, it&#8217;s tempting to immediately start pulling data.</span></p><p><span>But sometimes the most important question isn&#8217;t: &#8220;What analysis should I run?&#8221;</span></p><p><span>It&#8217;s: &#8220;What decision are we trying to make?&#8221;</span></p><p><span>Without understanding the decision, analysts can spend days producing work that doesn&#8217;t actually influence anything.</span></p><p><em><strong><span>What to do instead:</span></strong></em></p><p><span>Start with questions like:</span></p><ul><li><p><span>&#8220;What decision will this analysis inform?&#8221;</span></p></li><li><p><span>&#8220;What would you do differently depending on the result?&#8221;</span></p></li></ul><p><span>The goal isn&#8217;t just to answer questions&#8212;it&#8217;s to enable decisions.</span></p><h3><strong><span>2. Optimizing for Technical Correctness Instead of Business Relevance</span></strong></h3><p><span>Analysts love rigor. PMs love progress.</span></p><p><span>Sometimes analysts spend weeks building the perfect analysis when a directional answer would have been enough.</span></p><p><span>Meanwhile, the decision window closes.</span></p><p><em><strong><span>What to do instead:</span></strong></em></p><p><span>Match the depth of analysis to the importance and urgency of the decision.</span></p><p><span>Ask: &#8220;Do we need a precise answer, or do we need a useful answer?&#8221;</span></p><p><span>Perfect analyses delivered too late are rarely impactful.</span></p><h3><strong><span>3. Speaking in Metrics Instead of Narratives</span></strong></h3><p><span>Analysts spend so much time with data that it&#8217;s easy to assume the numbers speak for themselves.</span></p><p><span>They don&#8217;t.</span></p><p><span>PMs aren&#8217;t looking for charts&#8212;they&#8217;re looking for understanding.</span></p><p><em><strong><span>What to do instead:</span></strong></em></p><p><span>Move beyond: &#8220;Activation decreased by 3.2%.&#8221;</span></p><p><span>To: &#8220;Users are dropping off earlier in onboarding, which suggests the new flow may be creating friction.&#8221;</span></p><p><span>Help connect:</span></p><ul><li><p><span>What happened?</span></p></li><li><p><span>Why did it happen?</span></p></li><li><p><span>What should we do next?</span></p></li></ul><p><span>Insights are far more valuable than metrics.</span></p><h3><strong><span>4. Treating Requests as Tickets Instead of Partnerships</span></strong></h3><p><span>Sometimes analysts unintentionally position themselves as service providers:</span></p><p><span>&#8220;Tell me what you need, and I&#8217;ll deliver it.&#8221;</span></p><p><span>But the strongest analyst-PM relationships are collaborative, not transactional.</span></p><p><span>What to do instead: Challenge assumptions. Suggest alternatives. Push back when necessary.</span></p><p><em><strong><span>Instead of saying:</span></strong></em></p><p><span>&#8220;Sure, I&#8217;ll build that dashboard.&#8221;</span></p><p><em><strong><span>Try:</span></strong></em><span> &#8220;Before we build this, can we talk about the problem we&#8217;re trying to solve?&#8221;</span></p><p><span>Great analysts don&#8217;t just provide answers&#8212;they help frame better questions.</span></p><h3><strong><span>5. Disappearing Into Analysis and Returning Weeks Later</span></strong></h3><p><span>Nothing creates anxiety for PMs faster than silence. Without visibility, PMs don&#8217;t know:</span></p><ul><li><p><span>Is the analysis blocked?</span></p></li><li><p><span>Are assumptions changing?</span></p></li><li><p><span>Is the scope growing?</span></p></li></ul><p><span>What to do instead: Share progress early and often.</span></p><p><span>Even quick updates like:</span></p><ul><li><p><span>&#8220;I&#8217;m seeing something unexpected.&#8221;</span></p></li><li><p><span>&#8220;Initial results don&#8217;t support our hypothesis.&#8221;</span></p></li><li><p><span>&#8220;I may need another day to validate this.&#8221;</span></p></li></ul><p><span>These conversations prevent surprises and build trust.</span></p><div><hr></div><p>Ultimately, great products are rarely built by PMs or analysts in isolation.</p><p>The best decisions happen when product intuition and analytical thinking reinforce each other rather than compete with each other.</p><p>PMs don&#8217;t need analysts who simply answer questions.</p><p>Analysts don&#8217;t need PMs who simply assign tasks.</p><p>They both need partners who are willing to challenge assumptions, communicate openly, and stay focused on the outcome rather than the process.</p><p>Because when that happens, data stops being a reporting function and product stops being a guessing game. And that&#8217;s when teams do their best work.</p>]]></content:encoded></item><item><title><![CDATA[Regularization Explained for Data Science Interviews]]></title><description><![CDATA[Preventing Overfitting Without Sacrificing Performance]]></description><link>https://thepracticaldatascientist.substack.com/p/regularization-explained-for-data</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/regularization-explained-for-data</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 23 Jun 2026 14:01:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Jn2T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Jn2T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Jn2T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1581153,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/202997709?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Jn2T!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c9f02-773d-4982-bc19-e7f3a3be41ab_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>Introduction</h2><p>Over the past few issues of <em>The Practical Data Scientist</em>, we&#8217;ve explored increasingly powerful machine learning models and learned how to evaluate and tune them properly.</p><p>But regardless of the algorithm we use, one challenge remains:</p><blockquote><p>How do we prevent models from memorizing the training data?</p></blockquote><p>Suppose we train two models.</p><p>One achieves 99% training accuracy, while the other achieves 95%.</p><p>Does that automatically make the first model better?</p><p>Not necessarily.</p><p>A model that fits the training data too closely may end up learning noise rather than underlying patterns. As a result, it performs exceptionally well on the data it has already seen but struggles when presented with new observations.</p><p>This phenomenon, known as <strong>overfitting</strong>, is one of the most common problems in machine learning.</p><p>Ideally, we want models that capture meaningful relationships while ignoring random fluctuations in the data.</p><p>This is where <strong>regularization</strong> comes in.</p><p>Regularization introduces a penalty for model complexity, encouraging models to learn simpler and more generalizable patterns. Instead of focusing solely on fitting the training data, regularization helps strike a balance between accuracy and complexity.</p><p>In this issue, we&#8217;ll explore the intuition behind regularization, understand how Ridge (L2), Lasso (L1), and Elastic Net work, and discuss when to use each approach.</p><p>By the end of this issue, you&#8217;ll understand one of the most important ideas in machine learning:</p><blockquote><p><em>A model that fits the training data perfectly is not always the model that performs best.</em></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Idea Behind Regularization</h2><p>In the previous section, we saw that highly complex models can overfit the training data and struggle to generalize to unseen observations.</p><p>So how do we prevent this?</p><p>The answer is <strong>regularization</strong>.</p><p>At a high level, regularization discourages the model from becoming overly complex by penalizing large coefficients.</p><p>The intuition is simple:</p><blockquote><p>Simpler models tend to generalize better.</p></blockquote><p>Consider a linear model:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{y} = w_1x_1 + w_2x_2 + \\cdots + w_nx_n + b&quot;,&quot;id&quot;:&quot;ACJDNLNBYF&quot;}" data-component-name="LatexBlockToDOM"></div><p>where:</p><ul><li><p>(x_i) represents the input features</p></li><li><p>(w_i) represents the coefficients (weights)</p></li><li><p>(b) is the intercept</p></li></ul><p>The coefficients determine how strongly each feature influences the prediction.</p><p>Without any constraints, the model is free to assign very large values to these coefficients if doing so helps reduce training error. While this may improve performance on the training data, it can also make the model more sensitive to noise and increase the risk of overfitting.</p><p>To understand how regularization works, we first need to introduce the concept of a <strong>loss function</strong>.</p><p>A loss function measures how far the model&#8217;s predictions are from the true values. During training, the model tries to minimize this quantity.</p><p>Regularization modifies the loss function by adding a penalty term that discourages complexity.</p><p>Instead of minimizing only the prediction error, the model now tries to balance two objectives:</p><ul><li><p>Fit the data well</p></li><li><p>Keep the model reasonably simple</p></li></ul><p>Mathematically, the regularized loss becomes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathcal{L}_{\\text{regularized}}\n=\n\\mathcal{L}_{\\text{prediction}}\n+\n\\lambda \\times \\text{Penalty}&quot;,&quot;id&quot;:&quot;FKETYACYAC&quot;}" data-component-name="LatexBlockToDOM"></div><p>where:</p><ul><li><p>_prediction measures prediction error</p></li><li><p>lambda controls the strength of regularization</p></li><li><p>The penalty term discourages large coefficients</p></li></ul><p>By introducing this penalty, the model may sacrifice a small amount of training accuracy, but it often achieves much better performance on unseen data.</p><h4>Why Does Regularization Work?</h4><p>Large coefficients can make a model highly sensitive to small changes in the input data.</p><p>By shrinking the coefficients, regularization:</p><ul><li><p>Reduces model complexity</p></li><li><p>Lowers variance</p></li><li><p>Improves generalization</p></li><li><p>Helps prevent overfitting</p></li></ul><p>The tradeoff is that regularization may slightly increase bias. However, this increase in bias is often outweighed by a larger reduction in variance, leading to better overall performance.</p><p>Regularization does not try to fit the training data perfectly. Instead, it intentionally sacrifices a little training accuracy in exchange for better generalization.</p><p>This idea lies at the heart of the bias-variance tradeoff and is one of the reasons regularization is so effective.</p><p><em><strong>&#127908; Interview Q:</strong></em> What is the intuition behind regularization?</p><p><em><strong>A:</strong></em> Regularization penalizes model complexity, encouraging simpler models that generalize better to unseen data.</p><div><hr></div><h2>Ridge Regression (L2)</h2><p>One of the most common forms of regularization is <strong>Ridge Regression</strong>, also known as <strong>L2 regularization</strong>.</p><p>The central idea behind Ridge Regression is simple:</p><p><em><strong>Instead of eliminating features, Ridge shrinks their coefficients toward zero.</strong></em></p><p>Recall that in the previous section, we introduced the idea of adding a penalty term to the loss function. In Ridge Regression, this penalty is proportional to the <strong>sum of the squared coefficients</strong>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathcal{L}\n=\n\\mathcal{L}_{\\text{prediction}}\n+\n\\lambda \\sum_{i=1}^{n} w_i^2&quot;,&quot;id&quot;:&quot;YQGYUHIIBB&quot;}" data-component-name="LatexBlockToDOM"></div><p>where:</p><ul><li><p>_prediction is the original loss function</p></li><li><p>(w_i) are the model coefficients</p></li><li><p><em>lambda</em> controls the strength of regularization</p></li></ul><h4>What Does Ridge Regression Do?</h4><p>Ridge discourages large coefficients by making them more expensive. As the value of <em>lambda</em> increases:</p><ul><li><p>Coefficients become smaller</p></li><li><p>Model complexity decreases</p></li><li><p>Variance decreases</p></li><li><p>Generalization improves</p></li></ul><p>However, unlike Lasso Regression, Ridge does <strong>not</strong> force coefficients exactly to zero. Instead, it keeps all features in the model while reducing their influence.</p><h4>Why Does Squaring Matter?</h4><p>Because when the coefficients are squared:</p><ul><li><p>Large coefficients receive a much larger penalty</p></li><li><p>Small coefficients are penalized less severely</p></li></ul><p>This allows Ridge Regression to prevent any single feature from dominating the model.</p><h4>When Should You Use Ridge Regression?</h4><p>Ridge is particularly useful when:</p><ul><li><p>Most features are believed to contain useful information</p></li><li><p>Features are highly correlated</p></li><li><p>The goal is prediction rather than feature selection</p></li><li><p>You want a more stable model</p></li></ul><p>For example, if multiple features contain overlapping information, Ridge tends to distribute the weights across them rather than relying heavily on a single feature.</p><p>Ridge Regression reduces model complexity without removing features. You can think of Ridge as:</p><p><em><strong>&#8220;Keep all the features, but don&#8217;t let any one of them become too important.&#8221;</strong></em></p><p><em><strong>&#127908; Interview Q:</strong></em> What is the effect of L2 regularization?</p><p><em><strong>A:</strong></em> L2 regularization shrinks the coefficients toward zero, reducing model complexity and improving generalization, but it does not perform feature selection because the coefficients rarely become exactly zero.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Lasso Regression (L1)</h2><p>While Ridge Regression shrinks coefficients toward zero, <strong>Lasso Regression</strong>, also known as <strong>L1 regularization</strong>, can go one step further:</p><p><em><strong>It can eliminate them entirely.</strong></em></p><p>In Lasso Regression, the penalty term is proportional to the <strong>sum of the absolute values of the coefficients</strong>.</p><p>The regularized loss function becomes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathcal{L}\n=\n\\mathcal{L}_{\\text{prediction}}\n+\n\\lambda \\sum_{i=1}^{n} |w_i|&quot;,&quot;id&quot;:&quot;KRAHAHEPGN&quot;}" data-component-name="LatexBlockToDOM"></div><p>where:</p><ul><li><p>_prediction is the original loss function</p></li><li><p>(w_i) are the model coefficients</p></li><li><p><em>lambda</em> controls the strength of regularization</p></li></ul><h4>What Does Lasso Regression Do?</h4><p>Lasso encourages sparsity. As the value of (\lambda) increases:</p><ul><li><p>Some coefficients shrink to exactly zero</p></li><li><p>Unimportant features are removed</p></li><li><p>The model becomes simpler</p></li><li><p>Variance decreases and generalization improves</p></li></ul><p>Unlike Ridge Regression, which retains all features, Lasso automatically performs <strong>feature selection</strong>.</p><h4>Why Is Feature Selection Useful?</h4><p>In many real-world problems, only a subset of features may truly be important. By driving some coefficients to zero, Lasso:</p><ul><li><p>Removes irrelevant features</p></li><li><p>Produces simpler models</p></li><li><p>Improves interpretability</p></li><li><p>Reduces storage and computational requirements</p></li></ul><p>This makes Lasso especially attractive when working with high-dimensional datasets.</p><h4>When Should You Use Lasso Regression?</h4><p>Lasso is particularly useful when:</p><ul><li><p>You suspect that only a few features are truly important</p></li><li><p>Interpretability is important</p></li><li><p>You want the model to automatically perform feature selection</p></li><li><p>The dataset contains many features</p></li></ul><p>The key difference between Ridge and Lasso is: <em><strong>Ridge shrinks coefficients. Lasso can eliminate them.</strong></em></p><p>This ability to set coefficients exactly to zero is what makes Lasso a powerful tool for feature selection.</p><p><em><strong>&#127908; Interview Q:</strong></em>  Why does Lasso perform feature selection but Ridge does not?</p><p><em><strong>A:</strong></em> L1 regularization can drive some coefficients exactly to zero, effectively removing those features from the model. In contrast, L2 regularization shrinks coefficients toward zero but rarely makes them exactly zero.&#8221;</p><div><hr></div><h2>Elastic Net</h2><p>So far, we&#8217;ve seen that Ridge Regression shrinks coefficients while Lasso Regression can eliminate them entirely. But what if we want the benefits of both?</p><p>This is where <strong>Elastic Net</strong> comes in.</p><p>Elastic Net combines the penalties used in Ridge and Lasso, allowing the model to simultaneously shrink coefficients and perform feature selection.</p><p>The regularized loss function becomes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathcal{L}\n=\n\\mathcal{L}_{\\text{prediction}}\n+\n\\lambda_1 \\sum_{i=1}^{n} |w_i|\n+\n\\lambda_2 \\sum_{i=1}^{n} w_i^2\n&quot;,&quot;id&quot;:&quot;ZJDGVBFFLQ&quot;}" data-component-name="LatexBlockToDOM"></div><p>where:</p><ul><li><p>_prediction is the original loss function</p></li><li><p>The first penalty term corresponds to L1 regularization</p></li><li><p>The second penalty term corresponds to L2 regularization</p></li><li><p><em>lambda_1</em> and <em>lambda_2</em> control the strength of the two penalties</p></li></ul><h4>What Does Elastic Net Do?</h4><p>Elastic Net inherits properties from both Ridge and Lasso. As a result, it:</p><ul><li><p>Shrinks coefficients toward zero</p></li><li><p>Sets some coefficients exactly to zero</p></li><li><p>Performs feature selection</p></li><li><p>Reduces overfitting</p></li><li><p>Produces more stable models</p></li></ul><h4>Why Not Always Use Lasso?</h4><p>One limitation of Lasso is that when features are highly correlated, it often selects one feature and discards the others. For example, suppose we are predicting house prices using:</p><ul><li><p>House size in square feet</p></li><li><p>Number of bedrooms</p></li><li><p>Number of bathrooms</p></li></ul><p>These features are strongly related. Lasso may arbitrarily keep one feature and remove the others.</p><p>Elastic Net behaves differently.</p><p>Instead of choosing a single feature, it tends to keep groups of correlated features together, making the resulting model more stable and often more accurate.</p><h4>When Should You Use Elastic Net?</h4><p>Elastic Net is particularly useful when:</p><ul><li><p>The dataset contains many features</p></li><li><p>Features are highly correlated</p></li><li><p>You want both shrinkage and feature selection</p></li><li><p>Lasso appears unstable</p></li></ul><p>In practice, Elastic Net is often preferred over pure Lasso when working with real-world datasets containing correlated variables.</p><p>You can think of the three methods as follows:</p><ul><li><p><strong>Ridge:</strong> Keep all features, but shrink them.</p></li><li><p><strong>Lasso:</strong> Remove unnecessary features.</p></li><li><p><strong>Elastic Net:</strong> Combine the strengths of both.</p></li></ul><p><em><strong>&#127908; Interview Q:</strong></em> When would you choose Elastic Net over Lasso?&#8221;</p><p><em><strong>A:</strong></em> Elastic Net is preferred when features are highly correlated because it combines the shrinkage effect of Ridge with the feature selection ability of Lasso, resulting in more stable models.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Ridge vs Lasso vs Elastic Net</h2><p>We&#8217;ve now explored the three most common regularization techniques individually. While they all aim to reduce overfitting and improve generalization, they do so in different ways.</p><p>The image below summarizes their key differences.</p><ul><li><p><strong>Ridge Regression (L2)</strong> shrinks coefficients toward zero while keeping all features in the model. It is often a good choice when most features contain useful information and prediction accuracy is the primary goal.</p></li><li><p><strong>Lasso Regression (L1)</strong> goes one step further by driving some coefficients exactly to zero. This makes it particularly useful for feature selection and building more interpretable models.</p></li><li><p><strong>Elastic Net</strong> combines the strengths of both Ridge and Lasso. By performing shrinkage and feature selection simultaneously, it tends to work well when features are highly correlated.</p></li></ul><p>There is no universally best regularization technique.</p><p>The right choice depends on the problem, the nature of the data, and whether prediction accuracy or model interpretability is more important.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vStT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vStT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1487785,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/202997709?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vStT!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a93796-782b-4a68-8614-fdc3dda15297_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Rapid Fire Interview Questions and Final Takeaway</h2><ol><li><p><strong>What is regularization?</strong></p></li></ol><p>Regularization is a technique that penalizes model complexity to reduce overfitting and improve generalization.</p><ol start="2"><li><p><strong>Why do we need regularization?</strong></p></li></ol><p>Highly complex models can memorize noise in the training data. Regularization encourages simpler models that perform better on unseen data.</p><ol start="3"><li><p><strong>What is the difference between L1 and L2 regularization?</strong></p></li></ol><p>L2 regularization (Ridge) shrinks coefficients toward zero but keeps all features.</p><p>L1 regularization (Lasso) can shrink some coefficients exactly to zero, effectively performing feature selection.</p><ol start="4"><li><p><strong>Why does Lasso perform feature selection?</strong></p></li></ol><p>Because the L1 penalty can drive some coefficients to exactly zero, removing those features from the model.</p><ol start="5"><li><p><strong>When would you use Elastic Net?</strong></p></li></ol><p>Elastic Net is often preferred when features are highly correlated because it combines the shrinkage effect of Ridge with the feature selection ability of Lasso.</p><ol start="6"><li><p><strong>Does regularization always improve training accuracy?</strong></p></li></ol><p>No. Regularization may slightly reduce training accuracy, but it often improves performance on unseen data by reducing overfitting.</p><h4>Final Takeaway</h4><p>The goal of machine learning is not to fit the training data perfectly. The goal is to build models that generalize well.</p><p>Regularization helps achieve this by controlling model complexity and striking a balance between bias and variance.</p><p>If there&#8217;s one idea to remember, it&#8217;s this:</p><p><em><strong>A slightly simpler model that generalizes well is often more valuable than a complex model that memorizes the training data.</strong></em></p><div><hr></div><h2>What&#8217;s Next?</h2><p>Earlier in <em>The Practical Data Scientist</em>, we explored <strong>A/B testing</strong> and briefly introduced several techniques that help us move beyond simple experimentation.</p><p>In the next issue, we&#8217;ll begin taking a deeper dive into these methods with <strong>Causal Inference for Data Science Interviews</strong>.</p><p>We&#8217;ll explore one of the most important questions in data science:</p><p><em><strong>Did this action actually cause the outcome?</strong></em></p><p>We&#8217;ll start by understanding the difference between correlation and causation, discuss common sources of bias, and introduce the tools data scientists use to estimate causal effects when randomized experiments aren&#8217;t possible.</p><p>Over the coming issues, we&#8217;ll go deeper into techniques such as propensity score matching, difference-in-differences, instrumental variables, and other approaches that help answer causal questions in the real world.</p><p>Because finding patterns is useful&#8212;but understanding causes is powerful.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Cross Validation and Hyperparameter Tuning Explained]]></title><description><![CDATA[The Interview Guide to Model Selection and Generalization]]></description><link>https://thepracticaldatascientist.substack.com/p/cross-validation-and-hyperparameter</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/cross-validation-and-hyperparameter</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 16 Jun 2026 13:32:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0qHK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0qHK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0qHK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1449602,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/202056278?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0qHK!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b05cb83-9a48-4419-9e17-286c77de159a_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Over the past few issues of <em>The Practical Data Scientist</em>, we&#8217;ve explored several machine learning models, from Logistic Regression and Decision Trees to Random Forests and XGBoost.</p><p>But regardless of how powerful a model may be, one question remains: <em><strong>how do we know if it will perform well on unseen data?</strong></em></p><p>Suppose we train two models and obtain the following training accuracies:</p><ul><li><p>Model A: 99%</p></li><li><p>Model B: 95%</p></li></ul><p>Which model is better?</p><p>At first glance, Model A might seem like the obvious choice. However, training performance alone tells us very little about how the model will behave in the real world.</p><p>A model that performs exceptionally well on the training data may simply be memorizing it rather than learning patterns that generalize to new observations. This phenomenon, known as <strong>overfitting</strong>, is one of the most common challenges in machine learning.</p><p>To build models that generalize well, we need reliable ways to estimate their performance on unseen data and systematic approaches for choosing their settings.</p><p>This is where <strong>cross validation</strong> and <strong>hyperparameter tuning</strong> come into play.</p><p>In this issue, we&#8217;ll explore how to properly evaluate machine learning models, understand why a single train-test split is often not enough, and learn how to tune models to achieve better performance without falling into common pitfalls.</p><p>By the end of this issue, you&#8217;ll understand one of the most important ideas in machine learning:</p><p><em><strong>&#8220;A model is only as good as the way it is evaluated&#8221;</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Train, Validation, and Test Sets</h2><p>To understand cross validation and hyperparameter tuning, we first need to understand how data is typically divided during model development.</p><p>A machine learning model should not be evaluated on the same data it was trained on. Doing so can give an overly optimistic estimate of performance and make it difficult to detect overfitting.</p><p>For this reason, datasets are usually split into three parts:</p><h4>Training Set</h4><p>The training set is used to fit the model. During this phase, the algorithm learns patterns and relationships from the data.</p><p>Examples:</p><ul><li><p>Estimating coefficients in Logistic Regression</p></li><li><p>Building trees in Random Forests</p></li><li><p>Learning residuals in XGBoost</p></li></ul><p>The model has direct access to this data.</p><h4>Validation Set</h4><p>The validation set is used during model development. Its primary purposes are:</p><ul><li><p>Hyperparameter tuning</p></li><li><p>Model selection</p></li><li><p>Comparing different algorithms</p></li></ul><p>For example, we might use the validation set to answer questions like:</p><ul><li><p>Should the Random Forest have 100 trees or 500 trees?</p></li><li><p>What learning rate should we use for XGBoost?</p></li><li><p>Is Logistic Regression performing better than Random Forest?</p></li></ul><p>Importantly, the model does not learn directly from the validation set.</p><p>Instead, the validation set helps us make decisions about the model.</p><h4>Test Set</h4><p>The test set is used only once, after all model development is complete. Its purpose is to provide an unbiased estimate of how the final model will perform on unseen data.</p><p>Think of the test set as a final exam.</p><p>Just as students should not see the exam before taking it, the model should not use information from the test set while it is being developed.</p><p>The typical workflow looks like this:</p><pre><code><code>Dataset
   &#8595;
Training Set
   &#8595;
Validation Set
   &#8595;
Test Set
</code></code></pre><ul><li><p>Train on the training set</p></li><li><p>Tune on the validation set</p></li><li><p>Evaluate once on the test set</p></li></ul><p>Many beginners focus heavily on training accuracy. However, what ultimately matters is how well the model performs on data it has never seen before.</p><p><em><strong>&#127908; Interview Q:</strong></em> Why can&#8217;t we tune hyperparameters using the test set?</p><p><em><strong>A:</strong></em> Because the test set should remain untouched until the very end. Using it during model development can lead to overly optimistic performance estimates and poor generalization.</p><div><hr></div><h2>Cross Validation</h2><p>In the previous section, we saw how datasets are typically divided into training, validation, and test sets. But there is a problem with this approach.</p><p>Suppose we randomly split our data and obtain a validation accuracy of 92%.</p><p>How much confidence should we have in that number? What if we had chosen a different validation set?</p><p>A different split could produce:</p><ul><li><p>90% accuracy</p></li><li><p>94% accuracy</p></li><li><p>Or something entirely different</p></li></ul><p>In other words, a single train-validation split may give us a misleading estimate of model performance. This is where <strong>cross validation</strong> comes in.</p><h4>The Idea Behind Cross Validation</h4><p>Instead of relying on a single validation set, cross validation repeatedly trains and evaluates the model on different subsets of the data.</p><p>By averaging performance across multiple splits, we obtain a more reliable estimate of how well the model will generalize to unseen data.</p><h4>K-Fold Cross Validation</h4><p>The most common approach is <strong>K-Fold Cross Validation</strong>.</p><p>The process works as follows:</p><ol><li><p>Divide the dataset into <strong>K equal folds</strong></p></li><li><p>Use K&#8722;1 folds for training</p></li><li><p>Use the remaining fold for validation</p></li><li><p>Repeat the process K times so that every fold serves as the validation set exactly once</p></li><li><p>Average the performance across all K runs</p></li></ol><p>For example, in <strong>5-fold cross validation</strong>:</p><ul><li><p>Train on folds 1&#8211;4 and validate on fold 5</p></li><li><p>Train on folds 1&#8211;3 and 5, validate on fold 4</p></li><li><p>Continue until each fold has been used as the validation set</p></li></ul><p>The final score is the average across all five runs.</p><h4>Why Use Cross Validation?</h4><p>Cross validation offers several advantages:</p><ul><li><p>Makes better use of the available data</p></li><li><p>Reduces dependence on a single random split</p></li><li><p>Produces a more stable estimate of performance</p></li><li><p>Helps compare different models more fairly</p></li></ul><p>For these reasons, cross validation is widely used in machine learning competitions and real-world projects.</p><p>The most commonly used values are:</p><ul><li><p><strong>5-Fold Cross Validation</strong></p></li><li><p><strong>10-Fold Cross Validation</strong></p></li></ul><p>In practice, 5-fold cross validation provides a good balance between computational cost and reliable performance estimates.</p><p>Cross validation does not improve the model itself. Instead, it improves our confidence in the model&#8217;s estimated performance.</p><p><em><strong>&#127908; Interview Q:</strong></em> Why do we use cross validation instead of a single train-test split?&#8221;</p><p><em><strong>A:</strong></em> Cross validation reduces the dependence on a particular split and provides a more reliable estimate of how the model will perform on unseen data.&#8221;</p><div><hr></div><h2>Hyperparameter Tuning</h2><p>So far, we&#8217;ve focused on evaluating models. But once we have a reliable evaluation strategy, another question naturally arises:</p><p><em><strong>How do we choose the best version of a model?</strong></em></p><p>This is where <strong>hyperparameter tuning</strong> comes in.</p><p>Hyperparameters are settings that control how a machine learning algorithm learns. Unlike model parameters, they are not learned from the data and must be specified before training begins. </p><p>For example: <strong>Random Forest</strong></p><ul><li><p>Number of trees (<code>n_estimators</code>)</p></li><li><p>Maximum depth (<code>max_depth</code>)</p></li><li><p>Minimum samples per split (<code>min_samples_split</code>)</p></li></ul><p><strong>XGBoost</strong></p><ul><li><p>Learning rate</p></li><li><p>Maximum depth</p></li><li><p>Number of trees</p></li><li><p>Subsampling rate</p></li></ul><p>Different hyperparameter values can lead to very different models and, consequently, very different performance.</p><h4>Parameters vs Hyperparameters</h4><p>It is important to distinguish between parameters and hyperparameters.</p><p><strong>Parameters</strong></p><ul><li><p>Learned automatically from the data</p></li><li><p>Examples:</p><ul><li><p>Coefficients in Logistic Regression</p></li><li><p>Split points in Decision Trees</p></li><li><p>Leaf values in XGBoost</p></li></ul></li></ul><p><strong>Hyperparameters</strong></p><ul><li><p>Chosen before training</p></li><li><p>Control how the model learns</p></li><li><p>Examples:</p><ul><li><p>Learning rate</p></li><li><p>Maximum depth</p></li><li><p>Number of trees</p></li></ul></li></ul><h4>Grid Search</h4><p>One approach to hyperparameter tuning is <strong>Grid Search</strong>. The idea is simple:</p><ol><li><p>Define a set of possible values for each hyperparameter.</p></li><li><p>Train models using every possible combination.</p></li><li><p>Select the combination that produces the best validation performance.</p></li></ol><p>For example:</p><pre><code><code>max_depth: [3, 5, 7]
n_estimators: [100, 300, 500]</code></code></pre><p>Grid Search evaluates all nine combinations.</p><p><strong>Advantages</strong></p><ul><li><p>Simple to understand</p></li><li><p>Exhaustive search</p></li></ul><p><strong>Limitations</strong></p><ul><li><p>Computationally expensive</p></li><li><p>Doesn&#8217;t scale well with many hyperparameters</p></li></ul><h4>Random Search</h4><p>Instead of trying every combination, Random Search samples combinations randomly. For example, instead of evaluating all 100 possible combinations, we might randomly test 20.</p><p>Surprisingly, Random Search often performs just as well while requiring significantly less computation.</p><p><strong>Advantages</strong></p><ul><li><p>More efficient</p></li><li><p>Scales better to large search spaces</p></li></ul><p><strong>Limitations</strong></p><ul><li><p>Not guaranteed to find the absolute best combination</p></li></ul><p>The typical workflow looks like this:</p><pre><code><code>Choose Hyperparameters
          &#8595;
Train Model
          &#8595;
Cross Validation
          &#8595;
Evaluate Performance
          &#8595;
Select Best Combination
</code></code></pre><p>Hyperparameter tuning is not about finding the perfect model. It is about finding a model that generalizes well to unseen data.</p><p><em><strong>&#127908; Interview Q:</strong></em>  What is the difference between Grid Search and Random Search?</p><p><em><strong>A:</strong></em> Grid Search evaluates every possible combination of hyperparameters, while Random Search samples combinations randomly. Random Search is often more computationally efficient and performs surprisingly well in practice.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Common Pitfalls</h2><p>Cross validation and hyperparameter tuning are powerful tools, but they can also give misleading results if used incorrectly. Understanding these pitfalls is just as important as understanding the techniques themselves.</p><h4>Data Leakage</h4><p>Data leakage occurs when information from outside the training set unintentionally influences the model. Examples include:</p><ul><li><p>Using future information to predict the past</p></li><li><p>Scaling the entire dataset before splitting</p></li><li><p>Performing feature selection on the full dataset</p></li></ul><p>Data leakage can produce unrealistically high performance and often leads to poor results in production.</p><h4>Tuning on the Test Set</h4><p>The purpose of the test set is to estimate how the final model will perform on unseen data. If we repeatedly evaluate models on the test set and use those results to make decisions, the test set effectively becomes part of the training process.</p><p>This leads to overly optimistic performance estimates. The test set should only be used once, after all model development is complete.</p><h4>Overfitting to the Validation Set</h4><p>Trying too many models and hyperparameter combinations can cause us to unknowingly tailor our model to the validation data.</p><p>As a result, the validation score may no longer reflect true generalization performance. Cross validation helps reduce this problem, but it cannot eliminate it entirely.</p><h4>Stratified K-Fold Cross Validation</h4><p>For imbalanced datasets, a regular K-Fold split can produce folds with very different class distributions. This may lead to unstable evaluation metrics.</p><p><strong>Stratified K-Fold</strong> preserves the class proportions in each fold, making performance estimates more reliable. For classification problems with class imbalance, Stratified K-Fold is often preferred over standard K-Fold.</p><h4>Time Series Cross Validation</h4><p>Traditional cross validation randomly shuffles observations. This approach breaks the temporal order of time series data and can introduce leakage from the future.</p><p>Instead, time series problems use expanding or rolling windows. The model is always trained on past observations and validated on future observations.</p><h4>Important Insight</h4><p>The biggest mistakes in machine learning often come not from the algorithms themselves, but from how models are evaluated. A sophisticated model evaluated incorrectly can be far worse than a simple model evaluated properly.</p><p>Common interview questions include:</p><ul><li><p>What is data leakage?</p></li><li><p>Why shouldn&#8217;t we tune on the test set?</p></li><li><p>When would you use Stratified K-Fold?</p></li><li><p>Why is regular K-Fold inappropriate for time series data?</p></li></ul><p>Being able to answer these questions demonstrates a strong understanding of practical machine learning, not just theory.</p><div><hr></div><h2>Rapid Fire Interview Questions and Final Takeaway</h2><ol><li><p><em><strong>Why do we need a validation set?</strong></em></p></li></ol><p>The validation set is used for model selection and hyperparameter tuning, while the test set is reserved for the final evaluation.</p><ol start="2"><li><p><em><strong>Why can&#8217;t we tune hyperparameters using the test set?</strong></em></p></li></ol><p>Using the test set during model development can lead to overly optimistic performance estimates and poor generalization.</p><ol start="3"><li><p><em><strong>Why do we use cross validation?</strong></em></p></li></ol><p>Cross validation provides a more reliable estimate of model performance by reducing dependence on a single train-validation split.</p><ol start="4"><li><p><em><strong>What is the difference between parameters and hyperparameters?</strong></em></p></li></ol><p>Parameters are learned from the data during training, while hyperparameters are specified before training and control how the model learns.</p><ol start="5"><li><p><em><strong>Grid Search vs Random Search?</strong></em></p></li></ol><p>Grid Search evaluates every possible hyperparameter combination, while Random Search evaluates a random subset of combinations.</p><p>Random Search is often more computationally efficient.</p><ol start="6"><li><p><em><strong>When should you use Stratified K-Fold?</strong></em></p></li></ol><p>For classification problems with imbalanced classes, Stratified K-Fold ensures that each fold preserves the class distribution of the original dataset.</p><ol start="7"><li><p><em><strong>Why can&#8217;t we use regular K-Fold for time series data?</strong></em></p></li></ol><p>Randomly shuffling time series observations can introduce information from the future into the training set, leading to data leakage.</p><p>Time series data requires specialized validation strategies that preserve temporal order.</p><div><hr></div><h2>Final Takeaway</h2><p>No matter how sophisticated a machine learning algorithm may be, its success ultimately depends on how it is evaluated.</p><p>Cross validation and hyperparameter tuning help us build models that generalize to unseen data rather than simply memorizing the training set.</p><p>If there&#8217;s one idea to remember, it&#8217;s this: <em><strong>A model is only as good as the way it is evaluated.</strong></em></p><div><hr></div><h2>What&#8217;s Next?</h2><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll explore <strong>Regularization for Data Science Interviews</strong>.</p><p>We&#8217;ll dive into overfitting, the bias-variance tradeoff, and how techniques like <strong>Ridge (L2)</strong>, <strong>Lasso (L1)</strong>, and <strong>Elastic Net</strong> help machine learning models generalize better to unseen data.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Gradient Boosting and XGBoost for Data Science Interviews]]></title><description><![CDATA[How Sequential Learning Builds Some of the Most Powerful Machine Learning Models]]></description><link>https://thepracticaldatascientist.substack.com/p/gradient-boosting-and-xgboost-for</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/gradient-boosting-and-xgboost-for</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 09 Jun 2026 14:02:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Sf5Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Sf5Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Sf5Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1432285,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/201093785?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Sf5Q!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2fb61cf8-adbb-4591-b515-b1a91382fb51_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>In the last issue of <em>The Practical Data Scientist</em>, we explored Random Forests and saw how combining many decision trees can produce models that are more stable and less prone to overfitting.</p><p>Random Forests improve performance by averaging the predictions of many independent trees. This helps reduce variance and makes the overall model more reliable.</p><p>But what if we took a completely different approach?</p><p>Instead of building many trees independently and averaging their predictions, what if we built trees sequentially and allowed each new tree to focus on correcting the mistakes made by the previous ones?</p><p>This simple idea gave rise to <strong>Gradient Boosting</strong>, one of the most powerful concepts in machine learning.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Over the years, several boosting algorithms have been developed, but none have had a greater impact than <strong>XGBoost</strong>. From Kaggle competitions to production systems, XGBoost became the algorithm of choice for many structured data problems and helped popularize boosting across the machine learning community.</p><p>In this issue, we&#8217;ll explore the intuition behind Gradient Boosting, understand how XGBoost improves upon it, and discuss why these models have become staples of both real-world machine learning and data science interviews.</p><p>By the end of this issue, you&#8217;ll understand not just how boosting works, but why learning from mistakes can sometimes be more powerful than averaging predictions.</p><div><hr></div><h2>The Key Idea Behind Gradient Boosting</h2><p>In the previous issue, we saw that Random Forests improve predictions by averaging many independent trees.</p><p>Gradient Boosting takes a very different approach.</p><p>Instead of building trees independently, it builds them <strong>sequentially</strong>.</p><p>The key idea is simple: <em><strong>Each new tree tries to correct the mistakes made by the previous trees.</strong></em></p><p>Imagine we&#8217;re trying to predict house prices.</p><p>Suppose the true price of a house is:</p><pre><code><code>Actual Value = 100</code></code></pre><p>Our first tree predicts:</p><pre><code><code>Prediction = 80</code></code></pre><p>The model has made an error of:</p><pre><code><code>Residual = Actual - Prediction = 20</code></code></pre><p>Instead of starting from scratch, Gradient Boosting trains the next tree to learn this residual&#8212;the part that the first tree missed.</p><p>The second tree might predict:</p><pre><code><code>Residual Prediction = 15</code></code></pre><p>Now our updated prediction becomes:</p><pre><code><code>80 + 15 = 95</code></code></pre><p>We&#8217;re closer, but we&#8217;re still not perfect. The remaining error is:</p><pre><code><code>100 - 95 = 5</code></code></pre><p>A third tree is trained to learn this remaining error.</p><p>This process continues, with each new tree focusing on what previous trees failed to capture. As a result:</p><ul><li><p>Errors become progressively smaller</p></li><li><p>Predictions become increasingly accurate</p></li><li><p>The ensemble improves step by step</p></li></ul><h4>Random Forests vs Gradient Boosting</h4><p>Although both methods combine many trees, their philosophies are very different.</p><p><strong>Random Forests</strong></p><ul><li><p>Build trees independently</p></li><li><p>Reduce variance through averaging</p></li><li><p>Rely on the wisdom of crowds</p></li></ul><p><strong>Gradient Boosting</strong></p><ul><li><p>Build trees sequentially</p></li><li><p>Reduce errors through correction</p></li><li><p>Learn from previous mistakes</p></li></ul><p>Random Forests improve predictions through averaging. Gradient Boosting improves predictions through learning.</p><p>This seemingly small difference is what makes boosting one of the most powerful ideas in machine learning.</p><p><em><strong>&#127908; Interview Q:</strong></em> What is the intuition behind Gradient Boosting?</p><p><em><strong>A:</strong></em> Gradient Boosting builds trees sequentially, where each new tree learns from the errors made by previous trees.</p><div><hr></div><h2>How Gradient Boosting Works</h2><p>Now that we understand the intuition behind boosting, let&#8217;s see how the algorithm works in practice.</p><p>The process begins by training a simple decision tree on the original dataset. This first tree produces an initial prediction.</p><p>Since the prediction is not perfect, the model computes the errors, also known as <strong>residuals</strong>, which represent the part of the target that the current model failed to explain.</p><p>A new tree is then trained to predict these residuals. Instead of predicting the target directly, this tree learns what the previous model got wrong.</p><p>This process repeats multiple times:</p><ol><li><p>Train a tree</p></li><li><p>Calculate the residuals</p></li><li><p>Train the next tree on those residuals</p></li><li><p>Add the new predictions to the existing ones</p></li><li><p>Repeat until the model reaches the desired performance</p></li></ol><p>As more trees are added, the ensemble gradually improves its predictions.</p><h4>Additive Learning</h4><p>Unlike Random Forests, where predictions are averaged, Gradient Boosting builds an <strong>additive model</strong>. Each new tree contributes a small correction to the existing prediction. The final prediction can be thought of as:</p><pre><code><code>Final Prediction =
Tree 1
+ Tree 2
+ Tree 3
+ ...
+ Tree N</code></code></pre><p>Every tree is responsible for explaining what the previous trees missed.</p><h4>The Role of Learning Rate</h4><p>Not all trees contribute equally.</p><p>Gradient Boosting introduces a parameter called the <strong>learning rate</strong>, which controls how much each new tree influences the final prediction.</p><p>A smaller learning rate:</p><ul><li><p>Learns more slowly</p></li><li><p>Requires more trees</p></li><li><p>Often generalizes better</p></li></ul><p>A larger learning rate:</p><ul><li><p>Learns faster</p></li><li><p>Requires fewer trees</p></li><li><p>Can increase the risk of overfitting</p></li></ul><p>The learning rate acts like a step size, ensuring that the model improves gradually rather than making large corrections all at once.</p><p>Gradient Boosting does not try to build one perfect tree. Instead, it builds many simple trees, each making small improvements to the overall prediction. This gradual correction process is one of the reasons boosting algorithms are so powerful.</p><p><em><strong>&#127908; Interview Q:</strong></em>  What does the learning rate do in Gradient Boosting?</p><p><em><strong>A:</strong></em> The learning rate controls how much each new tree contributes to the ensemble. Smaller learning rates lead to slower learning but often better generalization.&#8221;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>What Makes XGBoost Special?</h2><p>Gradient Boosting provided the core idea, but XGBoost transformed that idea into one of the most successful machine learning algorithms ever created.</p><p>For many years, XGBoost dominated Kaggle competitions and became the go-to algorithm for structured and tabular data problems.</p><p>So what made it so special?</p><h4>Regularization</h4><p>Unlike traditional Gradient Boosting, XGBoost includes regularization to control model complexity and reduce overfitting.</p><p>This allows the model to achieve high predictive performance while maintaining better generalization.</p><h4>Built-in Handling of Missing Values</h4><p>Many machine learning algorithms require missing values to be imputed before training.</p><p>XGBoost can automatically learn how to handle missing values, making preprocessing simpler and more robust.</p><h4>Feature Subsampling</h4><p>Similar to Random Forests, XGBoost can randomly sample features during training.</p><p>This introduces diversity and helps reduce overfitting.</p><h4>Efficient and Scalable Training</h4><p>XGBoost was designed with speed and scalability in mind. It supports:</p><ul><li><p>Parallel computation</p></li><li><p>Efficient memory usage</p></li><li><p>Distributed training</p></li></ul><p>These improvements make it suitable for large datasets and production environments.</p><h4>Why Did XGBoost Dominate Kaggle Competitions?</h4><p>XGBoost offered several advantages:</p><ul><li><p>Strong predictive performance</p></li><li><p>Excellent handling of tabular data</p></li><li><p>Robustness to noisy features</p></li><li><p>Flexibility through extensive hyperparameters</p></li><li><p>Effective regularization</p></li></ul><p>As a result, it became the default starting point for many machine learning competitions and real-world applications.</p><p>XGBoost is not a completely different algorithm. It is an optimized implementation of Gradient Boosting that adds several practical improvements to improve speed, robustness, and generalization.</p><p><em><strong>&#127908; Interview Q: </strong></em>&#8220;What is the difference between Gradient Boosting and XGBoost?&#8221;</p><p><em><strong>A:</strong></em> &#8220;XGBoost is an optimized implementation of Gradient Boosting that adds regularization, feature subsampling, efficient training, and automatic handling of missing values.&#8221;</p><div><hr></div><h2>Advantages and Limitations of XGBoost</h2><p>XGBoost became one of the most popular machine learning algorithms because it combines strong predictive performance with several practical improvements over traditional Gradient Boosting. It is often one of the first algorithms data scientists try.</p><h4>Advantages</h4><p>XGBoost offers several important benefits:</p><ul><li><p>Excellent predictive performance</p></li><li><p>Handles complex non-linear relationships</p></li><li><p>Built-in regularization helps reduce overfitting</p></li><li><p>Supports missing values natively</p></li><li><p>Provides feature importance scores</p></li><li><p>Scales efficiently to large datasets</p></li><li><p>Frequently achieves state-of-the-art results on structured data</p></li></ul><p>These advantages made XGBoost the dominant algorithm in many machine learning competitions and production systems.</p><h4>Limitations</h4><p>Despite its strengths, XGBoost is not always the best choice. Some common limitations include:</p><ul><li><p>Larger number of hyperparameters</p></li><li><p>Longer training times compared to simpler models</p></li><li><p>Greater risk of overfitting if not properly tuned</p></li><li><p>Less interpretable than simpler models</p></li><li><p>Requires more experimentation to achieve optimal performance</p></li></ul><h4>When Should You Use XGBoost?</h4><p>XGBoost is often a strong choice when:</p><ul><li><p>Predictive performance is more important than interpretability</p></li><li><p>Relationships between features are complex and non-linear</p></li><li><p>You are willing to spend time tuning hyperparameters</p></li></ul><p>In practice, XGBoost is commonly used for:</p><ul><li><p>Fraud detection</p></li><li><p>Customer churn prediction</p></li><li><p>Credit risk modeling</p></li><li><p>Recommendation systems</p></li><li><p>Demand forecasting</p></li><li><p>Marketing response modeling</p></li></ul><p>XGBoost is often considered the &#8220;power tool&#8221; of machine learning. It can achieve exceptional performance, but that performance usually comes at the cost of increased complexity and tuning effort.</p><p><em><strong>&#127908; Interview Q: </strong></em>When would you choose XGBoost over simpler models?</p><p><em><strong>A:</strong></em> When predictive performance is the primary objective and I&#8217;m working with structured data, XGBoost is often an excellent choice because of its ability to capture complex patterns while maintaining strong generalization.&#8221;</p><div><hr></div><h2>Random Forest vs XGBoost</h2><p>Random Forests and XGBoost are both ensemble methods built on decision trees, but they improve performance in very different ways.</p><p>Random Forests reduce variance by averaging the predictions of many independent trees. XGBoost, on the other hand, builds trees sequentially, with each new tree correcting the mistakes made by the previous ones.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Bu35!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Bu35!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1461054,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/201093785?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Bu35!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd97755d-0a3c-4182-bc2b-863d7676731a_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Which One Should You Use?</h4><p>If you want:</p><ul><li><p>A robust model with minimal tuning</p></li><li><p>Faster training</p></li><li><p>Lower risk of overfitting</p></li></ul><p>then <strong>Random Forest</strong> is often an excellent choice.</p><p>If you want:</p><ul><li><p>The highest possible predictive performance</p></li><li><p>Greater flexibility</p></li><li><p>Better results on many tabular datasets</p></li></ul><p>then <strong>XGBoost</strong> is frequently the preferred option.</p><p>Neither algorithm is universally better. Random Forests are often easier to train and interpret, while XGBoost typically offers superior performance when properly tuned.</p><p>A good rule of thumb is: <em><strong>Start with Random Forests. Optimize with XGBoost.</strong></em></p><p><em><strong>&#127908; Interview Q: </strong></em>Random Forest or XGBoost&#8212;which would you choose?</p><p><em><strong>A:</strong></em> Random Forest is a great baseline because it is simple and robust. If maximizing predictive performance is critical and I have time for tuning, I would consider XGBoost.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Rapid Fire Interview Questions</h2><p><em><strong>Q: What is the difference between bagging and boosting?</strong></em></p><p><em><strong>A:</strong></em> Bagging builds trees independently and combines their predictions through averaging. Boosting builds trees sequentially, with each new tree correcting the mistakes of the previous ones.</p><p><em><strong>Q: What are residuals in Gradient Boosting?</strong></em></p><p><em><strong>A:</strong></em> Residuals are the errors made by the current model. Each new tree is trained to predict these residuals, gradually improving the overall prediction.</p><p><em><strong>Q: What does the learning rate do?</strong></em></p><p><em><strong>A:</strong></em> The learning rate controls how much each new tree contributes to the ensemble. Smaller learning rates lead to slower learning but often better generalization.</p><p><em><strong>Q: Why is XGBoost so powerful?</strong></em></p><p><em><strong>A:</strong></em> XGBoost combines the idea of Gradient Boosting with several practical improvements, including regularization, feature subsampling, efficient training, and automatic handling of missing values.</p><p><em><strong>Q: Why can boosting overfit?</strong></em></p><p><em><strong>A:</strong></em> Because each new tree focuses on correcting previous errors, adding too many trees or using an aggressive learning rate can cause the model to learn noise in the training data.</p><p><em><strong>Q:</strong></em> Which is better: Random Forest or XGBoost?</p><p><em><strong>A:</strong></em> Neither model is universally better. Random Forests are simpler and more robust, while XGBoost often achieves higher predictive performance when properly tuned.</p><div><hr></div><h2>Final Takeaway</h2><p>Random Forests improve predictions through averaging.Gradient Boosting improves predictions through correction.</p><p>While Random Forests rely on the wisdom of crowds, Gradient Boosting and XGBoost learn from their mistakes.</p><p>This simple idea&#8212;building models sequentially and correcting errors step by step&#8212;helped make XGBoost one of the most successful machine learning algorithms ever created.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll shift our focus from building models to evaluating and selecting them.</p><p>We&#8217;ll explore <strong>Cross Validation and Hyperparameter Tuning</strong>, covering concepts like train-validation-test splits, K-fold cross validation, and common tuning strategies that help machine learning models generalize to unseen data.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Random Forests for Data Science Interviews]]></title><description><![CDATA[How Bagging Transforms Unstable Decision Trees into Reliable Models]]></description><link>https://thepracticaldatascientist.substack.com/p/random-forests-made-simple</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/random-forests-made-simple</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 02 Jun 2026 14:01:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IGVo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!IGVo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!IGVo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1505850,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/200073882?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IGVo!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdedce035-d9c8-4631-b60e-0b6022c4090a_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>In the last issue of <em>The Practical Data Scientist</em>, we explored Decision Trees and saw why they became one of the most popular machine learning algorithms.</p><p>They are intuitive, easy to visualize, and capable of capturing complex non-linear relationships. They can automatically learn feature interactions and require very little preprocessing.</p><p>But decision trees have an important weakness: <em><strong>A single decision tree can be highly unstable.</strong></em></p><p>Imagine training a decision tree on a customer churn dataset. If you slightly change the training data&#8212;perhaps by removing a few observations or collecting a new batch of customers&#8212;the resulting tree may look completely different.</p><p>Different data can lead to:</p><ul><li><p>Different tree structures</p></li><li><p>Different decision rules</p></li><li><p>Different predictions</p></li></ul><p>This behavior is known as <strong>high variance</strong>.</p><p>Decision trees are also prone to overfitting. As the tree grows deeper, it may begin learning random noise and dataset-specific patterns instead of meaningful relationships. While a deep tree can capture complex patterns, it may not generalize well to unseen data.</p><p>As a result, a single decision tree often suffers from:</p><ul><li><p>High variance</p></li><li><p>Overfitting</p></li><li><p>Unstable predictions</p></li></ul><p>This naturally raises an important question: <em><strong>How do we keep the flexibility of decision trees while reducing their instability?</strong></em></p><p>The answer is surprisingly simple. Instead of relying on one tree, build many trees and combine their predictions.</p><p>This idea forms the foundation of <strong>Random Forests</strong>.</p><p>Random Forests are one of the most successful machine learning algorithms ever created. They are widely used for fraud detection, churn prediction, recommendation systems, risk modeling, and countless other real-world applications.</p><p>In this issue, we&#8217;ll explore the key ideas behind Random Forests, including bagging, bootstrap sampling, feature randomness, and ensemble learning. More importantly, we&#8217;ll focus on the intuition behind why combining many trees often produces a model that is more stable, more accurate, and less prone to overfitting than any individual tree.</p><p>&#128204; <em><strong>Pro Tip:</strong></em> <em><strong>&#8220;What problem do Random Forests solve?&#8221;</strong></em></p><p>&#8220;Random Forests primarily reduce the variance of decision trees, making predictions more stable and improving generalization.&#8221;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>The Wisdom of Crowds</h2><p>Suppose you walk into a room and show 1,000 people a jar filled with jellybeans.</p><p>You ask each person: <em>&#8220;How many jellybeans are in the jar?&#8221;</em></p><p>Most individual guesses will be wrong.</p><p>Some people will underestimate.<br>Some will overestimate.</p><p>But something interesting happens when you average all the guesses together. The average estimate is often surprisingly close to the true answer. </p><p>This phenomenon is known as the <strong>wisdom of crowds</strong>.</p><p>The idea is simple:</p><p><em><strong>A collection of imperfect opinions can often be more accurate than any single opinion.</strong></em></p><p>Random Forests apply this exact principle to machine learning.</p><p>Instead of relying on a single decision tree, Random Forests build many decision trees and combine their predictions. Each individual tree:</p><ul><li><p>Sees a slightly different version of the data</p></li><li><p>Learns slightly different patterns</p></li><li><p>Makes slightly different mistakes</p></li></ul><p>While some trees may make poor predictions on certain observations, those mistakes tend to cancel out when we aggregate predictions across hundreds of trees. As a result, the combined model is often:</p><ul><li><p>More stable</p></li><li><p>More accurate</p></li><li><p>Less prone to overfitting</p></li></ul><p>This process of building many models and combining their predictions is called <strong>ensemble learning</strong>.</p><p>Random Forests are one of the most successful examples of ensemble learning, and they rely on a specific technique called <strong>bagging</strong>, which we&#8217;ll explore next.</p><p>&#128204; <em><strong>Pro Tip: </strong></em>A strong interview explanation is - </p><p><em>&#8220;Random Forests work because averaging many diverse decision trees reduces variance and produces more stable predictions.&#8221;</em></p><p>That&#8217;s often more valuable than diving straight into implementation details.</p><div><hr></div><h2>What Is Bagging?</h2><p>Bagging stands for <strong>Bootstrap Aggregating</strong>.</p><p>It is an ensemble learning technique that improves model performance by combining the predictions of multiple models instead of relying on a single model.</p><p>The idea is simple:</p><p><em><strong>Rather than training one decision tree, we train many decision trees and aggregate their predictions.</strong></em></p><p>The process works as follows:</p><ol><li><p>Create multiple versions of the training dataset</p></li><li><p>Train a separate decision tree on each dataset</p></li><li><p>Generate predictions from all trees</p></li><li><p>Combine the predictions into a final result</p></li></ol><p><strong>For classification problems:</strong></p><ul><li><p>Each tree votes for a class</p></li><li><p>The majority vote becomes the final prediction</p></li></ul><p><strong>For regression problems:</strong></p><ul><li><p>Each tree produces a numerical prediction</p></li><li><p>The average becomes the final prediction</p></li></ul><h4>The Key Requirement for Bagging</h4><p>For bagging to be effective, the models in the ensemble cannot all be identical.</p><p>Imagine training 100 decision trees on exactly the same dataset using the same settings. You would end up with 100 identical trees making the same predictions. In that case, combining their predictions provides no benefit.</p><p>For bagging to work, each tree must learn something slightly different. This creates diversity within the ensemble.</p><p>The challenge then becomes: <em><strong>How do we create many different trees while starting from the same dataset?</strong></em></p><p>The solution is <strong>bootstrap sampling</strong>.</p><p>Instead of training every tree on the entire dataset, we train each tree on a different randomly generated sample of the data.</p><p>As a result:</p><ul><li><p>Each tree sees slightly different observations</p></li><li><p>Each tree learns slightly different patterns</p></li><li><p>Each tree produces slightly different predictions</p></li></ul><p>This diversity is one of the key reasons bagging works so well in practice.</p><p>We&#8217;ll see exactly how bootstrap sampling creates these different datasets in the next section.</p><p>&#128204; <em><strong>Pro Tip: </strong></em>A common interview question is:</p><p><em>&#8220;Why doesn&#8217;t training multiple trees on the same dataset improve performance?&#8221;</em></p><p>A strong answer is:</p><p><em>&#8220;Because the trees would be nearly identical. Bagging works because each model is trained on a different bootstrap sample, creating diversity within the ensemble.&#8221;</em></p><div><hr></div><h2>Bootstrap Sampling</h2><p>In the previous section, we saw that bagging requires multiple decision trees that are slightly different from one another.</p><p>This raises an important question:</p><p><em><strong>How do we create different training datasets when we only have one dataset?</strong></em></p><p>The answer is <strong>bootstrap sampling</strong>.</p><p>Bootstrap sampling is a technique where we repeatedly sample observations from the original dataset <strong>with replacement</strong>.</p><p>The phrase <em>with replacement</em> is important.</p><p>It means that after selecting an observation, we put it back into the dataset before drawing the next observation. As a result, the same observation can be selected multiple times. For example, suppose our original dataset contains five observations:</p><pre><code><code>A B C D E</code></code></pre><p>One bootstrap sample might look like:</p><pre><code><code>A B B D E</code></code></pre><p>Another bootstrap sample might be:</p><pre><code><code>A C D D E</code></code></pre><p>Notice that:</p><ul><li><p>Some observations appear multiple times</p></li><li><p>Some observations are missing entirely</p></li></ul><p>Even though each bootstrap sample has the same size as the original dataset, the composition is slightly different. This is exactly what we want.</p><p>By training each decision tree on a different bootstrap sample:</p><ul><li><p>Every tree sees a slightly different dataset</p></li><li><p>Every tree learns slightly different patterns</p></li><li><p>Every tree makes slightly different predictions</p></li></ul><p>These differences create the diversity needed for bagging to be effective.</p><h4>Why Sample With Replacement?</h4><p>If we sampled without replacement, every bootstrap sample would contain exactly the same observations, just in a different order.</p><p>The resulting trees would be very similar.</p><p>Sampling with replacement introduces variability, which helps create diverse trees and ultimately improves the performance of the ensemble.</p><p>Although each bootstrap sample contains the same number of rows as the original dataset, it does not contain all of the original observations.</p><p>On average:</p><ul><li><p>About 63% of the original observations appear in a bootstrap sample</p></li><li><p>About 37% are left out</p></li></ul><p>These left-out observations are called <strong>Out-of-Bag (OOB) samples</strong>, which can later be used to estimate model performance without needing a separate validation dataset.</p><p>We&#8217;ll revisit OOB samples when discussing Random Forest evaluation.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>How Random Forests Work</h2><p>At this point, we have all the building blocks needed to understand Random Forests.</p><p>A Random Forest combines:</p><ul><li><p>Bootstrap sampling</p></li><li><p>Decision trees</p></li><li><p>Random feature selection</p></li><li><p>Aggregated predictions</p></li></ul><p>The training process is surprisingly straightforward.</p><p><strong>Step 1: Create Bootstrap Samples</strong></p><p>Multiple bootstrap samples are generated from the original dataset.</p><p>Each sample contains:</p><ul><li><p>The same number of observations as the original dataset</p></li><li><p>A slightly different combination of observations</p></li></ul><p>Every decision tree will be trained on a different bootstrap sample.</p><p><strong>Step 2: Train a Decision Tree on Each Sample</strong></p><p>A separate decision tree is trained on each bootstrap sample.</p><p>Because every tree sees a slightly different dataset:</p><ul><li><p>The trees learn different patterns</p></li><li><p>The trees make different mistakes</p></li><li><p>The trees produce different predictions</p></li></ul><p>This diversity is critical to the success of the ensemble.</p><p><strong>Step 3: Randomly Select Features at Each Split</strong></p><p>This is the step that makes Random Forests different from ordinary bagged trees.</p><p>Instead of considering all features when creating a split, each tree only evaluates a random subset of features.</p><p>For example, if a dataset contains 20 features:</p><ul><li><p>A tree may only consider 4 or 5 randomly selected features at a split</p></li><li><p>Another tree may consider a different subset</p></li></ul><p>This introduces additional diversity into the forest.</p><p>We&#8217;ll explore why this matters in the next section.</p><p><strong>Step 4: Aggregate Predictions</strong></p><p>Once all trees have been trained, their predictions are combined.</p><p>For classification:</p><ul><li><p>Each tree votes for a class</p></li><li><p>The majority vote becomes the final prediction</p></li></ul><p>For regression:</p><ul><li><p>Each tree produces a numerical prediction</p></li><li><p>The average prediction becomes the final output</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iJHz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iJHz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1347072,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/200073882?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iJHz!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3dc17062-dfbf-4858-a3c3-40c440d72301_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Putting It All Together, random Forests do not rely on a single &#8220;best&#8221; tree.</p><p>Instead, they rely on the collective intelligence of many trees trained on slightly different data and slightly different features.</p><p>This combination produces a model that is often:</p><ul><li><p>More accurate</p></li><li><p>More stable</p></li><li><p>Less prone to overfitting</p></li></ul><p>than any individual decision tree.</p><p>&#128204; <em><strong>Pro Tip: </strong>Random Forests combine bootstrap sampling, decision trees, random feature selection, and prediction aggregation to create a low-variance ensemble model.</em></p><div><hr></div><h2>Why Random Feature Selection Matters</h2><p>So far, we&#8217;ve seen that Random Forests use bootstrap sampling to train each tree on a different version of the data. But Random Forests add one more layer of randomness.</p><p>When a decision tree creates a split, it normally evaluates <strong>all available features</strong> and chooses the best one. For example, if a dataset contains 20 features, a traditional decision tree considers all 20 features at every split.</p><p>Random Forests work differently. At each split, the algorithm randomly selects a subset of features and only evaluates those features when deciding how to split the data.</p><p>For example:</p><ul><li><p>Dataset contains 20 features</p></li><li><p>Random Forest randomly selects 4 features</p></li><li><p>The tree chooses the best split from those 4 features only</p></li></ul><p>The next split may use a completely different subset of features.</p><h4>Why Is This Necessary?</h4><p>Imagine a dataset where one feature is extremely predictive.</p><p>If every tree always has access to all features, most trees will repeatedly choose the same feature near the top of the tree. As a result:</p><ul><li><p>Trees become very similar</p></li><li><p>Trees make similar predictions</p></li><li><p>Trees make similar mistakes</p></li></ul><p>In other words, the forest loses diversity.</p><p>Random feature selection helps prevent this problem. By forcing trees to consider different subsets of features:</p><ul><li><p>Trees become more diverse</p></li><li><p>Correlation between trees decreases</p></li><li><p>The ensemble becomes more effective</p></li></ul><h4>Bagged Trees vs Random Forests</h4><p>This distinction is important:</p><p><strong>Bagged Trees</strong></p><ul><li><p>Use bootstrap sampling</p></li><li><p>Train multiple trees</p></li><li><p>Consider all features at each split</p></li></ul><p><strong>Random Forests</strong></p><ul><li><p>Use bootstrap sampling</p></li><li><p>Train multiple trees</p></li><li><p>Randomly select features at each split</p></li></ul><p>This additional randomness is what makes Random Forests significantly more powerful than ordinary bagged trees.</p><p>A Random Forest does not try to build the best individual tree. Instead, it tries to build many strong but diverse trees. The combination of bootstrap sampling and random feature selection creates the diversity needed for the ensemble to outperform any single tree.</p><p>&#128204; <em><strong>Pro Tip: </strong></em>A very common interview question is:</p><p><em>&#8220;What is the difference between bagging and Random Forests?&#8221;</em></p><p>A strong answer is:</p><p><em>&#8220;Random Forests use bagging, but they also introduce random feature selection at each split. This reduces correlation between trees and improves ensemble performance.&#8221;</em></p><div><hr></div><h2>Advantages and Limitations of Random Forests</h2><p>Random Forests became one of the most popular machine learning algorithms because they combine strong predictive performance with relatively simple training procedures.</p><p>They often perform well out of the box and require less tuning than many other machine learning models.</p><h4>Advantages</h4><p>Random Forests offer several important benefits:</p><ul><li><p>Reduce overfitting compared to a single decision tree</p></li><li><p>Lower variance through averaging</p></li><li><p>Capture complex non-linear relationships</p></li><li><p>Handle feature interactions automatically</p></li><li><p>Require little preprocessing</p></li><li><p>Do not require feature scaling</p></li><li><p>Work well on both classification and regression problems</p></li></ul><p>Because Random Forests combine many trees, they are typically more stable and reliable than individual decision trees.</p><p>This is one reason they are often used as a strong baseline model in machine learning projects.</p><h4>Limitations</h4><p>Despite their strengths, Random Forests are not perfect.</p><p>Some common limitations include:</p><ul><li><p>Less interpretable than a single decision tree</p></li><li><p>Larger memory footprint</p></li><li><p>Slower training and inference than simpler models</p></li><li><p>Can become difficult to explain to stakeholders</p></li><li><p>May struggle when extrapolation beyond the training data is required</p></li></ul><p>For example, while a single decision tree can be visualized and explained relatively easily, understanding the behavior of hundreds of trees simultaneously is much more challenging.</p><h4>When Should You Use Random Forests?</h4><p>Random Forests are often a strong choice when:</p><ul><li><p>You are working with structured/tabular data</p></li><li><p>Relationships are non-linear</p></li><li><p>Predictive performance is more important than model simplicity</p></li><li><p>You want a strong baseline before exploring more advanced models</p></li></ul><p>In practice, Random Forests are commonly used for:</p><ul><li><p>Fraud detection</p></li><li><p>Customer churn prediction</p></li><li><p>Credit risk modeling</p></li><li><p>Marketing response modeling</p></li><li><p>Demand forecasting</p></li></ul><p>Random Forests are often described as a &#8220;safe default&#8221; machine learning model.</p><p>They may not always be the best-performing algorithm, but they frequently provide strong results with relatively little effort.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Decision Tree vs Random Forest</h2><p>Decision Trees and Random Forests are closely related, but they solve very different problems.</p><p>A Decision Tree focuses on building a single model that learns patterns from the data. Random Forests take a different approach by combining many decision trees to produce a more stable and accurate prediction.</p><p>The table below summarizes the key differences:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!d4gJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!d4gJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1343443,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/200073882?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!d4gJ!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0effade6-d9a0-4d6f-91df-54255b143090_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Which One Should You Use?</h4><p>If interpretability is your primary goal, a Decision Tree may be the better choice. Since the entire model can be visualized, it is often easier to explain predictions to stakeholders.</p><p>If predictive performance is the priority, Random Forests are usually the stronger option. By averaging many diverse trees, they reduce overfitting and often generalize better to unseen data.</p><p>Random Forests do not replace Decision Trees. Instead, they build upon them.</p><p>You can think of a Random Forest as: <em>Many Decision Trees working together to produce a more reliable prediction.</em></p><p>This is one of the most important ideas in ensemble learning.</p><div><hr></div><h2>Rapid Fire Interview Questions</h2><ol><li><p><em><strong>What problem do Random Forests solve?</strong></em></p></li></ol><p>Random Forests primarily reduce the variance of decision trees.</p><p>By combining predictions from many diverse trees, they produce more stable predictions and improve generalization on unseen data.</p><ol start="2"><li><p><em><strong>What is bagging?</strong></em></p></li></ol><p>Bagging (Bootstrap Aggregating) is an ensemble learning technique that trains multiple models on different bootstrap samples of the data and combines their predictions.</p><p>Its primary goal is to reduce variance.</p><ol start="3"><li><p><em><strong>What is bootstrap sampling?</strong></em></p></li></ol><p>Bootstrap sampling is the process of randomly sampling observations <strong>with replacement</strong> from the original dataset.</p><p>Because sampling is performed with replacement:</p><ul><li><p>Some observations appear multiple times</p></li><li><p>Some observations are omitted</p></li></ul><p>This creates different training datasets for different trees.</p><ol start="4"><li><p><em><strong>Why does Random Forest use random feature selection?</strong></em></p></li></ol><p>If every tree always considered all features, many trees would look very similar.</p><p>Random feature selection:</p><ul><li><p>Increases diversity</p></li><li><p>Reduces correlation between trees</p></li><li><p>Improves ensemble performance</p></li></ul><ol start="5"><li><p><em><strong>What is the difference between bagging and Random Forests?</strong></em></p></li></ol><p>Bagging:</p><ul><li><p>Uses bootstrap sampling</p></li><li><p>Trains multiple trees</p></li><li><p>Uses all features at each split</p></li></ul><p>Random Forest:</p><ul><li><p>Uses bootstrap sampling</p></li><li><p>Trains multiple trees</p></li><li><p>Uses a random subset of features at each split</p></li></ul><p>You can think of Random Forests as:</p><p><em><strong>Random Forest = Bagging + Random Feature Selection</strong></em></p><ol start="6"><li><p><em><strong>Does Random Forest reduce bias or variance?</strong></em></p></li></ol><p>Random Forests primarily reduce variance.</p><p>Individual decision trees tend to have:</p><ul><li><p>Low bias</p></li><li><p>High variance</p></li></ul><p>Averaging many trees reduces variance while maintaining the flexibility of the underlying trees.</p><ol start="7"><li><p><em><strong>Do Random Forests require feature scaling?</strong></em></p></li></ol><p>No.</p><p>Random Forests split data based on feature thresholds rather than distances, so feature scaling typically has little impact on performance.</p><ol start="8"><li><p><em><strong>Why are Random Forests less interpretable than Decision Trees?</strong></em></p></li></ol><p>A Decision Tree consists of a single set of decision rules that can be visualized and explained.</p><p>A Random Forest may contain hundreds of trees, making it much harder to understand the exact reasoning behind an individual prediction.</p><ol start="9"><li><p><em><strong>What are Out-of-Bag (OOB) samples?</strong></em></p></li></ol><p>Out-of-Bag samples are observations that were not selected in a tree&#8217;s bootstrap sample.</p><p>These observations can be used to estimate model performance without requiring a separate validation dataset.</p><ol start="10"><li><p><em><strong>Why does a Random Forest usually outperform a Decision Tree?</strong></em></p></li></ol><p>A Random Forest reduces variance by averaging predictions from many diverse decision trees, which makes the model more stable and less prone to overfitting.</p><div><hr></div><h2>Final Takeaway</h2><p>Random Forests demonstrate one of the most powerful ideas in machine learning:</p><blockquote><p>Combining many models often works better than relying on a single model.</p></blockquote><p>By using bootstrap sampling, random feature selection, and prediction aggregation, Random Forests transform unstable decision trees into models that are more accurate, more stable, and less prone to overfitting.</p><p>If there&#8217;s one idea to remember, it&#8217;s this:</p><blockquote><p>Decision Trees are powerful. Random Forests make them reliable.</p></blockquote><div><hr></div><h2>What&#8217;s Next?</h2><p>I&#8217;m currently working on something a little different for the next issue of <em>The Practical Data Scientist</em>&#8212;potentially a collaboration with another creator in the data science space.</p><p>Nothing is finalized yet, but if it comes together, I think you&#8217;ll enjoy it.</p><p><em><strong>Until then, keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p><h4></h4>]]></content:encoded></item><item><title><![CDATA[Introduction to Tree-Based Models]]></title><description><![CDATA[A practical guide for data science interviews]]></description><link>https://thepracticaldatascientist.substack.com/p/introduction-to-tree-based-models</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/introduction-to-tree-based-models</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 26 May 2026 15:27:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YIuI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YIuI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YIuI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1739120,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/199275782?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YIuI!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82560de3-b330-482a-a0c9-92137fd03814_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the previous issues of <em>The Practical Data Scientist</em>, we explored:</p><ul><li><p>Logistic Regression</p></li><li><p>Evaluation Metrics</p></li><li><p>Model performance trade-offs</p></li></ul><p>Those topics form the foundation of machine learning interviews.</p><p>But many of the most powerful modern machine learning models are not linear models.</p><p>They are tree-based models.</p><p>Models like Random Forests, XGBoost, LightGBM and CatBoost have become some of the most widely used algorithms in practical machine learning because they:</p><ul><li><p>Handle non-linear relationships well</p></li><li><p>Capture feature interactions automatically</p></li><li><p>Require less feature engineering</p></li><li><p>Perform extremely well on structured/tabular data</p></li></ul><p>At the core of all these models is one simple idea: <em><strong>Decision Trees</strong></em></p><p>Understanding decision trees deeply makes it much easier to understand modern ensemble methods later.</p><p>In this issue, we&#8217;ll build that foundation by covering:</p><ul><li><p>How decision trees work</p></li><li><p>Entropy and information gain</p></li><li><p>Gini impurity</p></li><li><p>Tree splitting logic</p></li><li><p>Overfitting and pruning</p></li><li><p>Feature importance</p></li><li><p>Advantages and limitations of tree-based models</p></li><li><p>Common interview questions and pro tips</p></li></ul><p>As always, the focus will be on intuition, practical understanding, and interview-oriented thinking rather than memorizing formulas.</p><p><em><strong>&#128204; Pro Tip: </strong></em>Interviewers often use decision trees to test whether candidates truly understand machine learning intuition &#8212; not just equations and definitions.</p><div><hr></div><h2>Why Tree-Based Models Became Popular</h2><p>Tree-based models became extremely popular because they solve many practical problems that traditional linear models struggle with.</p><p>Unlike models such as logistic regression, tree-based models can naturally capture:</p><ul><li><p>Non-linear relationships</p></li><li><p>Complex feature interactions</p></li><li><p>Different behaviors across segments of data</p></li></ul><p>For example:</p><ul><li><p>The relationship between income and loan default may behave differently for different age groups</p></li><li><p>Customer churn patterns may depend on combinations of features rather than single variables</p></li></ul><p>Tree-based models handle these patterns automatically without requiring heavy feature engineering.</p><h4>Key Advantages</h4><p><strong>Handles Non-Linear Relationships: </strong>Decision trees do not assume a linear relationship between features and the target. This makes them flexible for many real-world problems.</p><p><strong>Captures Feature Interactions Automatically: </strong>Linear models often require manually creating interaction features. Tree-based models naturally capture interactions through splits. For example:</p><ul><li><p>A tree may first split on age</p></li><li><p>Then create different income rules within each age group</p></li></ul><p><strong>Requires Less Preprocessing: </strong>Tree-based models usually - </p><ul><li><p>Do not require feature scaling</p></li><li><p>Are less sensitive to outliers</p></li><li><p>Can handle numerical and categorical features</p></li></ul><p>This makes them very practical in real-world workflows.</p><p><strong>Strong Performance on Tabular Data: </strong>Tree-based ensemble models like Random Forests, XGBoost, LightGBM often dominate structured/tabular machine learning problems. This is one reason they appear frequently in:</p><ul><li><p>Kaggle competitions</p></li><li><p>Production ML systems</p></li><li><p>Data science interviews</p></li></ul><p>Compared to logistic regression:</p><ul><li><p>Logistic regression creates linear decision boundaries</p></li><li><p>Tree-based models create flexible non-linear decision boundaries</p></li></ul><p>This allows trees to model more complex patterns with less manual feature engineering.</p><p><em><strong>&#128204; Pro Tip: </strong></em>A strong interview answer is &#8220;Tree-based models became popular because they combine strong predictive performance with relatively low preprocessing requirements on structured data.&#8221;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>What Is a Decision Tree?</h2><p>A decision tree is a machine learning model that makes predictions by repeatedly splitting the data into smaller groups based on feature values.</p><p>The structure resembles a flowchart:</p><ul><li><p>Each split represents a decision rule</p></li><li><p>Each branch represents an outcome of that rule</p></li><li><p>Each final node (leaf node) represents a prediction</p></li></ul><p>The goal of a decision tree is simple: <em><strong>Split the data in a way that creates increasingly &#8220;pure&#8221; groups.</strong></em></p><p>For example, imagine building a model to predict customer churn. The tree might learn rules like:</p><ul><li><p>Customers with low engagement are more likely to churn</p></li><li><p>Among those users, customers with many support tickets churn even more frequently</p></li></ul><p>The tree keeps splitting the data step-by-step until the groups become sufficiently homogeneous.</p><h4>Components of a Decision Tree</h4><p><strong>Root Node: </strong>The first split of the tree. This usually represents the feature that best separates the data initially.</p><p><strong>Internal Nodes: </strong>Intermediate decision points where the data is split further. Example:</p><p><em>&#8220;Is account age &gt; 12 months?&#8221;</em></p><p><strong>Branches: </strong>The possible outcomes of a split. For example:</p><ul><li><p>Yes branch</p></li><li><p>No branch</p></li></ul><p><strong>Leaf Nodes: </strong>The final prediction outputs. For classification problems:</p><ul><li><p>Class 0</p></li><li><p>Class 1</p></li></ul><p>For regression problems:</p><ul><li><p>Numerical prediction values</p></li></ul><h4>Recursive Partitioning</h4><p>Decision trees use a process called recursive partitioning:</p><ol><li><p>Find the best split</p></li><li><p>Split the data</p></li><li><p>Repeat the process on each subgroup</p></li></ol><p>This continues until:</p><ul><li><p>The data becomes sufficiently pure</p></li><li><p>Or stopping criteria are reached</p></li></ul><h4>Why Trees Feel Intuitive</h4><p>One reason decision trees are popular is because they mimic human decision-making. For example:</p><ul><li><p>&#8220;If income is low and debt is high &#8594; higher default risk&#8221;</p></li><li><p>&#8220;If engagement is low and complaints are high &#8594; higher churn risk&#8221;</p></li></ul><p>This makes trees highly interpretable compared to many other machine learning models.</p><p><em><strong>&#128204;  Pro Tip: </strong></em>A strong interview explanation is:</p><p><em>&#8220;Decision trees recursively split the data to create increasingly pure groups.&#8221;</em></p><p>That demonstrates deeper understanding than simply saying:</p><p><em>&#8220;Trees make decisions using branches.&#8221;</em></p><div><hr></div><h2>How Decision Trees Learn</h2><p>Now that we understand what a decision tree is, the next question becomes:</p><p><em>How does the tree decide where to split?</em></p><p>At every step, the tree evaluates different possible splits and chooses the one that creates the &#8220;purest&#8221; groups. For example, in a churn prediction problem, the model may test rules such as:</p><ul><li><p>Is engagement score &lt; 20?</p></li><li><p>Is account age &gt; 12 months?</p></li><li><p>Is number of support tickets &gt; 5?</p></li></ul><p>The goal is to reduce uncertainty and create groups where most observations belong to a single class. Decision trees learn using a recursive process:</p><ul><li><p>Find the best split at the current node</p></li><li><p>Split the data into smaller groups</p></li><li><p>Repeat the process on each subgroup</p></li></ul><p>This continues until the groups become sufficiently pure or the model reaches stopping criteria. A &#8220;good&#8221; split is one that:</p><ul><li><p>Separates the classes clearly</p></li><li><p>Reduces impurity within each group</p></li></ul><p>For example, a node that originally contains a 50/50 mix of churners and non-churners may become much cleaner after splitting:</p><ul><li><p>One subgroup may contain mostly churners</p></li><li><p>Another subgroup may contain mostly non-churners</p></li></ul><p>Decision trees use a greedy learning approach. At each step, the model chooses the best split available at that moment instead of searching for the globally optimal tree. This makes trees computationally efficient, but it can sometimes lead to suboptimal solutions.</p><p>To measure split quality, decision trees use impurity metrics such as entropy, information gain, and Gini impurity. These concepts help quantify how &#8220;pure&#8221; or &#8220;mixed&#8221; a node is, which we&#8217;ll explore next.</p><p><em><strong>&#128204; Pro Tip: </strong></em>A strong interview answer is</p><p><em>&#8220;Decision trees recursively choose splits that maximize purity improvement at each step.&#8221;</em></p><div><hr></div><h2>Entropy and Information Gain</h2><p>To decide the best place to split the data, decision trees need a way to measure how &#8220;pure&#8221; or &#8220;mixed&#8221; a node is.</p><p>One common measure is called <strong>entropy</strong>.</p><p>Entropy measures the amount of uncertainty or randomness in a node.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;H(S) = - \\sum p_i \\log_2(p_i)&quot;,&quot;id&quot;:&quot;SCMDEBCEEI&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where:</p><ul><li><p><em>p_i</em> = proportion (probability) of class iii in the node</p></li></ul><p>For example:</p><ul><li><p>If 70% of observations belong to Class 1, then p_1=0.7</p></li></ul><h4>Intuition Behind Entropy</h4><ul><li><p>High entropy &#8594; highly mixed classes</p></li><li><p>Low entropy &#8594; mostly one class</p></li><li><p>Zero entropy &#8594; completely pure node</p></li></ul><p>For example:</p><ul><li><p>A node with 50% churners and 50% non-churners has high entropy</p></li><li><p>A node with 95% churners has much lower entropy</p></li></ul><p>The goal of a decision tree is to reduce entropy after every split.</p><h4>Information Gain</h4><p>Decision trees use a metric called <strong>Information Gain</strong> to evaluate how useful a split is.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;IG = H(Parent) - \\sum w_i H(Child_i)&quot;,&quot;id&quot;:&quot;ENYUXGOCFW&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where:</p><ul><li><p><em>w_i</em> = proportion of samples that go to child node i</p></li><li><p>H(<em>Child_i</em>) = entropy of child node i</p></li></ul><p>For example:</p><ul><li><p>If 30% of observations go left and 70% go right:</p><ul><li><p>w1=0.3</p></li><li><p>w2=0.7</p></li></ul></li></ul><p>Information Gain measures how much uncertainty is reduced after splitting the data.</p><p>A split with high information gain creates cleaner and more homogeneous groups. For example:</p><ul><li><p>Before splitting, a node may contain a mixed set of churners and non-churners</p></li><li><p>After splitting, the groups may become much more separated</p></li></ul><p>The tree chooses the split that produces the highest information gain.</p><p>At every step, the decision tree is essentially asking: <em>&#8220;Which split gives me the largest reduction in uncertainty?&#8221;</em></p><p>This process repeats recursively until the tree stops growing.</p><p><em><strong>&#128204;  Pro Tip: </strong></em>A strong interview explanation is:</p><p><em>&#8220;Entropy measures impurity, and information gain measures how much a split reduces that impurity.&#8221;</em></p><p>That demonstrates much stronger intuition than simply memorizing formulas.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Gini Impurity</h2><p>Another common way decision trees measure impurity is through <strong>Gini Impurity</strong>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Gini = 1 - \\sum p_i^2&quot;,&quot;id&quot;:&quot;FWIKYXUCYH&quot;}" data-component-name="LatexBlockToDOM"></div><p>where:</p><ul><li><p><em>p_i</em> = proportion of class (i) in the node</p></li></ul><p>Like entropy, Gini Impurity measures how mixed the classes are within a node.</p><h4>Intuition Behind Gini Impurity</h4><ul><li><p>Low Gini &#8594; purer node</p></li><li><p>High Gini &#8594; more mixed classes</p></li><li><p>Gini = 0 &#8594; completely pure node</p></li></ul><p>For example:</p><ul><li><p>A node containing only churners has Gini = 0</p></li><li><p>A node with a 50/50 class split has high impurity</p></li></ul><p>The decision tree tries to find splits that reduce Gini Impurity as much as possible.</p><h4>Gini vs Entropy</h4><p>Both entropy and Gini are impurity measures used to evaluate splits. In practice:</p><ul><li><p>They often produce very similar trees</p></li><li><p>Both aim to create purer groups after splitting</p></li></ul><p>Main differences:</p><ul><li><p>Gini is computationally simpler and slightly faster</p></li><li><p>Entropy comes from information theory and is sometimes considered more interpretable mathematically</p></li></ul><p>Most modern decision tree implementations are based on CART (Classification and Regression Trees), which typically use Gini Impurity by default.</p><p>The exact impurity metric is usually less important than the core intuition: <em>Decision trees learn by repeatedly creating cleaner and more homogeneous groups.</em></p><p><em><strong>&#128204;  Pro Tip: </strong></em>If an interviewer asks:</p><p><em>&#8220;What&#8217;s the difference between entropy and Gini?&#8221;</em></p><p>you usually do not need a deeply mathematical answer. A strong practical answer is:</p><p><em>&#8220;Both measure impurity. Gini is computationally simpler and often used in practice, while entropy is based on information theory.&#8221;</em></p><div><hr></div><h2>Overfitting in Decision Trees</h2><p>One of the biggest challenges with decision trees is overfitting.</p><p>Because trees can continue splitting the data repeatedly, they may eventually start memorizing the training data instead of learning generalizable patterns.</p><p>An overfit tree often:</p><ul><li><p>Achieves extremely high training accuracy</p></li><li><p>Performs poorly on unseen data</p></li><li><p>Learns noise instead of meaningful relationships</p></li></ul><p>This usually happens when the tree becomes very deep and creates overly specific decision rules.</p><p>For example, instead of learning broad churn behavior, the tree may begin memorizing very small customer segments that only exist in the training dataset.</p><p>Decision trees are especially prone to overfitting because they are highly flexible models. They can continue splitting until nodes become almost completely pure, even if those splits are not useful for generalization.</p><p>This leads to a high-variance problem:</p><ul><li><p>Small changes in training data can produce very different trees</p></li><li><p>Different splits may lead to entirely different model structures</p></li><li><p>Predictions can become unstable</p></li></ul><p>From a bias-variance perspective:</p><ul><li><p>Deep trees &#8594; low bias, high variance</p></li><li><p>Shallow trees &#8594; higher bias, lower variance</p></li></ul><p>The challenge is finding the right balance between:</p><ul><li><p>Learning meaningful patterns</p></li><li><p>Avoiding memorization of noise</p></li></ul><p>A very important machine learning insight is:</p><p><em>A perfectly accurate training model is not necessarily a good model.</em></p><p>What matters most is how well the model performs on unseen data.</p><p><em><strong>&#128204;  Pro Tip: </strong></em>A strong interview answer is:</p><p><em>&#8220;Decision trees overfit because they can continue splitting until they memorize the training data, which increases variance and hurts generalization.&#8221;</em></p><div><hr></div><h2>Controlling Tree Complexity</h2><p>Since decision trees can easily overfit, we need ways to control how complex the tree becomes.</p><p>Instead of allowing the tree to grow indefinitely, we can introduce constraints that help the model generalize better to unseen data. Some common techniques include:</p><ul><li><p>Limiting the maximum depth of the tree</p></li><li><p>Requiring a minimum number of samples before splitting</p></li><li><p>Requiring a minimum number of samples in leaf nodes</p></li><li><p>Pruning unnecessary branches</p></li></ul><p><strong>Max Depth: </strong>Max depth limits how deep the tree can grow. A shallow tree:</p><ul><li><p>Simpler model</p></li><li><p>Lower variance</p></li><li><p>Higher bias</p></li></ul><p>A very deep tree:</p><ul><li><p>More flexible</p></li><li><p>Higher variance</p></li><li><p>Greater risk of overfitting</p></li></ul><p>Choosing the right depth is part of balancing the bias-variance tradeoff.</p><p><strong>Minimum Samples for Split: </strong>This parameter requires a minimum number of observations before creating another split. Without this restriction, the tree may create highly specific branches for very small groups of data.</p><p><strong>Minimum Samples per Leaf: </strong>This controls the minimum number of observations allowed in a final leaf node. Larger leaf sizes usually:</p><ul><li><p>Produce smoother decision boundaries</p></li><li><p>Reduce overfitting</p></li><li><p>Improve stability</p></li></ul><p><strong>Pruning: </strong>Pruning removes branches that add little predictive value. The idea is simple:</p><p><em>Not every split improves generalization.</em></p><p>By removing weak or unnecessary branches, pruning helps create simpler and more robust trees.</p><p>A more complex tree is not always a better tree.</p><p>The goal is not to perfectly memorize the training data, but to learn patterns that generalize well to new data.</p><p><em><strong>&#128204;   Pro Tip: </strong></em>A strong interview answer is:</p><p><em>&#8220;We control tree complexity to reduce overfitting and improve generalization on unseen data.&#8221;</em></p><div><hr></div><h2>Advantages and Limitations of Tree-Based Models</h2><p>Tree-based models became popular because they are powerful, flexible, and relatively easy to use on structured data.</p><p>One of their biggest strengths is their ability to model complex patterns without requiring heavy preprocessing or feature engineering.</p><h4>Advantages</h4><p>Tree-based models:</p><ul><li><p>Handle non-linear relationships naturally</p></li><li><p>Capture feature interactions automatically</p></li><li><p>Require little preprocessing</p></li><li><p>Do not require feature scaling</p></li><li><p>Work well with both numerical and categorical features</p></li><li><p>Are relatively easy to interpret</p></li></ul><p>Unlike linear models, decision trees can create flexible decision boundaries and adapt to different segments of the data automatically. For example:</p><ul><li><p>Customer churn behavior may differ across age groups</p></li><li><p>Fraud patterns may depend on combinations of features rather than individual variables</p></li></ul><p>Tree-based models can learn these relationships directly from the data.</p><p>Another major advantage is interpretability. Since the model makes decisions through sequential splits, it is often easier to explain predictions compared to more complex machine learning models.</p><h4>Limitations</h4><p>Despite their strengths, tree-based models also have important limitations.</p><p>Decision trees:</p><ul><li><p>Can overfit easily</p></li><li><p>Are sensitive to small changes in data</p></li><li><p>Often have high variance</p></li><li><p>Use greedy splitting strategies</p></li><li><p>May not generalize well individually</p></li></ul><p>For example, a small change in the training data can sometimes produce a very different tree structure.</p><p>This instability is one reason individual decision trees are often replaced by ensemble methods such as:</p><ul><li><p>Random Forests</p></li><li><p>Gradient Boosting</p></li><li><p>XGBoost</p></li></ul><p>These models combine many trees together to improve stability and predictive performance.</p><p><strong>Important Insight: </strong>Individual decision trees are usually not the final goal. They are the foundation for many of the most powerful machine learning models used in production today.</p><p><em><strong>&#128204;</strong></em> <em><strong>Pro Tip: </strong></em>A strong interview answer is:</p><p><em>&#8220;Decision trees are powerful because they handle non-linear relationships and interactions naturally, but individual trees can be unstable and prone to overfitting.&#8221;</em></p><div><hr></div><h2>Feature Importance and Interpretability</h2><p>One reason tree-based models are popular is that they are often easier to interpret than many other machine learning models.</p><p>As the tree makes splits, it learns which features are most useful for separating the data. Features that:</p><ul><li><p>Reduce impurity significantly</p></li><li><p>Appear near the top of the tree</p></li><li><p>Are used frequently across splits</p></li></ul><p>are often considered more important.</p><h4>How Feature Importance Works</h4><p>In decision trees, feature importance is typically calculated using impurity reduction.</p><p>The intuition is simple: <em>Features that create better splits contribute more to the model.</em></p><p>For example:</p><ul><li><p>If &#8220;engagement score&#8221; consistently creates strong splits in a churn model</p></li><li><p>The model may assign high importance to that feature</p></li></ul><h4>Why Interpretability Matters</h4><p>Interpretability is especially valuable in domains where understanding model decisions is important. Examples include:</p><ul><li><p>Healthcare</p></li><li><p>Finance</p></li><li><p>Risk modeling</p></li><li><p>Marketing</p></li></ul><p>Unlike many black-box models, decision trees allow us to trace:</p><ul><li><p>Which features influenced a prediction</p></li><li><p>What decision path the model followed</p></li></ul><p>This makes tree-based models easier to explain to non-technical stakeholders.</p><h4>Important Limitation</h4><p>Feature importance should still be interpreted carefully.</p><p>Some importance methods can become biased toward:</p><ul><li><p>Features with many unique values</p></li><li><p>High-cardinality variables</p></li><li><p>Correlated features</p></li></ul><p>This means: <em>A feature with high importance is not always truly &#8220;causal&#8221; or inherently more valuable.</em></p><p><strong>Important Insight: </strong>Interpretability is one of the major reasons tree-based models remain widely used, even when more complex models exist.</p><p>In many real-world applications:</p><ul><li><p>Explainability matters</p></li><li><p>Stakeholder trust matters</p></li><li><p>Transparency matters</p></li></ul><p> <em><strong>&#128204; Pro Tip:  </strong>&#8220;Feature importance in trees is typically based on impurity reduction, but importance scores should still be interpreted carefully because they can be biased.&#8221;</em></p><div><hr></div><h2>Decision Trees vs Logistic Regression</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!sp8N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 424w, /__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 848w, /__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!sp8N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png" width="1456" height="1017" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1017,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1318367,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/199275782?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 424w, /__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 848w, /__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 1272w, /__u/substackcdn.com/image/fetch/$s_!sp8N!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f28fff9-0a78-4228-baa6-f929f7644de5_1501x1048.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Important Insight: </strong>Neither model is universally better. In practice:</p><ul><li><p>Logistic regression is often a strong interpretable baseline</p></li><li><p>Tree-based models are often better for capturing complex non-linear patterns</p></li></ul><p>Many real-world machine learning workflows start with logistic regression and later move toward tree-based ensemble methods if additional predictive power is needed.</p><p> <em><strong>&#128204; Pro Tip:  </strong>Logistic regression works well when relationships are relatively linear and interpretability is critical, while tree-based models are better at capturing complex non-linear interactions automatically.</em></p><div><hr></div><h2>Rapid Fire Interview Questions</h2><ol><li><p><strong>Why are decision trees considered non-linear models?</strong></p></li></ol><p>Decision trees create decision boundaries through sequential splits instead of fitting a single linear equation. This allows them to model:</p><ul><li><p>Complex patterns</p></li><li><p>Feature interactions</p></li><li><p>Segment-specific behavior</p></li></ul><p>A strong answer is:</p><p><em>&#8220;Trees create piecewise decision boundaries through recursive splits, which makes them non-linear.&#8221;</em></p><ol start="2"><li><p><strong>Why don&#8217;t decision trees require feature scaling?</strong></p></li></ol><p>Decision trees split based on feature thresholds rather than distances.</p><p>For example: <em>&#8220;Is income &gt; 50K?&#8221;</em></p><p>Since splits depend only on ordering, scaling usually does not affect the tree structure.</p><p>Distance-based models like KNN or K-Means require scaling much more than trees do.</p><ol start="3"><li><p><strong>Why are decision trees unstable?</strong></p></li></ol><p>Small changes in training data can lead to very different splits and tree structures. This happens because trees:</p><ul><li><p>Use greedy splitting</p></li><li><p>Depend heavily on early split decisions</p></li></ul><p>This instability contributes to high variance.</p><p>This is one of the major motivations behind Random Forests.</p><ol start="4"><li><p><strong>Entropy vs Gini &#8212; what&#8217;s the difference?</strong></p></li></ol><p>Both measure node impurity.</p><ul><li><p>Entropy is based on information theory</p></li><li><p>Gini is computationally simpler and commonly used in CART trees</p></li></ul><p>In practice:</p><ul><li><p>They often produce very similar trees</p></li></ul><p>You usually do not need a deeply mathematical explanation in interviews. Focus on intuition.</p><ol start="5"><li><p><strong>Why do decision trees overfit?</strong></p></li></ol><p>Decision trees can continue splitting until they memorize the training data. Deep trees:</p><ul><li><p>Low bias</p></li><li><p>High variance</p></li></ul><p>This hurts generalization on unseen data.</p><p><em>&#8220;Trees overfit because they can create highly specific rules for small groups of observations.&#8221;</em></p><ol start="6"><li><p><strong>What is pruning?</strong></p></li></ol><p>Pruning removes branches that contribute little predictive value. The goal is to:</p><ul><li><p>Reduce complexity</p></li><li><p>Improve generalization</p></li><li><p>Prevent overfitting</p></li></ul><p>Pruning is essentially a way of simplifying the model to improve performance on unseen data.</p><div><hr></div><h2>Final Takeaway</h2><p>Decision trees are one of the most important foundations in machine learning.</p><p>Individually, they are simple and intuitive models that help explain core concepts such as:</p><ul><li><p>Non-linear decision boundaries</p></li><li><p>Recursive splitting</p></li><li><p>Entropy and impurity</p></li><li><p>Overfitting and variance</p></li><li><p>Feature importance</p></li></ul><p>But their importance goes far beyond individual trees.</p><p>Modern machine learning systems often rely on powerful ensemble methods built on top of decision trees, including:</p><ul><li><p>Random Forests</p></li><li><p>Gradient Boosting</p></li><li><p>XGBoost</p></li><li><p>LightGBM</p></li></ul><p>Understanding decision trees deeply makes these advanced models much easier to learn later.</p><p>One of the biggest interview advantages of studying decision trees is that they force you to think about:</p><ul><li><p>Model behavior</p></li><li><p>Trade-offs</p></li><li><p>Generalization</p></li><li><p>Interpretability</p></li></ul><p>instead of simply memorizing algorithms.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll move from individual decision trees to ensemble methods by exploring Random Forests and Bagging. </p><p>We&#8217;ll cover how combining multiple trees improves performance, reduces overfitting, and forms the foundation of many modern machine learning systems.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Evaluation Metrics for Data Science Interviews]]></title><description><![CDATA[How to choose the right metric beyond accuracy]]></description><link>https://thepracticaldatascientist.substack.com/p/evaluation-metrics-for-data-science</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/evaluation-metrics-for-data-science</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 19 May 2026 14:00:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8R7Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8R7Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8R7Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1484917,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/197777965?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8R7Z!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6286fc38-99fb-4b3c-bee9-485d3434e9aa_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the previous issue of <em>The Practical Data Scientist</em>, we explored logistic regression &#8212; one of the most commonly discussed algorithms in data science interviews. But building a model is only half the problem.</p><p>The next challenge is understanding whether the model is actually good.</p><p>And surprisingly, this is where many candidates struggle in interviews.</p><p>A model with 99% accuracy can still completely fail in production. A model with lower accuracy might actually deliver far more business value. Understanding <em>why</em> this happens is what separates candidates who memorize machine learning from candidates who truly understand it.</p><p>In this issue, we&#8217;ll break down the most important evaluation metrics used in machine learning interviews:</p><ul><li><p>Accuracy</p></li><li><p>Precision</p></li><li><p>Recall</p></li><li><p>F1 Score</p></li><li><p>ROC-AUC</p></li><li><p>PR Curves</p></li><li><p>Threshold tuning</p></li></ul><p>More importantly, we&#8217;ll focus on the intuition behind these metrics, the trade-offs they represent, and how to connect them to real business problems during interviews.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>The Confusion Matrix</h2><p>Before discussing metrics like Precision, Recall, or ROC-AUC, we first need to understand the foundation behind all of them: the confusion matrix.</p><p>Almost every classification metric in machine learning is derived from four possible prediction outcomes:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!EoLB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 424w, /__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 848w, /__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!EoLB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png" width="1412" height="262" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:262,&quot;width&quot;:1412,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:41779,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/197777965?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 424w, /__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 848w, /__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EoLB!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f771ed4-123c-46ba-a95a-e83ff6ec6724_1412x262.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Here&#8217;s what each term means:</p><ul><li><p><strong>True Positive (TP):</strong><br>The model correctly predicts the positive class.</p></li><li><p><strong>True Negative (TN):</strong><br>The model correctly predicts the negative class.</p></li><li><p><strong>False Positive (FP):</strong><br>The model predicts positive when the actual class is negative.</p></li><li><p><strong>False Negative (FN):</strong><br>The model predicts negative when the actual class is positive.</p></li></ul><p>Most evaluation metrics are simply different ways of measuring the trade-offs between these four quantities.</p><p>For example:</p><ul><li><p>Precision focuses on False Positives</p></li><li><p>Recall focuses on False Negatives</p></li><li><p>Accuracy combines all four values</p></li></ul><p>Understanding the confusion matrix deeply makes evaluation metrics much easier to understand intuitively.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>Many candidates memorize metric formulas without understanding what the errors actually represent.</p><p>Strong interview answers usually focus on:</p><ul><li><p>What kind of mistakes the model is making</p></li><li><p>Which mistakes are more costly for the business</p></li><li><p>How the metric captures those trade-offs</p></li></ul><div><hr></div><h2>Accuracy</h2><p>Accuracy is one of the simplest and most commonly used evaluation metrics.</p><p>It measures the proportion of predictions the model got correct:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{Accuracy} = \\frac{TP + TN}{TP + TN + FP + FN}&quot;,&quot;id&quot;:&quot;FBROBTNESP&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where:</p><ul><li><p>TP = True Positives</p></li><li><p>TN = True Negatives</p></li><li><p>FP = False Positives</p></li><li><p>FN = False Negatives</p></li></ul><p>Accuracy answers the question: <em><strong>&#8220;Out of all predictions, how many did the model get right?&#8221;</strong></em></p><p>For balanced datasets, accuracy can be a useful metric. For example:</p><ul><li><p>Predicting pass/fail outcomes in a balanced exam dataset</p></li><li><p>Predicting whether a customer clicks or not when both classes are similarly distributed</p></li></ul><p>However, accuracy becomes misleading when dealing with imbalanced datasets.</p><p><em><strong>Why Accuracy Can Fail</strong></em></p><p>Imagine a fraud detection dataset where:</p><ul><li><p>99% of transactions are not fraud</p></li><li><p>1% are fraud</p></li></ul><p>A model that predicts: &#8220;Not fraud&#8221; for every transaction would still achieve:</p><p>Accuracy = 99%</p><p>Even though the model completely fails to detect fraud. This is why relying only on accuracy can lead to poor business decisions.</p><p>Accuracy is most useful when:</p><ul><li><p>Classes are balanced</p></li><li><p>False positives and false negatives have similar costs</p></li></ul><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>If an interviewer says:</p><blockquote><p>&#8220;The model has 99% accuracy&#8221;</p></blockquote><p>your first instinct should be:</p><blockquote><p>&#8220;What does the class distribution look like?&#8221;</p></blockquote><p>That usually signals an imbalanced dataset question.</p><div><hr></div><h2>Precision</h2><p>Precision measures how often the model is correct when it predicts the positive class.</p><p>The formula for precision is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{Precision} = \\frac{TP}{TP + FP}&quot;,&quot;id&quot;:&quot;NJXXVKXOXB&quot;}" data-component-name="LatexBlockToDOM"></div><p>Precision answers the question: <em><strong>&#8220;When the model predicts positive, how often is it actually correct?&#8221;</strong></em></p><p>A high precision model makes fewer false positive mistakes. For example:</p><ul><li><p>In spam detection, low precision means important emails may be incorrectly marked as spam</p></li><li><p>In marketing campaigns, low precision means promotions may be sent to users unlikely to respond</p></li></ul><h4>Precision Focuses on False Positives</h4><p>Precision becomes important when false positives are costly. Examples:</p><ul><li><p>Spam filtering</p></li><li><p>Ad targeting</p></li><li><p>Recommendation systems</p></li><li><p>Loan approval systems</p></li></ul><h4>Precision vs Conservativeness</h4><p>A model can increase precision by becoming more conservative about predicting the positive class. For example:</p><ul><li><p>Instead of predicting 100 positives, it may only predict 20 highly confident positives</p></li><li><p>Precision may improve, but recall may decrease</p></li></ul><p>This is why precision should usually not be analyzed in isolation.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>A strong interview answer is:</p><blockquote><p>&#8220;Precision matters when acting on a prediction has a cost, so we want predicted positives to be highly reliable.&#8221;</p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Recall</h2><p>Recall measures how many actual positive cases the model successfully identifies.</p><p>The formula for recall is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{Recall} = \\frac{TP}{TP + FN}&quot;,&quot;id&quot;:&quot;KASDTZNSAB&quot;}" data-component-name="LatexBlockToDOM"></div><p>Recall answers the question: <em><strong>&#8220;Out of all actual positive cases, how many did the model correctly capture?&#8221;</strong></em></p><p>A high recall model makes fewer false negative mistakes.</p><h4>Recall Focuses on False Negatives</h4><p>Recall becomes important when missing a positive case is costly or dangerous. Examples:</p><ul><li><p>Fraud detection</p></li><li><p>Disease diagnosis</p></li><li><p>Security systems</p></li><li><p>Predicting equipment failures</p></li></ul><p>For example:</p><ul><li><p>In fraud detection, a false negative means fraudulent activity goes undetected</p></li><li><p>In medical diagnosis, a false negative could mean failing to detect a disease</p></li></ul><h4>Recall vs Aggressiveness</h4><p>A model can improve recall by becoming more aggressive about predicting the positive class. For example:</p><ul><li><p>Lowering the prediction threshold may capture more true positives</p></li><li><p>But this can also increase false positives</p></li></ul><p>This is why improving recall often reduces precision.</p><p><strong>&#128204; </strong><em><strong>Pro Tip</strong></em></p><p>A strong interview answer is:</p><blockquote><p>&#8220;Recall matters when missing a positive case is more costly than raising false alarms.&#8221;</p></blockquote><div><hr></div><h2>Precision vs Recall Tradeoff</h2><p>Precision and recall are often in tension with each other. Improving one can reduce the other. For example:</p><ul><li><p>Lowering the prediction threshold makes the model more likely to predict the positive class</p></li><li><p>This usually increases recall because more actual positives are captured</p></li><li><p>But it can also increase false positives, reducing precision</p></li></ul><p>Similarly:</p><ul><li><p>Raising the threshold makes the model more conservative</p></li><li><p>Precision may improve</p></li><li><p>But recall may decrease because more positives are missed</p></li></ul><h4>Intuition</h4><ul><li><p><strong>High Precision:</strong><br>&#8220;When the model says positive, it is usually correct.&#8221;</p></li><li><p><strong>High Recall:</strong><br>&#8220;The model captures most of the actual positive cases.&#8221;</p></li></ul><p>In practice, the right balance depends on the business problem.</p><h4>Real-World Examples</h4><p><strong>Fraud Detection</strong></p><p>Recall is usually prioritized because missing fraud can be very costly.</p><p><strong>Spam Detection</strong></p><p>Precision is often prioritized because users do not want important emails incorrectly marked as spam.</p><p><strong>Medical Diagnosis</strong></p><p>High recall is critical because missing a disease can be dangerous.</p><h4>Threshold Tuning</h4><p>The prediction threshold directly controls the precision-recall tradeoff. For example:</p><ul><li><p>Lower threshold &#8594; higher recall, lower precision</p></li><li><p>Higher threshold &#8594; higher precision, lower recall</p></li></ul><p>This is why threshold selection is often a business decision rather than a purely technical one.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>One of the biggest mistakes candidates make is saying:</p><blockquote><p>&#8220;Precision is more important&#8221; or &#8220;Recall is more important.&#8221;</p></blockquote><p>Strong candidates instead say:</p><blockquote><p>&#8220;The right balance depends on the cost of false positives versus false negatives.&#8221;</p></blockquote><div><hr></div><h2>F1 Score</h2><p>The F1 Score combines precision and recall into a single metric. It is defined as the harmonic mean of precision and recall:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;F1 = 2 \\times \\frac{Precision \\times Recall}{Precision + Recall}&quot;,&quot;id&quot;:&quot;EJLHYGXWOZ&quot;}" data-component-name="LatexBlockToDOM"></div><p>F1 Score answers the question:</p><blockquote><p>&#8220;How well does the model balance precision and recall?&#8221;</p></blockquote><p>Unlike accuracy, the F1 Score becomes especially useful when dealing with imbalanced datasets. A high F1 Score means:</p><ul><li><p>The model captures positive cases effectively</p></li><li><p>The model also keeps false positives relatively low</p></li></ul><h4>Why Harmonic Mean?</h4><p>The harmonic mean penalizes situations where one metric is much lower than the other. For example:</p><ul><li><p>Precision = 1.0</p></li><li><p>Recall = 0.1</p></li></ul><p>The F1 Score will still be low because the model is performing poorly on recall. This prevents models from appearing strong when only one metric is high.</p><h4>When F1 Score Is Useful</h4><p>F1 Score is commonly used when:</p><ul><li><p>Classes are imbalanced</p></li><li><p>Both false positives and false negatives matter</p></li><li><p>A balance between precision and recall is important</p></li></ul><p>Examples:</p><ul><li><p>Fraud detection</p></li><li><p>Churn prediction</p></li><li><p>Medical diagnosis</p></li></ul><h4>Limitations of F1 Score</h4><p>F1 Score does not consider true negatives directly. This means:</p><ul><li><p>Two models with very different true negative performance can still have similar F1 Scores</p></li></ul><p>Because of this, F1 Score should not always be used alone.</p><p><strong>&#128204;  </strong><em><strong>Pro Tip: </strong></em>A very common interview question is:</p><blockquote><p>&#8220;Why do we use the harmonic mean instead of the regular average?&#8221;</p></blockquote><p>The answer:</p><blockquote><p>&#8220;Because the harmonic mean penalizes extreme imbalance between precision and recall.&#8221;</p></blockquote><div><hr></div><h2>ROC Curve and ROC-AUC</h2><p>The ROC Curve (Receiver Operating Characteristic Curve) measures how well a model separates the positive and negative classes across different classification thresholds.</p><p>The curve plots:</p><ul><li><p><strong>True Positive Rate (TPR)</strong> on the Y-axis</p></li><li><p><strong>False Positive Rate (FPR)</strong> on the X-axis</p></li></ul><p>Where:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{TPR} = \\frac{TP}{TP + FN}&quot;,&quot;id&quot;:&quot;LTTPUGHJUE&quot;}" data-component-name="LatexBlockToDOM"></div><p></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{FPR} = \\frac{FP}{FP + TN}&quot;,&quot;id&quot;:&quot;LCOCWRSFKQ&quot;}" data-component-name="LatexBlockToDOM"></div><p></p><h4>Intuition</h4><p>ROC-AUC measures the model&#8217;s ability to rank positive examples higher than negative examples. A higher ROC-AUC means:</p><ul><li><p>Better class separation</p></li><li><p>Better ranking ability</p></li><li><p>Stronger discrimination between classes</p></li></ul><p>General interpretation:</p><ul><li><p>ROC-AUC = 0.5 &#8594; Random guessing</p></li><li><p>ROC-AUC = 1.0 &#8594; Perfect separation</p></li></ul><h4>Why ROC Curves Matter</h4><p>Different classification thresholds produce different:</p><ul><li><p>Precision</p></li><li><p>Recall</p></li><li><p>False positive rates</p></li></ul><p>ROC curves allow us to evaluate model performance across all thresholds instead of relying on a single threshold like 0.5.</p><p>ROC-AUC evaluates ranking quality, not classification accuracy at a specific threshold.</p><p>This is why:</p><ul><li><p>A model can have a strong ROC-AUC</p></li><li><p>But still perform poorly at a badly chosen threshold</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WvUk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WvUk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1278361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/197777965?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WvUk!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b75ca20-7926-46c9-971d-b4386cfed967_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One important thing to notice in the ROC curve is how different classification thresholds change the model&#8217;s behavior.</p><p>As the threshold decreases:</p><ul><li><p>The model becomes more aggressive about predicting the positive class</p></li><li><p>Recall (TPR) increases</p></li><li><p>False positives also increase</p></li></ul><p>As the threshold increases:</p><ul><li><p>The model becomes more conservative</p></li><li><p>Precision usually improves</p></li><li><p>Recall decreases because fewer positives are predicted</p></li></ul><p>Each point on the ROC curve represents a different threshold value.</p><p>This is why ROC curves are useful:<br>they show model performance across <em>all possible thresholds</em> instead of evaluating performance at just one fixed threshold like 0.5.</p><h4>Limitations of ROC-AUC</h4><p>ROC-AUC can sometimes appear overly optimistic on highly imbalanced datasets.</p><p>This is because:</p><ul><li><p>False Positive Rate may remain small even when the number of false positives is large in absolute terms</p></li></ul><p>This is one reason Precision-Recall curves are often preferred for rare-event prediction problems.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>A strong interview explanation is:</p><blockquote><p>&#8220;ROC-AUC measures how well the model separates the classes independent of a specific threshold.&#8221;</p></blockquote><p>That shows deeper understanding than simply saying:</p><blockquote><p>&#8220;Higher AUC is better.&#8221;</p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>PR Curve vs ROC Curve</h2><p>Both ROC Curves and Precision-Recall (PR) Curves are used to evaluate classification models across different thresholds.</p><p>However, they focus on different aspects of model performance.</p><p>ROC Curves plot:</p><ul><li><p>True Positive Rate (Recall)</p></li><li><p>False Positive Rate</p></li></ul><p>ROC-AUC measures how well the model separates the classes overall.</p><p>PR Curves plot:</p><ul><li><p>Precision</p></li><li><p>Recall</p></li></ul><p>Instead of focusing on class separation broadly, PR Curves focus specifically on performance for the positive class.</p><h4>Why PR Curves Matter for Imbalanced Data</h4><p>On highly imbalanced datasets, ROC-AUC can sometimes appear overly optimistic.</p><p>This happens because:</p><ul><li><p>The False Positive Rate may remain small even when the model generates many false positives in absolute terms</p></li></ul><p>PR Curves expose this issue more clearly because precision directly accounts for false positives.</p><p>This makes PR Curves especially useful for:</p><ul><li><p>Fraud detection</p></li><li><p>Disease prediction</p></li><li><p>Rare-event modeling</p></li><li><p>Anomaly detection</p></li></ul><p>ROC Curves answer:</p><blockquote><p>&#8220;How well can the model separate the classes overall?&#8221;</p></blockquote><p>PR Curves answer:</p><blockquote><p>&#8220;When the model predicts positive, how reliable are those predictions while still capturing positives?&#8221;</p></blockquote><p>General rule:</p><ul><li><p>Balanced datasets &#8594; ROC-AUC is often sufficient</p></li><li><p>Imbalanced datasets &#8594; PR Curves are usually more informative</p></li></ul><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>Mentioning PR Curves for imbalanced datasets is a strong signal of practical machine learning experience.</p><p>A strong interview answer is:</p><blockquote><p>&#8220;ROC-AUC can look strong even when precision is poor on rare-event problems, which is why PR curves are often preferred for imbalanced datasets.&#8221;</p></blockquote><div><hr></div><h2>Threshold Tuning</h2><p>Most classification models output probabilities, not final class predictions.</p><p>To convert probabilities into class labels, we choose a classification threshold.</p><p>For example:</p><ul><li><p>Probability &gt; 0.5 &#8594; Positive class</p></li><li><p>Probability &#8804; 0.5 &#8594; Negative class</p></li></ul><p>However, the default threshold of 0.5 is not always optimal.</p><h4>Why Thresholds Matter</h4><p>Changing the threshold directly changes:</p><ul><li><p>Precision</p></li><li><p>Recall</p></li><li><p>False Positive Rate</p></li><li><p>False Negative Rate</p></li></ul><p>This means the threshold controls the type of mistakes the model makes.</p><p>A lower threshold makes the model more aggressive about predicting the positive class. Effects:</p><ul><li><p>Higher recall</p></li><li><p>More true positives</p></li><li><p>More false positives</p></li><li><p>Lower precision</p></li></ul><p>A higher threshold makes the model more conservative. Effects:</p><ul><li><p>Higher precision</p></li><li><p>Fewer false positives</p></li><li><p>Lower recall</p></li><li><p>More false negatives</p></li></ul><h4>Real-World Examples</h4><p><strong>Fraud Detection</strong></p><p>Lower thresholds are often preferred because missing fraud can be very costly.</p><p><strong>Spam Detection</strong></p><p>Higher thresholds may be preferred to avoid incorrectly classifying important emails as spam.</p><p><strong>Medical Diagnosis</strong></p><p>Thresholds are often chosen carefully to balance patient safety and unnecessary interventions.</p><p>Threshold selection is usually a business decision, not just a modeling decision.</p><p>The &#8220;best&#8221; threshold depends on:</p><ul><li><p>Cost of false positives</p></li><li><p>Cost of false negatives</p></li><li><p>Operational constraints</p></li><li><p>User experience</p></li></ul><p><strong>&#128204;  </strong><em><strong>Pro Tip: </strong></em>Many candidates assume the threshold is fixed at 0.5.</p><p>Mentioning threshold tuning immediately makes your answer more practical and business-oriented.</p><p>A strong interview answer is:</p><blockquote><p>&#8220;The optimal threshold depends on the business trade-off between false positives and false negatives.&#8221;</p></blockquote><div><hr></div><h2>Choosing the Right Metric</h2><p>One important distinction to make is that the metrics discussed in this issue are <strong>model performance metrics</strong>. These metrics evaluate how well the machine learning model performs statistically.</p><p>However, business success is often measured using entirely different metrics, such as:</p><ul><li><p>Revenue</p></li><li><p>Retention</p></li><li><p>Conversion rate</p></li><li><p>Customer lifetime value</p></li><li><p>Cost savings</p></li></ul><p>A model with strong ROC-AUC or F1 Score does not automatically guarantee strong business impact. That is why choosing the right business metric is equally important.</p><blockquote><p>I covered this topic in a <a href="/__u/thepracticaldatascientist.substack.com/p/defining-product-metrics-what-to">previous issue </a>of <em>The Practical Data Scientist</em> on selecting the right business metrics for machine learning problems.</p></blockquote><p>There is no universally &#8220;best&#8221; evaluation metric. The right metric depends entirely on:</p><ul><li><p>The business problem</p></li><li><p>The type of mistakes that matter</p></li><li><p>The operational costs of those mistakes</p></li></ul><h4>Different Problems Need Different Metrics</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HCGn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 424w, /__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 848w, /__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HCGn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png" width="1456" height="502" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:502,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:85704,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/197777965?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 424w, /__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 848w, /__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HCGn!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac604940-0e99-42c4-9f59-a1e5be58348d_1584x546.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Business Context Matters</h4><p>Two models can have:</p><ul><li><p>Similar accuracy</p></li><li><p>Similar ROC-AUC</p></li><li><p>Very different business impact</p></li></ul><p>For example:</p><ul><li><p>A fraud model with high recall may detect more fraud</p></li><li><p>But excessive false positives may overwhelm investigators</p></li></ul><p>This is why evaluation metrics should never be discussed without business context.</p><p>Many candidates in an interview would answer:</p><blockquote><p>&#8220;I would optimize for accuracy.&#8221;</p></blockquote><p>without first understanding:</p><ul><li><p>Class imbalance</p></li><li><p>Cost of mistakes</p></li><li><p>Operational impact</p></li></ul><p>Strong candidates instead ask:</p><ul><li><p>Which errors are more expensive?</p></li><li><p>What happens when the model is wrong?</p></li><li><p>What are the business constraints?</p></li></ul><p>Metrics are not just mathematical tools.</p><p>They are ways of translating business priorities into model behavior.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>This is one of the highest-signal interview sections.</p><p>The strongest answers usually:</p><ol><li><p>Identify the business objective</p></li><li><p>Discuss trade-offs</p></li><li><p>Explain which mistakes matter most</p></li><li><p>Justify the metric choice accordingly</p></li></ol><div><hr></div><h2>Common Interview Questions</h2><p>Here are some of the most commonly asked evaluation metric interview questions:</p><ol><li><p><strong>Why is accuracy misleading on imbalanced datasets?</strong></p></li></ol><p>Accuracy can appear very high even when the model completely fails on the minority class.</p><p>For example:</p><ul><li><p>If 99% of transactions are non-fraud</p></li><li><p>A model predicting &#8220;not fraud&#8221; for every case still achieves 99% accuracy</p></li></ul><p>This is why metrics like Precision, Recall, F1 Score, and PR-AUC are often more useful for imbalanced problems.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>Whenever you hear:</p><blockquote><p>&#8220;The model has very high accuracy&#8221;</p></blockquote><p>immediately think about:</p><ul><li><p>Class imbalance</p></li><li><p>Minority class performance</p></li><li><p>Cost of mistakes</p></li></ul><ol start="2"><li><p><strong>Precision vs Recall &#8212; which is more important?</strong></p></li></ol><p>Neither metric is universally better.</p><p>The right choice depends on the business problem and the cost of errors.</p><p>Examples:</p><ul><li><p>Fraud detection &#8594; prioritize Recall</p></li><li><p>Spam detection &#8594; prioritize Precision</p></li></ul><p><strong>&#128204; </strong><em><strong>Pro Tip:  </strong></em>Avoid absolute answers like:</p><blockquote><p>&#8220;Recall is better.&#8221;</p></blockquote><p>Strong candidates discuss trade-offs instead.</p><ol start="3"><li><p><em><strong>Why use F1 Score?</strong></em></p></li></ol><p>F1 Score balances Precision and Recall into a single metric.</p><p>It becomes useful when:</p><ul><li><p>Classes are imbalanced</p></li><li><p>Both false positives and false negatives matter</p></li></ul><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>Mention that F1 Score ignores true negatives directly.<br>That shows deeper understanding.</p><ol start="4"><li><p><em><strong>ROC-AUC vs PR-AUC</strong></em></p></li></ol><ul><li><p>ROC-AUC measures overall class separation ability</p></li><li><p>PR-AUC focuses more directly on positive-class performance</p></li></ul><p>PR-AUC is often preferred for highly imbalanced datasets because it captures precision-recall trade-offs more clearly.</p><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>Mentioning PR-AUC for rare-event prediction problems is a strong practical ML signal.</p><ol start="5"><li><p><em><strong>How do thresholds impact model performance?</strong></em></p></li></ol><p>Changing the threshold changes:</p><ul><li><p>Precision</p></li><li><p>Recall</p></li><li><p>False positive rate</p></li><li><p>False negative rate</p></li></ul><p>Lower thresholds:</p><ul><li><p>Increase recall</p></li><li><p>Reduce precision</p></li></ul><p>Higher thresholds:</p><ul><li><p>Increase precision</p></li><li><p>Reduce recall</p></li></ul><p><strong>&#128204; </strong><em><strong>Pro Tip: </strong></em>Mention that threshold tuning is usually a business decision, not just a modeling decision.</p><ol start="6"><li><p><em><strong>Which metric would you optimize for?</strong></em></p></li></ol><p>The correct answer depends on:</p><ul><li><p>Business objectives</p></li><li><p>Operational constraints</p></li><li><p>Cost of false positives</p></li><li><p>Cost of false negatives</p></li></ul><p>There is rarely one universally correct metric.</p><p><strong>&#128204; </strong><em><strong>Pro Tip:  </strong></em>The strongest candidates:</p><ol><li><p>Clarify the business goal</p></li><li><p>Discuss trade-offs</p></li><li><p>Justify the metric choice using business impact</p></li></ol><div><hr></div><h2>Final Takeaway</h2><p>Evaluation metrics are not just mathematical formulas.</p><p>They are tools for understanding:</p><ul><li><p>How a model behaves</p></li><li><p>What kinds of mistakes it makes</p></li><li><p>Whether those mistakes align with business goals</p></li></ul><p>One of the biggest differences between beginner and experienced data scientists is the ability to think beyond accuracy and reason about trade-offs.</p><p>In interviews, strong candidates usually:</p><ul><li><p>Explain metrics intuitively</p></li><li><p>Connect them to real-world business problems</p></li><li><p>Discuss false positives vs false negatives</p></li><li><p>Justify why a particular metric matters</p></li></ul><p>The goal is not to memorize formulas.</p><p>The goal is to understand what the model is optimizing for &#8212; and whether that aligns with the actual problem being solved.</p><div><hr></div><h2>What&#8217;s Next?</h2><p>In the next issue of <em>The Practical Data Scientist</em>, we&#8217;ll move beyond linear models and step into the world of tree-based models with Decision Trees </p><p>Decision trees are one of the most important foundations for understanding modern machine learning models like Random Forests and Gradient Boosting.</p><p>As always, the focus will be on intuition, practical understanding, and interview-oriented thinking rather than memorizing formulas.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Logistic Regression: What Interviewers Expect You to Know]]></title><description><![CDATA[The concepts, intuition, and interview questions you actually need to know]]></description><link>https://thepracticaldatascientist.substack.com/p/logistic-regression-what-interviewers</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/logistic-regression-what-interviewers</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 12 May 2026 14:02:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Vvs5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Vvs5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Vvs5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1461631,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/197167973?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Vvs5!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F273f7c1e-a8dc-42c5-ad6e-5d2e5ab278f6_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What Is Logistic Regression?</h2><p>Logistic regression is a supervised machine learning algorithm used for classification problems. It predicts the probability that an observation belongs to a particular class.</p><p>For example, it can be used to predict:</p><ul><li><p>Whether a customer will churn</p></li><li><p>Whether a transaction is fraudulent</p></li><li><p>Whether a user will click on an ad</p></li></ul><p>Unlike linear regression, which predicts continuous values, logistic regression predicts probabilities between 0 and 1.</p><p>It does this using the sigmoid function:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p = \\frac{1}{1 + e^{-z}}&quot;,&quot;id&quot;:&quot;FVRYRERAEI&quot;}" data-component-name="LatexBlockToDOM"></div><p>The predicted probability is then converted into a class label using a threshold (commonly 0.5). For example:</p><ul><li><p>Probability &gt; 0.5 &#8594; Class 1</p></li><li><p>Probability &#8804; 0.5 &#8594; Class 0</p></li></ul><p>One of the biggest advantages of logistic regression is interpretability, which is why it remains one of the most commonly asked algorithms in data science interviews.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>The Core Intuition</h2><p>Logistic regression first creates a weighted combination of input features:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;z = \\beta_0 + \\beta_1 x_1 + \\beta_2 x_2 + \\dots + \\beta_n x_n&quot;,&quot;id&quot;:&quot;MOLRIBONFM&quot;}" data-component-name="LatexBlockToDOM"></div><p>This value is then passed through the sigmoid function:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p = \\frac{1}{1 + e^{-z}}&quot;,&quot;id&quot;:&quot;PSIREHINJR&quot;}" data-component-name="LatexBlockToDOM"></div><p>The sigmoid function transforms the output into a smooth S-shaped curve, allowing the model to map predictions into probabilities.</p><p>Key intuition:</p><ul><li><p>Higher values of z push predictions closer to 1</p></li><li><p>Lower values of z push predictions closer to 0</p></li></ul><p>This makes logistic regression effective for binary classification problems while keeping the model simple and interpretable.</p><div><hr></div><h2>Key Interview Concept: Log-Odds</h2><p>One of the most important concepts in logistic regression interviews is log-odds.</p><p>Instead of modeling probabilities directly, logistic regression models the logarithm of the odds:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\log\\left(\\frac{p}{1-p}\\right) = \\beta_0 + \\beta_1 x_1 + \\dots + \\beta_n x_n&quot;,&quot;id&quot;:&quot;GEYTJQJBBJ&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here:</p><ul><li><p>p represents the probability of the positive class</p></li><li><p>p/(1&#8722;p)&#8203; represents the odds</p></li><li><p>Taking the logarithm converts the relationship into a linear equation</p></li></ul><p>But why do we use log-odds instead of probabilities directly?</p><p>Probabilities are bounded between 0 and 1, making them difficult to model with a linear equation. A linear model could otherwise produce invalid predictions like 1.5 or -0.2.</p><p>By converting probabilities into odds and then taking the logarithm, the output range becomes:</p><ul><li><p>Negative infinity to positive infinity</p></li></ul><p>This makes it suitable for linear modeling.</p><h4>Intuition</h4><ul><li><p>Probability = 0.5 &#8594; log-odds = 0</p></li><li><p>Probability &gt; 0.5 &#8594; positive log-odds</p></li><li><p>Probability &lt; 0.5 &#8594; negative log-odds</p></li></ul><p>Another reason this formulation is important is interpretability:</p><ul><li><p>A positive coefficient increases the odds of the positive class</p></li><li><p>A negative coefficient decreases the odds</p></li></ul><p>A common interview question is:</p><blockquote><p>&#8220;How do you interpret logistic regression coefficients?&#8221;</p></blockquote><p>In logistic regression:</p><ul><li><p>A one-unit increase in a feature changes the log-odds by the value of its coefficient</p></li><li><p>Exponentiating the coefficient, e&#946;e^{\beta}e&#946;, gives the odds ratio</p></li></ul><p>For example:</p><ul><li><p>If e^&#946; = 2, the odds are doubled</p></li><li><p>If e^&#946; = 0.5, the odds are reduced by half</p></li></ul><p>This impacts the probability p indirectly through the sigmoid function:</p><ul><li><p>Positive coefficients increase the probability of the positive class</p></li><li><p>Negative coefficients decrease it</p></li></ul><p>However, the change in probability is not constant. It depends on the current value of p.</p><p>For example:</p><ul><li><p>Increasing log-odds from 0 to 1 changes probability from 0.50 to 0.73</p></li><li><p>Increasing log-odds from 4 to 5 changes probability only from 0.98 to 0.99</p></li></ul><p>This happens because the sigmoid curve becomes flatter near 0 and 1.</p><div><hr></div><h2>How Logistic Regression Is Trained</h2><p>Logistic regression is trained by finding the coefficients that best separate the classes.</p><p>Instead of using Mean Squared Error (MSE), logistic regression uses a loss function called Log Loss (or Binary Cross-Entropy):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;L = - \\sum \\left[y \\log(p) + (1-y)\\log(1-p)\\right]&quot;,&quot;id&quot;:&quot;XPIVDQPSMS&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where:</p><ul><li><p>y is the actual class label</p></li><li><p>p is the predicted probability</p></li></ul><p>The loss function penalizes incorrect predictions:</p><ul><li><p>Predicting a high probability for the wrong class results in a large penalty</p></li><li><p>Correct and confident predictions result in lower loss</p></li></ul><p>The model then uses optimization algorithms such as Gradient Descent to minimize this loss and learn the best coefficients.</p><h4>Why Not Use MSE?</h4><p>Mean Squared Error is commonly used in linear regression, but it is not ideal for logistic regression because:</p><ul><li><p>The sigmoid function makes the optimization problem non-linear</p></li><li><p>MSE can lead to non-convex loss surfaces</p></li><li><p>Log Loss provides better probabilistic interpretation and optimization behavior</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p></li></ul><div><hr></div><h2>Important Interview Topics</h2><h4>Assumptions</h4><p>Logistic regression makes a few important assumptions:</p><ul><li><p><strong>Linear relationship with log-odds:</strong><br>The features should have a linear relationship with the log-odds of the target, not necessarily with the target itself.</p></li><li><p><strong>Independent observations:</strong><br>Each observation should be independent. Repeated or highly related observations can bias the model.</p></li><li><p><strong>Low multicollinearity:</strong><br>Features should not be highly correlated with each other, as this can make coefficients unstable and difficult to interpret.</p></li><li><p><strong>Limited influence of outliers:</strong><br>Extreme values can heavily influence the decision boundary and coefficient estimates.</p></li></ul><p>&#128204; <em><strong>Pro Tip: </strong></em>A common mistake is saying logistic regression assumes a linear relationship with the target variable.<br>The correct answer is:</p><blockquote><p>&#8220;Logistic regression assumes a linear relationship between the features and the log-odds of the target.&#8221;</p></blockquote><h4>Evaluation Metrics</h4><p>Accuracy alone is often misleading, especially with imbalanced datasets.</p><p>Common evaluation metrics include:</p><ul><li><p>Precision</p></li><li><p>Recall</p></li><li><p>F1 Score</p></li><li><p>ROC-AUC</p></li></ul><p>Interview tip:</p><ul><li><p>Fraud detection problems usually prioritize recall</p></li><li><p>Marketing targeting problems often prioritize precision</p></li></ul><p><em>I&#8217;ll cover evaluation metrics in much more detail in a separate article since they are extremely important for interviews.</em></p><p>&#128204; <em><strong>Pro Tip: </strong></em>When discussing metrics, always tie them back to business impact instead of giving textbook definitions only.</p><h4>Imbalanced Data</h4><p>In many real-world datasets, one class appears much less frequently than the other.</p><p>Examples:</p><ul><li><p>Fraud transactions</p></li><li><p>Rare diseases</p></li><li><p>Customer churn</p></li></ul><p>A model predicting only the majority class may still achieve high accuracy, making accuracy unreliable.</p><p>Common solutions:</p><ul><li><p>Oversampling the minority class</p></li><li><p>Undersampling the majority class</p></li><li><p>Using class weights</p></li><li><p>Adjusting the classification threshold</p></li></ul><p>&#128204; <em><strong>Pro Tip: </strong></em>If an interviewer mentions &#8220;99% accuracy,&#8221; immediately think about class imbalance and ask about the class distribution.</p><h4>Regularization</h4><p>Regularization helps prevent overfitting by penalizing large coefficients.</p><p>Two common types:</p><ul><li><p><strong>L1 Regularization (Lasso):</strong><br>Can shrink some coefficients to zero, effectively performing feature selection.</p></li><li><p><strong>L2 Regularization (Ridge):</strong><br>Reduces coefficient magnitudes smoothly without eliminating features entirely.</p></li></ul><p>Regularization is especially useful when:</p><ul><li><p>There are many features</p></li><li><p>Features are correlated</p></li><li><p>The model is overfitting training data</p></li></ul><p>&#128204; <em><strong>Pro Tip: </strong></em>A strong interview answer is:</p><blockquote><p>&#8220;L1 helps with feature selection, while L2 helps stabilize the model by shrinking coefficients.&#8221;</p></blockquote><h4>Decision Threshold</h4><p>Logistic regression outputs probabilities, which are converted into class labels using a threshold.</p><p>Default threshold:</p><ul><li><p>Probability &gt; 0.5 &#8594; Positive class</p></li><li><p>Probability &#8804; 0.5 &#8594; Negative class</p></li></ul><p>However, the threshold can be adjusted depending on business objectives.</p><p>Examples:</p><ul><li><p>Lower threshold &#8594; higher recall</p></li><li><p>Higher threshold &#8594; higher precision</p></li></ul><p>&#128204; <em><strong>Pro Tip:</strong></em> Many candidates forget that the threshold is configurable. Mentioning threshold tuning shows practical understanding beyond theory.</p><h4>Multicollinearity</h4><p>Multicollinearity occurs when features are highly correlated with each other.</p><p>Problems caused by multicollinearity:</p><ul><li><p>Unstable coefficient estimates</p></li><li><p>Difficulty interpreting feature importance</p></li><li><p>Increased variance in predictions</p></li></ul><p>Common ways to detect it:</p><ul><li><p>Correlation matrix</p></li><li><p>Variance Inflation Factor (VIF)</p></li></ul><p>Possible solutions:</p><ul><li><p>Remove correlated features</p></li><li><p>Combine related variables</p></li><li><p>Apply regularization techniques</p></li></ul><p>&#128204; <em><strong>Pro Tip: </strong></em>If interpretability matters, always mention multicollinearity because unstable coefficients can lead to misleading business conclusions.</p><div><hr></div><h2>When to Use Logistic Regression</h2><p>Logistic regression is often one of the best starting models for classification problems because it is simple, fast, and highly interpretable.</p><p>It works especially well when:</p><ul><li><p>The relationship between features and the target is relatively linear</p></li><li><p>Interpretability is important</p></li><li><p>The dataset is not extremely large or complex</p></li><li><p>Probability estimates are needed</p></li></ul><p>Common real-world applications include:</p><ul><li><p>Customer churn prediction</p></li><li><p>Fraud detection</p></li><li><p>Credit risk modeling</p></li><li><p>Ad click prediction</p></li><li><p>Medical diagnosis</p></li></ul><p>One major advantage of logistic regression is explainability. Unlike many complex machine learning models, its coefficients can be interpreted directly, making it popular in industries where transparency matters.</p><p>However, logistic regression may struggle when:</p><ul><li><p>Relationships are highly non-linear</p></li><li><p>There are complex feature interactions</p></li><li><p>The decision boundary is not approximately linear</p></li></ul><p>In such cases, models like decision trees, random forests, or gradient boosting may perform better.</p><p>&#128204; <em><strong>Pro Tip:  </strong></em>A strong interview answer is &#8220;Logistic regression is usually my first baseline model because it is interpretable, fast to train, and provides strong performance on many structured datasets.&#8221;</p><div><hr></div><h2>Common Interview Questions</h2><p>Here are some of the most commonly asked logistic regression interview questions:</p><ol><li><p><em><strong>Why not use linear regression for classification?</strong></em></p></li></ol><p>Linear regression can produce predictions outside the range of 0 and 1, making it unsuitable for probability estimation. Logistic regression solves this using the sigmoid function.</p><ol start="2"><li><p><em><strong>Why does logistic regression use log-odds?</strong></em></p></li></ol><p>Log-odds transform probabilities from a bounded range (0,1)(0,1)(0,1) to an unbounded range (&#8722;&#8734;,+&#8734;)(-\infty, +\infty)(&#8722;&#8734;,+&#8734;), allowing a linear relationship with the features.</p><ol start="3"><li><p><em><strong>How do you interpret logistic regression coefficients?</strong></em></p></li></ol><ul><li><p>Coefficients represent changes in log-odds</p></li><li><p>e^&#946; represents the odds ratio</p></li><li><p>Positive coefficients increase the probability of the positive class</p></li></ul><ol start="4"><li><p><em><strong>What happens when features are highly correlated?</strong></em></p></li></ol><p>Highly correlated features can lead to unstable coefficient estimates and poor interpretability. This issue is known as multicollinearity.</p><ol start="5"><li><p><em><strong>How do you handle imbalanced datasets?</strong></em></p></li></ol><p>Common approaches include:</p><ul><li><p>Class weights</p></li><li><p>Oversampling / undersampling</p></li><li><p>Threshold tuning</p></li><li><p>Using better evaluation metrics such as Precision, Recall, and F1-score</p></li></ul><ol start="6"><li><p><em><strong>What is regularization and why is it important?</strong></em></p></li></ol><p>Regularization prevents overfitting by penalizing large coefficients.</p><ul><li><p>L1 regularization can perform feature selection</p></li><li><p>L2 regularization shrinks coefficients smoothly</p></li></ul><ol start="7"><li><p><em><strong>What is the role of the decision threshold?</strong></em></p></li></ol><p>The threshold converts predicted probabilities into class labels. Adjusting the threshold changes the balance between precision and recall.</p><p>&#128204; <em><strong>Pro Tip:  </strong></em>In interviews, avoid giving one-line textbook answers. Explain:</p><ol><li><p>The intuition</p></li><li><p>The business impact</p></li><li><p>The trade-offs</p></li></ol><p>That usually separates strong candidates from average ones.</p><div><hr></div><h2>Final Takeaway</h2><p>Logistic regression is one of the most important algorithms for data science interviews because it combines:</p><ul><li><p>Statistics</p></li><li><p>Machine learning</p></li><li><p>Probability</p></li><li><p>Business interpretation</p></li></ul><p>Even though it is considered a &#8220;simple&#8221; model, interviewers often use it to evaluate how deeply you understand core machine learning concepts.</p><p>To perform well in interviews, focus on:</p><ul><li><p>Understanding intuition instead of memorizing formulas</p></li><li><p>Explaining coefficients clearly</p></li><li><p>Connecting evaluation metrics to business goals</p></li><li><p>Discussing trade-offs and limitations</p></li></ul><p>A candidate who can explain logistic regression clearly and intuitively usually demonstrates strong machine learning fundamentals.</p><div><hr></div><h2>What&#8217;s Next</h2><p>In the next issue, I&#8217;ll cover one of the most important topics in machine learning interviews:</p><ul><li><p>Precision vs Recall</p></li><li><p>F1 Score</p></li><li><p>ROC-AUC</p></li><li><p>Confusion Matrix</p></li><li><p>Threshold tuning</p></li><li><p>Choosing the right metric for business problems</p></li></ul><p>Understanding evaluation metrics is critical because building a model is only half the problem &#8212; knowing how to evaluate it correctly is what drives business impact.</p><p>This is also one of the most frequently tested areas in data science interviews.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Practical Data Scientist! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Linear Regression (Part 2): How to Approach It in Data Science Interviews]]></title><description><![CDATA[A structured approach to interpreting results, evaluating models, and answering linear regression questions with clarity.]]></description><link>https://thepracticaldatascientist.substack.com/p/linear-regression-part-2-how-to-approach</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/linear-regression-part-2-how-to-approach</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 05 May 2026 14:02:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uVUk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uVUk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 424w, /__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 848w, /__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uVUk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png" width="1456" height="970" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:970,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2445096,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/196491829?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 424w, /__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 848w, /__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uVUk!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca6cb9dc-1450-4699-a85b-d3d06f6374b2_1796x1196.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Introduction: Why Linear Regression Shows Up in Interviews</strong></h2><p>Linear regression is one of the most commonly tested topics in data science interviews.</p><p>Not because companies expect you to build production models from scratch, but because it tests something more fundamental:</p><p>&#128073; <em>Do you understand how models work, and can you interpret what they&#8217;re telling you?</em></p><p>In Part 1, we focused on how linear regression works&#8212;how the line of best fit is determined, what residuals represent, and the assumptions the model relies on.</p><p>But in interviews, the focus shifts.</p><p>You&#8217;re rarely asked to derive formulas or write complex code. Instead, you&#8217;re expected to:</p><ul><li><p>interpret model outputs</p></li><li><p>explain assumptions clearly</p></li><li><p>reason about when the model works (and when it doesn&#8217;t)</p></li></ul><p>This is where many candidates struggle. They know how to fit the model, but find it harder to explain what the results actually mean.</p><p>In this article, we&#8217;ll focus on exactly that&#8212;how to approach linear regression questions in interviews with clarity and structure.</p><div><hr></div><h2><strong>A Structured Way to Answer Linear Regression Questions</strong></h2><p>When faced with a linear regression question in an interview, it&#8217;s easy to jump straight into explaining coefficients or metrics.</p><p>A better approach is to stay structured.</p><p>A simple framework you can use is:</p><p>&#128073; <strong>Model &#8594; Interpretation &#8594; Assumptions &#8594; Limitations</strong></p><h4><strong>1. Start with the Model</strong></h4><p>Briefly explain what the model is doing.</p><p>For example:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{y} = \\beta_0 + \\beta_1 x&quot;,&quot;id&quot;:&quot;GFMPPKGOIW&quot;}" data-component-name="LatexBlockToDOM"></div><p>You don&#8217;t need to go deep into the math&#8212;just show that you understand that the model is capturing a relationship between input and output.</p><h4><strong>2. Interpret the Results</strong></h4><p>Next, focus on what the model is telling you.</p><ul><li><p>What does the coefficient mean?</p></li><li><p>What is the direction of the relationship?</p></li><li><p>How does a change in input affect the output?</p></li></ul><p>This is usually the most important part of your answer.</p><p>We&#8217;ll cover this in detail in the next section, where we break down how to interpret coefficients, p-values, and model outputs with concrete examples.</p><h4><strong>3. Call Out Assumptions</strong></h4><p>Then, connect back to Part 1.</p><p>Mention key assumptions such as:</p><ul><li><p>linearity</p></li><li><p>independence</p></li><li><p>constant variance</p></li></ul><p>You don&#8217;t need to list all of them&#8212;just show that you know the model relies on certain conditions.</p><h4><strong>4. Discuss Limitations</strong></h4><p>Finally, go beyond the model.</p><ul><li><p>Does correlation imply causation?</p></li><li><p>Are there missing variables?</p></li><li><p>Could the relationship be non-linear?</p></li></ul><p>This shows maturity in thinking and is often what differentiates strong candidates.</p><h4><strong>Putting It Together</strong></h4><p>Instead of giving a fragmented answer, this structure helps you build a clear narrative:</p><p>&#128073; <em>Here&#8217;s what the model does &#8594; here&#8217;s what it means &#8594; here&#8217;s when it works &#8594; here&#8217;s where it can fail</em></p><p>&#128204; <strong>Pro tip:</strong> In interviews, clarity and structure matter more than depth. A well-organized answer with clear reasoning is often more valuable than a technically dense one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2><strong>Interpreting Coefficients: What the Model Is Telling You</strong></h2><p>Once we&#8217;ve fitted the model, the next step is to understand what it is actually telling us.</p><p>Recall the model:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{y} = \\beta_0 + \\beta_1 x&quot;,&quot;id&quot;:&quot;RXVJGCPYSQ&quot;}" data-component-name="LatexBlockToDOM"></div><p></p><p>Here, &#946;0&#8203; is the intercept and &#946;1&#8203; is the coefficient (or slope). These are the two most important outputs of the model.</p><h4><strong>What Does the Coefficient Mean?</strong></h4><p>The coefficient &#946;1&#8203; tells us how much the output changes for a one-unit change in the input.</p><p>In our example:</p><p>&#128073; If &#946;1&#8203;=0.05, it means:</p><blockquote><p>For every additional unit increase in TV advertising spend, sales increase by <strong>0.05 units</strong>, on average.</p></blockquote><p>The key phrase here is <strong>&#8220;on average&#8221;</strong>. This is not a guaranteed outcome for every observation&#8212;it is the model&#8217;s best estimate across the data.</p><h4><strong>Direction and Magnitude</strong></h4><p>There are two things to focus on when interpreting a coefficient:</p><ul><li><p><strong>Direction: </strong>A positive coefficient means the variables move in the same direction. A negative coefficient means they move in opposite directions.</p></li><li><p><strong>Magnitude: </strong>The size of the coefficient tells you how strong the relationship is. Larger values indicate a stronger impact.</p></li></ul><p>However, magnitude should always be interpreted in context. A coefficient of 0.05 may be meaningful or negligible depending on the scale of the data.</p><h4><strong>What About the Intercept?</strong></h4><p>The intercept &#946;0&#8203; represents the predicted value of y when x = 0.</p><p>In this case:<br>&#128073; It is the predicted sales when TV spend is zero.</p><p>While mathematically necessary, the intercept is not always meaningful in real-world scenarios&#8212;especially if x = 0 is not a realistic value.</p><h4><strong>Multiple Variables: A Key Shift in Interpretation</strong></h4><p>So far, we&#8217;ve looked at a single variable. In practice, we often use multiple variables:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{y} = \\beta_0 + \\beta_1 x_1 + \\beta_2 x_2 + \\dots&quot;,&quot;id&quot;:&quot;CSOEZNSOCF&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, the interpretation changes slightly.</p><p>&#128073; Each coefficient now represents the effect of that variable <strong>holding all other variables constant</strong>.</p><p>For example:</p><blockquote><p>The coefficient for TV spend tells us how sales change with TV spend, assuming radio and newspaper spend stay the same.</p></blockquote><p>This is a subtle but important point&#8212;and one that is often tested in interviews.</p><h4><strong>Important Caveat: Correlation vs Causation</strong></h4><p>A coefficient does <strong>not</strong> imply causation.</p><p>Just because TV spend is associated with higher sales does not mean it is the sole cause. There could be:</p><ul><li><p>Other influencing factors</p></li><li><p>Confounding variables</p></li><li><p>Reverse relationships</p></li></ul><p>Linear regression captures relationships&#8212;not necessarily cause-and-effect.</p><p>&#128204; <strong>Pro tip:</strong> In interviews, a strong answer includes both interpretation and caution. Explaining what the coefficient means <em>and</em> what it doesn&#8217;t mean is a key differentiator.</p><div><hr></div><h2><strong>Statistical Significance and Model Fit: P-values and R&#178;</strong></h2><p>Once we understand what the coefficients mean, the next step is to evaluate how reliable those estimates are and how well the model fits the data. This is where <strong>p-values</strong> and <strong>R&#178;</strong> come in.</p><h4><strong>P-values: Is the Relationship Real?</strong></h4><p>In linear regression, each coefficient is associated with a hypothesis test:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;H_0: \\beta_1 = 0 \\quad \\text{vs} \\quad H_1: \\beta_1 \\neq 0&quot;,&quot;id&quot;:&quot;LPDSCTXIZS&quot;}" data-component-name="LatexBlockToDOM"></div><p>The p-value tells us how likely it is to observe a coefficient as large as the one we estimated, assuming the true effect is actually zero.</p><p>&#128073; In simple terms:</p><ul><li><p>A <strong>small p-value</strong> suggests that the relationship is unlikely to be due to random chance</p></li><li><p>A <strong>large p-value</strong> suggests that the relationship may not be meaningful</p></li></ul><p>A common threshold is 0.05, but this should not be treated as a strict rule.</p><p>However, it&#8217;s important to interpret p-values carefully:</p><ul><li><p>A small p-value does <strong>not</strong> mean the effect is large</p></li><li><p>It does <strong>not</strong> imply causation</p></li><li><p>It only tells us whether there is evidence of a relationship</p></li></ul><h4><strong>R&#178;: How Well Does the Model Explain the Data?</strong></h4><p>While p-values focus on individual coefficients, <strong>R&#178;</strong> looks at the model as a whole.</p><p>R&#178; measures the proportion of variance in the outcome that is explained by the model:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;R^2 = 1 - \\frac{\\sum (y_i - \\hat{y}_i)^2}{\\sum (y_i - \\bar{y})^2}&quot;,&quot;id&quot;:&quot;PCNAOHUVYN&quot;}" data-component-name="LatexBlockToDOM"></div><p>&#128073; Intuition:</p><ul><li><p>R&#178; = 0 &#8594; model explains nothing</p></li><li><p>R&#178; = 1 &#8594; model explains everything</p></li></ul><p>For example:</p><ul><li><p>R&#178; = 0.6 means the model explains 60% of the variation in sales</p></li></ul><h4><strong>How to Use Them Together</strong></h4><p>P-values and R&#178; answer different questions:</p><ul><li><p><strong>P-value</strong> &#8594; Is there evidence of a relationship?</p></li><li><p><strong>R&#178;</strong> &#8594; How much of the outcome does the model explain?</p></li></ul><p>A model can have:</p><ul><li><p>Significant coefficients but low R&#178; &#8594; relationship exists but weak explanatory power</p></li><li><p>High R&#178; but insignificant variables &#8594; potential overfitting or redundancy</p></li></ul><h4><strong>Why This Matters</strong></h4><p>Neither metric should be used in isolation.</p><p>A strong analysis looks at:</p><ul><li><p>Whether the relationships are statistically meaningful</p></li><li><p>Whether the model explains enough variation to be useful</p></li></ul><p>&#128204; <strong>Pro tip:</strong> In interviews, a strong answer doesn&#8217;t stop at &#8220;the p-value is small&#8221; or &#8220;R&#178; is high.&#8221; Explain what each tells you and how you would use both to assess the model.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2><strong>Common Pitfalls in Linear Regression</strong></h2><p>Linear regression is simple to use, but easy to misuse. Many mistakes don&#8217;t come from fitting the model, but from how the results are interpreted.</p><p>One of the most common pitfalls is confusing <strong>correlation with causation</strong>. A positive coefficient between TV spend and sales does not mean that increasing TV spend will necessarily cause sales to increase. There could be other factors influencing both variables. Linear regression captures relationships in the data&#8212;it does not prove cause and effect.</p><p>Another common issue is <strong>ignoring multicollinearity</strong>. When multiple input variables are highly correlated (for example, TV and radio spend increasing together), the model struggles to separate their individual effects. This can lead to unstable or misleading coefficients, even if the overall model seems to fit well.</p><p>It&#8217;s also easy to <strong>overinterpret R&#178;</strong>. A high R&#178; may give the impression that the model is strong, but it only tells you how well the model explains variation in the data&#8212;not whether the relationships are meaningful or reliable. Similarly, a low R&#178; does not necessarily mean the model is useless, especially in domains with inherently noisy data.</p><p>Another mistake is <strong>ignoring model assumptions</strong>. As we saw earlier, linear regression relies on assumptions like linearity and constant variance. If these assumptions are violated, the model may produce biased or unreliable results, even if the output looks reasonable.</p><p>Finally, there is the risk of <strong>overfitting or oversimplifying</strong>. A model with too many variables may fit the training data well but fail to generalize, while an overly simple model may miss important patterns in the data.</p><p>&#128204; <strong>Pro tip:</strong> In interviews, calling out even one or two of these pitfalls&#8212;especially correlation vs causation or multicollinearity&#8212;can strongly differentiate your answer. It shows that you understand not just how to use the model, but how it can fail.</p><div><hr></div><h2><strong>When to Use (and Not Use) Linear Regression</strong></h2><p>Linear regression is a powerful and widely used model&#8212;but it works best in the right situations.</p><p>At its core, it is designed to capture <strong>simple, linear relationships</strong> between variables. When that assumption holds, it can be both effective and highly interpretable.</p><h4><strong>When to Use Linear Regression</strong></h4><p>Linear regression works well when:</p><ul><li><p>The relationship between input and output is approximately <strong>linear</strong></p></li><li><p>You want a model that is <strong>easy to interpret and explain</strong></p></li><li><p>You are looking to understand how variables are related, not just make predictions</p></li><li><p>The number of features is relatively small and well-behaved</p></li></ul><p>In many real-world scenarios, especially as a baseline model, linear regression performs surprisingly well and provides valuable insights.</p><h4><strong>When Not to Use Linear Regression</strong></h4><p>Linear regression may not be the right choice when:</p><ul><li><p>The relationship between variables is <strong>non-linear</strong></p></li><li><p>There are <strong>complex interactions</strong> between features</p></li><li><p>The data has strong <strong>multicollinearity</strong></p></li><li><p>You are trying to capture <strong>causal relationships</strong> without proper controls</p></li><li><p>The assumptions of the model are clearly violated</p></li></ul><p>In these cases, more flexible models may provide better performance&#8212;but often at the cost of interpretability.</p><h4><strong>The Tradeoff</strong></h4><p>Choosing linear regression is often a tradeoff between:</p><ul><li><p><strong>Simplicity and interpretability</strong></p></li><li><p><strong>Flexibility and predictive power</strong></p></li></ul><p>Linear regression sits on the side of simplicity. It may not always be the most accurate model, but it is often the easiest to understand and communicate.</p><p>&#128204; <strong>Pro tip:</strong> In interviews, explicitly stating <em>when you would use linear regression and when you wouldn&#8217;t</em> is a strong signal. It shows that you can make thoughtful modeling decisions, not just apply techniques.</p><div><hr></div><h2><strong>Closing Thoughts</strong></h2><p>Linear regression is one of the simplest models in data science, but as we&#8217;ve seen, using it effectively requires more than just fitting a line.</p><p>In Part 1, we focused on how the model works&#8212;how it finds the line of best fit, what residuals represent, and the assumptions it relies on. In this part, we shifted to interpretation&#8212;understanding coefficients, evaluating model fit, and recognizing where the model can mislead.</p><p>Together, these pieces form a complete picture:<br>&#128073; <em>how the model works, what it tells you, and when you can trust it.</em></p><p>The goal is not to memorize formulas or metrics, but to build intuition. Once you understand what the model is doing and how to interpret its outputs, you can use it more confidently&#8212;whether in interviews or real-world problems.</p><p>Linear regression is often the starting point, but the thinking behind it&#8212;structured reasoning, careful interpretation, and awareness of limitations&#8212;is what carries forward to more complex models.</p><div><hr></div><h2><strong>What&#8217;s Next</strong></h2><p>In the next issue, we&#8217;ll shift focus to logistic regression for interviews&#8212;covering the questions that actually matter and how to answer them clearly.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Linear Regression (Part 1): Why It Still Works and How It Actually Works]]></title><description><![CDATA[A practical guide to understanding the line of best fit, model assumptions, and why this simple model remains powerful]]></description><link>https://thepracticaldatascientist.substack.com/p/linear-regression-part-1-why-it-still</link><guid isPermaLink="false">https://thepracticaldatascientist.substack.com/p/linear-regression-part-1-why-it-still</guid><dc:creator><![CDATA[Gowthami Peri]]></dc:creator><pubDate>Tue, 28 Apr 2026 14:03:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XZUh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XZUh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 424w, /__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 848w, /__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XZUh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png" width="1446" height="959" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:959,&quot;width&quot;:1446,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1709077,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 424w, /__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 848w, /__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XZUh!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946bba38-4e89-4786-bd7d-9b468f60bd22_1446x959.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>Introduction</strong></h1><p>Linear regression is one of the simplest models in data science&#8212;and one of the most widely used.</p><p>In a world of increasingly sophisticated models&#8212;gradient boosting, deep learning, and large language models&#8212;it&#8217;s natural to ask:</p><p>&#128073; <em>Why does linear regression still matter?</em></p><p>The answer is simple: <strong>it works surprisingly well in many real-world situations</strong>.</p><p>Linear regression is fast, interpretable, and often provides a strong baseline. More importantly, it forces you to think clearly about relationships in your data&#8212;what drives what, and by how much. This clarity is something more complex models often trade off for performance.</p><p>In practice, many problems don&#8217;t need complexity. A well-specified linear model can capture the majority of the signal, especially when relationships are approximately linear or when interpretability is critical.</p><p>This is also why linear regression shows up frequently in interviews. It&#8217;s not just about the model&#8212;it&#8217;s about whether you understand how models work, what assumptions they rely on, and how to reason about them.</p><p>In this article, we&#8217;ll focus on the foundations:</p><ul><li><p>what linear regression is</p></li><li><p>how the <strong>line of best fit</strong> is determined</p></li><li><p>what assumptions the model makes</p></li></ul><p>The goal is not to memorize formulas, but to build an intuitive understanding of how the model works and when it can be trusted.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2><strong>What Is Linear Regression? (An Intuitive View)</strong></h2><p>To make this concrete, let&#8217;s start with a simple example.</p><p>We&#8217;ll use a small dataset where we try to understand how <strong>TV advertising spend relates to sales</strong>. Each row represents a different market, with how much was spent on TV ads and the resulting sales.</p><p>Let&#8217;s first load the data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;df296270-5cdd-4d9c-bc5a-bcb686cb7d7c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import pandas as pd

url = "https://raw.githubusercontent.com/selva86/datasets/master/Advertising.csv"

df = pd.read_csv(url, index_col=0)
df.head()</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!w6fu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 424w, /__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 848w, /__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 1272w, /__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!w6fu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png" width="494" height="338" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad20e519-1f13-43d3-9276-a1a765854dda_494x338.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:338,&quot;width&quot;:494,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:31741,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 424w, /__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 848w, /__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 1272w, /__u/substackcdn.com/image/fetch/$s_!w6fu!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad20e519-1f13-43d3-9276-a1a765854dda_494x338.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Now let&#8217;s visualize the relationship between TV spend and sales:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e46502b9-1f8e-4167-8eba-8007e0f1e4c9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import matplotlib.pyplot as plt

plt.scatter(df['TV'], df['sales'])
plt.xlabel('TV Spend')
plt.ylabel('Sales')
plt.title('TV Spend vs Sales')
plt.show()</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!46Z6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 424w, /__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 848w, /__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 1272w, /__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!46Z6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png" width="1136" height="918" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:918,&quot;width&quot;:1136,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:126664,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 424w, /__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 848w, /__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 1272w, /__u/substackcdn.com/image/fetch/$s_!46Z6!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8651306e-c56f-42a6-a700-bdfd153c87ac_1136x918.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Each point in this plot represents an observation. On the x-axis, we have how much was spent on TV advertising, and on the y-axis, we have the resulting sales.</p><p>At a glance, you can already see a pattern:<br>&#128073; As TV spend increases, sales tend to increase as well.</p><p>But this relationship isn&#8217;t perfectly clean. The points don&#8217;t lie on a straight line&#8212;they&#8217;re scattered around. This is because real-world data always has some amount of noise and variability.</p><p>This is where linear regression comes in.</p><p>&#128073; <strong>Linear regression tries to capture this relationship with a simple line.</strong></p><p>The idea is straightforward:</p><ul><li><p>We assume that sales can be explained (at least partially) by TV spend</p></li><li><p>We try to draw a line that best represents the overall trend in the data</p></li></ul><p>This line won&#8217;t pass through every point, but it should summarize the general direction of the relationship.</p><p>In other words:</p><p>&#128073; Linear regression helps us move from a scattered set of points to a <strong>simple, interpretable relationship</strong> between variables.</p><p>This gives us two powerful capabilities:</p><ul><li><p><strong>Understanding</strong> &#8594; how changes in TV spend affect sales</p></li><li><p><strong>Prediction</strong> &#8594; estimating sales for a given level of TV spend</p></li></ul><p>In the next section, we&#8217;ll make this idea more precise by understanding what we mean by the <strong>line of best fit</strong>&#8212;and how the model actually finds it.</p><div><hr></div><h2><strong>The Line of Best Fit</strong></h2><p>From the previous section, we saw that the data forms a pattern&#8212;but it&#8217;s scattered. The goal of linear regression is to summarize this pattern with a single line.</p><p>&#128073; This line is called the <strong>line of best fit</strong>.</p><h4><strong>What Does &#8220;Best Fit&#8221; Mean?</strong></h4><p>The line of best fit is the line that <strong>best represents the relationship</strong> between the input (TV spend) and the output (sales). But what does &#8220;best&#8221; actually mean?</p><p>It means:<br>&#128073; The line that minimizes the difference between the actual values and the predicted values.</p><p>These differences are called <strong>residuals</strong>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{Residual} = y_i - \\hat{y}_i&quot;,&quot;id&quot;:&quot;KTUQZBFLOF&quot;}" data-component-name="LatexBlockToDOM"></div><p>Since some residuals are positive and some are negative, we square them and minimize the total:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\min \\sum (y_i - \\hat{y}_i)^2&quot;,&quot;id&quot;:&quot;IWKVOGHZTI&quot;}" data-component-name="LatexBlockToDOM"></div><p>&#128073; This is called <strong>least squares regression</strong>.</p><h4><strong>Fitting the Line in Code</strong></h4><p>Let&#8217;s now fit a linear regression model to our data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f2d00304-3284-4a26-8a32-040f7ce6ac35&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from sklearn.linear_model import LinearRegression

X = df[['TV']]
y = df['sales']

model = LinearRegression()
model.fit(X, y)</code></pre></div><p>This gives us a line of the form:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{y} = \\beta_0 + \\beta_1 x&quot;,&quot;id&quot;:&quot;IDAQUNOYCB&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where:</p><ul><li><p>&#946;0&#8203; = intercept</p></li><li><p>&#946;1&#8203; = slope</p></li></ul><p>Let&#8217;s print them:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2c7f4420-7f2c-4e1d-8e7b-deb16bdfcba1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">print("Intercept:", model.intercept_)
print("Slope:", model.coef_[0])</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Mw8s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Mw8s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png" width="248" height="46" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/930b76ed-e899-43e8-b133-7b7440d87806_248x46.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:46,&quot;width&quot;:248,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9455,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 424w, /__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 848w, /__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Mw8s!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F930b76ed-e899-43e8-b133-7b7440d87806_248x46.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Now let&#8217;s overlay the fitted line on top of our data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dd703d3b-f2ff-41ec-8267-95f0b145dc99&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">y_pred = model.predict(X)

plt.scatter(df['TV'], df['sales'], label='Actual Data')
plt.plot(df['TV'], y_pred, color='red', label='Best Fit Line')
plt.xlabel('TV Spend')
plt.ylabel('Sales')
plt.title('Linear Regression Fit')
plt.legend()
plt.show()</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!VIzD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 424w, /__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 848w, /__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 1272w, /__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!VIzD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png" width="616" height="470" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:470,&quot;width&quot;:616,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:43695,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 424w, /__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 848w, /__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 1272w, /__u/substackcdn.com/image/fetch/$s_!VIzD!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9032bd4c-4f02-402a-8acd-fe17c8fb9483_616x470.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This red line is the model&#8217;s best estimate of how sales change with TV spend.</p><p>&#128073; For any given TV spend, the line gives us a predicted sales value.</p><p>It does two things:</p><ul><li><p>Captures the overall trend in the data</p></li><li><p>Smooths out the noise</p></li></ul><p>You&#8217;ll notice:</p><ul><li><p>Some points lie above the line</p></li><li><p>Some lie below</p></li></ul><p>That&#8217;s expected. No model is perfect.</p><p>The key idea is:<br>&#128073; The line minimizes the total squared error across all points.</p><h4><strong>What This Line Represents</strong></h4><p>This red line is the model&#8217;s best estimate of how sales change with TV spend.</p><p>&#128073; For any given TV spend, the line gives us a predicted sales value.</p><p>It does two things:</p><ul><li><p>Captures the overall trend in the data</p></li><li><p>Smooths out the noise</p></li></ul><p>You&#8217;ll notice:</p><ul><li><p>Some points lie above the line</p></li><li><p>Some lie below</p></li></ul><p>That&#8217;s expected. No model is perfect.</p><p>The key idea is:<br>&#128073; <strong>The line minimizes the total squared error across all points.</strong></p><p>This simple line is doing a lot of work:</p><ul><li><p>It summarizes the relationship between variables</p></li><li><p>It enables prediction</p></li><li><p>It forms the foundation for understanding more complex models</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/subscribe"><span>Subscribe now</span></a></p></li></ul><div><hr></div><h2><strong>Understanding Residuals: What the Model Gets Wrong</strong></h2><p>Even though we&#8217;ve fitted a line, it doesn&#8217;t perfectly match every data point.</p><p>Some points lie above the line, and some lie below it. This difference between the actual value and the predicted value is called a <strong>residual</strong>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{Residual} = y_i - \\hat{y}_i&quot;,&quot;id&quot;:&quot;ZMNVPWATYW&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where:</p><ul><li><p>yi&#8203; is the actual value</p></li><li><p>y^&#8203;i&#8203; is the predicted value from the model</p></li></ul><p>In simple terms, the residual captures how far off the model&#8217;s prediction is for a given observation.</p><p>Residuals help us understand how well the model is performing. When the residuals are small, the predictions are close to the actual values, indicating a good fit. Larger residuals suggest that the model is not capturing the relationship well for those observations.</p><p>To better understand this, it&#8217;s useful to visualize residuals:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a4ef6951-c119-4ce2-841d-13e1bccdff15&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">residuals = y - y_pred

plt.scatter(df['TV'], residuals)
plt.axhline(y=0, color='red', linestyle='--')
plt.xlabel('TV Spend')
plt.ylabel('Residuals')
plt.title('Residual Plot')
plt.show()</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!MEr5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 424w, /__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 848w, /__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!MEr5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png" width="592" height="451" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:451,&quot;width&quot;:592,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:37000,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 424w, /__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 848w, /__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MEr5!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10ad733d-339c-47d8-af73-dcae7fa46630_592x451.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In a good model, residuals should be scattered randomly around zero without any clear pattern. This suggests that the model has captured the underlying relationship in the data.</p><p>When the model is not appropriate, residual plots often show patterns. For example:</p><ul><li><p>A curved pattern suggests the relationship is not linear</p></li><li><p>A widening or funnel shape suggests changing variance</p></li><li><p>Clusters may indicate missing variables or segments</p></li></ul><p>These patterns are signals that the model may need to be improved or replaced.</p><p>This is why residuals are more than just errors&#8212;they act as a diagnostic tool. They help you assess whether your model is appropriate and whether there are patterns the model is failing to capture.</p><p>&#128204; <strong>Pro tip:</strong> In interviews, mentioning residuals and what they reveal about the model is a strong signal. It shows that you understand not just how to fit a model, but how to evaluate it.</p><div><hr></div><h2><strong>Model Assumptions: When Linear Regression Works</strong></h2><p>So far, we&#8217;ve seen how linear regression fits a line and how residuals help us evaluate the model. But for these results to be reliable, the model relies on a few key assumptions.</p><p>These assumptions allow us to interpret the results correctly. When they don&#8217;t hold, the model may still produce outputs&#8212;but those outputs can be misleading.</p><h4><strong>Key Assumptions</strong></h4><p>Linear regression makes a few core assumptions about the relationship between inputs and outputs, as well as the behavior of the errors:</p><ul><li><p><strong>Linearity:</strong><br>The relationship between the input (TV spend) and output (sales) should be approximately linear. If the true relationship is curved or more complex, a straight line will systematically miss patterns in the data.</p></li><li><p><strong>Independence:</strong><br>Observations should be independent of each other. For example, sales in one market should not directly influence sales in another. This assumption is often violated in time-series data or when there are clusters (e.g., stores in the same region).</p></li><li><p><strong>Constant Variance (Homoscedasticity):</strong><br>The spread of residuals should be roughly the same across all values of the input. If variability increases with TV spend (a funnel shape), it indicates that the model&#8217;s errors are not consistent.</p></li><li><p><strong>Normality of Residuals (for inference):</strong><br>Residuals should be approximately normally distributed. This assumption is less important for prediction, but matters when interpreting statistical significance (e.g., p-values, confidence intervals).</p></li><li><p><strong>No Strong Multicollinearity (for multiple variables):</strong><br>When using multiple features, they should not be highly correlated with each other. If they are, it becomes difficult to isolate the individual effect of each variable.</p></li></ul><p>Rather than checking each assumption separately, we can visualize them together using a <strong>diagnostic plot</strong>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f700d941-7cce-483c-9c34-377d203bb482&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import matplotlib.pyplot as plt
import seaborn as sns
import scipy.stats as stats

residuals = y - y_pred

fig, axes = plt.subplots(2, 2, figsize=(12, 10))

# Residuals vs X (Linearity)
axes[0, 0].scatter(df['TV'], residuals)
axes[0, 0].axhline(y=0, color='red', linestyle='--')
axes[0, 0].set_title('Residuals vs TV Spend')

# Residuals vs Predictions (Variance)
axes[0, 1].scatter(y_pred, residuals)
axes[0, 1].axhline(y=0, color='red', linestyle='--')
axes[0, 1].set_title('Residuals vs Predictions')

# Histogram (Normality)
sns.histplot(residuals, kde=True, ax=axes[1, 0])
axes[1, 0].set_title('Residual Distribution')

# QQ Plot (Normality)
stats.probplot(residuals, dist="norm", plot=axes[1, 1])
axes[1, 1].set_title('QQ Plot')

plt.tight_layout()
plt.show()</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!EGpg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 424w, /__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 848w, /__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_webp, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!EGpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png" width="1209" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1209,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:121075,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thepracticaldatascientist.substack.com/i/195582852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_424, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 424w, /__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_848, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 848w, /__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_1272, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 1272w, /__u/substackcdn.com/image/fetch/$s_!EGpg!, /__u/thepracticaldatascientist.substack.com/w_1456, /__u/thepracticaldatascientist.substack.com/c_limit, /__u/thepracticaldatascientist.substack.com/f_auto, /__u/thepracticaldatascientist.substack.com/q_auto:good, /__u/thepracticaldatascientist.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c7c806c-5608-4dee-b271-e5040ee04917_1209x983.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Each part of this plot helps diagnose a different assumption:</p><ul><li><p><strong>Top-left:</strong> checks whether the relationship is linear</p></li><li><p><strong>Top-right:</strong> checks whether variance is constant</p></li><li><p><strong>Bottom-left &amp; bottom-right:</strong> check whether residuals follow a normal distribution</p></li></ul><p>A well-behaved model will show residuals that are randomly scattered around zero, with no clear patterns and a roughly symmetric distribution.</p><p>These assumptions determine whether you can trust your model.</p><p>If they hold:</p><ul><li><p>The model is stable and interpretable</p></li><li><p>Coefficients reflect meaningful relationships</p></li><li><p>Statistical measures (like p-values) are reliable</p></li></ul><p>If they are violated:</p><ul><li><p>The model may miss important patterns</p></li><li><p>Coefficients may be misleading</p></li><li><p>Predictions may be unreliable</p></li></ul><p>This is why checking assumptions is just as important as fitting the model itself.</p><div><hr></div><h2><strong>Closing Thoughts</strong></h2><p>Linear regression is often introduced as a simple model&#8212;but as we&#8217;ve seen, there&#8217;s more to it than just fitting a line.</p><p>From understanding the <strong>line of best fit</strong> to analyzing <strong>residuals</strong> and checking <strong>assumptions</strong>, each step plays a role in determining whether the model is actually capturing a meaningful relationship. Without these checks, it&#8217;s easy to rely on a model that looks reasonable on the surface but fails to represent the data accurately.</p><p>The goal of this first part was to build that foundation&#8212;how the model works and when it can be trusted. Once you have that clarity, the next step becomes much more interesting:</p><p>&#128073; <em>What is the model actually telling us?</em></p><p>In Part 2, we&#8217;ll focus on interpreting the results&#8212;breaking down coefficients, p-values, and R-squared, and understanding how to use them to make decisions.</p><div><hr></div><h2><strong>What&#8217;s Next</strong></h2><p>In the next issue, we&#8217;ll have a Part 2 where the focus will be on interpreting the results&#8212;breaking down coefficients, p-values, and R-squared, and understanding how to use them to make decisions.</p><p><em><strong>Keep building, keep learning&#8212;wishing you the best in your data journey.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/p/linear-regression-part-1-why-it-still?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thepracticaldatascientist.substack.com/p/linear-regression-part-1-why-it-still?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thepracticaldatascientist.substack.com/p/linear-regression-part-1-why-it-still?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p></p>]]></content:encoded></item></channel></rss>