<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Data Espresso]]></title><description><![CDATA[Data engineering updates and commentary to accompany your afternoon espresso.]]></description><link>https://dataespresso.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!md3X!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png</url><title>Data Espresso</title><link>https://dataespresso.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 18:19:07 GMT</lastBuildDate><atom:link href="/__u/dataespresso.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Mahdi Karabiben]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[dataespresso@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[dataespresso@substack.com]]></itunes:email><itunes:name><![CDATA[Mahdi Karabiben]]></itunes:name></itunes:owner><itunes:author><![CDATA[Mahdi Karabiben]]></itunes:author><googleplay:owner><![CDATA[dataespresso@substack.com]]></googleplay:owner><googleplay:email><![CDATA[dataespresso@substack.com]]></googleplay:email><googleplay:author><![CDATA[Mahdi Karabiben]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[☕ #15: The uncomfortable future of internal data tools, and the bigger problem for BI]]></title><description><![CDATA[What utility do internal data tools actually serve in an AI-Agent-powered world? And why is BI, with its nuanced history, a completely different story?]]></description><link>https://dataespresso.substack.com/p/15-the-uncomfortable-future-of-internal</link><guid isPermaLink="false">https://dataespresso.substack.com/p/15-the-uncomfortable-future-of-internal</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 24 Jun 2026 14:50:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!md3X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>First off, I promise no World Cup references in this edition&#8212;except to say that <a href="https://www.globalchalice.com/">the Neo4j-powered Global Chalice</a> is absolutely worth checking out if you enjoy mixing data with <s>soccer</s> football.</p><p>The first part of this edition will cover a broad topic: internal data tools and how their position and value are (very) quickly evolving as AI disrupts software in general. The second part will then zoom into the implications of this shift for one category of internal tooling: BI, and my perspective on how &amp; why the acronym should shift from business intelligence to business intent. This connects with my recent article in <span>Towards Data Science:&nbsp;</span><em><a href="https://towardsdatascience.com/bi-is-dead-long-live-bi/">"BI is Dead, Long Live BI</a></em><span>."</span></p><p>So, without further ado, let&#8217;s talk data &amp; analytics while the espresso is still hot.</p><div><hr></div><h2>Internal tools, anyone?</h2><p>Historically, and specifically during the Modern Data Stack era, data teams were particularly spoiled in terms of the footprint of their internal tooling. The data stack was usually a collection of 20+ tools and open-source packages, moving, transforming, and interacting with data across a lengthy (and often unnecessarily complex) path.</p><p>This tool sprawl resulted in a bloated stack with many moving pieces, but also many rough edges where a gap between the features of two tools remains open or a critical piece of functionality remains elusive. An ideal solution to circumvent this would be developing lean in-house tools that fill these specific gaps, but realistically, no one had time for that&#8212;we were all, collectively, jumping from one broken dbt model to another while updating a query filter and adding a new Fivetran connector to a spaghetti-like Airflow DAG.</p><p><a href="/__u/dataespresso.substack.com/p/espresso-11-ai-as-a-last-mile-enabler">Around this time last year, I wrote about how AI can be a true enabler for these &#8220;missing last-mile pieces&#8221;</a> because it writes software that&#8217;s good enough for internal usage, where bugs are usually acceptable and downstream dependencies are limited (so tech debt is cheap). I used the example of a dbt SLA tracker tool. The tool aimed to recreate <a href="https://medium.com/airbnb-engineering/visualizing-data-timeliness-at-airbnb-ee638fdf4710">Airbnb&#8217;s brilliant data timeliness UI</a>, which features sophisticated lineage and neat visual styling to show which specific jobs/steps are causing SLA misses for a given data pipeline.</p><p>The result last year was (barely) acceptable. I had functioning lineage at the dbt-model-level with nice coloring based on model run metadata, but it was clunky and lacked all other functionality. Getting to that result took a few hours of vibe coding with Replit + Claude Sonnet 3.7.</p><p>This year, around the same time, knowing the massive leap AI models had made in coding, I tried the exact same exercise with Claude Code + Claude Opus 4.8 Max. The results were simply mind-blowing. With just the Airbnb blog post and my ideas as a starting point, it built an end-to-end tool with a very sophisticated (yet elegant) and intuitive UI that covers all key features discussed in the blog post. I was able to add rather complex functionalities (running via Docker, refreshing data via a button, providing YAML config, etc.) with just one prompt at a time, without any significant corrections or back-and-forth. (The full GitHub project is available <a href="https://github.com/mahdiqb/dbt_timeline_analysis">here</a>.)</p><ul><li><p>The overview page for a given dbt project:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bA7F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 424w, /__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 848w, /__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bA7F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png" width="1456" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Report view &#8212; every tracked dataset, 30 days of landing times at a glance&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Report view &#8212; every tracked dataset, 30 days of landing times at a glance" title="Report view &#8212; every tracked dataset, 30 days of landing times at a glance" srcset="/__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 424w, /__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 848w, /__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bA7F!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe39f93c2-093b-48ad-b0bb-b7b5b4b55922_2880x1800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Overview page</figcaption></figure></div><ul><li><p>The timeline page for a particular run:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!i4LN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 424w, /__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 848w, /__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!i4LN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png" width="1456" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Timeline view &#8212; the full lineage Gantt with the bottleneck path highlighted&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Timeline view &#8212; the full lineage Gantt with the bottleneck path highlighted" title="Timeline view &#8212; the full lineage Gantt with the bottleneck path highlighted" srcset="/__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 424w, /__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 848w, /__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 1272w, /__u/substackcdn.com/image/fetch/$s_!i4LN!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a963b52-ad0b-447b-bb04-70c3d849827d_2880x1800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Timeline view</figcaption></figure></div><ul><li><p>The other pages of the app exhibit the same level of &#8220;sharpness&#8221;, all courtesy of Claude Opus 4.8 raw capabilities.</p></li></ul><p>This result, given the current capabilities of state-of-the-art models, isn&#8217;t surprising (and partly made Anthropic&#8217;s near-1-trillion valuation somewhat reasonable). But it made me realize an uncomfortable fact about where the future of internal tools lies.</p><p>Claude built a fantastic UI, sure, with a brilliantly engineered backend (I checked!). But at the end of the day, my actual goal of &#8220;<em>which models are slowing my pipeline and what are my SLA trends?</em>&#8221; becomes totally answerable without <em>any</em> of these mechanics. They turn into redundant code that serves no purpose other than vanity &#8220;<em>oh cool UI!</em>&#8221; dynamics. Claude can just look at the raw dbt runs metadata, figure out everything worth figuring out (either on its own or based on my guidance), tell me about it, and&#8212;in 95% of cases&#8212;make the most optimal code and process improvements directly.</p><p>Internal tools were built for a world where humans needed to find the insight on their own, making every visual cue helpful. But <em><strong>AI agents don&#8217;t need a UI</strong></em>; they extract insights directly from the raw data. So, as access to agents gets democratized, internal tools ecosystem... quo vadis?</p><h2>The BI link</h2><p>This realization is even more pertinent when we think about BI tooling specifically, which for decades (disappointingly) served the sole purpose of &#8220;<em>this data says X, and I can&#8217;t tell you anything beyond that or why you should care</em>&#8221;. If Claude can look at the data and surface everything worth surfacing, what&#8217;s the benefit of maintaining an expensive and demanding BI platform?</p><p>The nuance here, however, is that running too quickly to the &#8220;<em>ok let&#8217;s replace BI with AI chatbots</em>&#8221; path misses the whole point of BI. The key question all BI tools ultimately should aim to help answer isn&#8217;t &#8220;what&#8217;s the data answer to this specific question?&#8221;, but instead &#8220;<em><strong>based on the data, what questions should I be asking?</strong></em>&#8221;.</p><p>Right now, dashboards suffer from what I call the &#8220;closed-window problem&#8221;: they only show you what someone already decided to measure. They can&#8217;t surface the patterns no one thought to ask about. To actually answer &#8220;<em><strong>what should I be looking at?</strong></em>&#8221;, we need to stop optimizing the query layer and start building for the <em><strong>business intent layer</strong></em> (defining the business outcome, and letting the agent watch the data autonomously).</p><p>This exact shift (moving the acronym from Business Intelligence to Business Intent) is the focus of <a href="https://towardsdatascience.com/bi-is-dead-long-live-bi/">my latest article in Towards Data Science</a>. In it, I dive into the future of BI in an AI-Agent-powered world: where it&#8217;s currently headed, why intent is the missing piece, and what the new model actually looks like in practice.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Hope you enjoyed this edition of Data Espresso! If you found it useful, feel free to share it with fellow data folks in your network.</p><p>Feedback is always appreciated as usual, so feel free to share your thoughts in the comments or reach out directly &#8211; I&#8217;d love to hear your take on the edition&#8217;s topics!</p><p>Until next time, stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[☕ #14: The Data Team's Survival Kit for the Next Era of Data]]></title><description><![CDATA[Six pillars for a world where AI agents are the main data platform consumer: decluttering the stack, adopting product thinking, building an AI-ready context layer, and more.]]></description><link>https://dataespresso.substack.com/p/14-the-data-teams-survival-kit-for</link><guid isPermaLink="false">https://dataespresso.substack.com/p/14-the-data-teams-survival-kit-for</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 18 Mar 2026 14:15:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!md3X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>We are officially settling into 2026, and I hope the start of the year has been kind to you all!</p><p>This edition will be a deep dive into just one topic. I&#8217;ve spent the last few weeks reflecting on where data teams stand today, and I believe we are at a critical crossroads: We are paralyzed by tech debt and &#8220;model sprawl&#8221; (thanks MDS!) right when the business needs us to be more agile than ever. So, instead of a news digest, I&#8217;m sharing a survival kit for the next era of data&#8212;covering the strategic pivots we need to make to escape the &#8220;service trap&#8221; we&#8217;re in today and build the missing foundations for an era in which AI agents will be the main (only?) data consumer.</p><p>So, without further ado, let&#8217;s talk data &amp; analytics while the espresso is still hot.</p><div><hr></div><h2>The Survival Kit: 6 Imperatives for the Agent Era</h2><p>The dust from the Modern Data Stack era is finally settling, and it has left us staring at a very uncomfortable reality: <strong>the architecture we spent the last five years building is fundamentally incompatible with the future we are now living in.</strong></p><p>For half a decade, we (data teams) optimized for volume. We scaled our warehouses, treated every problem as a nail that needed a dbt hammer, and bought a specialized SaaS tool for every minor step in the pipeline. We built an ecosystem that was easier to manage than Hadoop, sure, but we optimized for <strong>human eyeballs</strong>&#8212;building dashboards that, let&#8217;s be honest, often went unread.</p><p>But the new primary consumer of data isn&#8217;t a human analyst anymore; it&#8217;s the <strong>AI Agent</strong>. And unlike a human, an agent doesn&#8217;t care about your star schema or your &#8220;best-of-breed&#8221; lineage tool if it lacks the unified context to reason about the data.</p><p>Yet, instead of moving fast to prep for this shift, teams are bogged down managing thousands of fragile models and a sprawling vendor list. We were caught off guard: We built a stack for 2021, and it is failing us in 2026. (For a group of people who work with predictions daily, we turned out to be terrible at predicting our own future.)</p><p><strong>And so, fellow data person, it is now time to pivot.</strong></p><p>In this month&#8217;s deep dive, I am sharing a <strong>survival kit for 2026 and beyond</strong>, consisting of 6 strategic imperatives that data teams need to adopt to re-align the data stack with the new era of use cases:</p><ol><li><p><strong>Put the Stack on a Diet:</strong> The era of buying a specialized SaaS tool for every feature is over. If your platform offers it natively, use it&#8212;consolidation is the only way to lower the &#8220;cognitive tax&#8221; enough to move fast.</p></li><li><p><strong>True Decoupling (Storage is Yours):</strong> To be future-proof, your central gravity must shift to Open Table Formats (Iceberg FTW!). Your data should live in a neutral state that any engine (Snowflake, Spark, or something completely new) can plug into without a migration.</p></li><li><p><strong>Stop Being a Service, Start Being a Product:</strong> The &#8220;ticket-taking&#8221; model is the wrong path forward. You cannot survive by answering every ad-hoc question; trust is won through focused excellence on a few high-value Data Products, not mediocre ubiquity.</p></li><li><p><strong>The Context Library:</strong> An AI agent is context-blind. We need a semantic layer that captures business logic, not just metrics&#8212;and the best way to build it is to have AI write the initial documentation for you.</p></li><li><p><strong>From &#8220;What Happened?&#8221; to &#8220;What Now?&#8221;:</strong> We must stop using data just to validate past decisions. The goal is now <strong>Navigation</strong>&#8212;automating the response to a change in a given metric. This requires shifting from the passive BI model to automated AI-powered processes that perform data-driven actions.</p></li><li><p><strong>The New Persona:</strong> &#8220;I write SQL&#8221; is no longer a career. The new data role moves <strong>Upstream</strong> (enforcing governance with engineers) and <strong>Downstream</strong> (acting as a PM to ensure data drives workflows).</p></li></ol><p>I unpack the nuance of all these points (and more) in <a href="https://towardsdatascience.com/the-data-teams-survival-guide-for-the-next-era-of-data/">the full article on Towards Data Science</a>. If you are trying to figure out how to position your team&#8212;or your own career&#8212;for the next 12 months and beyond, this is the full blueprint.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Hope you enjoyed this edition of Data Espresso! If you found it useful, feel free to share it with fellow data folks in your network.</p><p>Feedback is always appreciated as usual, so feel free to share your thoughts in the comments or reach out directly&#8212;I&#8217;d love to hear your take on the edition&#8217;s topics!</p><p>Until next time, stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[☕ #13: The Semantic Layer Renaissance, The Death of the “Shopping List” Architecture, and The Big/Small Metadata Question]]></title><description><![CDATA[A recipe for an AI-ready semantic layer, the new dynamic of the data stack's tools and categories, and a counter-thesis to distributed metadata.]]></description><link>https://dataespresso.substack.com/p/13-the-semantic-layer-renaissance</link><guid isPermaLink="false">https://dataespresso.substack.com/p/13-the-semantic-layer-renaissance</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 10 Dec 2025 14:32:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!md3X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>We meet again after a long(-ish) break, mainly due to a busy few months on both the personal and professional sides. As you might&#8217;ve noticed if we&#8217;re connected on LinkedIn, I left my role at Sifflet and joined Neo4j to work on graph analytics (something I&#8217;ve always been passionate about!) - this won&#8217;t impact the content of the newsletter: I&#8217;ll keep discussing the overall data and analytics space and how to build data products &amp; platforms at scale. (But maybe we&#8217;ll have a graph-related &#8220;twist&#8221; here and there, since I&#8217;m getting more familiar with graph analytics use cases.)</p><p>Now let&#8217;s briefly go through what we&#8217;ll cover in this month&#8217;s edition: we&#8217;ll first discuss the semantic layer renaissance (really, it&#8217;s back in style!) and why we need a new approach (other than defining aggregations and joins in YAML) for an AI-ready semantic layer, then we&#8217;ll go through how the data stack is shifting into a &#8220;hub and spike&#8221; model instead of a &#8220;shopping list&#8221; architecture, and finally we&#8217;ll end with a note on DuckDB&#8217;s new release: DuckLake. </p><p>So, without further ado, let&#8217;s talk data &amp; analytics while the espresso is still hot.</p><div><hr></div><h2><strong>The Semantic Layer Renaissance (and why it&#8217;s actually different this time)</strong></h2><p>Back in 2021/2022, during the peak of the Modern Data Stack craze (R.I.P.), the semantic layer felt a bit like&#8230; &#8220;unnecessary&#8221; homework.</p><p>In theory it was an essential component, sure, and I was a strong believer in its value - but it ultimately consisted of thousands of YAML lines just to define how <code>customers</code> joins with <code>orders</code>. The promise was &#8220;metric consistency,&#8221; but deep down, many of us knew it was overkill: The use cases weren&#8217;t there yet, and all the semantic layer was doing was adding friction between a BI tool (where analytics lived at the time) and a data warehouse. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!y3nH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 424w, /__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 848w, /__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!y3nH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png" width="1400" height="748" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:748,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 424w, /__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 848w, /__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y3nH!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe50f41-2e5a-45b1-ae03-12d7e92c2c95_1400x748.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Why the push for semantic layers in 2022 was difficult to justify</figcaption></figure></div><p>The friction was unnecessary because human analysts are smart - they have implicit context. They know not to count test accounts in churn figures. They know who to ask if a number looks weird. Building a massive configuration layer just to save them a <code>JOIN</code> often didn&#8217;t feel like high ROI.</p><p><strong>But the landscape has shifted.</strong></p><p>We are moving from a world of <strong>Analytics</strong> (humans looking at charts) to <strong>Agents</strong> (AI taking actions). And unlike human data analysts, an LLM is context-blind: it doesn&#8217;t know your business nuance. So if you feed it raw SQL snippets or simple metric definitions, it <em>will</em> hallucinate.</p><p>This is where the need for a &#8220;new&#8221; semantic layer comes in: it&#8217;s no longer about optimizing for SQL generation; it&#8217;s about optimizing for <strong>reasoning</strong>. It shifts the job from defining the <em>Math</em> (&#8221;Here is the formula for Churn&#8221;) to defining the <em>Meaning</em> (&#8221;Here is <em>why</em> we track Churn, <em>who</em> owns it, and <em>what</em> business goal it impacts&#8221;).</p><p>A couple of weeks ago, I wrote a deep dive on Medium about this shift. If you are building for AI, or just tired of writing YAML for dashboards that nobody uses, give it a read: <strong><a href="https://blog.dataengineerthings.org/semantic-layer-for-ai-beyond-sql-aae652837a5a">Building a Semantic Layer for the AI Era: Beyond SQL Generation</a>.</strong></p><h2>The Death of the &#8220;Shopping List&#8221; Architecture: Why You Just Need a Platform (and a Few Exceptions)</h2><p>For years, we treated data architecture diagrams as a shopping list - the standard advice was to buy (or build, for the brave souls) a specialized &#8220;Best-of-Breed&#8221; tool for every single area of the stack: A data catalog, an orchestrator, a data observability tool, etc. But in 2025, with data platforms like Snowflake and Databricks rebundling the stack (via features that cover everything from orchestration to metadata management), the &#8220;baseline&#8221; requirements for governance, observability, and execution are now built-in - especially given how AI is minimizing the surface area of human involvement in data work. And so if you&#8217;re just starting out, buying a dedicated catalog before there&#8217;s a concrete need for it is definitely not the best place to invest money or resources.</p><p>Instead, we are moving toward a <strong>&#8220;Hub and Spike&#8221;</strong> architecture. The &#8220;Hub&#8221; is your massive, all-encompassing platform that handles the boring, day-to-day utility work. The &#8220;Spikes&#8221; are the specialized tools you buy <em>only</em> when you hit a specific, wall-sized problem that the platform can&#8217;t solve - like orchestrating a job across an on-prem legacy component and Salesforce, or governing 20 years of legacy Oracle data.</p><p>This marks a subtle but critical shift: <strong>Yesterday, we bought tools to build a </strong><em><strong>Stack</strong></em><strong>. Today, we buy tools to fix a </strong><em><strong>Spike</strong></em><strong>.</strong> You no longer buy a tool just to &#8220;fill a category&#8221; (e.g., &#8220;we need a catalog&#8221;); you buy it to solve a bounded, painful edge case that your central hub can no longer contain. In this context, my recommendation (for most cases) is the following: start with the platform, and earn the right to buy the specialist.</p><h2>The Duck Is Out of the Lake: What if Metadata Doesn&#8217;t Need to Be Distributed?</h2><p>DuckDB disrupted the data industry a few years ago with a simple yet brilliant observation: <strong>most &#8220;Big Data&#8221; isn&#8217;t actually big.</strong> They proved that for 99% of use cases, data fits on a single machine, so using a distributed cluster makes no sense (to an extent). Now, with the introduction of <strong>DuckLake</strong>, they are applying that exact same philosophy to the Data Lakehouse architecture.</p><p>While the new wave of distributed table formats (hi Apache Iceberg &#128075;) treat metadata as &#8220;Big Data&#8221; (scattering it across thousands of distributed files in object storage), DuckLake asks the obvious question: <strong>Why distribute the metadata if it doesn&#8217;t need to be distributed?</strong> Their approach keeps the <em>data</em> distributed (Parquet in S3) but centralizes the <em>metadata</em> in a standard, ACID-compliant database. It recognizes that while your row count might be massive, your catalog operations are almost always &#8220;small data.&#8221;</p><p>This adds a much-needed nuance to the stack, challenging the &#8220;distributed everything&#8221; default that dominated the Iceberg/Delta/Hudi conversation. It&#8217;s a healthy correction towards simplicity that aligns better with how most teams actually operate. If you want to dive deeper into why centralized metadata might be the future (again), MotherDuck&#8217;s CEO Hannes Mu&#776;hleisen gave <a href="https://www.youtube.com/watch?v=DxwDaoUijTc">a fantastic talk</a> breaking it down.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Hope you enjoyed this edition of Data Espresso! If you found it useful, feel free to share it with fellow data folks in your network.</p><p>Feedback is always appreciated as usual, so feel free to share your thoughts in the comments or reach out directly &#8211; I&#8217;d love to hear your take on the edition&#8217;s topics!</p><p>Until next time, stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #12: Data modeling for data products, a Spotify Wrapped for everything, and building things that matter]]></title><description><![CDATA[A modern playbook for data modeling in a product-driven world, how AI can power a supercharged Spotify Wrapped, and a two-step formula for building valuable data products.]]></description><link>https://dataespresso.substack.com/p/espresso-12-data-modeling-for-data</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-12-data-modeling-for-data</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Fri, 01 Aug 2025 14:01:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!md3X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>This month, we&#8217;re diving into the big comeback of data modeling (and how to adapt it to today&#8217;s data-product-driven world), the data product we all need (analytics and insights combining data from all the apps we use), and how to find data problems worth solving. So, without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><div><hr></div><h2>Data Modeling for Data Products: A Practical Guide</h2><p>For the last couple of years, it feels like we&#8217;ve all time-traveled back to 1999. Nearly everyone in the data space is talking about Kimball, Inmon, and Data Vault again. This isn&#8217;t just nostalgia; it&#8217;s a very reasonable (and much-needed) reaction to the chaos of the last decade, where we threw endless compute and thousands of dbt models at every problem without a coherent strategy. After data budgets got much tighter in 2022, we all collectively realized that we needed a process to clean up the mess.</p><p>But here&#8217;s the tricky part: while we absolutely need the discipline of modeling, we can&#8217;t just copy-paste the old playbook. The slow, rigid, waterfall approach that defined the data warehousing era is not a great fit for modern data teams trying to efficiently build and ship data products at scale.</p><p>So, how do we bridge that gap? How do we take the timeless principles of good modeling and adapt them for a decentralized, product-driven world?</p><p>I decided to write down my full thoughts on this in <a href="https://blog.dataengineerthings.org/data-modeling-for-data-products-a-practical-guide-2db003cc7e72">a long-form article for Data Engineering Things</a>. It&#8217;s a practical guide that lays out a modern playbook for data modeling, covering (among other things):</p><ul><li><p>A "Go Wide, then Go Deep" strategy for balancing big-picture business alignment with the focused work of building a specific data product.</p></li><li><p>Why decentralized ownership isn&#8217;t just a buzzword, but a prerequisite for making this approach work.</p></li><li><p>The specific tools and frameworks (Metric Trees, the Semantic Layer, etc.) that help turn these principles into practice.</p></li></ul><p>It&#8217;s my take on how we move forward: keeping the good parts of modeling without the old-school baggage. If you&#8217;re grappling with these questions on your own team, I think you&#8217;ll find it useful. (You can read the full article <a href="https://blog.dataengineerthings.org/data-modeling-for-data-products-a-practical-guide-2db003cc7e72">here</a>.)</p><h2>A Spotify Wrapped for Everything Else</h2><p>Every year, Spotify Wrapped comes out and provides users with a wide range of analytics about the music they listened to throughout the year - and unsurprisingly, people love it. It&#8217;s a genuinely great data product that gives you a fun, insightful look at your own habits. But it&#8217;s also a rare exception. We use tens (hundreds?) of apps and services on a daily basis, yet we actually get back very little data from them.</p><p>What were my most productive hours according to my calendar and code commits? How did my reading habits on Kindle correlate with my running activity on Strava? What does my Uber history say about my social life? Some apps do sprinkle high-level metrics here and there, but there are mountains of personal data that are yet to be used for anything other than ads.</p><p>Getting the answers today would mean wrestling with a dozen different APIs and trying to stitch it all together myself. Unfortunately, but unsurprisingly, nobody has time for that.</p><p>This is where I think the current AI wave gets interesting. Connecting the dots across various services and generating useful insights about our habits and who we are is a very interesting "last-mile" problem for AI. The idea of a personal AI agent that can securely connect to the APIs of all the tools I use, pull the data, and just <em>tell me</em> something interesting about myself feels... actually useful.</p><p>It&#8217;s not about creating more dashboards. It&#8217;s about getting a coherent narrative out of the fragmented data of our own lives. That&#8217;s a data product I&#8217;d actually want to use.</p><p>(This is after all a data engineering newsletter, so I can&#8217;t talk about Spotify Wrapped without mentioning <a href="https://engineering.atspotify.com/2021/2/how-spotify-optimized-the-largest-dataflow-job-ever-for-wrapped-2020">the fascinating article about the data engineering magic behind it</a>.)</p><h2>Finding Data Problems that are Worth Solving</h2><p>For a myriad of reasons, data teams built a reputation for being disconnected (to an extent) from the business. The narrative is that we get excited about the tech and the new tools - and sometimes lose sight of whether we&#8217;re actually having an impact. (i.e. we end up building things that are technically impressive but practically useless.)</p><p>Recently I came across a great article by Sven Balnojan called <a href="https://svenbalnojan.medium.com/why-internal-data-teams-build-the-wrong-things-8ff5d3129a9d">"Why Internal Data Teams Build the Wrong Things"</a>, in which he provides 15 extremely relevant and useful learnings to ensure you don&#8217;t walk the wrong path.</p><p>But IMO we can distill things even further. In my experience, breaking out of this cycle and building things that are valuable for the business can come down to a simple, two-step path.</p><p>First, you have to find the low-hanging fruit to win some initial momentum<strong>.</strong> There&#8217;s always a team that is obviously starved for data. A classic example is the marketing team trying to figure out which campaigns are actually working. Building them a solid attribution model is a clear, contained project with an obvious business value. It&#8217;s a quick win that builds trust and opens new doors with other stakeholders who may have yet more interesting use cases.</p><p>But the real magic happens with the second step: truly understanding the business by diving into the details, and finding <em>new</em> (relevant) data use cases. This means going beyond shipping what&#8217;s obvious/requested, and instead watching how the business operates and identifying gaps where the data can shine.</p><p>A perfect example is product analytics. Most product teams have a tool like Amplitude or Mixpanel that gives them great high-level metrics. But the real gold (the raw, granular event stream) is often just sitting untouched in an S3 bucket or a production database. The high-value data product here isn&#8217;t about replacing the existing tool, but building a much richer analytics layer on top of that raw data. By modeling it properly, the data team can unlock granular, feature-level insights that are impossible to get otherwise. This is how you empower the product team to go from asking "How many active users do we have?" to "How does the adoption of our new search filter impact 30-day retention for our enterprise customers?" while also building a foundation for a myriad of new vertical use cases for product data.</p><p>These are the kinds of opportunities you only find when you go looking for them - by observing how the business actually runs and identifying the real-world friction that data can solve.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Hope you enjoyed this edition of Data Espresso! If you found it useful, feel free to share it with fellow data folks in your network.</p><p>Feedback is always appreciated as usual, so feel free to share your thoughts in the comments or reach out directly &#8211; I&#8217;d love to hear your take on the edition&#8217;s topics!</p><p>Until next time, stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #11: AI as a "last-mile" enabler, dbt Fusion engine, and navigating the thin layer between fact and BS]]></title><description><![CDATA[How AI is finally democratizing the Data Platform&#8217;s &#8220;last-mile&#8221; layer, dbt&#8217;s new Fusion engine, and why a healthy dose of skepticism is always needed when it comes to data.]]></description><link>https://dataespresso.substack.com/p/espresso-11-ai-as-a-last-mile-enabler</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-11-ai-as-a-last-mile-enabler</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Mon, 02 Jun 2025 13:03:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>This month, we&#8217;re diving into how AI is finally democratizing the Data Platform&#8217;s &#8220;last-mile&#8221; layer, dbt&#8217;s new Fusion Engine, and why a healthy dose of skepticism is always needed when it comes to data. So, without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><div><hr></div><h2>How AI is finally paving the data platform&#8217;s "last mile" for everyone</h2><p>We&#8217;ve all seen the incredible advancements in the data space over the past years &#8211; powerful warehouses with limitless scale, slick ELT tools, and dbt bringing software engineering rigor to data pipelines. Yet, that final layer of polish, the seamless &#8220;last mile&#8221; experience seen in Big Tech tools (like <a href="https://medium.com/airbnb-engineering/visualizing-data-timeliness-at-airbnb-ee638fdf4710">Airbnb&#8217;s data timeliness UIs</a> or <a href="https://netflixtechblog.com/notebook-innovation-591ee3221233">Netflix&#8217;s integrated notebooks</a>), still feels out of reach for most data teams, often a luxury only Big Tech could afford. Even with all the capabilities of the Modern Data Stack, data teams are still swamped with firefighting and ad-hoc requests, leaving little room for building these bespoke experience layers.</p><p>But what if that&#8217;s changing?</p><p>In my latest article, <strong>"<a href="https://medium.com/data-science-collective/how-ai-is-finally-democratizing-the-data-platforms-last-mile-layer-8b311aa3c107">How AI is Finally Democratizing the Data Platform&#8217;s Last-Mile Layer</a>"</strong>, I dive into how AI is emerging as a powerful &#8220;last-mile enabler&#8221;. I experienced this firsthand via a personal experiment that I discuss in the article: I persisted the artifacts of a dbt project using the &#8216;<a href="https://hub.getdbt.com/brooklyn-data/dbt_artifacts/latest/">dbt-artifacts</a>&#8217; package, and then using Replit's AI, I built a basic dbt run execution timeline visualizer in under one hour &#8211; a task that previously might have seemed daunting (especially when using something as &#8220;delicate&#8221; as D3.js). Sure, it wasn&#8217;t nearly as sophisticated as Airbnb&#8217;s internal platform, but it <em>worked</em>. I had a functional app showing dbt model lineage, run statuses, and execution times, built by one person, in less time than it takes to watch a movie.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_nXB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 424w, /__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 848w, /__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_nXB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png" width="1456" height="707" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:707,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:436141,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://dataespresso.substack.com/i/164420965?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 424w, /__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 848w, /__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_nXB!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea238afb-5758-4334-9bd6-aeaa309293dc_3000x1456.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">My very basic (but functional) dbt run timeline UI</figcaption></figure></div><p>This isn&#8217;t just about building cool internal tools faster. It&#8217;s about:</p><ul><li><p><strong>Democratizing Excellence:</strong> AI is lowering the barrier of the experience layer, allowing teams of all sizes to achieve the kind of platform sophistication once reserved for tech giants.</p></li><li><p><strong>Boosting Productivity:</strong> Enabling data professionals to focus more on insights by spending less time wrestling with clunky interfaces, hunting for metadata, or navigating complex access procedures.</p></li><li><p><strong>A Richer Ecosystem:</strong> Potentially fostering more collaboration and a vibrant open-source landscape for these &#8220;last mile&#8221; components (imagine all the dbt packages waiting to be built!).</p></li></ul><p>Of course, AI isn&#8217;t a magic wand (for now at least), and I touch upon the necessary cautions. But the potential for AI to help us finally build truly complete, user-centric data experiences is immense.</p><p>Want to explore the Big Tech examples, the &#8220;how-to&#8221; with AI, and what this means for the future of data platforms? <a href="https://medium.com/data-science-collective/how-ai-is-finally-democratizing-the-data-platforms-last-mile-layer-8b311aa3c107">Check out the full post on Medium here</a>.</p><h2><strong>New dbt engine, who dis?</strong></h2><p>So, dbt Labs <a href="https://www.getdbt.com/blog/dbt-launch-showcase-2025-recap">finally shared their detailed roadmap for the post-SDF-acquistion world</a>, and it&#8217;s pretty exciting stuff! The star of the show was the new Rust-based dbt Fusion engine. This is a big deal for quite a few reasons: Immense performance improvements compared to dbt Core (up to 30x faster parsing), smarter SQL understanding for things like real-time code validation, and state-aware orchestration that can actually cut down your warehouse compute costs.</p><p>The new engine will co-exist with the Python-based dbt Core engine (<a href="https://github.com/dbt-labs/dbt-core/blob/main/docs/roadmap/2025-05-new-engine-same-language.md">which dbt Labs will continue to maintain</a>), but will understandably have the more restrictive ELv2 license (you can use it internally but can&#8217;t just take it and build a competing SaaS).</p><p>Maintaining the two engines separately is a smart strategy. The dbt <em>language</em> stays consistent, which is key for &#8220;universal&#8221; dbt concepts. If you're already running dbt Core, you can continue as is, or you can choose to adopt the ELv2 Fusion engine for the performance boost without worrying about migration costs.</p><p>dbt Labs announced a wide range of dbt Cloud features that augment the new engine and provide impressive quality-of-life improvements to the dbt experience, like a supercharged VS Code extension and new/improved components like dbt Canvas, dbt Insights, and an expanded dbt Catalog.</p><p>In my opinion, dbt Labs pulled off something pretty impressive. Throughout the early years of dbt, the &#8220;distance&#8221; between dbt Core and dbt Cloud was relatively minimal, making it difficult to justify the cost of the &#8220;upgrade&#8221; from Core to Cloud. However, with this release (and the additional updates made throughout the past two years to areas such as the semantic layer), the value proposition of moving to dbt Cloud is stronger than ever. They&#8217;ve finally created that necessary distance between Core and Cloud to warrant internal discussions about upgrading for most data teams.</p><p>Although the open-source offering is still part of the picture, the big focus on dbt Cloud features means teams using dbt Core (or even the new Fusion engine) will inevitably start to wonder why they&#8217;re not using SQLMesh instead, which continues to champion a more &#8220;compelling&#8221;, feature-rich, open-source-first narrative.</p><h2>A shot of skepticism: thinking critically about our data conclusions</h2><p>As data professionals, we&#8217;re immersed in extracting insights and building data-driven solutions (or at least we like to think so). But this very closeness to the data can sometimes lead to a subtle trap: we might become a bit <em>too</em> quick to believe we&#8217;re drawing the right conclusions, simply because the data &#8220;said so&#8221;.</p><p>I recently read a fantastic book (which I bought a few years ago and forgot about) that serves as an effective reminder of the dark side of data: <strong>&#8220;<a href="https://www.goodreads.com/book/show/48889983-calling-bullshit">Calling Bullshit: The Art of Skepticism in a Data-Driven World</a>&#8221;</strong> by Carl T. Bergstrom and Jevin D. West. While it tackles broader themes of misinformation, its lessons are incredibly relevant for us data folks. (And it&#8217;s a fun read!)</p><p>The book highlights how easy it is for anyone&#8212;even experts&#8212;to be misled if we&#8217;re not rigorously questioning our inputs and interpretations. For us in the data world, this means constantly asking:</p><ul><li><p>Are we truly asking the right questions to begin with?</p></li><li><p>Have we considered all potential confounding factors or biases in the data?</p></li><li><p>Are we diligently distinguishing between mere correlation and actual causation?</p></li></ul><p>Even with sophisticated tools and vast datasets, the critical thinking we apply&#8212;the art of healthy skepticism towards our own findings&#8212;is as important as ever. If you&#8217;re looking to sharpen your internal &#8220;bullshit detector&#8221; and ensure your data narratives are as robust as your pipelines, this book is a worthwhile addition to your reading list.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Hope you enjoyed this edition of Data Espresso! If you found it useful, feel free to share it with fellow data folks in your network.</p><p>Feedback is always appreciated as usual, so feel free to share your thoughts in the comments or reach out directly &#8211; I&#8217;d love to hear your take on the edition&#8217;s topics!</p><p>Until next time, stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #10: A new ice(berg) age, revisiting old designs, and thriving on constraints]]></title><description><![CDATA[The inevitable Apache Iceberg era, the often-overlooked benefits of frequently revisiting design decisions, and the upsides of working in constraint-heavy environments.]]></description><link>https://dataespresso.substack.com/p/espresso-10-a-new-iceberg-age-revisiting</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-10-a-new-iceberg-age-revisiting</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 16 Apr 2025 13:31:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>This month, we&#8217;re diving into the inevitable Apache Iceberg era, the often-overlooked benefits of frequently revisiting design decisions, and the upsides of working in constraint-heavy environments. So, without further ado, let&#8217;s talk data engineering while the espresso (or cappuccino <a href="https://dreambeanscoffee.ie/why-italians-dont-order-cappuchinos-after-11am/">if you&#8217;re feeling adventurous</a>) is still hot.</p><h2><strong>How to navigate an Ice(berg) age</strong></h2><p>Apache Iceberg is all the rage in the data space. After a half-decade showdown between Hudi, Delta Lake, and Iceberg, the latter has emerged as the de facto modern table format, becoming the central piece of the increasingly composable data stack. While there&#8217;s been plenty of great content about Iceberg&#8217;s ascent, in this edition of Data Espresso, I want to offer my perspective on two questions many data teams are grappling with:</p><h3>1. <em>&#8220;I bet on Hudi or Delta Lake, should I migrate to Iceberg?&#8221;</em></h3><p>If you built your data platform around Apache Hudi or Delta Lake, you might be wondering how Iceberg&#8217;s dominance will impact your data platform in the coming years, and whether a migration is warranted.</p><p>Here&#8217;s a heuristic I&#8217;ve observed: once a data technology gets <a href="https://www.icebergsummit2025.com/">its own dedicated conference</a>, you know it&#8217;s going to dominate the discourse (and feature prominently in data stack diagrams) for at least a few years. This "<em>de facto standard</em>" status translates directly into ecosystem focus. While Hudi and Delta Lake aren&#8217;t disappearing overnight, the reality is that new tools and integrations across the data landscape will likely prioritize Iceberg compatibility first.</p><p>This shift will result in some friction for Hudi/Delta Lake shops. It&#8217;s not necessarily that existing Hudi/Delta integrations will break, but rather that the <em>Iceberg</em> path will likely receive new features, performance enhancements, and smoother integrations more rapidly. You might find yourself waiting longer for support or missing out on optimizations readily available to Iceberg users &#8212; and this gap is likely to widen over the next few years. A good historical parallel is the Parquet vs. ORC situation: While ORC remained in use, Parquet consistently had broader and earlier support across the ecosystem, making it the smoother choice for many.</p><p>So, does migration make sense for <em>you</em>? I highly recommend a structured evaluation:</p><ul><li><p>Assess the tangible risks of sticking with your current format (e.g., integration limitations, potentially slower adoption of new query engines, developer friction).</p></li><li><p>Estimate the potential benefits of switching to Iceberg (e.g., access to specific features, performance gains, simplified architecture, broader community support).</p></li><li><p>Factor in the cost and effort of migration.</p></li><li><p>Crucially, run a Proof of Concept (POC) with Iceberg on your specific workloads.</p></li></ul><p>It&#8217;s also worth noting that Iceberg has rapidly closed the feature gap with Hudi and Delta Lake, thanks to significant contributions from numerous &#8220;<em>big tech</em>&#8221; companies. A use case where Hudi/Delta might have been the clear winner a couple of years ago might be well-served, or even better served, by Iceberg today.</p><p>Migrating now, if the evaluation points that way, means doing it on your terms &#8211; defining a timeline that minimizes disruption while positioning your platform to leverage Iceberg&#8217;s momentum. However, if your assessment clearly shows minimal risk and limited benefit in switching <em>right now</em>, your engineering resources are likely better invested elsewhere &#8212; don&#8217;t migrate just for the sake of it.</p><h3>2. <em>&#8220;I&#8217;m not currently using a modern table format, should I add Iceberg to my stack?&#8221;</em></h3><p>This really breaks down into two sub-questions:</p><ul><li><p><em><strong>Should I add any modern table format to my platform right now? (either Iceberg or Hudi/Delta Lake) </strong></em><br>This is the classic "it depends", and the answer needs to be based on the value the tool would bring to your specific use cases. Considering the capabilities of modern table formats (and their increasingly available managed offerings like Iceberg Tables on AWS, GCP BigLake, etc.):</p><ul><li><p>How specifically would it benefit <em>your</em> platform and increase the business value derived from your data?</p></li><li><p>Can it streamline your architecture, perhaps removing redundant layers or simplifying data pipelines?</p></li><li><p>Does it unlock new use cases by allowing different compute engines (Spark, Trino, etc.) to seamlessly work with the same data?</p></li><li><p>Does it solve specific pain points you have <em>today</em> (e.g., schema evolution issues, partition management overhead, concurrent write conflicts)?</p></li></ul></li><li><p><em><strong>If I decide to adopt a modern table format, should it be Iceberg?</strong></em><br>In most cases today, the answer is likely <em>yes</em>. Iceberg has achieved feature parity (or near enough) for many common use cases, and its widespread adoption means it&#8217;s the safest bet for future-proofing and ensuring broad compatibility across the data ecosystem.</p></li></ul><p>Work through these kinds of questions honestly. The goal is to ensure your decision stems from a genuine, explainable need that Iceberg addresses, not just because it&#8217;s the technology <em>du jour</em>.</p><h3>If you decide to adopt Iceberg:</h3><p>It&#8217;s a fantastic technology that truly enables a more composable architecture (think true separation of storage and compute). However, be prepared for a journey. Expect some rough edges, especially with tooling maturity around specific engines or complex migration scenarios. Plan for a multi-phase rollout to de-risk the transition and gradually realize the benefits within your specific environment.</p><h2>An exercise that always pays off: Revisiting design decisions</h2><p>Within the data space, the focus is always on the new cool thing &#8212; be it a paradigm like data mesh or a technology like Iceberg. This focus, combined with a natural engineering tendency to build cool things with cool new tech (worrying about business value is never fun, after all), makes it easy to overlook significant improvement opportunities (in cost, maintenance, or performance) that aren&#8217;t necessarily "cool."</p><p>I recently came across <a href="https://medium.com/gumgum-tech/switching-from-snowpipe-to-data-lake-ingestion-for-simplicity-and-cost-savings-c661d3087c10">a great article by GumGum</a> detailing their switch from Snowpipe to a Data Lake ingestion pattern using Snowflake External Tables over S3 data. The result? A staggering <strong>60% cut</strong> in their data ingestion costs. This might sound counterintuitive, since Snowpipe is often positioned as the go-to for easy and efficient Snowflake ingestion. But for GumGum, the combination of their scale, data structure (already nicely partitioned in S3!), and access patterns meant the Snowpipe approach &#8211; which involved copying data into raw, internal Snowflake tables &#8211; created unnecessary overhead and cost (think expensive table scans for processing and retention on unpartitioned raw data). By switching to querying External Tables directly (and so leveraging the S3 partitioning), they eliminated Snowpipe&#8217;s compute costs, avoided data duplication in Snowflake storage, and drastically cut down query times. It&#8217;s a prime example of how the &#8216;best&#8217; approach is highly contextual and definitely not static.</p><p>As we chase the cutting edge, the tools and platforms we already use are constantly evolving: New features, pricing models, or even entirely new tools might make previously optimal design patterns suboptimal today. That&#8217;s why I believe every data team should schedule regular (every six months for example) review sessions to revisit their past design decisions and identify areas of improvement based on tech progress or changing business needs. Such sessions are a great opportunity to take a hard look at your current architecture:</p><ul><li><p>What are the major cost drivers?</p></li><li><p>Where are the performance bottlenecks?</p></li><li><p>What are the biggest maintenance headaches?</p></li><li><p>Have the capabilities of your core tools (or viable alternatives) changed significantly? Are you leveraging their new features?</p></li><li><p>Have your business requirements or data volumes shifted?</p></li></ul><p>This deliberate review often uncovers areas where significant gains can be made, sometimes with surprisingly simple changes, driven by technological progress or evolving business needs.</p><h2>Out of the comfort zone: The joy of working in a constraint-heavy environment</h2><p> I recently watched <a href="https://www.youtube.com/watch?v=0ML7ZLMdcl4&amp;ab_channel=AIEngineer">a fantastic talk by John Crepezzi from Jane Street about how they built their AI coding assistant</a>. What struck me was the sheer ingenuity required to operate within their unique context: a codebase predominantly in a niche functional language (OCaml) and an environment lacking some tools many engineers take for granted (like having <a href="https://github.com/janestreet/iron">an in-house version control system</a> instead of a Git-based system).</p><p>On the surface, this type of environment could frustrate many engineers (myself included), especially if constraints feel arbitrary or lack clear business or technical justification. (As in: "Is this limitation truly necessary, or just historical baggage?")</p><p>However, if you get the chance to operate within well-justified, albeit tight, constraints, you might find that it pushes you to elevate your game by having to be creative and solving complex problems. You&#8217;re forced to:</p><ul><li><p>Think deeply about the core problem.</p></li><li><p>Navigate ambiguity with more rigour.</p></li><li><p>Be incredibly deliberate about architectural choices, minimizing irreversible "one-way door" decisions.</p></li><li><p>Often, find simpler, more fundamental solutions when fancy off-the-shelf options aren&#8217;t available or suitable.</p></li></ul><p>I&#8217;ve personally worked in a couple of similar high-constraint environments (my investment banking days come to mind - so fun). While the roadblocks were certainly frustrating in the moment, looking back, those experiences were incredibly valuable learning grounds. They force a level of resourcefulness, experimentation, and rapid iteration that you don&#8217;t always encounter in less-constrained settings.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Hope you enjoyed this edition of Data Espresso! If you found it useful, feel free to share it with fellow data folks in your network.</p><p>I always appreciate hearing your feedback. What&#8217;s your take on the Iceberg wave, revisiting design decisions, or working with constraints? Share your thoughts in the comments or reach out directly &#8211; I&#8217;d love to discuss!</p><p>Until next time, stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #9: Real-time analytics' big moment and managing fewer data layers]]></title><description><![CDATA[Hello data friends,]]></description><link>https://dataespresso.substack.com/p/espresso-9-real-time-analytics-big</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-9-real-time-analytics-big</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Thu, 13 Feb 2025 15:03:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>This month, we&#8217;re diving into how (and why) streaming is becoming more accessible in the data space and the fate of the data platform&#8217;s raw/bronze layer. So, without further ado, let&#8217;s talk data engineering while the espresso (or cappuccino <a href="https://dreambeanscoffee.ie/why-italians-dont-order-cappuchinos-after-11am/">if you&#8217;re feeling adventurous</a>) is still hot.</p><h2><strong>Will 2025 be the year real-time analytics finally goes mainstream?</strong></h2><p>Since the Hadoop era and the early days of Big Data, one prediction kept coming back year after year: <em><strong>&#8220;Next year will be the year of streaming.&#8221;</strong></em> There was a constant expectation that eventually, most (all?) data pipelines would evolve to real-time patterns instead of batch processing. Yet, despite the hype, real-time analytics has largely remained limited to tech giants and niche industries with highly specialized streaming needs.</p><p>By this time, the general consensus in the data space is &#8220;<em>you don&#8217;t need streaming</em>&#8221; / &#8220;<em>batch is (mostly) always enough</em>&#8221;. However, I believe 2025 might finally be <em><strong>the year </strong></em>streaming makes sense for common data use cases. This year, two key factors might allow real-time analytics to break out of its niche and finally hit the mainstream.</p><h3><strong>Streaming&#8217;s unfulfilled promise</strong></h3><p>Streaming is exciting. The ability to process and analyze data at scale in real time, unlocking insights and enabling immediate action, has been a core promise of data platforms ever since the Hadoop era. But then the realities of its technical challenges, coupled with the lack of concrete and valuable-enough use cases, often lead data teams to the catch-all <em>&#8220;But do we really need it?&#8221;</em> counter-argument.</p><p>Technologies like Apache Storm and Apache Flink solved many of the technical hurdles, but the fundamental constraints remained: Building and maintaining streaming pipelines required complex infrastructure, specialized skill sets, and significant financial investment. As streaming for pure data movement gained traction in the software engineering world, with patterns like event-based architectures and a meteoric rise for Apache Kafka, the data world remained stuck with batch processing, and architectures like the one below became the standard:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LfWF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 424w, /__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 848w, /__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LfWF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png" width="1400" height="926" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:926,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 424w, /__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 848w, /__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LfWF!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9bd800f-56ce-4e31-b49d-5cd88e1d6d78_1400x926.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Sample (highly-abstracted) data movement and transformation architecture</figcaption></figure></div><p>This architecture often resulted in data arriving at the data platform&#8217;s landing zone (typically an object store) in real-time but only being moved and transformed in batches, often on a daily schedule. Two arguments typically justified such architectures:</p><ul><li><p>Ingesting and transforming data in real-time within the warehouse is more complex and expensive than batch processing.</p></li><li><p>The value generated by real-time use cases doesn&#8217;t justify the added cost and complexity.</p></li></ul><p>However, these arguments are rapidly losing their validity. <a href="https://medium.com/towards-data-science/will-2025-be-the-year-real-time-analytics-finally-goes-mainstream-74556ab7cd8c">In my latest article on Towards Data Science</a>, I discuss two major shifts that could make 2025 a tipping point for real-time analytics.</p><h2>Out of the comfort zone: One layer too many</h2><p>Tech companies have spent the last few years championing "efficiency" &#8212; often by streamlining their organizational structure and removing <em>layers</em> (i.e., middle management) to cut costs. The data world is experiencing its own efficiency push, with a focus on building business-value-driven data pipelines and minimizing the number of tools within the stack &#8212; but there's one area that should be the center of our efficiency discussions: <em>data layers</em>.</p><p>While most high-level architectures showcase the classic three-layer medallion approach (bronze/raw &#8594; silver/curated &#8594; gold/consumption-ready), many data teams in practice maintain <em>far</em> more &#8220;implicit&#8221; layers. Sometimes, data is even replicated multiple times <em>within</em> each layer due to ad-hoc use cases and lack of governance. So, as we strive to reduce costs and increase velocity, maybe it&#8217;s time to challenge our traditional layering approach?</p><p><a href="https://www.infoq.com/articles/rethinking-medallion-architecture/?utm_source=substack&amp;utm_medium=email">A recent article</a> by Adam Bellemare and Thomas Betts dives deep into this very question, arguing that it's time to rethink the bronze layer of the data platform by shifting it left. They propose several compelling options for building data products <em>before</em> the data even reaches the data platform's landing zone, rather than simply dumping everything into object storage and leaving data teams to untangle the complexity. If you're working on a data platform and looking to optimize costs or streamline your architecture, this article is a must-read.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to data folks in your entourage.</p><p>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Navigating Your Career Transition in Tech: A Practical Roadmap]]></title><description><![CDATA[A practical guide to a successful career pivot in tech: from making the decision to thriving in your new role.]]></description><link>https://dataespresso.substack.com/p/navigating-your-career-transition</link><guid isPermaLink="false">https://dataespresso.substack.com/p/navigating-your-career-transition</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Tue, 24 Dec 2024 16:53:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At the start of the year, while contemplating a transition to product management after seven years in data engineering, I struggled to find resources that could help with the decision-making process. It struck me that, although we engineers excel at writing technical content about the technologies we work with, we rarely focus on what ultimately matters most: key career decisions and transitions.</p><p>One of the conversations that helped me the most during that time was with Anis, a good friend who had made his own successful transition from engineering to sales a few years back. Anis shared many valuable insights that felt like they should be readily available online, not confined to a one-on-one conversation. So, a few months after my own transition, we had another conversation, this time about writing this very article: a roadmap with practical guidance for engineers looking to make a career pivot in tech. In this post, we'll share our personal experiences &#8211; Anis's journey into sales and mine into product management &#8211; to give you a behind-the-scenes look at what these transitions entail. We'll cover:</p><ul><li><p><strong>Why Change?</strong> Is a career pivot right for you?</p></li><li><p><strong>De-Risking the Change:</strong> How to research, gather information, and prepare for a smooth transition.</p></li><li><p><strong>Two Journeys:</strong> Real-world examples of moving from engineering to product and engineering to sales.</p></li><li><p><strong>Navigating the New Role:</strong> Tips for thriving in your first few months.</p></li><li><p><strong>If It Doesn't Work Out:</strong> Strategies for adjusting and finding your path.</p></li><li><p><strong>Key Takeaways:</strong> Lessons learned and resources that helped us along the way.</p></li></ul><p>Our goal is to provide a candid, practical guide for anyone in tech considering a career change. Let's dive in.</p><h2><strong>Part 1 - Is a Career Pivot Right for You?</strong></h2><p>Before we dive into the "how" of a career change, let's address the crucial "why." Switching roles, even within tech, requires time, effort, and a willingness to embrace the unfamiliar. So, how do you determine if it's the right move?</p><p><strong>Start with Self-Reflection:</strong></p><ul><li><p><strong>Envision Your Ideal Day:</strong> Imagine your perfect workday. What are you doing? Who are you working with? What impact are you making?</p></li><li><p><strong>Identify the Gaps:</strong> Honestly compare that ideal day to your current reality. What's missing? What aspects of your job do you genuinely enjoy, and what do you dislike? Are there recurring patterns?</p></li><li><p><strong>Revisit Your Motivations:</strong> Think back to why you chose your current career path. What's changed since then? Have your priorities shifted?</p></li><li><p><strong>Consider Your Long-Term Goals:</strong> Where do you see yourself in 5 or 10 years? Does your current trajectory align with that vision?</p></li></ul><p><strong>Gather External Insights:</strong></p><p>Introspection is crucial, but don't stop there. Seek diverse perspectives to gain a more complete picture:</p><ul><li><p><strong>Learn from Others' Experiences:</strong> Read books, articles, and blogs by people who've made similar transitions.</p></li><li><p><strong>Tap into Your Network:</strong> Talk to people doing the job you're considering. Ask about their daily routines, challenges, and rewards.</p></li><li><p><strong>Consult Your Mentors:</strong> If you have mentors, solicit their advice. They may offer valuable insights based on their own experiences.</p></li><li><p><strong>Sounding Board:</strong> While your friends and family might not grasp the intricacies of your industry, they can still offer valuable support and help you process your thoughts.</p></li></ul><p><strong>Three Potential Realizations:</strong></p><p>This process of reflection and information gathering will likely lead you to one of three conclusions:</p><ol><li><p><strong>You love your current role:</strong> Fantastic! You're already on the right path.</p></li><li><p><strong>You need to tweak your current role:</strong> Perhaps a new project, a different team, or a shift in responsibilities is all it takes.</p></li><li><p><strong>You need a fundamental change:</strong> This is also great news! Recognizing the need for something different is the first step towards finding a more fulfilling career path.</p></li></ol><blockquote><p><strong>Remember</strong>: Change isn't just a risk; it's an opportunity for growth, a chance to discover a path that truly aligns with your passions and goals.</p></blockquote><h2><strong>Part 2 - De-Risking Your Career Pivot: Preparation is Paramount</strong></h2><p>So, you've done the soul-searching and you're seriously considering a career change. Now it's time to strategize. Approaching this transition thoughtfully will minimize risk and maximize your chances of success.</p><p><strong>1. Understand the Terrain:</strong></p><ul><li><p><strong>Explore Potential Paths:</strong> Research roles that align with your interests and skills. Analyze the pros and cons of each.</p></li><li><p><strong>Analyze the Market:</strong> What's the current job market like for your target roles? What are the projected trends for the next 6-12 months? Are there specific industries or companies that look particularly promising?</p></li><li><p><strong>Assess Your Skills:</strong> Honestly identify the skills required for your desired role. Which do you already possess? What are the gaps you need to bridge?</p></li></ul><p><strong>2. Bridge the Gap:</strong></p><ul><li><p><strong>Targeted Learning:</strong> Focus on acquiring the necessary skills. This could involve online courses, workshops, certifications, or even self-directed learning through books and online resources.</p></li><li><p><strong>Expand Your Network:</strong> Connect with people already working in your desired field, attend industry events and meetups, and join relevant online communities.</p></li><li><p><strong>Gain Practical Experience:</strong> Seek opportunities to gain hands-on experience in your target role. Could you shadow someone at your current company? Could you take on a relevant side project or volunteer your time with a non-profit to apply your new skills?</p></li></ul><blockquote><p><strong>Pro Tip:</strong> Explore whether your current company offers personal development programs. These programs can often provide valuable resources and support for your learning journey.</p></blockquote><h2><strong>Part 3 - Two Journeys: Engineering to Product &amp; Sales</strong></h2><p>Let's get into the specifics. We'll share our own experiences transitioning from engineering, highlighting the challenges, surprises, and key learnings.</p><h3><strong>Mahdi's Journey: From Data Engineering to Product Management</strong></h3><p>My transition to product management felt like a natural progression. I've always been passionate about problem-solving and understanding user needs. As a data engineer, I found myself increasingly drawn to the "why" behind the features we were building rather than just the "how."</p><p><strong>The Itch:</strong> It began as a subtle feeling. While I enjoyed coding, I felt a stronger pull towards product strategy and the bigger picture. I was that engineer constantly asking, "Why are we building this?" and "Who is this for?"</p><p><strong>Finding My Niche:</strong> I realized my technical background could be a significant asset in a product role. I focused on products in the data space, where my data engineering experience gave me a unique edge.</p><p><strong>The Mindset Shift:</strong> Moving from engineering to product requires a significant shift in perspective:</p><ul><li><p><strong>Trusting Your Team:</strong> As an engineer, I was used to having complete control over technical implementation. As a PM, I had to learn to trust my engineering team's expertise while still providing valuable input and guidance.</p></li><li><p><strong>Becoming a Translator:</strong> A big part of the PM role is bridging the gap between technical and non-technical stakeholders. I had to become fluent in both "languages."</p></li><li><p><strong>Embracing Ambiguity:</strong> Product management often involves navigating uncertainty. I had to get comfortable making decisions with incomplete information and iterating based on feedback.</p></li></ul><p><strong>The Essentials:</strong> I quickly realized I needed to master these core PM skills:</p><ul><li><p><strong>Prioritization:</strong> Learning to prioritize features and make tough calls about what to build (and what <em>not</em> to build) is crucial.</p></li><li><p><strong>Metrics &amp; Measurement:</strong> Understanding how to define and track key metrics is essential for measuring product success.</p></li><li><p><strong>Collaboration:</strong> Working effectively with designers, marketers, and other stakeholders is a daily necessity.</p></li></ul><h3><strong>Anis's Journey: From Engineering to Sales</strong></h3><p>My transition into sales was intentional, yet it still came with its share of uncertainties. While I enjoyed my engineering work, I found myself more excited about explaining our projects than actually building them. I got a real kick out of transforming complex technical ideas into relatable stories. I craved more of that&#8212;more conversations, human connection, and opportunities to help people solve problems.</p><p>This inclination started during university. I loved presenting our projects&#8212;not just explaining them, but truly <em>selling</em> them. I found ways to make them relevant, demonstrating their value beyond the classroom. That passion for communication and connection led me to a part-time business development role at a startup. It was my first real taste of sales. Some days were exhilarating; others were tough. I learned firsthand about rejection, how to handle it, and how to move forward. I discovered that I thrived on the ups and downs, the competition, and the challenge of working with people&#8212;not machines.</p><p>After graduation, I knew I didn't want to abandon my engineering background entirely, but I also didn't want to be confined to purely technical roles. I yearned for something different. That's when I pursued a Master's in International Business Development. It wasn't just about the degree; it was a strategic step to facilitate my transition. I learned about negotiation, networking, and operating in the business world. But the most valuable lessons came from the alumni network&#8212;hearing their stories and seeing how they navigated similar career changes. Their experiences gave me the confidence to believe I could do the same.</p><p><strong>Key Differences:</strong></p><ul><li><p><strong>Human vs. Machine:</strong> Sales is fundamentally about people. It's about understanding their needs, building trust, and earning respect. My engineering background helped me connect with technical customers, but true success in sales came from listening, asking insightful questions, and demonstrating a genuine desire to help.</p></li><li><p><strong>Shared Frameworks:</strong> Some engineering principles translated surprisingly well. KISS&#8212;Keep It Simple and Stupid&#8212;became a guiding principle. The simpler I made things for my customers, the better the results. Sales isn't about overwhelming people with complexity; it's about making their lives easier.</p></li><li><p><strong>Networking and Branding:</strong> In sales, networking isn't just a skill&#8212;it's a lifeline. But it's not solely about revenue. It's about cultivating relationships, building a personal brand, and opening doors for the future. Over time, I've seen these connections blossom into friendships, mentorships, and unexpected opportunities.</p></li></ul><p><strong>The Learning Curve:</strong></p><p>The transition wasn't a walk in the park. There were stumbles, and I had to learn quickly.</p><ul><li><p><strong>Patience:</strong> Building relationships and closing deals takes time. I learned that the hard way. Success comes from consistently showing up and proving your trustworthiness.</p></li><li><p><strong>Adaptability:</strong> People are unpredictable. Unlike machines, there's no easy debugging. Every conversation is unique, and I had to learn to read situations, adapt, and respond in real-time.</p></li><li><p><strong>Embrace Mistakes:</strong> Early on, I mistakenly thought sales was about working harder. It's not. It's about working smarter. I learned to listen more, talk less, and focus on the customer's goals&#8212;not my own.</p></li></ul><p>Sales has taught me that it's not about having the perfect pitch. It's about showing up, being present, and helping others achieve their goals. My engineering background remains a valuable asset&#8212;it helps me understand my customers' technical challenges and offer solutions that resonate. But now I also have the tools to connect, to inspire, and to lead conversations.</p><blockquote><p><strong>Pro Tip:</strong> If you're contemplating a similar change, my advice is this: don't overthink it. Take the leap. Your skills are more transferable than you might realize, and the journey will teach you things you can't yet imagine. You might surprise yourself. I know I did.</p></blockquote><h2><strong>Part 4 - Navigating Your New Role: Mastering the First Few Months</strong></h2><p>You've made the leap! Now comes the critical phase of settling into your new role. These first few months are crucial for setting yourself up for long-term success.</p><p><strong>1. Embrace the Learning Curve:</strong></p><ul><li><p><strong>Be a Sponge:</strong> Absorb as much information as possible. Ask questions, seek out resources, and learn from your colleagues. It's important to filter this information, though. Prioritize ruthlessly, focusing on knowledge directly relevant to your immediate tasks and goals.</p></li><li><p><strong>Seek Feedback:</strong> Don't hesitate to ask for feedback, both positive and constructive. It's the fastest way to identify areas for improvement.</p></li><li><p><strong>Set Realistic Expectations:</strong> You won't become an expert overnight. Give yourself time to adjust, learn the ropes, and define both realistic and ambitious goals to track your progress.</p></li></ul><p><strong>2. Build Relationships:</strong></p><ul><li><p><strong>Connect with Your Team:</strong> Get to know your colleagues on both a professional and personal level. These relationships will be vital to your success.</p></li><li><p><strong>Find a Mentor:</strong> Seek out an experienced person in your new role who can offer guidance, support, and valuable insights.</p></li><li><p><strong>Network Internally:</strong> Build relationships with people in other departments. This will broaden your understanding of the company and foster more effective collaboration.</p></li></ul><p><strong>3. Reflect and Adjust:</strong></p><ul><li><p><strong>Regular Check-ins:</strong> Schedule time to reflect on your progress. What's going well? What needs improvement?</p></li><li><p><strong>Course Correction:</strong> Don't be afraid to make adjustments along the way. If something isn't working, try a different approach. Be agile and adapt to new information.</p></li></ul><h2><strong>Part 5 - When Things Don't Go as Planned: Strategies for Recalibration</strong></h2><p>Despite your best efforts and intentions, there's always a chance that a new role might not be the perfect fit. And that's perfectly okay! It's not a failure; it's a valuable learning opportunity. Here's how to navigate this situation:</p><p><strong>1. Open Communication is Key:</strong></p><ul><li><p><strong>Talk to Your Manager:</strong> Be honest about your challenges and concerns. They might be able to offer support, suggest solutions, or help you adjust your role.</p></li><li><p><strong>Seek Feedback from Colleagues:</strong> Get their perspectives. They may offer insights into team dynamics or company culture that you haven't fully grasped.</p></li></ul><p><strong>2. Analyze and Diagnose:</strong></p><ul><li><p><strong>Identify the Root Cause:</strong> Try to pinpoint the underlying reasons why things aren't working out. Is it the role itself, the team dynamics, the company culture, or something else entirely?</p></li><li><p><strong>Reflect on Your Decision:</strong> Revisit the factors that led you to make the jump. What did you overlook or underestimate? This reflection will be invaluable for future career decisions.</p></li></ul><p><strong>3. Explore Your Options:</strong></p><ul><li><p><strong>Consider an Internal Transfer:</strong> If you like the company but not the specific role, explore an internal transfer. Perhaps you could return to a previous role or move to a different department.</p></li><li><p><strong>Reconnect with Your Previous Employer:</strong> If you maintained a good relationship with your previous company, reaching out to them might be an option.</p></li><li><p><strong>Explore External Opportunities:</strong> If internal options aren't feasible, start looking for new opportunities elsewhere.</p></li></ul><p><strong>4. Give it Time, But Set Limits:</strong></p><ul><li><p><strong>Allow for Adjustment:</strong> Give yourself and the role/team sufficient time before making any drastic decisions. It takes time to adapt, and initial challenges might resolve themselves.</p></li><li><p><strong>Define a Timeframe:</strong> If you decide to try and make it work, set a clear timeframe. This prevents you from staying in an unfulfilling situation indefinitely.</p></li></ul><blockquote><p><strong>Remember:</strong> Careers are rarely linear. Be open to pivoting or re-pivoting as needed. Each experience, even the challenging ones, provides valuable insights that inform your future choices and bring you closer to your ideal career path.</p></blockquote><h2><strong>Part 6 - Key Takeaways, Common Fears, and Resources</strong></h2><p>Career transitions can be intimidating, but they can also be incredibly rewarding. Here are some of the most important lessons we've learned along the way:</p><p><strong>Key Takeaways:</strong></p><ul><li><p><strong>Embrace Change as an Opportunity:</strong> View change as a chance for growth, self-discovery, and finding a path that truly aligns with your goals.</p></li><li><p><strong>Preparation is Paramount:</strong> Thorough research and preparation significantly increase your chances of a successful transition.</p></li><li><p><strong>Lifelong Learning is Essential:</strong> The tech landscape is constantly evolving. Never stop learning, and continuously adapt to stay relevant.</p></li><li><p><strong>Pivoting is a Strength:</strong> Don't be afraid to change course if things aren't working out. It's a sign of strength and adaptability, not failure.</p></li><li><p><strong>Networking is Invaluable:</strong> Build and nurture your professional network. It can open doors to new opportunities and provide crucial support.</p></li></ul><p><strong>Common Fears &amp; Misconceptions:</strong></p><ul><li><p><strong>"I'm not qualified":</strong> Imposter syndrome is common, but don't let it hold you back. Focus on your transferable skills, your willingness to learn, and your past successes.</p></li><li><p><strong>"It's too late to change":</strong> It's never too late to pursue a career path that aligns with your passions and goals.</p></li><li><p><strong>"I'll lose my seniority":</strong> While you might experience a temporary adjustment in title or compensation, your experience and skills remain valuable assets.</p></li></ul><p><strong>Resources That Helped Us:</strong></p><ul><li><p><strong>Sales &amp; Negotiation (Anis):</strong></p><ul><li><p><strong>"Never Split the Difference"</strong> by Chris Voss</p></li><li><p><strong>"SPIN Selling"</strong> by Neil Rackham</p></li><li><p><strong>&#8220;Cold Calling Sucks (And That's Why It Works)&#8221;</strong> by Armand Farrokh and Nick Cegelski</p></li><li><p><strong>&#8220;Selling with&#8221;</strong> by Nate Nasralla</p></li></ul></li><li><p><strong>Product Management (Mahdi):</strong></p><ul><li><p><strong>"Inspired: How to Create Tech Products Customers Love&#8221;</strong> by Marty Cagan</p></li><li><p><strong>"Monetizing Innovation: How Smart Companies Design the Product Around the Price"</strong> by Madhavan Ramanujam, Georg Tacke</p></li><li><p><strong>Lenny&#8217;s Podcast and Lenny&#8217;s Newsletter</strong> by Lenny Rachitsky</p></li></ul></li><li><p><strong>Mindset &amp; Personal Development:</strong></p><ul><li><p><strong>(Anis) "The Mamba Mentality"</strong> by Kobe Bryant</p></li><li><p><strong>(Anis) "Next Play"</strong> by Coach K</p></li><li><p><strong>(Mahdi) "Principles: Life and Work"</strong> by Ray Dalio</p></li><li><p><strong>(Mahdi) "Range: Why Generalists Triumph in a Specialized World"</strong> by David Epstein</p></li></ul></li></ul><p><strong>Conclusion:</strong></p><p>Ultimately, your career journey is unique. Don't be afraid to forge your own path, embrace change, and seek out opportunities that excite and challenge you. We hope that sharing our experiences and the resources that helped us along the way will empower you to navigate your own career transition with confidence and achieve your professional goals. Remember that career paths are dynamic, not linear.</p><p><strong>Best of luck,</strong></p><p><strong>Anis &amp; Mahdi</strong></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>We'd love to hear from you! Share your own experiences with career transitions in the comments below. What challenges did you face? What advice would you give to others?</p>]]></content:encoded></item><item><title><![CDATA[Espresso #8: Why the data space struggles with standardization + Reflecting on turning 30]]></title><description><![CDATA[We&#8217;re tackling the barriers to standardization in the data space, some personal wisdom I&#8217;ve picked up in my 20s (it&#8217;s been a ride!), and the 100-page gems of The Do Book Co.]]></description><link>https://dataespresso.substack.com/p/espresso-8-why-the-data-space-struggles</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-8-why-the-data-space-struggles</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Mon, 11 Nov 2024 14:31:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello data friends,</p><p>Grab your espresso (or cappuccino <a href="https://dreambeanscoffee.ie/why-italians-dont-order-cappuchinos-after-11am/">if you&#8217;re feeling adventurous</a>) and let&#8217;s dive into this month&#8217;s brew! We&#8217;re tackling the barriers to standardization in the data space, some personal wisdom I&#8217;ve picked up in my 20s (it&#8217;s been a ride!), and the 100-page gems of The Do Book Co. So, without further ado, let&#8217;s talk data engineering (and personal updates) while the espresso is still hot.</p><h2>Barriers to standardization in the data space</h2><p>Maxime Beauchemin, one of the pioneers (<a href="https://medium.com/free-code-camp/the-rise-of-the-data-engineer-91be18f1e603">THE pioneer?</a>) of the data engineering field, <a href="https://preset.io/blog/why-data-teams-keep-reinventing-the-wheel/">recently wrote</a> about why data teams keep reinventing the wheel by building their transformation layer from scratch. In his article (which I highly recommend), he offers a proposal relying on &#8220;<strong>Parametric Pipelines&#8221; </strong>and<strong> &#8220;Unified Models&#8221; </strong>to build reusable and generic data components/assets that are also flexible enough to accommodate the inevitable specificities of every business.</p><p>Maxime&#8217;s article navigates the nuances of one key part of the data stack (transformation) and why it lacks standards, but I think he touched on a topic that applies to the stack as a whole: We're great at standardizing the <em>tools</em> (thanks dbt &amp; Airflow), but everything else feels like the Wild West - which is a shame since we universally agree that all data teams do <em>slightly different versions of the same job</em>.</p><p>I think this lack of standardization outside of the tools themselves is mainly due to two barriers we still need to overcome:</p><ul><li><p><strong>Every tool generates its own metadata (and doesn&#8217;t share it)</strong>: Probably the thing that hurt data teams the most in the past five years is that every tool within the Modern Data Stack wanted to store its own metadata and  build features on top of it. This meant that every tool knew <em>a bit</em> about the data and the pipelines, but no tool had the full story (not even the data catalog) - making it very difficult to centralize all of the stack&#8217;s metadata in one place and build insights on top of it (or even define standards for managing it). The ideal scenario in this area was to have a standard representation of metadata that tools within the stack can implement and then leverage to exchange metadata in an automated manner, allowing for the rise of standardized ways to manage the data itself. I still believe we&#8217;ll reach that state eventually, but the road ahead is, unfortunately, long.</p></li><li><p><strong>We don&#8217;t write enough YAML (seriously - hear me out)</strong>: Although the &#8220;<em>data engineers are in fact YAML engineers</em>&#8221; reflection started out of (valid) frustration with YAML, I think we&#8217;re still paying a high price for not doing (most) things in code. Even though the Modern Data Stack (RIP) brought with it a lot of advancements, the focus remained on the UI, and many technical workflows (essential for automating and standardizing things) were simply ignored. By moving more things (dashboards, tooling configuration, metrics, data observability, etc.) to code/YAML, we dramatically shorten the path toward standardization and open new industry-wide collaboration doors. The &#8220;Post-Modern&#8221; Data Stack (PMDS?) is fortunately taking things in the right direction, and I believe we&#8217;ll overcome this barrier in the next few years.</p></li></ul><h2>Out of the comfort zone: 5 learnings from my twenties</h2><p>This September I turned 30 and decided that it was a good moment to reflect on the biggest learnings of my twenties (which were, in a nutshell, a rollercoaster). The end result felt like something worth sharing (more on <a href="https://medium.com/swlh/your-twenties-are-meant-for-fun-after-all-or-are-they-9d228cc734e0">that</a> below), but these are the five learnings I&#8217;m extremely grateful for:</p><ul><li><p><strong>Figure out how you&#8217;re wired</strong>: What makes you <em>you</em>?</p></li><li><p><strong>It&#8217;s all about resilience (Luctor et Emergo)</strong>: You&#8217;ll inevitably face an insurmountable mountain, and you need to be ready for it.</p></li><li><p><strong>Curiosity and open-mindedness are the path forward</strong>: The combination of these two characteristics, whether in a personal or professional context, ensures that you&#8217;re always on the right track toward becoming a better version of yourself.</p></li><li><p><strong>Take measured risks</strong>: Understand what risk is and figure out how much of it you want in your life.</p></li><li><p><strong>Learn from your losses, and celebrate your wins</strong>: As you progress in life, you&#8217;ll inevitably make many wrong turns and yet more right ones. Acknowledge the outcome of these turns and make them count.</p></li></ul><p>Want to jump into the details? <a href="https://medium.com/swlh/your-twenties-are-meant-for-fun-after-all-or-are-they-9d228cc734e0">I wrote a full article about it!</a> (Isn&#8217;t that what you do when you turn 30?)</p><h2>Out of office: The 100-page gems of The Do Book Co</h2><p>On a recent trip to Vienna (highly recommended!) I stumbled upon <a href="https://thedobook.co/">The Do Book Co</a>&#8217;s &#8220;pocket guides&#8221; at a cool concept store called <a href="https://www.calienna.com/">Calienna</a>. These tiny books (around 100 pages long on average) are all written by subject matter experts and come in an ideal format: long enough to cover a complex topic and deep-dive into it, but not a big time commitment like a 400-page book.</p><p>The first two I read are:</p><ul><li><p><a href="https://thedobook.co/products/do-deal-negotiate-better-find-hidden-value-enrich-relationships?Format=Paperback">Do Deal</a> by Richard Hoare &amp; Andrew Gummer: Fantastic advice and insights that cover all the stages of negotiations (in any given context).</p></li><li><p><a href="https://thedobook.co/products/do-start-how-to-create-and-run-a-business-that-doesnt-run-you?Format=Paperback">Do Start</a> by Dan Kieran: This one was so good that it deserved <a href="https://x.com/MahdiKarabiben/status/1831073327646933403">a Twitter/X post</a>. If you're interested in working at startups or starting your own, this book is a must-read. Dan covers all the key areas of entrepreneurship while providing valuable insights and telling the inspiring story of <a href="https://unbound.com/">Unbound</a>.</p></li></ul><p>They have a ton of other guides on all sorts of topics, so check them out!</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.</p><p>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #7: Data modeling in a post-Modern Data Stack world, a reverse-MDS journey, and the data team's tricky first steps]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Tuesday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/espresso-7-data-modeling-in-a-post</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-7-data-modeling-in-a-post</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Tue, 06 Aug 2024 13:30:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,</p><p>In this edition, we will discuss data modeling techniques that you can leverage in a post-<em>Modern Data Stack</em> world, Notion&#8217;s very interesting <em>&#8220;reverse-Modern-Data-Stack&#8221;</em> journey, and the tricky first steps of building a data team. So, without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><h2><strong>Data modeling techniques for the post-</strong><em><strong>Modern Data Stack </strong></em><strong>world</strong></h2><p>Over the past few years, as the Modern Data Stack (MDS) introduced new patterns and standards for moving, transforming, and interacting with data, dimensional data modeling gradually became a relic of the past. In its place, data teams relied on One-Big-Tables (OBT) and stacking layer upon layer of dbt models to tackle new use cases. However, these approaches led to unfortunate situations in which data teams became a cost center with unscalable processes. So, as we enter a &#8220;post-modern&#8221; data stack era, defined by the pursuit to reduce costs, tidy up data platforms, and limit model sprawl, data modeling is witnessing a resurrection.</p><p>This transition puts data teams in front of a dilemma: should we revert back to strict data modeling approaches that were defined decades ago for a completely different data ecosystem, or can we introduce new principles that are defined based on today&#8217;s technology and business problems?</p><p>In my opinion, the answer is (drumroll) somewhere in the middle. It's important to move away from model sprawl, but without impacting the delivery speed of new data products. How? <a href="https://towardsdatascience.com/data-modeling-techniques-for-the-post-modern-data-stack-03fc2e4a210c">In my latest article on Towards Data Science</a>, I discuss different techniques to achieve this. (Spoiler: it starts with defining the right standards.)</p><h2>Fresh off the press: Notion&#8217;s reverse-MDS journey</h2><p>As data professionals, we&#8217;re all familiar with the data lake &#8594; Modern Data Stack (MDS) path: Data teams realize that they&#8217;re spending too much engineering time managing the data platform itself due to its complexity (thanks for the trauma, Hadoop) and decide to migrate to SaaS tools that offer a streamlined experience and minimal maintenance needs.</p><p>In <a href="https://www.notion.so/blog/building-and-scaling-notions-data-lake">a recent blog post</a>, the Notion data team discussed their rather unique journey in the opposite direction. Instead of migrating from a data lake to an MDS platform, they replaced their shiny Snowflake- and Fivetran-based platform with a data lake architecture using Hudi, Spark, and Kafka &#8212; with very positive results.</p><p>We should note here that their use case is unique to their product since around 90% of their upserts are updates, which Snowflake hates while Hudi excels at. Their journey is, however, a great reminder that the data architecture should be defined first and foremost around your specific business needs, and the tooling that&#8217;s &#8220;cool&#8221; or &#8220;modern&#8221; is not always the best fit for your particular use cases.</p><h2>Out of the comfort zone: The data team&#8217;s tricky first steps</h2><p>As I transitioned from engineering to product (a journey that I&#8217;ll discuss in an upcoming blog post), I started listening to <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Lenny Rachitsky&quot;,&quot;id&quot;:1849774,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/afba5161-65bb-4d99-8d6b-cce660917fa1_1540x1540.png&quot;,&quot;uuid&quot;:&quot;7e82349b-939c-4d2a-b100-e289d6c96cd7&quot;}" data-component-name="MentionToDOM"></span>&#8217;s fantastic <a href="https://www.lennysnewsletter.com/podcast">podcast</a> on a weekly basis (which I highly recommend even if you don&#8217;t work in product). In <a href="https://www.youtube.com/watch?v=D4PDb_C8Dww&amp;ab_channel=Lenny%27sPodcast">a recent episode</a>, his guest was Jessica Lachs (VP of Analytics and Data Science at DoorDash). The whole episode was filled with valuable insights, but I personally appreciated the discussion around a very important yet overlooked topic: the data team&#8217;s &#8220;delicate&#8221; first steps.</p><p>Data teams have a rather awkward positioning in most companies: They don&#8217;t directly generate value, and so it&#8217;s easy to head in the wrong direction with just one wrong step. Decisions like building an unnecessarily complex platform early on, focusing on the wrong business goals, or adopting a reactive &#8220;data team as a support team&#8221; mindset will ultimately make it very difficult for the data team to scale and showcase its value.</p><p>Outside of the podcast episode, which covers many other topics like defining effective metrics and the ideal data team structure, Jessica published <a href="https://review.firstround.com/starting-an-analytics-org-from-scratch-lessons-from-a-decade-at-doordash/">a blog post</a> that covers this topic in particular. If you&#8217;re leading a data team in its early days, this is a must-read.</p><h2>Out of office</h2><p>Until we meet again for another espresso, and if you still have room for one addition to your summer reading list, I highly recommend <a href="https://www.goodreads.com/book/show/51115322-the-art-of-rest">The Art of Rest</a> by Claudia Hammond. If, like myself, you feel &#8220;guilty&#8221; when resting, then this book will definitely change your perspective.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.</p><p>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #6: From data mess to data mesh, Spotify's data platform, and the post-MDS world]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Monday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/espresso-6-from-data-mess-to-data</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-6-from-data-mess-to-data</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Mon, 08 Apr 2024 14:00:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!byzV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,<br><br>In this edition, we will talk about a path to successfully scaling your data platform, Spotify&#8217;s exceptional data culture, and the state of the data world after we declared the <em>Modern Data Stack (MDS)</em> dead. So without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><h2>Navigating your data platform&#8217;s growing pains: a path from data mess to data mesh</h2><p>When working on software components, developers can leverage a wide range of frameworks, design patterns, and principles to scale their products and seamlessly adjust their architecture to support new use cases and handle increasing usage and complexity. This allows software engineering teams to ensure optimized performance and reliability as their platform (and its value) grows in scale.</p><p>Data teams, however, are not so fortunate. While the initial months of a data platform&#8217;s lifecycle are often marked by the excitement of tackling complex technical challenges and the joy of delivering a first wave of data products, what often follows is a daunting spiral of mounting complexity, rising costs, and diminishing returns.</p><p>Unlike other problems that we need to navigate as data teams, our scalability struggles are inherently different from the ones faced by software teams. In the data world, these struggles come in the form of unavoidable technical complexity (like mixing a multitude of patterns to move and transform data across an ever-growing list of systems) and the data platform&#8217;s unique positioning within the company (since eventually every business unit gets connected to it either directly or indirectly).</p><p>So, <a href="/__u/benn.substack.com/p/the-problem-was-the-product">in this post-MDS world</a>, where data teams are thoroughly scrutinized over their spending and continuously asked to showcase their value, it is more important than ever to define standards and tenets for successfully scaling a data platform.</p><p>In <a href="https://towardsdatascience.com/navigating-your-data-platforms-growing-pains-a-path-from-data-mess-to-data-mesh-c16df72f5463">my latest article</a>, I offer five principles (complete with actionable strategies) to navigate the tricky state between being a new data team and building a large-scale data platform that generates substantial business value.</p><h2>Fresh off the press: Spotify&#8217;s data platform</h2><p>Among the big tech behemoths, some companies stand out from the rest when it comes to data - think Airbnb, Netflix, Uber, and Spotify. These companies play a key role in driving the data field forward by constantly rethinking their approaches to data and, more importantly, by being open about their data work (whether by open-sourcing the tools they build or sharing the learnings of their journey).</p><p>I&#8217;m personally a big fan of how the Spotify data teams approach their data projects, whether it&#8217;s <a href="https://engineering.atspotify.com/2021/02/how-spotify-optimized-the-largest-dataflow-job-ever-for-wrapped-2020/">to build the largest Dataflow job ever</a> or to scale and democratize both <a href="https://medium.com/spotify-insights/analytics-engineering-at-spotify-f165180a6722">analytics engineering</a> and <a href="https://medium.com/spotify-insights/visual-analytics-at-spotify-3d4221d8686">data visualization</a>. Last week, Spotify published <a href="https://engineering.atspotify.com/2024/04/data-platform-explained/">the first article</a> in a new series that will discuss the different aspects of its data platform and data endeavors. I recommend following along the journey since it would definitely present many valuable learnings that  can be applied when working on data projects.</p><h2>Out of the comfort zone: The post-Modern-Data-Stack world</h2><p>In the past few months, one of the hottest topics in the data world has been <em>the death of the modern data stack</em>. I believe that enough ink has been spilled on this topic, but I also want to reflect on the Modern Data Stack era and how it transformed the data field.</p><p>When I started my career in early 2017, the Hadoop ecosystem was all the rage. Companies were spending enormous resources on building data lakes that had a terrible return on investment and were very taxing to maintain. The problem wasn&#8217;t the data itself but the systems, which were simply too complex and, in most cases, barely working.</p><p>Fast forward to 2024. Last week, while reading <a href="https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2024">dbt Labs&#8217; annual state of analytics engineering report</a>, I was pleasantly surprised by the following chart:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!byzV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 424w, /__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 848w, /__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 1272w, /__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!byzV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png" width="1456" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:235705,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 424w, /__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 848w, /__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 1272w, /__u/substackcdn.com/image/fetch/$s_!byzV!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50b44eeb-3f93-47f3-950c-322599a3c8d1_1842x1022.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The takeaway is that we reached a state where the biggest two challenges that data teams face (poor data quality and ambiguous ownership) are related to people and processes and not to technical complexity or the data platform itself. This is, in a way, a direct acknowledgment of the modern data stack&#8217;s role in solving one of the Hadoop era&#8217;s biggest problems: today&#8217;s data platform just works. Gone are the days of messy configuration, navigating chaotic error logs across different systems, and spending endless engineering time on building and maintaining the platform.</p><p>The Modern Data Stack, for all its flaws, solved the data platform&#8217;s technical hurdles. Now it&#8217;s time to solve the business ones.</p><div><hr></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.<br><br>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #5: Design docs, dbt at scale, and how to move up the ladder]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Monday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/espresso-5-design-docs-dbt-at-scale</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-5-design-docs-dbt-at-scale</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Mon, 12 Jun 2023 13:55:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,<br><br>In this edition, we will talk about design docs for data pipelines, setting standards and foundations for scalable dbt usage, and the core principle to keep in mind when working towards your next promotion. So without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><h2>Writing design docs for data pipelines</h2><p>Over the past few years, adopting software engineering best practices has become a common theme within the data engineering space. From dbt&#8217;s software-engineering-inspired capabilities to the rise of data observability, data engineers are getting increasingly accustomed to the software engineer&#8217;s toolset and principles.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This shift had a major impact on how we design and build data pipelines. It made data pipelines more robust (since we moved away from hard-coded business logic and complex SQL queries to modular dbt models and macros) and drastically lowered the number of &#8220;<em>Hey can you check this table?</em>&#8221; Slack messages (via automated data quality monitoring and alerting).</p><p>These changes helped us move the industry in the right direction &#8212; but there are still areas in which we still have much to learn from our software engineering counterparts.</p><p>Based on <a href="https://www.getdbt.com/blog/analytics-engineering-next-step-forwards/#:~:text=Complexity%20in%20the%20dbt%20ecosystem">recent data released by dbt Labs</a>, around 20% of dbt projects have more than 1,000 models (and 5% have more than 5,000). These numbers highlight one fundamental problem in our data pipelines: we&#8217;re not intentional enough with how we design them. We stack layers upon layers of ad-hoc models and use-case-specific transformations and end up with ten models that represent the same logical entity &#8220;with subtle differences&#8221;.</p><p>In <a href="https://towardsdatascience.com/writing-design-docs-for-data-pipelines-d49550f95580">my latest article</a>, I discuss one artifact that can help us design (and build) robust foundations for our data platforms: <strong>design docs</strong>.</p><h2>Fresh off the press: our <strong>dbt journey at Zendesk</strong></h2><p>At Zendesk, we started our dbt journey more than a year ago (and have been running it in production for over six months) - and today, more than half of our data pipelines are dbt jobs.</p><p>The most important pillar to which we owe the success of our approach is setting the right foundations from day one and defining standards that make sense within the context of our use cases and architecture.</p><p>Last month, I published <a href="https://zendesk.engineering/dbt-at-zendesk-part-i-setting-foundations-for-scalability-34b55e6a6aa1">an article in the Zendesk engineering blog</a> that discusses the core foundations of our dbt setup (and more) - it&#8217;s intended to be the first article of a three-part series, so stay tuned for more technical articles on our dbt implementation.</p><h2>Out of the comfort zone: What got you here won&#8217;t get you there.</h2><p>Last year I participated in the fantastic <a href="https://www.lidr.co/en/ignite">Ignite tech lead mentoring program</a> (which I highly recommend), and one of the key learnings that stuck with me is &#8220;<em><strong>What got you here won&#8217;t get you there.</strong></em>&#8221; This was a very elegant way to phrase a principle that I&#8217;ve been applying throughout my career, and I think it&#8217;s important to keep it in mind as you progress throughout yours.</p><p>Every vertical (or horizontal) transition in your career introduces you to a new job requiring a specific skill set that differs from your previous role. Senior, Staff, Tech Lead, or Engineering Manager are all very different hats that require some preparation before wearing them. Some companies merely rely on the &#8220;years of experience&#8221; metric to determine your readiness to change hats, but what really matters is how comfortable you are with the skill set of the new job.</p><p>As you think about your career goals, consider the role that you currently aim to achieve and the skills that it requires, then use this information as a map to define the areas on which you want to focus in your development and the experiences you want to gain.</p><h2>A sound to code to</h2><p>Ever since the <a href="https://www.youtube.com/watch?v=OzYxJV_rmE8&amp;ab_channel=HBO">Succession</a> series finale, I&#8217;ve been mostly relying on Nicholas Britell&#8217;s beautiful, grandiose, and melancholic <a href="https://music.youtube.com/playlist?list=OLAK5uy_lOyxaZgusDUT9c5NQK5176AvBKQYTruNA">season 4 score</a> as my coding playlist - with a focus on <a href="https://music.youtube.com/watch?v=K2l-GKWBbtw&amp;list=OLAK5uy_lOyxaZgusDUT9c5NQK5176AvBKQYTruNA">Andante Risoluto</a> (unsurprisingly).</p><div><hr></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.<br><br>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Espresso #4: Data orchestration philosophies, and why Coalesce matters]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Friday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/espresso-4-data-orchestration-philosophies</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-4-data-orchestration-philosophies</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Fri, 14 Oct 2022 15:00:21 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,<br><br>It has been a while since our last espresso, and the data space has evolved a lot in the meantime - but that gives us yet more topics to talk about.<br><br>In this edition, we will talk about the philosophies behind two different data orchestration approaches and why Coalesce (dbt Labs&#8217; yearly conference) matters. So without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><h2>Airflow vs. Dagster: it&#8217;s all about philosophy</h2><p>Ever since it was open-sourced by Airbnb back in 2015, Apache Airflow established itself as the de-facto standard for orchestration within the data space. Thanks to its feature-rich User Interface (UI), its ability to manage a wide range of operations, and particularly its no-nonsense and intuitive approach to organizing workflows via DAGs and tasks, it quickly eclipsed existing orchestrators like Spotify&#8217;s Luigi (<strong><a href="https://developer.spotify.com/community/news/2012/09/24/hello-world/">which was open-sourced in 2012</a></strong>) and the Hadoop ecosystem&#8217;s <strong><a href="https://oozie.apache.org/">Oozie</a></strong>.</p><p>Airflow is used today by data engineering teams around the world for an ever-expanding list of use cases, supported by custom operators, in-house abstractions, and a myriad of hacks to leverage <strong><a href="https://airflow.apache.org/docs/apache-airflow/stable/concepts/xcoms.html">some of its aging features</a></strong>. On the other hand, the way we interact with data today is very different from how things were seven years ago:</p><ul><li><p>We no longer just want to run Spark jobs. Instead, we think about the quality, state, and lineage of our heterogeneous data assets.</p></li><li><p>We can no longer tolerate weeks-long development cycles to generate new data assets. Instead, data pipelines should be written, tested, and deployed as efficiently and as fast as possible.</p></li><li><p>We can no longer rely on a small centralized data engineering team that builds and maintains all the DAGs. Instead, we aim for self-service capabilities and automation that would allow a larger set of contributors to build data assets and push them to production.</p></li><li><p>Finally, the data stack we built for the above is moving fast and changing old patterns. And so we no longer think about tasks, but we think about dbt models, Airbyte connectors, metrics, and a whole ecosystem of capabilities that are the modern incarnation of 2015&#8217;s Airflow <em>operators</em> we once had to write from scratch.</p></li></ul><p>With the above in mind, it&#8217;s definitely time to ask the question: Is Airflow still the undisputed go-to orchestrator for data pipelines, or is Dagster, the new orchestrator that&#8217;s built with all the previous points in mind, the better option?</p><p>In <a href="https://www.restack.io/docs/airflow-vs-dagster">my latest article</a>, published on <a href="https://www.restack.io/">Restack&#8217;s</a> blog, I go into the details of how Dagster differs from Airflow and whether data teams that have already invested a lot in Airflow should make the jump. But most importantly, I explain why &#8220;Airflow vs. Dagster&#8221; is not a technical question at all - it's just a matter of philosophy.</p><h2>From binge-reading to binge-watching</h2><p>I personally believe that the year&#8217;s most important Modern Data Stack conference isn&#8217;t Databricks&#8217; Data + AI summit or Snowflake&#8217;s summit - instead, it&#8217;s dbt Labs&#8217; <a href="https://coalesce.getdbt.com/">Coalesce</a>. This is not only because dbt is the tool that opened the door to <a href="https://towardsdatascience.com/building-an-end-to-end-open-source-modern-data-platform-c906be2f31bd">the third wave of data technologies</a>, but also because at Coalesce everyone within the data community <em><strong>belongs</strong></em>.</p><p>If you never attended Coalesce before, I totally recommend doing so this year. You&#8217;ll learn quite a lot about the Modern Data Stack and how fellow data practitioners are doing more with third-wave data technologies. You&#8217;ll see how fun and engaging the dbt Slack is. You&#8217;ll experience how welcoming and diverse the data community is. And most importantly, <em><strong>you&#8217;ll feel that you belong - because you do</strong></em>.</p><p>Throughout the next couple of weeks, let&#8217;s binge-watch Coalesce talks (there are <em>way too many</em> interesting ones) instead of binge-reading data content. And hey - it&#8217;s our yearly opportunity to tackle the endless debate: is Kimball's dimensional data modeling approach still relevant? (even though we all know that the only right answer is &#8220;<em>mostly yes, but it depends</em>&#8221;).</p><h2>A sound to code to</h2><p>After watching <a href="https://www.youtube.com/watch?v=JtqIas3bYhg">Cyberpunk Edgerunners</a>, I&#8217;ve been listening to most of <a href="https://music.youtube.com/playlist?list=PLl-vhnGPY7cpGmDrC1P5qzY9QebtwuYAe&amp;feature=share">its soundtrack</a> on repeat for the past couple of weeks (it&#8217;s really that good) - and I can confirm that it offers a productivity boost similar to <a href="https://cyberpunk.fandom.com/wiki/Cyberpunk_2077_Cyberware">Night City&#8217;s finest cyberware upgrades</a>. My personal favorite is - unsurprisingly - &#8220;<a href="https://youtu.be/h4VJGNNSQnw">I Really Want to Stay At Your House</a>&#8221;.</p><div><hr></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.<br><br>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data Espresso! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Espresso #3: Data quality, unbundle to rebundle, and navigating data content]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Wednesday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/espresso-3-data-quality-unbundle</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-3-data-quality-unbundle</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 09 Mar 2022 16:00:54 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,<br><br>In this edition, we will talk about data quality, the future of orchestration in a fragmented modern data stack, and how to navigate the endless stream of data content. So without further ado, let&#8217;s talk data engineering while the espresso is still hot.</p><h2>Having trust issues with your data?</h2><p>A statement that I keep running into variations of online is how &#8220;<em>you can&#8217;t create value with ML if you don&#8217;t have good data to begin with</em>&#8221;. (I.e., <em>garbage in, garbage out</em>).</p><p>I personally believe that neglecting data quality and building data products without first ensuring robust data validation is one of the key issues that eventually lead to failed data initiatives. But why did data quality only gain attention rather recently?</p><p>For many years (think 2010 to 2016) data engineers were building pipelines without software engineering best practices - the goal was to deliver as much data as possible, as fast as possible. This didn&#8217;t cause any immediate issues at the time, because &#8220;<em>big data</em>&#8221; was still a secondary factor for decision-making at most companies. But now that data is a first-class citizen everywhere and metrics are consumed by a wide range of users/teams, we frequently find ourselves trying to understand why two dashboards present different values for the same metric, or how a failure in one pipeline would impact our end-users. Now, we&#8217;re paying our overdue debts for building data pipelines without having data quality in mind. So how should you tackle data trust issues?</p><p>Within the Modern Data Stack, you&#8217;ll find dedicated companies that concentrate on solving this particular issue. While this may be the way to go if you don&#8217;t actually have any data engineers within your company (you should though!), I find that it adds unnecessary complexity and costs to solve a rather simple problem - for most cases.</p><p>Data quality, at its core, can be achieved by being able to answer three main questions:</p><ul><li><p>What types of checks do you want to implement for your pipelines? (schema checks, data checks, etc.)</p></li><li><p>Which operations should you implement to ensure that your data meets the standards defined by answering the first question? (row count, handling nulls, deduplicating the data, recasting fields, ensuring default values, etc.)</p></li><li><p>What action should happen when one of your checks fails? (raising a warning or an error based on the type of the check and the severity of the issue, sending alerts, etc.)</p></li></ul><p>For example, if you have an architecture that relies on Spark-based pipelines, you can implement the checks and actions related to them as Spark-based nodes within your orchestrated pipelines, with the aim of running these checks as close to the source as possible (to minimize the impact on downstream processes). Or better yet, you can leverage a tool like Soda SQL or Great Expectations, which are open-source projects that will simplify your data quality checks (you&#8217;ll have plenty of pre-defined checks out of the box) without adding much complexity to your stack.</p><p>Data quality is a key pillar of having a successful data-driven strategy, and yet as a problem, it&#8217;s actually <em><strong>not that complex</strong></em>. Adding yet-one-more-vendor to your stack solely for data quality is unnecessary in most cases because there&#8217;s no hidden complexity behind the core problem.</p><h2>Fresh off the press</h2><p>A couple of weeks ago, the unbundle vs. rebundle debate took Data Twitter by storm.</p><p>First, Gorkem Yurtseven published <a href="https://blog.fal.ai/the-unbundling-of-airflow-2/">The Unbundling of Airflow</a> on Features &amp; Labels&#8217; blog, arguing that Airflow (and workflow engines in general) is being unbundled into separate tools each focusing on one part of the data stack, which would eventually make the orchestration engine redundant.</p><p>Then, Nick Schrock published <a href="https://dagster.io/blog/rebundling-the-data-platform">Rebundling the Data Platform</a> on Dagster&#8217;s blog, countering with the point that the &#8220;next thing&#8221; wouldn&#8217;t be giving up on workflow engines, but instead evolving them so that they can orchestrate <em><strong>software-defined assets </strong></em>- which Dagster now supports.</p><p>With the modern data stack being extremely fragmented in its current state, the debate about how to <em>connect</em> all these tools is just getting started.</p><h2>Out of the comfort zone: navigating the endless stream of data content</h2><p>A few weeks ago I came across a tweet about how it&#8217;s extremely hard to keep up with all the topics/discussions in the data community - and I immediately related to it:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://twitter.com/tayloramurphy/status/1491574303783071745?s=20&amp;t=1p9AeUZmNzbalgUdu_oqVA&quot;,&quot;full_text&quot;:&quot;genuinely distressed about the volume of data content I don't have time to consume (let alone write) and the number of great conversations happening in the community. I need to spend time with my family people! stop being so great!!!&quot;,&quot;username&quot;:&quot;tayloramurphy&quot;,&quot;name&quot;:&quot;data &#127345;&#65039;oi&quot;,&quot;profile_image_url&quot;:&quot;&quot;,&quot;date&quot;:&quot;Thu Feb 10 00:46:40 +0000 2022&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:0,&quot;retweet_count&quot;:1,&quot;like_count&quot;:37,&quot;impression_count&quot;:0,&quot;expanded_url&quot;:{},&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p> This is an extremely beneficial and healthy environment for the field itself since it&#8217;s boosting innovation and opening new possibilities every single day - but it can also get very overwhelming for us humans of data.</p><p>Personally, I try to stick to two rules when it comes to navigating data content:</p><ol><li><p>Streamline your sources: Multiple people in the data space try to post content on a daily basis - while this is done with good intentions, I genuinely believe that as humans we simply can&#8217;t deliver thoughtful insights every single day. If you&#8217;re following many data &#8220;influencers&#8221; on LinkedIn for example, you&#8217;ll find yourself frequently consuming short pieces of content throughout the day that are written mostly because the writer wants to share content, not because there&#8217;s a genuine thought or reflection. Consuming these &#8220;bites&#8221; of content drains your precious energy and points you in dozens of directions in one single scrolling session. Instead of merely following people who post &#8220;daily&#8221;, I recommend seeing what the community is discussing on Twitter once or twice per day (personally I usually try to avoid threads), where the 280 character limit enforces everyone to go straight to the point - and then you can make your own deep dives into topics that interest you.</p></li><li><p>Streamline your thoughts:  This second rule is a direct result of the first one and the fact that data engineering is a vast and ever-evolving field. By limiting your sources and consuming long-form content only when the topic genuinely interests you, you&#8217;ll be able to build knowledge in areas that matter to you. Instead of knowing a bit of everything, the aim should be to know a bit of everything and a lot about a few things. This would allow you to formulate your own thoughts and opinions instead of only consuming content, and will help you get a better understanding of the bigger picture. If you&#8217;re interested in the concept of a &#8220;metrics layer&#8221; for example, you can spend time going through articles and white papers by major tech companies who built their own internal metrics platforms, to learn the &#8220;why&#8221;s, &#8220;how&#8221;s, and the lessons they learned along the way.</p></li></ol><p>My point is that trying to stay up to date with everything happening within data engineering / the data stack is a futile effort that won&#8217;t give you long-term knowledge, whereas focusing on a few areas that interest you and building meaningful and deep knowledge in them is a path towards widening your expertise and potential.</p><p>But again, this is the approach that works for me - and it won&#8217;t necessarily work for everyone.</p><h2>A sound to code to</h2><p>During the last few weeks I&#8217;ve been the most productive when playing Polo &amp; Pan&#8217;s latest album, <a href="https://music.youtube.com/playlist?list=OLAK5uy_md42pYz4E7fDP6IXdWfYmoyJ1v0g7JOKU">Cyclorama</a>. With <em><a href="https://music.youtube.com/watch?v=V9r1TOtilQQ&amp;feature=share">Attrape-r&#234;ve</a> </em>as a definite standout.</p><div><hr></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.<br><br>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Espresso #2: Open table formats, real-time at scale, and what happens behind closed doors]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Wednesday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/espresso-2-open-table-formats-real</link><guid isPermaLink="false">https://dataespresso.substack.com/p/espresso-2-open-table-formats-real</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 02 Feb 2022 15:54:27 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,</p><p>If this is your first time reading Data Espresso, I recommend going through the first two posts (<a href="/__u/dataespresso.substack.com/p/so-whats-data-espresso">post #0</a> and <a href="/__u/dataespresso.substack.com/p/our-first-espresso-open-source-modern">post #1</a>) to get familiar with the newsletter&#8217;s concept and the motivation behind it.</p><p>As mentioned in the previous edition, the newsletter is still  not in its finalized format and so certain sections can be changed in the future - but enough with the introductions, let&#8217;s talk data engineering while the espresso is still hot.</p><h2>Why open table formats are a game-changer</h2><p>One word that characterized the second wave of data systems (the Hadoop ecosystem and co.) is <em>compromise</em> - to have horizontal scalability in storage and compute we had to let go of certain features that were a given in existing systems, like ACID transactions and schema enforcement. And although most of the compromises were necessary (as explained by <a href="https://en.wikipedia.org/wiki/CAP_theorem">the CAP theorem</a> for example), some of them weren&#8217;t.</p><p>The de-facto metadata store of the Hadoop ecosystem, the Hive metastore, is where most of the compromises happened. To manage distributed and partitioned file-based tables, the Hive metastore was designed in a way completely disconnected from the data itself - it tracks directories (for both tables and partitions) instead of data files, and so most of the features that are offered by a typical RDBMS are effectively unattainable since the metastore doesn&#8217;t directly manage the data files.</p><p>This design choice in the Hive table format meant that Hive tables were problematic to manage (since the schema isn&#8217;t verified for every data file), hard to trust (since there&#8217;s no table transaction log), and also inefficient when they reach a certain scale (since we can&#8217;t optimize the queries below the partition level and each query necessitates listing the data files).</p><p>Due to such issues, many companies struggled with file-based tables. Additionally, advancements in distributed query engines and the metadata management space were slowed down by the inefficiencies of the Hive metastore - pushing companies to opt for a scalable data warehouse for their analytics architecture instead of relying on a file-based system.</p><p>But that all changed in 2019. In one year, Databricks open-sourced its Delta project under the name Delta Lake, and both Uber and Netflix submitted their own table formats, Hudi and Iceberg respectively, to the Apache Software Foundation. </p><p>Each one of these new table formats, in its own way, solved most of the issues related to the Hive metastore with a set of features and key strengths that differentiate it from the other formats - but most importantly, their introduction meant that data lakes are no longer inferior to data warehouses when it comes to table management and query optimization. <em><strong>The data lake is dead, long live the <a href="https://databricks.com/blog/2020/01/30/what-is-a-data-lakehouse.html">lakehouse</a>!</strong></em></p><p>To familiarize yourself with open table formats, and determine which one suits your needs the most, the following blogs would be a great place to start:</p><ul><li><p>Iceberg:</p><ul><li><p><a href="https://tabular.io/blog/iceberg-metadata-indexing/">Metadata Indexing in Iceberg</a> by Ryan Blue (Iceberg creator)</p></li><li><p><a href="https://medium.com/expedia-group-tech/a-short-introduction-to-apache-iceberg-d34f628b6799">A Short Introduction to Apache Iceberg</a> by Christine Mathiesen (Expedia Group)</p></li><li><p><a href="https://www.dremio.com/resources/webinars/apache-iceberg-an-architectural-look-under-the-covers-2/">Apache Iceberg &#8211; An Architectural Look Under the Covers</a> by Jason Hughes</p></li></ul></li><li><p>Delta Lake:</p><ul><li><p><a href="https://medium.com/adobetech/massive-data-processing-in-adobe-experience-platform-using-deltalake-de5ceef2c150">Massive Data Processing in Adobe Experience Platform Using DeltaLake</a> by the Adobe Experience Platform team</p></li><li><p><a href="https://engineering.salesforce.com/engagement-activity-delta-lake-2e9b074a94af">Engagement Activity Delta Lake</a> by the Salesforce Engineering team</p></li><li><p><a href="https://databricks.com/blog/2019/08/21/diving-into-delta-lake-unpacking-the-transaction-log.html">Diving Into Delta Lake: Unpacking The Transaction Log</a> on the Databricks blog</p></li></ul></li><li><p>Hudi:</p><ul><li><p><a href="https://hudi.apache.org/blog/2021/07/21/streaming-data-lake-platform/">Apache Hudi - The Data Lake Platform</a> on the Apache Hudi blog</p></li><li><p><a href="https://medium.com/slalom-build/data-lakehouse-building-the-next-generation-of-data-lakes-using-apache-hudi-41550f62f5f">Data Lakehouse: Building the Next Generation of Data Lakes using Apache Hudi</a> by Ryan D'Souza &amp; Brandon Stanley</p></li></ul></li></ul><h2>Fresh off the press</h2><h4><a href="https://medium.com/@ZhenzhongXu/the-four-innovation-phases-of-netflixs-trillions-scale-real-time-data-infrastructure-2370938d7f01">The Four Innovation Phases of Netflix&#8217;s Trillions Scale Real-time Data Infrastructure</a></h4><p>Zhenzhong Xu, who led the Stream Processing Platform team at Netflix, dives deep into the different phases that the Netflix real-time data infrastructure went through, with their respective challenges and learnings.</p><h4><a href="https://towardsdatascience.com/data-to-engineers-ratio-a-deep-dive-into-50-top-european-tech-companies-58abc23e36ca">Data to engineers ratio: A deep dive into 50 top European tech companies</a></h4><p>A very interesting analysis by Mikkel Dengs&#248;e, the head of data science at Monzo, in which he compares the data engineers ratio at 50 tech companies from different sectors.</p><h2>Out of the comfort zone</h2><p>From the outside, it&#8217;s easy to romanticize Silicon Valley and the idea of launching startups that turn into tech giants making the world a better place - but we all know that things don&#8217;t happen that way.</p><p>If you watched <a href="https://www.imdb.com/title/tt1285016/">The Social Network</a> then you probably already know that &#8220;<em><strong>you don't get to 500 million friends without making a few enemies&#8221;</strong></em>, and the power struggle that we witness in the movie isn&#8217;t an exception or something specific to Facebook (um, Meta), but it&#8217;s a story that&#8217;s rather familiar in Silicon Valley. Some of these stories that happen behind closed doors were documented by award-winning journalists via page-turner books. Below are my recommendations:</p><ul><li><p><a href="https://www.goodreads.com/en/book/show/44573628-super-pumped">Super Pumped: The Battle for Uber</a> by Mike Isaac: this one is my personal favorite because it not only compellingly tells the story behind Uber, but also delves into the dark side of Silicon Valley unicorns and the &#8220;hustlin&#8217;&#8221; that happens behind the scenes.</p></li><li><p><a href="https://www.goodreads.com/book/show/50772888-no-filter">No Filter: The Inside Story of Instagram</a> by Sarah Frier: even though Kevin Systrom and Mike Krieger (Instagram&#8217;s founders) agreed to sell the company to Facebook early on, life didn&#8217;t get any easier for them and their small team. Sarah Frier masterfully tells the story of a company within a company, and how the decisions of a small group of engineers can change the world we live in.</p></li><li><p><a href="https://www.goodreads.com/en/book/show/56470423-an-ugly-truth">An Ugly Truth: Inside Facebook's Battle for Domination</a> by Sheera Frenkel and Cecilia Kang: Yes, it&#8217;s Facebook/Meta again, and this time you get to discover the dynamics between Mark Zuckerberg and Sheryl Sandberg, how the Cambridge Analytica scandal unfolded, and what led to the company&#8217;s fall from grace.</p></li><li><p><a href="https://www.goodreads.com/book/show/37976541-bad-blood">Bad Blood: Secrets and Lies in a Silicon Valley Startup</a> by John Carreyrou: Last but not least is the story of the unicorn that never was. Theranos was probably one of the biggest scams of the 21st century, and the story of Elizabeth Holmes&#8217; startup is filled with lessons about what can go wrong when aiming to change the world - but most importantly, it shows that most people aren&#8217;t as smart as they think they are.</p></li></ul><h2>A sound to code to</h2><p><a href="https://music.youtube.com/playlist?list=OLAK5uy_mVCuxlsZbnB75VO4F0PVzTLo3gEXoQoQc">Bonobo&#8217;s new album, Fragments</a>, is all you need for a productive coding session.</p><div><hr></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.<br><br>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[Our first Espresso: Open-source, Modern Data Stack, and Range]]></title><description><![CDATA[Make yourself an espresso and join me for a short break on a Wednesday afternoon &#9749;]]></description><link>https://dataespresso.substack.com/p/our-first-espresso-open-source-modern</link><guid isPermaLink="false">https://dataespresso.substack.com/p/our-first-espresso-open-source-modern</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Wed, 12 Jan 2022 18:14:50 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello fellow data enthusiasts,</p><p>First of all, I&#8217;m glad that you&#8217;re joining me on this ride - hopefully this will be a mutual learning experience. The newsletter is still going through the first cycles of its life, so probably this is not the finalized format and certain sections can be changed. Throughout the first few issues, you&#8217;ll basically see the newsletter go through adolescence.</p><p>As stated in my initial post, my aim for this newsletter is that it will be more of an exchange of ideas, so you&#8217;re definitely encouraged to reply to the emails with your thoughts and feedback.</p><h2>Let&#8217;s talk data engineering</h2><h4>Can open-source survive in the modern data stack?</h4><p>If we look back to the previous wave of data technologies, the wave of the Hadoop ecosystem, NoSQL databases, and the term &#8220;Big Data&#8221;, we&#8217;ll notice that open-source was the standard. The data technologies realm was one of the first battles that open-source won, and this in turn was one of the most important factors that led to the rapid pace of innovation during the Hadoop era. The whole ecosystem thrived on the open-source model, with multiple teams of engineers from the biggest tech giants contributing to the same projects under the umbrella of open-source foundations.</p><p>In contrast, if we look at the ecosystem of the modern data stack, we&#8217;ll notice that things are not so similar. We&#8217;ve somehow taken a step back in a way and decided to rethink open-source-based business models. This is the result of multiple factors, one of which is cloud providers offering managed versions of open-source projects that effectively render the companies built on top of these projects redundant.</p><p>The consequence of this is either closed-source business models (which are back in style) or stricter licenses similar to <a href="https://www.elastic.co/pricing/faq/licensing">what Elastic did in 2021</a> - which would offer open-source-based companies a fighting chance against cloud providers.</p><p>With the tweaked licensing and the continued support of the tech communities, open-source is still a viable option for data companies. The best proof of that is Airbyte, which disrupted the data integration space with its open-source approach and just recently announced <a href="https://www.businesswire.com/news/home/20211217005648/en/Airbyte-Closes-150-Million-Series-B-Funding-Round-Led-by-Altimeter-Capital-and-Coatue-Management">a $150 Million series B funding round</a>, at a $1.5 billion valuation. (<a href="https://airbyte.com/blog/a-new-license-to-future-proof-the-commoditization-of-data-integration">after also opting for a stricter license</a>)</p><p>So, yes, open-source will survive (and thrive) in this era, but we&#8217;ll potentially witness a slower rate of innovation due to the fragmented ecosystem and the closed-source model that multiple modern data stack companies opted for.</p><h2>Fresh off the press</h2><h4>Life with dbt</h4><p>When considering adding a new tool to your stack, nothing is more helpful than the feedback of other engineering teams that are already using it. Well, if you&#8217;re considering using dbt (and you should be), the data engineering team at Devoted published <a href="https://tech.devoted.com/one-year-of-dbt-b2e8474841ca">a very insightful piece</a> on their first year using it and how powerful it can be when integrated with the different components of your stack. <em>(the post is written by <a href="https://www.linkedin.com/in/adam-boscarino-47267313/">Adam Boscarino</a> and <a href="https://www.linkedin.com/in/jasonbrownstein/">Jason Brownstein</a>.)</em></p><h4>A look back at 2021</h4><p>2021 was an eventful year when it comes to the Modern Data Stack, and in <a href="https://towardsdatascience.com/trends-that-shaped-the-modern-data-stack-in-2021-4e2348fee9a3">an excellent article published on Towards Data Science</a>, Salma Bakouk (co-founder &amp; CEO of <a href="https://www.siffletdata.com/">Sifflet</a>) goes through the trends that shaped it.</p><h4>&#8230; And a look ahead to 2022</h4><p><a href="https://towardsdatascience.com/the-future-of-the-modern-data-stack-in-2022-4f4c91bb778f">In an article also on Towards Data Science</a>, Prukalpa Sankar (co-founder of <a href="https://atlan.com/">Atlan</a>) discusses six ideas that will continue to shape the Modern Data Stack in 2022. </p><h2>Out of the comfort zone</h2><p>One of the best books I&#8217;ve ever read is <strong><a href="https://www.goodreads.com/en/book/show/41795733-range">Range: Why Generalists Triumph in a Specialized World</a> </strong>by <a href="https://www.goodreads.com/author/show/7164089.David_Epstein">David Epstein</a>, for a very simple reason - it proved something that I always believed in: you should always start gathering knowledge, expertise, and skills horizontally before specializing in a particular subfield.</p><p>If you just graduated from university and you want to start a career in a technology field (not necessarily data engineering), it might be tempting to jump headfirst into learning the most in-demand technology/tool of that field. While technically that might indeed help you land your first job, your long-term strategy should be built around widening your knowledge and expertise <strong>horizontally </strong>as much as possible, before making vertical deep dives into a specific technology/trend.</p><p>The reasoning behind this is that by first familiarizing yourself with concepts and paradigms, you&#8217;ll be able to:</p><ul><li><p>Learn and master the technologies themselves faster</p></li><li><p>Recognize patterns more easily and know when and how to leverage the right technology for a specific use-case</p></li><li><p>Have the flexibility to switch between technologies without a significant learning curve</p></li></ul><p>For example, if you want to get into data engineering, instead of starting with a deep dive into writing Apache Spark jobs, focus on learning concepts like distributed computing, massively parallel processing (MPP), MapReduce, ETL/ELT, and data modeling. This will first help you learn Spark in a shortened amount of time (because you&#8217;ll be comfortable with the concepts behind it), better understand how Spark implements specific concepts, and more importantly, recognize when Spark isn&#8217;t the right tool for a given problem. In the long term, this will give you one (very important) additional possibility throughout your career:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!abyU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 424w, /__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 848w, /__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!abyU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif" width="320" height="204.8" data-attrs="{&quot;src&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:160,&quot;width&quot;:250,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:482110,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 424w, /__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 848w, /__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 1272w, /__u/substackcdn.com/image/fetch/$s_!abyU!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F3670b850-533c-4063-8d7e-79dd8ce6edd4_250x160.gif 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Yes, you&#8217;ll be able to pivot from one technology/field to another one without much trouble - and in a field that&#8217;s moving extremely fast like data engineering, this is a very valuable asset.</p><p>A great book that can help you understand a wide range of data-related concepts is <strong><a href="https://www.goodreads.com/book/show/23463279-designing-data-intensive-applications">Designing Data-Intensive Applications</a> </strong>by <a href="https://www.goodreads.com/author/show/7969625.Martin_Kleppmann">Martin Kleppmann</a>.</p><h2>A sound to code to</h2><p>Asking a developer whether they like to listen to music while coding is like asking someone whether they&#8217;re a cat person or a dog person: it&#8217;s something that feels like a personality tell, but you&#8217;re never quite sure what it actually means.</p><p>Personally, if I&#8217;m not working on a complex topic, I find myself more productive when listening to certain types of music - and the aim of this section is to share with you some of the records that I find the most suitable for coding sessions.</p><p>Since this is the first issue of the newsletter, my recommendation is actually the artist I listen to the most when working:<a href="https://www.youtube.com/channel/UCbcUt4SAP587pHR4v9O60aA"> Jamie xx</a> (yes, <a href="https://music.youtube.com/channel/UCQ1Xqlk89CLZ9aAFyo-1THg">the xx</a>&#8217;s discreet DJ). I won&#8217;t burden you with unnecessary details - I&#8217;d just recommend that you give his 2015 album, <a href="https://music.youtube.com/playlist?list=OLAK5uy_nRRUoAJ4wFbgImMxhWo7975oZIQeW-9vw">In Colour</a>, a listen during your next coding session. It&#8217;ll either turn into a dance session or you&#8217;ll finish what you&#8217;re working on in half the expected time: it&#8217;s a win-win.</p><div><hr></div><p>If you enjoyed this issue of Data Espresso, feel free to recommend the newsletter to people in your entourage.<br><br>Your feedback is also very welcome, and I&#8217;d be happy to discuss one of this issue&#8217;s topics in detail and hear your thoughts on it.</p><p>Stay safe and caffeinated &#9749;</p>]]></content:encoded></item><item><title><![CDATA[So, what's Data Espresso?]]></title><description><![CDATA[Before you open this, go make yourself an espresso. You'll need it.]]></description><link>https://dataespresso.substack.com/p/so-whats-data-espresso</link><guid isPermaLink="false">https://dataespresso.substack.com/p/so-whats-data-espresso</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Thu, 23 Dec 2021 17:08:43 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey there &#128075; &#9749;</p><p>Welcome to Data Espresso, a bi-weekly newsletter&nbsp;in which I&#8217;ll discuss various topics related to data engineering and technology in general.</p><h2>Who am I?</h2><p>My name is Mahdi Karabiben, and I have been working with data engineering technologies on a daily basis since 2017. My (rather short) career has already allowed me to work with data in very different scenarios: I first started my career at <a href="https://democracyinternational.com/">Democracy International</a>, where I worked on visualizing electoral data; after that, I joined the data-marketing firm <a href="https://numberly.com/en/">Numberly</a> to build data pipelines for AdTech use cases; I then transitioned into the financial sector, where I worked on data projects at <a href="https://www.ca-cib.com/">CA-CIB</a> and <a href="https://www.factset.com/">FactSet</a>; and now I'm building data products at <a href="/__u/www.zendesk.com/">Zendesk</a> and experimenting with open-source projects in my free time.</p><p>I&#8217;m someone who has always been very passionate about data, but I also like building human connections. And those are the two pillars of this newsletter.</p><h2>Why am I starting a newsletter?</h2><p>Before COVID, I used to take an afternoon coffee break with a colleague of mine during which we discussed topics that range from the latest news, trends, and advancements in the data engineering world to the latest TV show that one of us binge-watched. When I look back at those conversations now, it&#8217;s clear to me that those coffee breaks were one of the highlights of my typical workday.</p><p>My aim with Data Espresso is to create a similar experience for you as a subscriber. Each issue would consist of the following sections:</p><ul><li><p>3 brief observations related to data engineering (trends, commentary on certain technologies, updates, etc.)</p></li><li><p>The most interesting article that I came across during the past two weeks.</p></li><li><p>A sound to code to: This is mostly the result of my 4 years as a member of the radio club back in university - Every issue will come with a song that I personally enjoy coding to.</p></li><li><p>Out of the comfort zone: This is a section that won&#8217;t have anything to do with data engineering or technology as a whole.</p></li></ul><p>As you can see, the idea consists of mixing technology with a human/personal aspect - like in an actual discussion between two people who happen to have the same job, while drinking an espresso.</p><h2>Why subscribe?</h2><p>Subscribe to get full access to the newsletter and <a href="/__u/dataespresso.substack.com/archive">website</a>. Never miss an update.</p><h3>Stay up-to-date</h3><p>You won&#8217;t have to worry about missing anything. Every new edition of the newsletter goes directly to your inbox.</p><h3>Join the crew</h3><p>Be part of a community of people who share your interests.</p><p>To find out more about the company that provides the tech for this newsletter, visit <a href="/__u/www.substack.com/">Substack.com</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/dataespresso.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Coming soon]]></title><description><![CDATA[This is Data Espresso, a newsletter about Data engineering updates and deep dives to accompany your morning espresso..]]></description><link>https://dataespresso.substack.com/p/coming-soon</link><guid isPermaLink="false">https://dataespresso.substack.com/p/coming-soon</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Fri, 17 Dec 2021 10:20:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!md3X!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F80458ee1-e940-4d9c-8e1a-619db61121eb_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This is Data Espresso</strong>, a newsletter about Data engineering updates and deep dives to accompany your morning espresso..</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://dataespresso.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/dataespresso.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[Fully Automating Your ML Pipelines With the AWS CI/CD Tools]]></title><description><![CDATA[A guide to building an automated MLOps pipeline by leveraging the trusted DevOps toolset.]]></description><link>https://dataespresso.substack.com/p/fully-automating-your-ml-pipelines-with-the-aws-ci-cd-tools-bfc337aa77c3</link><guid isPermaLink="false">https://dataespresso.substack.com/p/fully-automating-your-ml-pipelines-with-the-aws-ci-cd-tools-bfc337aa77c3</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Sun, 18 Jul 2021 23:18:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LR6N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LR6N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LR6N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg" data-attrs="{&quot;src&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/a777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!LR6N!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fa777289b-1978-4ee0-b090-a49c9b58f826_2600x1788.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><p>A guide to building an automated MLOps pipeline by leveraging the trusted DevOps toolset.</p><p><a href="https://towardsdatascience.com/fully-automating-your-ml-pipelines-with-the-aws-ci-cd-tools-bfc337aa77c3?source=rss-7cda12823b7a------2">Continue reading on Towards Data Science &#187;</a></p>]]></content:encoded></item><item><title><![CDATA[Running Apache Superset at Scale]]></title><description><![CDATA[A set of recommendations and starting points to efficiently run Superset at scale]]></description><link>https://dataespresso.substack.com/p/running-apache-superset-at-scale-1539e3945093</link><guid isPermaLink="false">https://dataespresso.substack.com/p/running-apache-superset-at-scale-1539e3945093</guid><dc:creator><![CDATA[Mahdi Karabiben]]></dc:creator><pubDate>Mon, 22 Mar 2021 15:49:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qKYS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qKYS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 424w, /__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 848w, /__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_webp, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qKYS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png" data-attrs="{&quot;src&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/d3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_424, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 424w, /__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_848, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 848w, /__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_1272, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qKYS!, /__u/dataespresso.substack.com/w_1456, /__u/dataespresso.substack.com/c_limit, /__u/dataespresso.substack.com/f_auto, /__u/dataespresso.substack.com/q_auto:good, /__u/dataespresso.substack.com/fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3ed4991-97f3-4f91-a7d2-f5d6c5d11645_2600x1593.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><p>A set of recommendations and starting points to efficiently run Superset at scale</p><p><a href="https://towardsdatascience.com/running-apache-superset-at-scale-1539e3945093?source=rss-7cda12823b7a------2">Continue reading on Towards Data Science &#187;</a></p>]]></content:encoded></item></channel></rss>