<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Data Marks]]></title><description><![CDATA[A newsletter dedicated to talking about data science, experimentation and how to pass the mark.]]></description><link>https://datamarks.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!xdKG!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fdatamarks.substack.com%2Fimg%2Fsubstack.png</url><title>Data Marks</title><link>https://datamarks.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 00:36:08 GMT</lastBuildDate><atom:link href="/__u/datamarks.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Eltsefon Mark]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[datamarks@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[datamarks@substack.com]]></itunes:email><itunes:name><![CDATA[Eltsefon Mark]]></itunes:name></itunes:owner><itunes:author><![CDATA[Eltsefon Mark]]></itunes:author><googleplay:owner><![CDATA[datamarks@substack.com]]></googleplay:owner><googleplay:email><![CDATA[datamarks@substack.com]]></googleplay:email><googleplay:author><![CDATA[Eltsefon Mark]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[How to stay relevant as a data scientist]]></title><description><![CDATA[Everything is changing super fast, and the obvious question is:]]></description><link>https://datamarks.substack.com/p/how-to-stay-relevant-as-a-data-scientist</link><guid isPermaLink="false">https://datamarks.substack.com/p/how-to-stay-relevant-as-a-data-scientist</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 29 Jul 2026 11:42:47 GMT</pubDate><content:encoded><![CDATA[<p>Everything is changing super fast, and the obvious question is:</p><div class="pullquote"><p><strong>How can you make sure you stay relevant?</strong></p></div><p>The traditional career path used to be clear.</p><p>You started as a junior data scientist, became senior, moved into management, managed larger teams, and eventually aimed for a director or VP position.</p><p>I don&#8217;t think this path is as obvious anymore.</p><p>Companies are becoming flatter, some middle-management layers are getting thinner.</p><p>Strong individual contributors are becoming more valuable. </p><p>We have all seen former CTOs move into IC roles at AI companies. For example, Workday&#8217;s former CTO recently joined Anthropic as a Member of Technical Staff.</p><p>This doesn&#8217;t mean that management is disappearing.</p><p>It means that middle layers are shrinking and companies are betting on builders.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>We have entered the era of builders</h2><p>Today, your ability to build and deliver something matters more than ever.</p><p>It is no longer enough to know how to train a model, write SQL, create a dashboard, or explain different statistical methods.</p><p>&#8594; Can you take an ambiguous problem and turn it into something useful?</p><p>&#8594; Can you understand what the business really needs?</p><p>&#8594; Can you bring people together, make decisions, build a solution, launch it, measure its impact, and improve it?</p><p>&#8594; Do you understand how to accelerate business growth - with or without AI?</p><p>This is what companies desperately need.</p><p>The strongest data scientists will be the ones who can combine data, engineering, product thinking, and AI to move a project from an idea to a real outcome , not just complete a small part of it.</p><p>Your technical skills still matter. But technical knowledge without ownership will become less valuable.</p><h2>Get as close to AI as possible</h2><p>One of the clearest moves you can make today is to get close to AI.</p><p>But being close to AI doesn&#8217;t necessarily mean becoming an AI researcher or training large language models ( though it&#8217;s a great place to be right now) </p><p>For a data scientist, it can mean:</p><ul><li><p>Building AI-powered products</p></li><li><p>Designing evaluations for the AI features/agents</p></li><li><p>Measuring their impact on users and the business</p></li><li><p>Creating experimentation systems for AI products</p></li></ul><p>If you receive a strong offer for a role closely connected to AI, especially at a company where AI is the core product, I would go for it.</p><p>But don&#8217;t follow the word &#8220;AI&#8221; blindly.</p><p>A role where you only maintain dashboards for an AI team may help you less than a non-AI role where you own important problems and build things from start to finish.</p><p>The real question is not whether &#8220;AI&#8221; appears in the job title.</p><p>It is whether the role gives you real ownership and helps you understand how AI is changing products, companies, and the way people work.</p><h2>Learn to work with AI, not only use it</h2><p>Almost everyone is using ChatGPT or another AI tool now. That alone is not a competitive advantage.</p><p>The advantage comes from learning how to reorganise your work around AI.</p><p>Can you use it to explore unfamiliar codebases, analyse data, build prototypes, write documentation, challenge your thinking, and complete projects that were previously too large for one person?</p><p>Do you know how to give AI the right context, connect different tools, design effective workflows, and understand where human judgment is still needed?</p><p>AI gives individuals much more leverage. In some cases, one person can now do work that previously required a small team.</p><p>But this also raises expectations.</p><p>If everyone can produce code and analysis faster, the value moves towards choosing the right problem, judging the quality of the output, and making the right decision.</p><p>As AI becomes stronger, your judgment becomes more important&#8212;not less.</p><h2>Your company is not your identity</h2><p>You are not your job title, and you are not the company you work for.</p><p>Companies change. Teams are reorganised. Managers leave. Strategies shift. Roles that look safe today may disappear tomorrow.</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/how-to-stay-relevant-as-a-data-scientist">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What to learn when you're already a senior DS]]></title><description><![CDATA[There comes a point in a data science career when learning becomes harder to define.]]></description><link>https://datamarks.substack.com/p/what-to-learn-when-youre-already</link><guid isPermaLink="false">https://datamarks.substack.com/p/what-to-learn-when-youre-already</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 01 Jul 2026 12:02:04 GMT</pubDate><content:encoded><![CDATA[<p>There comes a point in a data science career when learning becomes harder to define.</p><p>At the beginning, the path is obvious. </p><p>Learn SQL. Learn statistics. Get comfortable with Python. Understand experiments, models, data pipelines and production systems.</p><p>Then you become senior.</p><p>You can still find gaps.</p><p>In fact, you can find an unlimited number of them: </p><ol><li><p>causal inference</p></li><li><p>LLM evaluation</p></li><li><p>Bayesian methods</p></li><li><p>Spark</p></li><li><p>Kubernetes</p></li><li><p>a new vector database released last Tuesday</p></li><li><p>another LLM framework</p></li></ol><p>There is always another subject that can make you feel slightly behind.</p><p>But another technical course or simulator rarely change the kind of work you are trusted with. </p><blockquote><p>Obviously abmarks.io is an exception and it will change the kind of work you are trusted with</p></blockquote><p>It might make you better at solving a particular problem. It will not necessarily make you better at finding the right problem, getting people to act on your work or helping your company make better decisions.</p><p>That is the real shift after senior.</p><div><hr></div><p>6 month ago, a colleague asked me:</p><p>&#8220;I am senior now. Great. What is left to study?&#8221;</p><p>I have been thinking about that question ever since. Here is what I think matters next.</p><h2>Learn to choose the problem</h2><p>A competent data scientist can answer a question.</p><p>A strong senior data scientist notices that the question is wrong.</p><p>&#8220;Can we predict which users will churn?&#8221; sounds reasonable. But the useful question might be:</p><blockquote><p>Which users can we realistically retain, with which intervention, and at what cost?</p></blockquote><p>The second question changes everything. It affects the target, the evaluation metric, the data you need and whether a model is useful at all.</p><p>At this level, the work is no longer just about choosing methods, validating data, and defining metrics. It is about judgment, priorities, and making decisions when the answer is not obvious.</p><p>A simple habit helps: before starting an analysis, write down five things.</p><ul><li><p><strong>What decision will this work support?</strong></p></li><li><p><strong>Who will make it?</strong></p></li><li><p><strong>What are the available choices?</strong></p></li><li><p><strong>What evidence could change the decision?</strong></p></li><li><p><strong>What happens if we are wrong?</strong></p></li></ul><p>When these questions have vague answers, better modelling will not save the project.</p><h2>Understand the whole system</h2><p>Senior data scientists usually understand models.</p><p>The next step is understanding everything around the model.</p><ul><li><p>Where did the data come from? </p></li><li><p>What behaviour generated it? </p></li><li><p>How does the output enter a product or workflow? </p></li><li><p>Who monitors it? </p></li><li><p>What happens when it fails quietly?</p></li></ul><p>The model is only one part of the work. Often, it is not even the most important part.</p><p>Choose one important metric or model in your company and trace it from beginning to end:</p><ul><li><p>how the raw event is produced;</p></li><li><p>how it is transformed;</p></li><li><p>where assumptions enter;</p></li><li><p>how the result is presented;</p></li><li><p>which action follows;</p></li><li><p>who notices when it is wrong.</p></li></ul><p>You will probably learn more from that exercise than from a general course on &#8220;advanced machine learning&#8221;.</p><h2>Learn to make decisions with incomplete information</h2><p>Earlier in a career, ambiguity often feels like a problem that someone else should remove.</p><p>Later, ambiguity is the work.</p><p>You will rarely have the sample size you wanted, a perfectly stable metric and six months to investigate. You will have three imperfect sources, a deadline on Friday and a stakeholder asking what you recommend.</p><p>The skill is not creating certainty. It is making uncertainty clear.</p><p>A useful senior answer often sounds like this:</p><blockquote><p>Here is what we know. Here is what we are assuming. Here is the risk in each option. Given that, this is what I would do.</p></blockquote><p>Keep an assumption log for important projects. Separate reversible decisions from expensive ones. Agree on failure criteria before seeing the results. Say what additional information is worth collecting, and what is unlikely to change the decision.</p><p>Knowing when to stop analysing is part of statistical judgment too.</p><h2>Learn to influence without becoming unbearable</h2><p>Data does not speak for itself.</p><p>Good analysis is not enough. Someone still has to explain why it matters.</p><p>Many technically strong data scientists discover this late. They produce correct work and assume that correctness should be enough. Then a weaker analysis wins because it is attached to a clearer story, a trusted person or a decision people already understand.</p><p>The most valuable learning rarely comes from a famous person giving a broad lecture. It usually comes from two experienced people discussing a real problem and being honest about what happened..</p><p>Influence does not have to mean manipulation. Usually it means doing the unglamorous work:</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/what-to-learn-when-youre-already">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[My own AI stack ]]></title><description><![CDATA[Most people still use AI one prompt at a time.]]></description><link>https://datamarks.substack.com/p/my-own-ai-stack</link><guid isPermaLink="false">https://datamarks.substack.com/p/my-own-ai-stack</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 17 Jun 2026 12:36:11 GMT</pubDate><content:encoded><![CDATA[<p>Most people still use AI one prompt at a time.</p><p>Open ChatGPT. Explain the problem. Paste some context. Get an answer. Copy it somewhere. Close the tab.</p><p>Next day, do it again.</p><p>That was me for a long time. The answers were useful, but the process was still annoying. Every session started from zero, and I was the one carrying the context.</p><p>This works for small tasks.</p><p>It breaks down for real data science work, because real DS work has memory.</p><p>You need to remember:</p><p>&#8594; which table has the correct metric</p><p>&#8594; which stakeholder asked for which cut</p><p>&#8594; which analysis version was trusted</p><p>&#8594; which caveat matters</p><p>&#8594; which meeting changed the direction </p><p>&#8594; which experiment had logging issues</p><p>&#8594; which dashboard is outdated but still used by leadership.</p><p>A lot of data science is not writing code. It is carrying context.</p><p>So my goal became simple:</p><p><strong>Stop re-explaining my work to AI.</strong></p><p>Below is the setup I use across personal and work projects.</p><p></p><h3><strong>1. Second Brain &#8212; the foundation</strong></h3><p><strong>The most important part of my setup is a structured workspace.</strong></p><p>I have seen many people and companies already call this kind of system a Second Brain, so why not keep the same name?</p><p>It is a set of files that describe my active projects, stakeholders, metric definitions, useful links, decisions, notes, and recurring workflows.</p><p>I organize it roughly using PARA:</p><ul><li><p>Projects</p></li><li><p>Areas</p></li><li><p>Resources</p></li><li><p>Archives</p></li></ul><p>It lives in Google Drive, so it survives across machines and sessions. The key part: my AI coding assistant can read it when I start working.</p><p>So instead of saying:</p><blockquote><p>Here is the background, here are the tables, here is the metric definition, here is what happened last week</p></blockquote><p>I can say:</p><blockquote><p>Continue the campaign performance analysis. Check if the WoW change still holds after excluding market B.</p></blockquote><p>That is a very different interaction. The AI is no longer starting from a blank page.</p><h3>What I keep inside</h3><p>For each active project, I usually keep:</p><ul><li><p>project goal</p></li><li><p>current status</p></li><li><p>main stakeholders</p></li><li><p>important links</p></li><li><p>table names</p></li><li><p>metric definitions</p></li><li><p>known caveats</p></li><li><p>decisions already made</p></li><li><p>open questions</p></li><li><p>next steps</p></li></ul><p>This is boring.</p><p>Boring is good.</p><p>Most productivity systems fail because they try to be too clever. This one works because it is just structured context.</p><p>The compounding effect is real. Every time I save a note, capture a decision, or update a project file, the next session gets better. After a few weeks of feeding it, you stop noticing the setup cost &#8212; you just describe what you need and it happens.</p><p><strong>One thing I got wrong early</strong>: I kept stuffing everything into one big context file. It grew to 700+ lines and actually made the agent <em>slower</em> and less focused. What works better is a layered approach &#8212; a slim top-level file that points to deeper context, loaded only when relevant.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/p/my-own-ai-stack?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Data Marks! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/p/my-own-ai-stack?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/datamarks.substack.com/p/my-own-ai-stack?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h3><strong>2. Claude Code &#8212; the hands</strong></h3><p>Second Brain is the memory. Claude Code is the hands.</p><p>For me as a data scientist, this means: writing and iterating on SQL, building analysis scripts, generating reports, editing documents, creating presentations. It&#8217;s not a chatbot &#8212; it&#8217;s a coding partner that lives in my terminal and can reach into my file system.</p><p>The way I use it most: I&#8217;ll describe an analysis I want to run &#8212; say, &#8220;I need to understand the week-over-week trend in campaign performance across these three markets, broken down by format&#8221; &#8212; and Claude Code will write the query, run it, iterate if something looks off, and produce a structured output. I steer, it executes.</p><p>Where this really shines is when the analysis isn&#8217;t a one-liner. Anything that involves multiple tables, joins, validation steps, or formatting into a deliverable &#8212; that&#8217;s where the autonomous loop saves me hours. I used to spend a full afternoon on what&#8217;s now a 20-minute conversation.</p><p>I also use it for things that aren&#8217;t strictly &#8220;data science&#8221; &#8212; writing documentation, drafting comms, structuring my thoughts into something coherent. It&#8217;s become the default interface for anything that requires sustained thinking plus output.</p><p></p><h3><strong>3. Wispr Flow &#8212; capturing thoughts before they disappear</strong></h3><p>This one sounds trivial but genuinely changed my workflow. </p><p>Wispr Flow is voice-to-text.</p><p>I speak, it writes.</p><p>Why does a data scientist need voice-to-text? Because my best thinking doesn&#8217;t happen at the keyboard.</p><p>I use Wispr to:</p><ul><li><p><strong>Save thoughts in transit.</strong> Walking between meetings, on the tube, making coffee &#8212; ideas come and I used to lose them. Now I speak them into a note and they&#8217;re captured. I have a running &#8220;thoughts&#8221; doc that I periodically feed back into my Second Brain.</p></li><li><p><strong>Write richer prompts.</strong> When I&#8217;m describing a complex analysis to Claude, speaking naturally produces way more context than typing. I&#8217;ll ramble for 30 seconds about what I actually want, and the prompt comes out better than anything I&#8217;d type.</p></li><li><p><strong>Draft first versions of documents.</strong> Instead of staring at a blank page, I talk through what I want to say. The structure is messy, but the ideas are there &#8212; then Claude cleans it up.</p></li><li><p><strong>Process meetings in real-time.</strong> Right after a meeting, while context is fresh, I&#8217;ll dictate my takeaways and action items. Thirty seconds of talking captures what would otherwise evaporate by the time I sit back down.</p></li></ul><p>The speed difference matters more than you&#8217;d think. Speaking is 3-4x faster than typing, and for some reason it&#8217;s also more <em>honest</em> &#8212; I say what I actually think instead of self-editing as I type.</p><div><hr></div><h3><strong>4. MyClaw &#8212; the always-on watcher</strong></h3><p>MyClaw is my always-on agent. It monitors things I care about and surfaces them without me having to go looking.</p><p>The setup: I&#8217;ve configured it to watch specific chat channels, track mentions of topics I care about, and generate periodic summaries. Think of it as a personalized news feed for my work context &#8212; but instead of me scrolling through hundreds of messages, it distills what matters.</p><p>I get:</p><ul><li><p>A daily digest of what happened in key channels while I was offline or in meetings</p></li><li><p>Alerts when specific topics come up (project names, metric names, stakeholder names)</p></li><li><p>Weekly summaries that help me draft my own updates</p></li></ul><p>It&#8217;s not glamorous. It&#8217;s not even particularly &#8220;smart.&#8221; But it solves a real problem: I was spending 30-40 minutes a day just <em>reading</em> to stay in the loop. Now the staying-in-the-loop happens in the background, and I get a 2-minute summary instead.</p><div><hr></div><h3><strong>5. NotebookLM &#8212; the research partner for messy, multi-document problems</strong></h3><p>NotebookLM is Google&#8217;s tool for reasoning over a collection of documents. You upload sources &#8212; PDFs, docs, web pages &#8212; and it lets you ask questions <em>across</em> them, synthesize, compare, find contradictions.</p><p>I use it for two very different purposes:</p><p><strong>Work: synthesizing domain research.</strong> When I&#8217;m scoping a new analysis area or trying to understand a complex business question, I&#8217;ll throw 10-20 relevant documents into NotebookLM &#8212; past analyses, strategy docs, meeting notes, industry reports &#8212; and have it build me a structured understanding. It finds connections I&#8217;d miss reading sequentially, and it cites everything so I can verify.</p><p><strong>Personal: my Global Talent visa application.</strong> This is the use case that sold me on the tool completely. The UK Global Talent visa requires you to demonstrate &#8220;exceptional talent or exceptional promise&#8221; in your field. The application involves pulling together evidence from multiple sources &#8212; publications, reference letters, proof of impact, immigration guidelines, successful application examples.</p><p>I uploaded everything: </p><p>&#8594; the official government guidance</p><p>&#8594;  successful application examples I found</p><p>&#8594; my own CV </p><p>&#8594; reference letter drafts</p><p>&#8594; proof of work documents</p><p>&#8594; relevant immigration forum threads. </p><p>Then I used NotebookLM as my research partner &#8212; asking it things like &#8220;what evidence do I have that maps to criterion X?&#8221;, &#8220;what gaps exist in my application compared to successful examples?&#8221;, &#8220;how should I frame this achievement in the context of the guidance?&#8221;</p><p>It&#8217;s essentially a research assistant that holds a hell of a context at once and lets you have a conversation with it. For any multi-document problem where the answer lives <em>between</em> documents rather than in any single one, it&#8217;s an awesome tool.</p><div><hr></div><h3><strong>6. Zapier / n8n &#8212; the glue and the automations</strong></h3><p>The last piece is automation. Zapier (and more recently n8n for things I want more control over) connects everything together and handles the recurring stuff.</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/my-own-ai-stack">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AI patience]]></title><description><![CDATA[You have heard about AI and how it can boost your productivity , your coding, everything and you have kicked off your journey.]]></description><link>https://datamarks.substack.com/p/ai-patience</link><guid isPermaLink="false">https://datamarks.substack.com/p/ai-patience</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 03 Jun 2026 11:35:26 GMT</pubDate><content:encoded><![CDATA[<p>You have heard about AI and how it can boost your productivity , your coding, everything and you have kicked off your journey.</p><p>You installed the tools. You watched the tutorials. You set up the workflows people on Twitter swear by. And then you actually tried to use AI for real work - your work - and it gave you something wrong. Not slightly wrong. Confidently wrong.</p><p>So you corrected it. It went in circles. You corrected it again. It hallucinated a table that doesn&#8217;t exist. You tried one more time, then closed the tab and thought: </p><p><em>I could&#8217;ve done this in half the time myself.</em></p><p>You&#8217;re not wrong. You probably could have.</p><p>But here&#8217;s what nobody tells you: <strong>that moment - the one where you want to quit - is the most expensive moment to quit.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The New Hire</h2><p>Imagine you hired someone brilliant but brand new. Day one, they don&#8217;t know your systems. They don&#8217;t know that when your team says &#8220;the dashboard,&#8221; they mean one specific dashboard. They don&#8217;t know that the column called <code>revenue</code> in your table actually means <em>gross</em> revenue and there&#8217;s a totally different field for net. They don&#8217;t know that your manager hates bullet points.</p><p>Would you fire that person after three days because they asked obvious questions? Of course not. You&#8217;d invest in onboarding them. You&#8217;d correct their mistakes, explain the context, and expect them to pick it up over time.</p><p>AI is that new hire. Except most people fire it after day two.</p><h2>Why the First Week Feels Broken</h2><p>The thing that makes AI frustrating at the start is the same thing that makes it powerful later: <strong>it adapts to context.</strong> But context doesn&#8217;t exist on day one. You have to build it.</p><p>When you first use AI on your real work, it&#8217;s operating with zero knowledge of:</p><ul><li><p>Your specific data (what tables exist, what columns mean, what&#8217;s stale)</p></li><li><p>Your conventions (how your team names things, what &#8220;done&#8221; looks like)</p></li><li><p>Your preferences (how you like explanations structured, what level of detail you want)</p></li><li><p>Your domain (the unwritten rules that everyone on your team knows but nobody&#8217;s documented)</p></li></ul><p>So it guesses. And guesses from a general-purpose model applied to a specific-purpose problem will be wrong. A lot. That&#8217;s not a bug. That&#8217;s the starting line.</p><p></p><h2>The Mistake That Costs People Months</h2><p>Here&#8217;s what most people do:</p><ol><li><p>Try AI on a real task</p></li><li><p>Get a mediocre or wrong output</p></li><li><p>Conclude &#8220;AI isn&#8217;t there yet for my work&#8221;</p></li><li><p>Go back to doing it manually</p></li><li><p>Try again in six months, repeat from step 1</p></li></ol><p>Every time they restart, they&#8217;re back at zero. No context was saved. No corrections were captured. The AI they come back to is exactly as clueless as the one they left.</p><p>Meanwhile, the people who are getting scary-good results from AI did something different. They didn&#8217;t have better tools or more technical skill. They just <strong>didn&#8217;t quit after step 2.</strong></p><p></p><h2>The Flywheel Nobody Sees</h2><p>Here&#8217;s what &#8220;sticking with it&#8221; actually looks like in practice:</p><p><strong>You finish the task.</strong> Even when it&#8217;s painful. Even when you&#8217;re correcting the AI five times. Even when doing it yourself would&#8217;ve been faster. You finish because the output isn&#8217;t the point - the <em>corrections</em> are the point.</p><p>Every time you say &#8220;No, not that table - use this one,&#8221; that&#8217;s a piece of knowledge. Every time you say &#8220;That&#8217;s the wrong metric definition, here&#8217;s the right one,&#8221; that&#8217;s a lesson. But a lesson only has value if you write it down somewhere the AI can find it next time.</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/ai-patience">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[ABMarks is live!]]></title><description><![CDATA[This day has finally come.]]></description><link>https://datamarks.substack.com/p/abmarks-is-live</link><guid isPermaLink="false">https://datamarks.substack.com/p/abmarks-is-live</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Thu, 07 May 2026 11:35:45 GMT</pubDate><content:encoded><![CDATA[<p>This day has finally come.</p><p>The product I&#8217;ve been working on for quite a while is now live.</p><p>I&#8217;ll share more about how I realized there was a real need for this product. The A/B testing and causal inference domain is unique when it comes to the learning curve.</p><p>The foundations are everywhere: t-tests, z-tests, what A/B testing is, statistics basics.</p><p>But when it comes to truly understanding experimentation and working on real projects, so many questions start to appear &#8212; and most answers are scattered across years of practice, discussions with other practitioners, standalone blog posts, or academic papers that are very hard to find.</p><p>Questions like:</p><ul><li><p>How does randomization actually work in A/B testing?</p></li><li><p>How can we run many experiments in parallel?</p></li><li><p>How do we measure long-term impact?</p></li><li><p>Is there a way to bundle multiple experiments for one feature and get a more precise answer?</p></li><li><p>What actually happens when randomization and analysis units are different?</p></li><li><p>Why do we need geo testing or cluster testing?</p></li></ul><p>And there&#8217;s an endless number of questions like these.</p><p>I couldn&#8217;t find one place where I could learn all of this, practice it, and start applying it confidently.</p><p>Now there is one: <a href="https://abmarks.io/">ABmarks</a>.</p><p>If you go through every lesson and every exercise, you&#8217;ll be ready for experimentation interviews, ready to deliver real insights from experiments, ready to work with complex experimental designs, and &#8212; what I think is most important &#8212; ready to validate assumptions, results, and experiments properly.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>For subscribers of my newsletter, I&#8217;m sharing a discount for the first two weeks:</p><p>LAUNCH2026</p><p>Meanwhile, I&#8217;ve already started working on a fully practical module for the platform &#8212; where people with zero experience can get everything they need to become successful in experimentation, and people with strong exposure can sharpen their skills by working on a variety of real-world designs.</p>]]></content:encoded></item><item><title><![CDATA[AI DS Workflow]]></title><description><![CDATA[If you are a data scientist using AI tools, you might have noticed that sometimes the AI feels like a junior analyst who needs constant hand-holding, and other times it feels like a superpower.]]></description><link>https://datamarks.substack.com/p/ai-ds-workflow</link><guid isPermaLink="false">https://datamarks.substack.com/p/ai-ds-workflow</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Thu, 30 Apr 2026 11:45:54 GMT</pubDate><content:encoded><![CDATA[<p>If you are a data scientist using AI tools, you might have noticed that sometimes the AI feels like a junior analyst who needs constant hand-holding, and other times it feels like a superpower.</p><p>The difference usually comes down to how you interact with it.</p><p>I have tried to accumulate the main hints that can make your &#8220;AI data science&#8221; workflow flow significantly better.</p><h2>The Core Workflow Patterns</h2><h3>1. Your Context File is Everything</h3><p>The single most impactful move you can make is to maintain a small, always-updated context file that the AI can read at the start of a session. You can call it AI_CONTEXT.md or PROJECT.md.</p><p>A good context file ensures the AI knows your role, domain, and constraints. It stops you from repeating yourself each session and ensures outputs match your standards for SQL dialects, naming conventions, tone, and statistical assumptions.</p><h4>How to start:</h4><p>Include your role and domain, your data stack at a high level (warehouse, notebook, BI), your default assumptions (time zones, definitions, experiment conventions), and a few things you repeatedly ask for.</p><p>Example bullets to include:</p><ul><li><p>I am a data scientist working on growth questions.</p></li><li><p>Default SQL dialect: Snowflake. Prefer CTEs; avoid vendor-specific functions unless asked.</p></li><li><p>Lead with the business implication, then show the analysis.</p></li><li><p>Challenge my assumptions before agreeing.</p></li><li><p>When I say &#8216;experiment&#8217;, assume: primary metric, guardrails, segmentation, power/significance checks.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>2. Be Painfully Specific About What You Want</h3><p>Treat the AI like briefing a new teammate. Taking 30 seconds to provide context upfront saves 15 minutes of back-and-forth.</p><p>&#10060; Bad:</p><p>&#8220;Help me with this experiment.&#8221;</p><p>&#9989; Better:</p><p>&#8220;Evaluate experiment X. Focus on segment Y. Show metric Z by country and platform. Include significance, effect sizes, and a sanity-check section (SRM, sample ratio, missingness). Return a 6-bullet executive summary plus a table.&#8221;</p><h3>3. Ask for the Approach Before the Answer</h3><p>This catches wrong assumptions early before the AI writes any code.</p><h5>Example:</h5><p>Before you write any SQL or code, propose: &#8226; The dataset(s) I likely need &#8226; The join keys and grain you will assume &#8226; The edge cases to verify &#8226; 2&#8211;3 validation queries you would run first</p><h3>4. Prefer Multi-Turn Collaboration Over One-Shot Prompts</h3><p>Do not restart from scratch for every question. Build a thread where each turn carries context you do not have to restate. The compounding effect is real.</p><ul><li><p>Step 1  &#8220;Help me identify candidate tables and grains.&#8221;</p></li><li><p>Step 2  &#8220;Now draft the query and sanity checks.&#8221;</p></li><li><p>Step 3  &#8220;Add segment breakdowns and moving averages.&#8221;</p></li><li><p>Step 4  &#8220;Turn results into a readout and caveats.&#8221;</p></li></ul><h3>5. Specify Output Format </h3><p>If you need something pasteable, say so explicitly.</p><p><strong>Return</strong>: </p><ol><li><p>A 5-bullet executive summary </p></li><li><p>A markdown table with: metric, control, test, delta, delta_pct, p_value</p></li><li><p>A &#8220;Checks&#8221; section (SRM, missing data, outliers)</p></li></ol><h3>6. Do Not Fight the Model &#8212; Guide It</h3><p></p>
      <p>
          <a href="/__u/datamarks.substack.com/p/ai-ds-workflow">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What actually means Analytical Rigor? ]]></title><description><![CDATA[As AI continues to automate the mechanical parts of data analysis&#8212;writing SQL, building dashboards, and even running basic models&#8212;a critical question emerges: what is our role as human data scientists?]]></description><link>https://datamarks.substack.com/p/what-actually-means-analytical-rigor</link><guid isPermaLink="false">https://datamarks.substack.com/p/what-actually-means-analytical-rigor</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Fri, 03 Apr 2026 10:11:04 GMT</pubDate><content:encoded><![CDATA[<p>As AI continues to automate the mechanical parts of data analysis&#8212;writing SQL, building dashboards, and even running basic models&#8212;a critical question emerges: what is our role as human data scientists?</p><p>I think an essential part of the answer lies in one core competency: <strong>Analytical Rigor.</strong></p><p>When you check on LinkedIn, 80% of data roles will have this requirement everywhere, but what does it actually mean?</p><h2>What is Analytical Rigor?</h2><p>Let&#8217;s start with what it is not. </p><p>Analytical rigor is not about applying the most complex math possible. It is not about achieving absolute certainty or perfection.</p><p>Instead, it is about ensuring that the entire analytical process is sound, logical, and defensible. It involves several key practices:</p><ul><li><p><strong>Solving the Right Problem:</strong> Finding the right questions and answering them. You can execute a flawless analysis, but if it doesn&#8217;t address the core business challenge, it is useless.</p></li></ul><blockquote><p><strong>Example</strong>: A product manager asks for a "quick data pull" of how many users clicked the new 'Share' button this week. Instead of just running the SQL, a rigorous analyst asks: "What decision will this number drive?" If the goal is to measure feature success, the analyst might suggest looking at the retention of users who shared, rather than just a raw click count.</p></blockquote><p></p><ul><li><p><strong>Applying the Best Methodology:</strong> Using the highest quality measurement and analytic techniques available for the specific problem.</p></li></ul><blockquote><p><strong>Example</strong>: Choosing a causal inference method like Difference-in-Differences when you need to understand the impact of a marketing campaign, rather than just looking at simple pre/post correlations.</p></blockquote><p></p><ul><li><p><strong>Relying on Trustworthy Data</strong>: Verifying that the underlying data is accurate, reliable, and that the processes generating it are sound.</p></li></ul><blockquote><p><strong>Example</strong>: Before analyzing user engagement, checking for duplicate tracking events or missing data from a recent app release.</p></blockquote><p></p><ul><li><p><strong>Validating Results</strong>: Cross-checking your findings. This means looking at the problem from multiple angles to see if the conclusions hold up.</p></li></ul><blockquote><p><strong>Example</strong>: You see a sudden spike in daily active users (DAU) in your internal dashboard. To validate this, you triangulate: do you see a corresponding spike in server load logs? Do third-party analytics tools (like Google Analytics or Mixpanel) show the same trend? If the internal dashboard says traffic doubled but server load remained flat, you likely have a logging bug, not a growth spurt.</p></blockquote><p></p><ul><li><p><strong>Acknowledging Constraints</strong>: Clearly calling out assumptions, caveats, and unknowns.</p></li></ul><blockquote><p><strong>Example</strong>: Stating upfront, &#8220;This analysis assumes that user behavior during the holiday season is representative of the rest of the year, which may not be true.&#8221;</p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Ultimately, your analysis should be reproducible and robust. If someone else (or you, six months from now) runs the same analysis on the same data, they should reach the same conclusions.</p><div class="pullquote"><p>If you want to develop your analytical rigor, especially in experimentation, join the waitlist for my new A/B testing simulator. It covers 90% of real-world A/B testing scenarios, helping you master design, implementation, and measurement. Join the <a href="https://tally.so/r/jaLaEa">waitlist</a> now!</p></div><h2>Why Does It Even Matter?</h2><p>Rigor is what turns a quick data pull into a decision that actually moves the business. Without it, you are just handing someone a number. Therefore, it provides several crucial benefits:</p><ul><li><p><strong>Improving Your Likelihood of Success</strong>: Teams don't have to rely on instinct. They have something concrete to act on &#8212; and that changes the quality of the decision entirely.</p></li></ul>
      <p>
          <a href="/__u/datamarks.substack.com/p/what-actually-means-analytical-rigor">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Evolution of the Data Scientist role]]></title><description><![CDATA[Something remarkable is happening in data science right now.]]></description><link>https://datamarks.substack.com/p/the-evolution-of-the-data-scientist</link><guid isPermaLink="false">https://datamarks.substack.com/p/the-evolution-of-the-data-scientist</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Thu, 19 Mar 2026 11:42:10 GMT</pubDate><content:encoded><![CDATA[<p>Something remarkable is happening in data science right now.</p><p>Not long ago, a data scientist&#8217;s day was consumed by writing SQL queries, tuning models, and running experiments. Today, thanks to AI, a single data scientist can do the work of an entire cross-functional team. They are writing production code, building end-to-end systems, driving business decisions, and solving complex problems without waiting on anyone else.</p><p>This isn&#8217;t just a trend at a few big tech companies; it is happening everywhere. AI tools are democratizing capabilities that traditionally required armies of engineers and product managers.</p><p>For data scientists today, the question is no longer if our jobs will change, but how we will adapt.</p><h2>Why the Old Way is Breaking Down</h2><p>In the past, the data science workflow was linear and slow. A data scientist found a problem, wrote queries, built a model, and then handed it off to an engineer to put into production. Next came designing an experiment, running it, analyzing the results, and deciding whether to scale up or pivot. Every single handoff introduced friction and delays.</p><p>Now, AI tools are eliminating those bottlenecks. A data scientist can use an AI agent to identify an issue, trace it back to the source code, and deploy a fix&#8212;all without waiting for an engineering ticket to be picked up.</p><p>This doesn&#8217;t mean AI is replacing humans. It means AI is amplifying us. When AI can write SQL, generate charts, and draft boilerplate code, the most valuable skills shift from syntax memorization to big-picture thinking and strategic execution. We are moving from merely analyzing data to actively building solutions.</p><p>Based on this shift, I see three distinct archetypes emerging for how data scientists are adopting AI.</p><h2>1. The Systems Data Scientist</h2><p>These are the foundational builders. They ensure that a company&#8217;s data infrastructure is organized, clean, and structured so that AI tools can actually understand and utilize it effectively.</p><p>If an AI agent is fed messy, unstructured data, it will confidently deliver bad answers. The Systems Data Scientist prevents this. They create the rules, semantic layers, and governance structures that allow AI agents to operate quickly and accurately.</p><p>What they do:</p><p>&#8226;Structure and organize data pipelines so AI models can easily ingest and interpret the information.</p><p>&#8226;Design robust testing frameworks to validate that AI tools are generating accurate and reliable outputs.</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/the-evolution-of-the-data-scientist">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AI experimentation for DS]]></title><description><![CDATA[AI is reshaping data science.]]></description><link>https://datamarks.substack.com/p/ai-experimentation-for-ds</link><guid isPermaLink="false">https://datamarks.substack.com/p/ai-experimentation-for-ds</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Thu, 26 Feb 2026 13:12:01 GMT</pubDate><content:encoded><![CDATA[<p>AI is reshaping data science. It&#8217;s not a &#8220;nice to have&#8221; anymore - it&#8217;s quickly becoming part of the job.</p><p>If you want to stay sharp, you need to become an AI experimentalist.</p><p>Over time, I&#8217;ve realized there are 6 habits that will help you succeed with AI in the long run.</p><div><hr></div><h2>1. Schedule 30 minutes for AI practice</h2><p>If you don&#8217;t practice, you won&#8217;t get good at this.</p><p>Block <strong>30 minutes a day</strong>, <strong>1 hour a week</strong>, or whatever cadence works. Think of it like an <strong>AI gym session</strong>: low pressure, high reps.</p><p>The key: don&#8217;t just passively read about AI - <strong>use it</strong>.</p><p>Try something small:</p><ul><li><p>rewrite a query</p></li><li><p>debug an error</p></li><li><p>summarize a notebook</p></li><li><p>generate a chart explanation</p></li><li><p>test a new tool</p></li></ul><p>Once you&#8217;ve built some momentum, experiment with <strong>tool chaining</strong>: multiple AI tools working together for one outcome. For example:</p><ul><li><p>analyze data &#8594; gather context &#8594; draft slides &#8594; refine narrative</p></li></ul><p>Over time, you&#8217;ll build instincts for when AI is helpful - and when it&#8217;s guessing.</p><div><hr></div><h2>2. Use AI to learn, not just to execute</h2><p>AI isn&#8217;t just there to write code. Use it as a tutor.</p><p>Ask it to:</p><ul><li><p><strong>Explain a pipeline.</strong> </p><ul><li><p>&#8220;Read through this pipeline thoroughly and give me a summary of what it does, which data sources it uses, and how the output is structured.&#8221;<br>It almost always beats reading the pipeline code line by line.</p></li></ul></li><li><p><strong>Understand the metric definition.</strong> </p><ul><li><p>&#8220;Trace the data lineage for this signal and tell me where it originally comes from.&#8221;<br>It will follow the joins, upstream tables, and transformations &#8212; and often give you a clear answer.</p></li></ul></li><li><p><strong>Explain some of your teammates works.</strong> </p><ul><li><p>Point it to an existing notebook or query and ask it to walk through the methodology.</p></li></ul></li></ul><p>And don&#8217;t forget the follow-ups.</p><blockquote><p>&#8220;What assumptions are you making?&#8221;</p></blockquote><p>That one question alone can surface hidden problems. </p><div><hr></div><h2>3. Give context, not just instructions</h2><p>I believe this is the most common advice - and at the same time, the most important.</p><p>AI is smart, but it doesn&#8217;t know your business.</p><p>That&#8217;s your biggest lever.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>You wouldn&#8217;t throw a new hire or even an experienced one into a project with zero context and expect amazing output. Same thing here.</p><p>If you ask AI for &#8220;conversion rate,&#8221; it has no idea what that means in your company unless you explain it.</p><p>So give context:</p><ul><li><p>Describe your tables and how they connect</p></li><li><p>Share the business processes behind the data and what you&#8217;re actually working on</p></li><li><p>Define your metrics (numerator, denominator, filters, time grain)</p></li></ul><p>The more context you provide, the better the output. Vague prompts = vague answers.</p><p><strong>Pro tip: use suggestive language over rigid commands.</strong> </p>
      <p>
          <a href="/__u/datamarks.substack.com/p/ai-experimentation-for-ds">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Product Data Science technical assignments in 2026]]></title><description><![CDATA[There are still technical tasks for data scientists that are given to solve offline.]]></description><link>https://datamarks.substack.com/p/product-data-science-technical-assignments</link><guid isPermaLink="false">https://datamarks.substack.com/p/product-data-science-technical-assignments</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Thu, 12 Feb 2026 12:45:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4hfc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4a60cdf-0514-40cd-9602-f11f4adbbff2_1507x383.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There are still technical tasks for data scientists that are given to solve offline. I really thought that nobody is doing it even in 2025, but I was wrong.</p><p>The main difference now is time. You are given a short window &#8212; sometimes 2 hours, sometimes a couple of days if the goal is to prepare a small deck with insights.</p><p>There is no question whether to use AI or not. You should. That&#8217;s real work today. The real question is: what do you ask it? What do you solve yourself? What trade-offs do you make under time pressure?</p><p>That&#8217;s what I want to share here.</p><p>These are real examples that my mentees and I received over the last 6 months, and how I would approach them in general along with some hints.</p><h1>Part 1: SQL Section</h1><p>You are usually given either:</p><ul><li><p>A data structure and a file</p></li><li><p>Or access to a database (dummy schema, not production-level)</p></li></ul><p>The warm-up task is simple. Something like: aggregate by <code>user_id</code>.</p><p>Real example:</p><pre><code><code>SELECT 
      user_id, 
      COUNT(*) as total_usage
FROM activity
GROUP BY user_id
ORDER BY total_usage DESC
LIMIT 10</code></code></pre><p>Then it gets deeper. You&#8217;ll need ranking functions.</p><p>Example question:</p><blockquote><p>Write a SQL query that ranks users by their total number of log-ins within each acquisition channel.</p></blockquote><p>Or with extra complexity:</p><blockquote><p>In addition return only the top 3 users per channel.</p></blockquote><p>My answer </p><pre><code><code>with total_data as (
select u.user_id,
       acquisition_channel,
       sum(log_in) total_log_ins
from users u left join
user_activity_daily uad on u.user_id = uad.user_id
group by 1,2),
total_data_rn as (
select *, rank() over(partition by acquisition_channel order by total_log_ins desc) as rn,
100.0 * PERCENT_RANK() OVER (
    PARTITION BY acquisition_channel
    ORDER BY total_log_ins
  ) pct_rank

 from total_data)
select * from total_data_rn where rn &lt;= 3</code></code></pre><p>The last SQL task is often a gaps-and-islands problem.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><blockquote><p>Find each user's longest streak of consecutive days where they purchase at least once.</p></blockquote><p>How I approached it:</p><pre><code><code>  WITH active_days AS (
  SELECT
    user_id,
    activity_date AS dt,
    row_number() OVER (PARTITION BY user_id ORDER BY activity_date) AS rnk
  FROM user_activity_daily
  WHERE purchase &gt;= 1
  order by 1,2
) , grps as (
  select *, date(dt, '-' || rnk || ' days') AS grp from active_days ) ,
streaks as (
select user_id,
       min(dt) as streak_start_date,
       max(dt) as streak_end_date,
       count(1) as total_days 
from grps 
       group by user_id,grp
  ) , top_streaks as (
  select *,
         row_number() over(partition by user_id order by total_days desc) as rn 
  from streaks)
select user_id,
       total_days as longest_streak,
       streak_start_date,
       streak_end_date 
from top_streaks where rn = 1  order by total_days desc limit 15</code></code></pre><p>To summarize, to pass the SQL section you just need:</p><ul><li><p>Understand JOINs and be comfortable with them</p></li><li><p>Know when and how to use window functions</p></li><li><p>Use CTEs</p></li></ul><p>That&#8217;s it. Simple? Yes.</p><p>But solving it fast under time pressure takes practice. Sometimes even with AI.</p><h2>Part 2: Anomaly Detection</h2><p>Very often you&#8217;ll get something like:</p><blockquote><p>As a data scientist, you'll need to monitor metrics and identify when something unexpected happens. Find and detect the anomaly in data</p></blockquote><p>Sometimes they tell you what to focus on, something like &#8216;Check if there are any anomalies in revenue data&#8217; or it can be in general like on the example above.</p><p>In both of these cases your output should include:</p><ul><li><p>What metric is anomalous</p></li><li><p>When it happened</p></li><li><p>How big the change is</p></li><li><p>How it relates to other metrics</p></li><li><p>Visualization</p></li><li><p>Hypothesis about possible causes</p></li></ul><h3>How to start ? </h3>
      <p>
          <a href="/__u/datamarks.substack.com/p/product-data-science-technical-assignments">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What have been helping me along my career]]></title><description><![CDATA[There are endless pieces of advice on how to stay efficient &#8212; morning routines, productivity hacks, new frameworks every week.]]></description><link>https://datamarks.substack.com/p/what-have-been-helping-me-along-my</link><guid isPermaLink="false">https://datamarks.substack.com/p/what-have-been-helping-me-along-my</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 28 Jan 2026 12:49:27 GMT</pubDate><content:encoded><![CDATA[<p>There are endless pieces of advice on how to stay efficient &#8212; morning routines, productivity hacks, new frameworks every week.</p><p>Most of them are either vague, too general, or just didn&#8217;t work for me.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Over time, I found a few simple habits that actually do. Nothing fancy, no magic tricks &#8212; just small things that compound.</p><p>Here&#8217;s what helped me stay sharp and useful.</p><h2>Keep a brag doc</h2><p>I&#8217;ve shared this already on LinkedIn, but it&#8217;s still the most powerful hint that helps me along the way.</p><p>Make a private Google Doc.</p><p>Drop in:</p><ul><li><p>wins</p></li><li><p>feedback</p></li><li><p>projects you shipped</p></li><li><p>moments when you helped someone out</p></li><li><p>interesting challenges</p></li></ul><p>You will forget this stuff faster than you expect. I personally forget it in a week max.</p><p>Then performance review season comes, or impostor syndrome kicks in &#8212; and suddenly your brain is empty.</p><p>That doc becomes pure gold.</p><p>Just try it for at least 3 months and you&#8217;ll see the difference.</p><div><hr></div><h2>If you actually want to remember things &#8212; teach them</h2><p>That&#8217;s my favourite.</p><p>Who actually remembers how algorithm X works if you don&#8217;t prepare for interviews?<br>Not me, at least.</p><p>Start teaching.</p><p>Nothing locks knowledge in like explaining it to another human, especially right after you learned it yourself.</p><p>It works for everything:</p><ul><li><p>fundamentals and math</p></li><li><p>ML concepts</p></li><li><p>interview prep</p></li><li><p>honestly, almost anything</p></li></ul><p>Teaching forces you to organize your thoughts. It shows you where you&#8217;re fuzzy. Stuff moves from short-term memory into long-term storage.</p><div><hr></div><h2>Run demos. Even if nobody asked for them.</h2><p>If your team doesn&#8217;t do demos or knowledge sharing &#8212; just start.</p><p>Keep it casual. Internal only. Let people show what they&#8217;re working on. Most teams have more projects than people, so this quickly creates shared context.</p><p>I like having at least one short presentation per project:</p><ul><li><p>first inside the team</p></li><li><p>later for a wider audience</p></li><li><p>optionally summarized somewhere internal</p></li></ul><p>A few nice side effects:</p><ul><li><p>everyone knows who&#8217;s doing what</p></li><li><p>random connections pop up (&#8220;oh, we tried something similar&#8221;)</p></li><li><p>onboarding gets way easier (new folks can skim slides instead of pinging 10 people)</p></li></ul><p>Small habit, big payoff.</p><p>And like every habit, it&#8217;s hard to maintain &#8212; but it&#8217;s really powerful.</p><div><hr></div><h2>Build</h2><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/p/what-have-been-helping-me-along-my?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Data Marks! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/p/what-have-been-helping-me-along-my?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/datamarks.substack.com/p/what-have-been-helping-me-along-my?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p>Most data people get dragged into endless meetings, docs, exploratory analysis, etc.</p><p>If you don&#8217;t build, your skills become rusty. Maybe in 2 months, maybe in 5 months, maybe in a year &#8212; but the outcome is inevitable.</p><p>You can build a small system for manual file uploads, or new data pipelines to automate part of your job. Or a self-service dashboard to analyze your meetings.</p><p>I personally build something once a month as a rule. It may be very simple, but it helps me keep my skills clean.</p><p>And now with vibe coding, it&#8217;s easier than it&#8217;s ever been.</p><div><hr></div><h2>Random Coffee</h2><p>You come to a new place. Your circle is limited: you have your team, your XFNs, your stakeholders &#8212; and that&#8217;s mostly it.</p><p>But what happens if you randomly reach out to a Data Engineer in a completely different team? Or a Data Scientist from a totally different company?</p><p>Most interesting connections don&#8217;t come from formal networking events or work projects. They come from random space.</p><p>A simple loop:</p><ul><li><p>find people doing work you&#8217;re curious about (coworkers, LinkedIn, Twitter, meetups, Substack, GitHub &#8212; anywhere)</p></li><li><p>send a short, friendly message</p></li><li><p>find a bit of common ground</p></li><li><p>suggest a quick coffee or virtual chat</p></li></ul><p>No pitch. No agenda. Just curiosity.</p><p>Something like:<br><em>&#8220;Hey &#8212; I saw your work on X. I&#8217;m also into Y. Want to grab a 20-minute coffee sometime?&#8221;</em></p><p>You&#8217;ll be surprised how often people say yes.</p><p>Over time, this compounds:</p><ul><li><p>you learn what&#8217;s actually happening across teams and companies (not just what&#8217;s in Slack or blogs)</p></li><li><p>you start seeing patterns</p></li><li><p>you connect people who should probably know each other</p></li><li><p>opportunities surface earlier</p></li><li><p>your name starts popping up in places you&#8217;re not even in</p></li></ul><p>I&#8217;ve met a lot of great and interesting people this way.</p><div><hr></div><p>Don&#8217;t limit yourself to just my approach. Do your own experiments.</p><p>Something that works for me may not work for you.</p><p>Explore.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Best Practices for the experiments]]></title><description><![CDATA[We all use experiments for data-driven decision-making, whether it&#8217;s product changes, marketing tweaks, or research.]]></description><link>https://datamarks.substack.com/p/best-practices-for-the-experiments</link><guid isPermaLink="false">https://datamarks.substack.com/p/best-practices-for-the-experiments</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 14 Jan 2026 12:38:28 GMT</pubDate><content:encoded><![CDATA[<p>We all use experiments for data-driven decision-making, whether it&#8217;s product changes, marketing tweaks, or research. But even experienced people fall into subtle statistical (and not only) traps that quietly ruin their results.</p><p>Here I will share some of the best practices to ensure your experiments yield reliable, actionable insights.</p><div><hr></div><h2><strong>1. Always Focus on a Measurable Hypothesis</strong></h2><p>Before you start, clearly define what you want to test. A strong experiment begins with a specific, testable hypothesis that can drive future improvements. Your train of thoughts should go from the general question to the specific hypothesis. For example:</p><ul><li><p><strong>General question:</strong> &#8220;Does changing the price affect sales?&#8221;</p></li><li><p><strong>Refined question:</strong> &#8220;Does lowering the monthly price from $29 to $25 increase conversions?&#8221;</p></li><li><p><strong>Specific hypothesis:</strong> &#8220;Lowering the price to $25 will increase conversions by at least 10% because the users are susceptible to the price.&#8221;</p></li></ul><p></p><p>A good hypothesis answers three things:</p><ul><li><p><strong>What are you changing?</strong></p></li><li><p><strong>What metric should move?</strong></p></li><li><p><strong>How much do you expect it to change?</strong></p></li></ul><p>The best hypotheses answer an additional one: </p><ul><li><p>Why should it work? </p></li></ul><p>If you can&#8217;t answer those, you&#8217;re not running an experiment -  you&#8217;re just vibes-testing.</p><div><hr></div><h2><strong>2. Don&#8217;t Rerun Experiments Until You &#8220;Win&#8221;</strong></h2><p>Running the same experiment multiple times until you get a statistically significant result is a form of p-hacking.</p><p>If your experiment failed:</p><ul><li><p>Figure out <em>why</em></p></li><li><p>Improve the hypothesis</p></li><li><p>Fix the design</p></li><li><p>Then try again</p></li></ul><p>Fishing for a better p-value doesn&#8217;t make the result real.</p><div><hr></div><h2><strong>3. Don&#8217;t Cherry-Pick Dates or Stop Experiments Early</strong></h2><p>Selecting dates or stopping experiments based on when results look favorable introduces bias and reduces reproducibility.</p><p><strong>Best practices:</strong></p><ul><li><p>Define objective criteria for experiment duration and data inclusion before starting.</p></li><li><p>Set clear pre-defined stopping rules (especially for detecting <em>degradation</em>)</p></li><li><p>Avoid peeking at results unless using really principled adaptive methods.</p></li></ul><div><hr></div><h2><strong>4. Don&#8217;t Equate &#8220;Not Statistically Significant&#8221; with &#8220;Neutral&#8221;</strong></h2><p>A &#8220;not statistically significant&#8221; result does not mean there is no effect, it only means the experiment didn&#8217;t detect an effect with high confidence.</p><p><strong>Why this matters:</strong></p><ul><li><p>The absence of evidence is not evidence of absence.</p></li><li><p>Small samples and noisy data can hide real changes</p></li><li><p>Tiny negative undetected effects can accumulate over time. Run enough of them, and you&#8217;ll see real degradation.</p><ul><li><p>Simple rule here helps. Check that the lower boundary of your confidence interval is above your threshold for regression. For example, to be 95% sure you haven&#8217;t lost more than 0.5%, the lower bound should be above -0.5%.</p></li></ul></li></ul><div><hr></div><h2><strong>5. </strong>Don&#8217;t just &#8220;explore&#8221; the data</h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>When it comes to the measurement of the experiment, don&#8217;t just &#8220;explore&#8221; the data.</p><p>Come prepared with a question you know can be answered.</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/best-practices-for-the-experiments">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Happy 2026 ]]></title><description><![CDATA[This year changed a lot for me.]]></description><link>https://datamarks.substack.com/p/happy-2026</link><guid isPermaLink="false">https://datamarks.substack.com/p/happy-2026</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Fri, 02 Jan 2026 12:47:02 GMT</pubDate><content:encoded><![CDATA[<p>This year changed a lot for me. I started sharing my ideas and thoughts on LinkedIn, and I created a newsletter. We&#8217;ve already grown to more than 3,000 people. I never thought so many people would be interested in what I think and write. I&#8217;m very grateful to all of you.</p><p>I learned three main lessons this year.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><ol><li><p><strong>AI is useful, cool and scary at the same time</strong></p></li></ol><p>I realized that without AI it no longer makes sense to work the old way, especially in tech. Without it, you are much slower. It&#8217;s not just about asking a chatbot to write code. It&#8217;s about planning architecture, writing clear specs, and setting up your process so AI fits into it easily.</p><p>My advice: keep using AI at work and try a lot more &#8220;vibe-coding.&#8221; It&#8217;s fun, and it helps you better understand the technologies that shape our world.</p><ol start="2"><li><p><strong>Don&#8217;t be afraid to try new things</strong></p></li></ol><p>You never know what you will find there.</p><p>I hadn&#8217;t played football for years. I started again, went out more, met many great people, got into better shape, and began enjoying life much more.</p><p>I also returned to poker this year. I earned around $5,000, met some great people who work in investing, and we may even work together on projects. And again, I enjoy my life more.</p><p>I started writing too, even though I&#8217;m shy and I worry about what people think about me and my ideas. Now many people read me: 37k LinkedIn followers and 3k newsletter subscribers &#8212; and I enjoy my life more.</p><p>The main point is simple: when you try something new, life becomes more interesting.</p><p>Don&#8217;t be afraid to start.</p><ol start="3"><li><p><strong>Consistency is a king.</strong></p></li></ol><p>If you are very talented but not consistent, you won&#8217;t be successful. But if you&#8217;re fairly average and consistent, you can achieve a lot, but don&#8217;t expect results tomorrow or next week &#8212; they will come with time.</p><p>Remember, 95% of people lack it. If you stay consistent in 2026, you&#8217;re already ahead of almost all of them.</p><div><hr></div><p>I wish all of you to try something new, enjoy your life, and let data help you in it. Thank you all for supporting me and a lot more interesting and useful content happening in 2026!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The right mindset to get a data scientist interview]]></title><description><![CDATA[If you&#8217;re preparing for data science interviews and feeling frustrated, exhausted, or questioning your self-worth after a rejection - this post is for you.]]></description><link>https://datamarks.substack.com/p/the-right-mindset-to-get-a-data-scientist</link><guid isPermaLink="false">https://datamarks.substack.com/p/the-right-mindset-to-get-a-data-scientist</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Fri, 12 Dec 2025 12:31:29 GMT</pubDate><content:encoded><![CDATA[<p>If you&#8217;re preparing for data science interviews and feeling frustrated, exhausted, or questioning your self-worth after a rejection - this post is for you.</p><p>One of the biggest mindset shifts that helped me (and many people I&#8217;ve worked with) is this:</p><p><strong>Getting a data scientist job is not a judgment of you as a person.<br>It&#8217;s a game.</strong></p><p>And like any serious game, it has rules, strategies, losses, and progression systems. Once you start seeing it this way, everything becomes clearer - and a lot less emotionally draining.</p><p>Let&#8217;s break it down.</p><div><hr></div><h2>1. It&#8217;s a Game. Learn the rules and the optimal strategy</h2><p>Every game has mechanics. Interviews are no different.</p><ul><li><p>There are <strong>levels</strong> (screening, technical, case study, behavioral).</p></li><li><p>There are <strong>win conditions</strong> (clear communication, correct assumptions, trade-offs).</p></li><li><p>There are <strong>constraints</strong> (time, ambiguity, incomplete information).</p></li><li><p>There is a <strong>meta</strong> (what companies currently value, what interviewers expect).</p></li></ul><p>Many candidates fail not because they&#8217;re &#8220;bad at data science,&#8221; but because they <strong>don&#8217;t know the rules of the game they&#8217;re playing</strong>.</p><p>You wouldn&#8217;t start a chess tournament without knowing how pieces move.<br>Yet many people walk into interviews without understanding:</p><ul><li><p>What kind of signals interviewers are looking for</p></li><li><p>How answers are evaluated</p></li><li><p>What &#8220;good enough&#8221; actually means</p></li></ul><p>Treat preparation as <strong>strategy design</strong>, not brute-force grinding.</p><div><hr></div><h2>2. You will lose, a lot, and that&#8217;s normal</h2><p>In hard games, losing is part of progression.</p><p>You&#8217;ll fail interviews.<br>You&#8217;ll get rejected after &#8220;great conversations.&#8221;<br>You&#8217;ll be ghosted.<br>You&#8217;ll feel like you did everything right and still lost.</p><p>That doesn&#8217;t mean you&#8217;re bad.<br>It means you&#8217;re playing a <strong>high-difficulty game</strong>.</p><p>The key question after every loss is not:</p><blockquote><p>&#8220;What&#8217;s wrong with me?&#8221;</p></blockquote><p>But:</p><blockquote><p>&#8220;What did this run teach me?&#8221;</p></blockquote><ul><li><p>Did I freeze under pressure?</p></li><li><p>Did I struggle explaining my thinking?</p></li><li><p>Was my SQL not optimal?</p></li><li><p>Did I misunderstand the business context?</p></li></ul><p>If you extract <strong>one concrete improvement</strong> from each loss, you are leveling up, even if it doesn&#8217;t feel like it.</p><div><hr></div><h2>3. Interviewers Are Not Final Bosses (They&#8217;re NPCs With a Job)</h2><p>It&#8217;s easy to see interviewers as enemies standing between you and the offer.</p><p>They&#8217;re not.</p><p>They&#8217;re usually:</p><ul><li><p>Tired</p></li><li><p>Busy</p></li><li><p>Slightly stressed</p></li><li><p>Trying to assess you fairly with limited time</p></li></ul><p>Most interviewers <strong>want you to do well</strong>. Your success makes their job easier.</p><p>They are not there to destroy you, they&#8217;re there to understand:</p><ul><li><p>How you think</p></li><li><p>How you communicate</p></li><li><p>How you approach uncertainty</p></li></ul><p>When you remember this, nerves don&#8217;t disappear, but they become manageable.</p><p>You&#8217;re not being attacked.<br>You&#8217;re being observed.</p><div><hr></div><h2>4. Show how you think, not just the final answer</h2><p>Many candidates treat interviews like speedruns:</p><blockquote><p>&#8220;Solve the task. Give the correct answer. Done.&#8221;</p></blockquote><p>That&#8217;s a mistake.</p><p>In real data science:</p><ul><li><p>Problems are ill-defined</p></li><li><p>Goals change</p></li><li><p>There is rarely one correct answer</p></li></ul><p>Interviewers care deeply about:</p><ul><li><p>What assumptions you make</p></li><li><p>Why you choose one approach over another</p></li><li><p>What trade-offs you consider</p></li><li><p>How you react when information is missing</p></li></ul><p>Sometimes <strong>there is no &#8220;right&#8221; answer</strong>, only a well-reasoned one.</p><p>Think out loud.<br>Let them see your decision-making process.</p><p>That&#8217;s often the real quest.</p><div><hr></div><h2>5. Know your strengths and weaknesses before  the &#8220;Raid&#8221;</h2><p>Before you charge into interviews, you need self-awareness.</p><p>Ask yourself:</p><ul><li><p>Am I strong in SQL but weak in statistics?</p></li><li><p>Am I good technically but struggle with business framing?</p></li><li><p>Do I panic under time pressure?</p></li><li><p>Do I over-focus on perfect solutions?</p></li></ul><p>One of the biggest mistakes I see:</p><blockquote><p>Grinding LeetCode endlessly because it <em>feels</em> productive</p></blockquote><p>Solving 100 problems in a week won&#8217;t help if:</p><ul><li><p>Your real weakness is communication</p></li><li><p>Or framing ambiguous questions</p></li><li><p>Or explaining trade-offs</p></li></ul><p>Preparation should be <strong>targeted</strong>, not random.</p><p>Good players don&#8217;t just grind, they train intentionally.</p><div><hr></div><h2>6. Don&#8217;t Play on &#8220;Hard Mode&#8221; When You Don&#8217;t Have To</h2><p>Not all interviews are equally difficult.<br>Not all companies expect the same depth.<br>Not all roles test the same skills.</p><p>Yet many candidates:</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/the-right-mindset-to-get-a-data-scientist">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Interpretation does matter]]></title><description><![CDATA[Interpreting experiment results, whether you&#8217;re building a product, running a marketing campaign, or doing research, can be tricky.]]></description><link>https://datamarks.substack.com/p/interpretation-does-matter</link><guid isPermaLink="false">https://datamarks.substack.com/p/interpretation-does-matter</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 12 Nov 2025 12:41:42 GMT</pubDate><content:encoded><![CDATA[<p>Interpreting experiment results, whether you&#8217;re building a product, running a marketing campaign, or doing research, can be tricky. It&#8217;s easy to fall into common traps that mess up your conclusions or waste your time. Let&#8217;s talk about some of the biggest mistakes people make, why they matter, and how you can avoid them.</p><div><hr></div><h2><strong>1. Statistical Significance: What It Really Means</strong></h2><p>A lot of folks think &#8220;not statistically significant&#8221; means &#8220;no effect.&#8221; That&#8217;s not true! If your confidence interval includes zero, it just means you can&#8217;t be sure there&#8217;s an effect. But you also can&#8217;t be sure there isn&#8217;t one: there could be a big effect hiding in there.</p><ul><li><p><strong>Why this matters:</strong> If your confidence intervals are wide, you might be missing large effects simply because your data isn&#8217;t precise enough.</p></li><li><p><strong>What to do:</strong> Don&#8217;t just look for whether zero is in the interval. Consider the range of possible effects. If you want to be sure you&#8217;re not causing harm, check that the lower bound of your confidence interval is above any value you&#8217;d consider risky.</p></li></ul><div><hr></div><h2><strong>2. Small Experiments = Big Problems</strong></h2><p>Running an experiment with too few participants is like trying to predict the weather by looking out the window for five minutes. Small sample sizes lead to noisy results and wide confidence intervals, making it hard to detect real changes.</p><ul><li><p><strong>Why this matters:</strong> You could miss something big, or get fooled by random noise.</p></li><li><p><strong>What to do:</strong> Estimate the size of the effect you care about and calculate how many participants you need to reliably detect it (this is called a &#8220;power calculation&#8221;). Don&#8217;t rely on default sample sizes, tailor your experiment to your specific needs.</p></li></ul><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>3. Health Check Failures: Red Flags</strong></h2><p>Health checks are designed to catch problems in your experiment setup, like unbalanced groups or unexpected data issues. If a health check fails, it&#8217;s a warning sign.</p><ul><li><p><strong>Why this matters:</strong> Ignoring these warnings can mess up your results.</p></li><li><p><strong>What to do:</strong> Investigate and fix the underlying issue before drawing conclusions. Only in rare, well-understood cases should you proceed despite a failed health check.</p></li></ul><div><hr></div><h2><strong>4. Don&#8217;t Dismiss Results Because of Pre-Experiment Differences</strong></h2><p>Sometimes, the groups in your experiment differ even before the experiment starts. This can happen by chance, especially with random assignment.</p><ul><li><p><strong>Why this matters:</strong> It&#8217;s tempting to ignore results if you see these differences, but that&#8217;s not necessary. Statistical methods can adjust for these imbalances.</p></li><li><p><strong>What to do:</strong> Use methods that account for pre-experiment differences, and remember that confidence intervals already factor in the possibility of random imbalances.</p></li></ul><div><hr></div><h2><strong>5. Beware of P-Hacking</strong></h2><p>P-hacking happens when you keep searching through different metrics, time periods, or subgroups until you find something statistically significant. This practice increases the risk of finding &#8220;false positives&#8221;, results that look real but are actually just due to chance.</p><ul><li><p><strong>Why this matters:</strong> The more you look, the more likely you are to find something that isn&#8217;t actually there.</p></li><li><p><strong>What to do:</strong> Decide what you&#8217;re going to measure before you start, and stick to it. If you do explore, treat those findings as ideas for future experiments, not as proof.</p></li></ul><div><hr></div><h2><strong>6. A/A Tests and False Positives</strong></h2><p>An A/A test is when you split your sample into two groups but don&#8217;t apply any treatment. Even in these cases, you might see some statistically significant results just by chance.</p><ul><li><p><strong>Why this matters:</strong> This is a normal part of statistical testing and doesn&#8217;t mean your system is broken.</p></li><li><p><strong>What to do:</strong> Expect a small percentage of false positives, especially if you&#8217;re looking at many metrics.</p></li></ul><h2><strong>7. Filtering on Post-Experiment Variables Can Bias Results</strong></h2>
      <p>
          <a href="/__u/datamarks.substack.com/p/interpretation-does-matter">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How Not to Turn Your 1:1 into a 0:0]]></title><description><![CDATA[1:1s can be the most rewarding meetings - but they can also be wasted time.]]></description><link>https://datamarks.substack.com/p/how-not-to-turn-your-11-into-a-00</link><guid isPermaLink="false">https://datamarks.substack.com/p/how-not-to-turn-your-11-into-a-00</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Fri, 17 Oct 2025 11:35:13 GMT</pubDate><content:encoded><![CDATA[<p><strong>1:1s can be the most rewarding meetings - but they can also be wasted time.</strong><br>It&#8217;s on you to make them useful.</p><p>Whether you&#8217;re new or experienced in your role, your 1:1s with your manager can make a big difference in your growth and day-to-day happiness.</p><p>Below are practical tips on how to run 1:1s well so you get more value from them.</p><div><hr></div><h2><strong>1. Own Your 1:1: It&#8217;s Your Time</strong></h2><ul><li><p><strong>Remember:</strong> The 1:1 is <em>for you</em>. Use it to bring up what matters most to <em>you</em>.</p></li><li><p><strong>Prepare an agenda:</strong> Keep a running doc where you jot down topics as they come up. This could be anything from technical blockers, career questions, to feedback on team processes.</p></li><li><p><strong>Prioritize: </strong>When too many topics accumulate, choose the ones that matter most. It&#8217;s okay to ask to extend the meeting or schedule follow-up time.</p></li></ul><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>2. Go Beyond Status Updates</strong></h2><ul><li><p><strong>Avoid repeating what&#8217;s already covered in team meetings.</strong> Use this time to dive deeper: data strategy, experiment design, metrics, assumptions.</p></li><li><p><strong>Talk about challenges: </strong>Bring up where you&#8217;re stuck (modeling issues, ambiguous results, debugging) and ask for guidance.</p></li><li><p><strong>Share wins and learnings:</strong> Celebrate what&#8217;s working, and reflect on what didn&#8217;t.</p></li></ul><div><hr></div><h2><strong>3. Track Follow-Ups</strong></h2><ul><li><p><strong>Write down action items:</strong> After each 1:1, note down what you and your manager agreed to do. Tag who owns each follow-up.</p></li><li><p><strong>Check back:</strong> Review these notes before your next 1:1 to keep things moving.</p></li></ul><div><hr></div><h2><strong>4. Ask for Feedback and Support</strong></h2><ul><li><p><strong>Proactively seek feedback:</strong> Don&#8217;t wait for review cycles. Ask how you can improve your analyses, presentations, or impact.</p></li><li><p><strong>Clarify expectations:</strong> If you&#8217;re unsure about priorities or what &#8220;success&#8221; looks like, ask directly.</p></li><li><p><strong>Request resources:</strong> Need mentorship on causal inference? Want to learn a new tool? Ask for support or connections.</p></li></ul><div><hr></div><h2><strong>5. Talk About Career Growth</strong></h2><ul><li><p><strong>Discuss your goals:</strong> Share your interests&#8212;whether it&#8217;s leading projects, deepening technical skills, or exploring new domains.</p></li><li><p><strong>Ask for opportunities:</strong> Let your manager know what kinds of projects or skills you want to develop.</p></li></ul><div><hr></div><h2><strong>6. Lead with Vulnerability</strong></h2><ul><li><p><strong>Model openness:</strong> Be open about the areas you struggle with, your uncertainties, or what feedback you fear.</p></li></ul><h2>7. Use Smart Cadence &amp; Format (bonus one)</h2>
      <p>
          <a href="/__u/datamarks.substack.com/p/how-not-to-turn-your-11-into-a-00">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[When intuition wins over A/B test?]]></title><description><![CDATA[When making choices that affect products, policies, or strategies, it&#8217;s important to understand not just what the data says, but how much trust we can place in it.]]></description><link>https://datamarks.substack.com/p/when-intuition-wins-over-ab-test</link><guid isPermaLink="false">https://datamarks.substack.com/p/when-intuition-wins-over-ab-test</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 01 Oct 2025 11:35:12 GMT</pubDate><content:encoded><![CDATA[<p>When making choices that affect products, policies, or strategies, it&#8217;s important to understand not just <em>what</em> the data says, but <em>how much trust</em> we can place in it. Not all evidence is created equal&#8212;some approaches offer more reliable insights about what will happen if we take a particular action. This guide outlines a spectrum of evidence types, from the least to the most trustworthy, and offers practical questions to help you assess each one.</p><div><hr></div><h3><strong>1. Intuition and Experience</strong></h3><ul><li><p><strong>What it is:</strong> Relying on personal judgment, gut feelings, or past experiences to make a decision. The other name for that is HiPPO (highest-paid person&#8217;s opinion)</p></li><li><p><strong>When it&#8217;s used:</strong> Often the default when time is short or data is unavailable.</p></li><li><p><strong>Risks:</strong> Highly subjective; may be and is influenced by biases or overconfidence. Not suitable for high-impact or complex decisions.</p></li><li><p><strong>Consider:</strong> Is the potential downside of being wrong significant? If so, go up one step</p></li></ul><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong>2. Correlations</strong></h3><ul><li><p><strong>What it is:</strong> Noticing that two things tend to happen together (e.g., users who see more notifications also engage more), therefore correlated.</p></li><li><p><strong>When it&#8217;s used:</strong> Early exploration or when data is limited.</p></li><li><p><strong>Risks:</strong> Associations can be misleading&#8212;just because two things move together doesn&#8217;t mean one causes the other. Hidden factors may be at play. <a href="https://www.tylervigen.com/spurious-correlations?utm_source=chatgpt.com">Here </a>you can find all the weird correlation examples you can find and imagine now, you will make the business decisions based on that!</p></li><li><p><strong>Consider:</strong> What other explanations could there be for this pattern? Could a third factor be influencing both?</p></li></ul><div><hr></div><h3><strong>3. Statistical Adjustments</strong></h3><ul><li><p><strong>What it is:</strong> Using statistical models to account for differences between groups (e.g., regression, matching, classification).</p></li><li><p><strong>When it&#8217;s used:</strong> When you have data on relevant characteristics and want to estimate the effect of an action.</p></li><li><p><strong>Risks:</strong> Only as good as the variables included; unmeasured factors can still bias results. A lot of quality checks and diagnostics should be made for this kind of design.</p></li><li><p><strong>Consider:</strong> Have you included all important variables? What might you be missing? The main question here is how could two users with the same features get different treatment statuses and why?</p></li></ul><div><hr></div><h3><strong>4. Quasi-experiments</strong></h3><ul><li><p><strong>What it is:</strong> Taking advantage of real-world changes that mimic random assignment (sometimes called &#8220;natural experiments&#8221;). These changes allow you to compare groups that are similar except for the action or change you&#8217;re interested in, helping to estimate cause-and-effect relationships.</p></li><li><p><strong>Examples of methods:</strong></p><ul><li><p>Differences-in-differences. </p></li><li><p>Synthetic control methods.</p></li><li><p>Regression discontinuity </p></li><li><p>Instrumental variables</p></li></ul></li><li><p><strong>When it&#8217;s used:</strong> When true randomization isn&#8217;t possible, but some external process creates comparable groups.</p></li><li><p><strong>Risks:</strong> These methods depend on specific assumptions (e.g., that nothing else changed at the same time or parallel trends assumption).</p></li><li><p><strong>Consider:</strong> What must be true for this approach to give a fair answer? Are those conditions likely met?</p></li></ul><div><hr></div><h3><strong>5. Meta-analysis of quasi-experiments</strong></h3>
      <p>
          <a href="/__u/datamarks.substack.com/p/when-intuition-wins-over-ab-test">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What actually speed up your A/B Test]]></title><description><![CDATA[We all want to speed up our A/B tests and run them as fast as possible.]]></description><link>https://datamarks.substack.com/p/what-actually-speed-up-your-ab-test</link><guid isPermaLink="false">https://datamarks.substack.com/p/what-actually-speed-up-your-ab-test</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 17 Sep 2025 11:58:31 GMT</pubDate><content:encoded><![CDATA[<p>We all want to speed up our A/B tests and run them as fast as possible. That means we&#8217;ll get insights faster, scale the solutions that drive impact quicker, and ultimately create business impact sooner.</p><p>But how exactly can we do it?</p><p>Let&#8217;s remind ourselves of the formula for sample size:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9Zyg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_424, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_webp, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 424w, /__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_848, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_webp, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 848w, /__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_1272, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_webp, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_1456, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_webp, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9Zyg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png" width="182" height="59" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:59,&quot;width&quot;:182,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3792,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://datamarks.substack.com/i/173534783?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_424, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_auto, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 424w, /__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_848, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_auto, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 848w, /__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_1272, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_auto, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9Zyg!, /__u/datamarks.substack.com/w_1456, /__u/datamarks.substack.com/c_limit, /__u/datamarks.substack.com/f_auto, /__u/datamarks.substack.com/q_auto:good, /__u/datamarks.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb58b905e-8e8e-4f88-a9d7-321ae6e861c4_182x59.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p>There are only three levers to reduce the sample size: </p><ul><li><p>Change the variance</p></li><li><p>Change the error level</p></li><li><p>Change the Minimum Detectable Effect (MDE)</p></li></ul><p>Let&#8217;s dive into the methods that actually move the needle.</p><p></p><h2>1. Change the Minimum Detectable Effect (MDE) directly</h2><p>The smaller the effect you want to detect, the larger the sample you need.</p><ul><li><p>If you want to catch <strong>tiny changes</strong>, you&#8217;ll wait months.</p></li><li><p>If you&#8217;re okay only detecting <strong>big shifts</strong>, you&#8217;ll finish much faster.</p></li></ul><p>&#128073; Example: detecting a 0.1% uplift in conversion might need millions of users. But if your MDE is 1%, you&#8217;ll wrap things up 100x sooner.</p><p>For <strong>startups</strong>, you don&#8217;t need to detect a 0.001% effect&#8212;your growth should be steep, and there are plenty of low-hanging fruits. Speed matters more than ultra-precision.</p><p>For <strong>mature companies</strong>, it&#8217;s the opposite: you&#8217;re looking for smaller uplifts because exponential growth is no longer possible, and small improvements compound significantly.</p><p><strong>Rule of thumb:</strong> don&#8217;t aim for microscopic wins unless they really matter to the business.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>2. Increase Your Error Level</h2><p>Statisticians hate this one, but it works.</p><ul><li><p>Most teams default to <strong>95% significance (&#945; = 0.05)</strong>.</p></li><li><p>But what if you&#8217;re okay with <strong>90% (&#945; = 0.10)</strong>?</p></li></ul><p>&#128073; Use this when:</p><ul><li><p>The decision isn&#8217;t mission-critical.</p></li><li><p>You&#8217;re running early explorations.</p></li><li><p>You can re-run later for more precision.</p></li></ul><p>Remember, it&#8217;s a trade-off: more speed, slightly higher risk of false positives.</p><p><strong>Rule of thumb:</strong> only use this if there are no better options left.</p><p></p><h2>3. Use Another Metric</h2><p>Some metrics are <strong>naturally noisy</strong> (e.g., revenue per user). Others are <strong>cleaner</strong> (e.g., conversion rate, click-through).</p><ul><li><p>If you want speed, choose a <strong>lower-variance proxy metric</strong> that&#8217;s still tightly linked to your business goal.</p></li><li><p>Example: Instead of testing impact directly on <strong>LTV</strong> (which is long-term, high variance), test on <strong>7-day retention</strong> or <strong>ARPU at 14 days</strong>.</p></li></ul><p>&#128073; The trick: pick a leading indicator that balances <strong>signal</strong> and <strong>speed</strong>.</p><p>This is why companies need a <strong>metric hierarchy</strong>&#8212;a clear understanding of which metrics are leading indicators and which ones are lagging. Without it, you&#8217;ll end up measuring what feels intuitive rather than what&#8217;s both <em>fast</em> and <em>useful</em>.</p><p></p><h2>4. Decrease Variance</h2><p>This is where the real magic happens. Lower variance means you need fewer users for the same level of confidence.</p><p>Here are the tools in your toolbox:</p><h3>a) Outlier Removal</h3><p>Extreme values can balloon variance (especially in revenue data). By trimming the top/bottom 1&#8211;5%, you tighten distributions without losing much signal.</p><p>But be careful&#8212;if you define &#8220;outliers&#8221; incorrectly, you might accidentally cut out a meaningful customer segment and bias your results.</p><h3>b) Stratification</h3><p>Instead of lumping everyone together, break users into groups (e.g., by country, traffic source, or device). This ensures balanced comparisons and reduces noise.</p><p>There are two main flavors:</p><ul><li><p><strong>Pre-stratification</strong>: ensuring randomization is balanced across groups before the test starts (e.g., equal split of mobile vs. desktop).</p></li><li><p><strong>Post-stratification</strong>: reweighting results after the test to correct imbalances that appeared by chance.</p></li></ul><p>&#128073; Caveat: stratification works best for small sample sizes. As your experiment grows, randomization naturally balances groups, and the benefit of stratification fades toward zero.</p><h3>c) CUPED (Controlled Pre-Experiment Data)</h3><p>CUPED is one of those techniques that feels like magic the first time you see it in action. It was introduced by Microsoft researchers, and it works by using <strong>pre-experiment data</strong> to reduce noise in your experiment.</p><p>Here&#8217;s the idea:</p><ul><li><p>For most metrics, user behavior before the experiment strongly predicts behavior during the experiment.</p></li><li><p>Example: If a user spent $100 last month, they&#8217;re more likely to spend a lot this month too, regardless of which variant they&#8217;re in.</p></li><li><p>That prior information is <em>not</em> related to the treatment, but it explains part of the variance in outcomes.</p></li></ul><p>CUPED uses this historical data as a <strong>covariate adjustment</strong> in the analysis. Essentially, it says:</p><blockquote><p>&#8220;Let&#8217;s control for what we already know about a user so we can focus only on the treatment effect.&#8221;</p></blockquote><p>&#128073; Why it helps: By explaining away the variance that comes from user history, the leftover &#8220;unexplained variance&#8221; is smaller. That means fewer users are needed to reach significance.</p><p>In practice, CUPED can reduce variance by <strong>20&#8211;40%</strong> depending on the metric and how predictive past behavior is.</p><h3>d) VWE (Variance-Weighted Estimator)</h3>
      <p>
          <a href="/__u/datamarks.substack.com/p/what-actually-speed-up-your-ab-test">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[7 reasons why your experiments fail]]></title><description><![CDATA[When you first start working with experiments, it feels simple:]]></description><link>https://datamarks.substack.com/p/7-reasons-why-your-experiments-fail</link><guid isPermaLink="false">https://datamarks.substack.com/p/7-reasons-why-your-experiments-fail</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 27 Aug 2025 11:55:17 GMT</pubDate><content:encoded><![CDATA[<p>When you first start working with experiments, it feels simple:<br>pick a metric &#8594; randomize users &#8594; apply a statistical test &#8594; get a result &#8594; make a decision.</p><p>But we all know it&#8217;s a lot harder than it looks at first glance.</p><p>Here are 7 of the most common reasons why your experimentation might not be working properly:</p><h3>1. Decentralized experiments = chaos</h3><p>If every team runs tests on their own, you&#8217;ll quickly face:</p><ul><li><p>Conflicting experiments affecting the same users.</p></li><li><p>Overlapping metrics and cannibalization of them</p></li><li><p>Teams cherry-picking results to look good.</p></li></ul><p>Solutions: </p><ul><li><p>Centralization.</p></li></ul><p> <br>A single A/B platform, shared methodology, validated test designs, and a knowledge base of past experiments. This ensures transparency, consistency, and less wasted effort.</p><div><hr></div><h3>2. Skipping experiment design</h3><p>Some people underestimate the importance of experiment design. But without it, you don&#8217;t know:</p><ul><li><p>How long to run your test.</p></li><li><p>Whether you&#8217;ll detect meaningful effects.</p></li><li><p>The actual power of your experiment.</p></li></ul><p>Good design means calculating <strong>MDE (minimum detectable effect)</strong>, <strong>sample size</strong>, defining the <strong>randomization unit</strong>, understanding the <strong>metrics</strong>, and much more. Without this, you risk stopping too early or wasting weeks on inconclusive data.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datamarks.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Marks is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>3. Peeking too early</h3><p>A classic mistake. You see significance, get excited, and want to stop right away. But stopping a test the moment you &#8220;see a win&#8221; increases false positives <strong>tenfold.</strong><br></p><p>Solutions:</p><ul><li><p>General rule: don&#8217;t stop before the pre-determined duration.</p></li><li><p>Only stop early for strong <em>negative</em> signals (e.g. revenue tanks).</p></li><li><p>If your experimentation system is mature, use <strong>sequential testing methods</strong> (like mSPRT) to control error rates.</p></li></ul><div><hr></div><h3>4. Ignoring multiple testing</h3><p>Many data scientists forget that multiple testing is not only about having multiple variants, but it&#8217;s also about metrics. </p><p>We rarely measure just one metric. But if you track 10+ KPIs without corrections, you&#8217;ll almost always find false &#8220;wins.&#8221;</p><p>Solutions: </p>
      <p>
          <a href="/__u/datamarks.substack.com/p/7-reasons-why-your-experiments-fail">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What I actually Do as a Data Scientist.]]></title><description><![CDATA[It&#8217;s always interesting to peek into a data scientist&#8217;s calendar.]]></description><link>https://datamarks.substack.com/p/what-i-actually-do-as-a-data-scientist</link><guid isPermaLink="false">https://datamarks.substack.com/p/what-i-actually-do-as-a-data-scientist</guid><dc:creator><![CDATA[Eltsefon Mark]]></dc:creator><pubDate>Wed, 13 Aug 2025 11:54:17 GMT</pubDate><content:encoded><![CDATA[<p>It&#8217;s always interesting to peek into a data scientist&#8217;s calendar. </p><p>Is it really just wall-to-wall modeling and number crunching? </p><p>Spoiler: not even close. </p><p>In reality, about 70% of my day ( and not only mine ) is spent in discussions with cross-functional partners, aligning on priorities, clarifying metrics, and making sure the work we ship actually moves the needle.</p><p>When I share my schedule here, I&#8217;ll skip most of the &#8220;Discuss X&#8221; or &#8220;Discuss Y&#8221; items (trust me, there are many) and focus on the more distinctive, high-impact activities. Think causal inference puzzles, data sleuthing, experiment design&#8230; and a little table tennis.</p><p>Here&#8217;s what my week actually looks like - from Monday morning coffee to Friday afternoon wrap-up.</p><h3><strong>MONDAY &#8212; Strategy + Sizing</strong></h3><p><strong>8:00 AM &#8212; Breakfast + Inbox Scan</strong><br>Over coffee and avocado toast, review invites draft for the in-person advertiser event. Strategy isn&#8217;t just &#8220;who to invite&#8221; &#8212; it&#8217;s balancing reach, potential revenue uplift, and relationship-building.</p><p><strong>9:00 AM &#8212; Invitation Strategy Deep Dive</strong><br>Segment advertisers by spend potential, engagement, and likelihood to convert post-event. Build a quick simulation to see which maximizes ROI without overfilling the venue.</p><p><strong>10:30 AM &#8212; Opportunity Sizing</strong><br>Run vertical-level revenue models for multiple ad categories. Compare projections to historical campaign lift to validate the numbers.</p><p><strong>12:00 PM &#8212; Lunch with Colleagues</strong><br>Swap notes on last quarter&#8217;s campaign results and &#8212; more importantly &#8212; where our next offsite will be.</p><p><strong>1:00 PM &#8212; Data Engineering Huddle</strong><br>Explain to DE why we need a new metric in the vertical datatable. Clarify definitions to avoid &#8220;two teams, two different metrics&#8221; syndrome.</p><p><strong>3:00 PM &#8212; PM Discussion on Uplift Results</strong><br>Share last quarter&#8217;s experiment outcomes. Brainstorm how to fold the uplift data into next quarter&#8217;s planning, adjusting spend and targeting based on measured impact.</p><p><strong>4:30 PM &#8212; Campaign Analysis Across Regions</strong><br>Evaluate marketing Threads campaigns worldwide. Identify which creative angles work best per market &#8212; and which just drain budget.</p><p></p><h3><strong>TUESDAY &#8212; Data &amp; Causal Inference Lab</strong></h3><p><strong>8:00 AM &#8212; Coffee + Prep</strong><br>List open questions for the day: &#8220;Why does PSM + regression adjustment disagree with DiD?&#8221; tops the list.</p><p><strong>9:00 AM &#8212; Causal Method Comparison</strong><br>Run side-by-side results for PSM+RegAdj vs. DiD. Diagnose differences: sample imbalance? Parallel trends violated? Different covariate coverage?</p>
      <p>
          <a href="/__u/datamarks.substack.com/p/what-i-actually-do-as-a-data-scientist">
              Read more
          </a>
      </p>
   ]]></content:encoded></item></channel></rss>